Transferrin fusions and uses thereof

By integrating a therapeutic transgene into the transferrin locus using CRISPR/Cas9, the method addresses the limitations of ERT by achieving efficient and sustained expression of transferrin fusions, improving treatment efficacy for lysosomal storage disorders.

WO2025227229A1PCT designated stage Publication Date: 2025-11-06UNIVERSITE LAVAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2025/050607
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-01
Filing Date
2025-04-28
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Current enzyme replacement therapies for lysosomal storage disorders, such as enzyme replacement therapy (ERT), are not curative, require lifelong intravenous injections, and fail to efficiently reach target tissues, including the brain, due to low efficacy and limited tissue penetration.

Method used

A method for expressing a polypeptide of interest as a transferrin fusion via targeted modification of the transferrin locus in a cell, using CRISPR/Cas9-based genome editing to integrate a therapeutic transgene into the transferrin gene, allowing for high-level expression and secretion of the fusion protein, including across the blood-brain barrier.

Benefits of technology

The method enables robust and widespread production and secretion of active therapeutic enzymes, effectively treating lysosomal storage disorders like Hurler syndrome and Pompe disease by enhancing tissue penetration and maintaining enzyme activity in serum and various tissues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2025050607_06112025_PF_FP_ABST
    Figure CA2025050607_06112025_PF_FP_ABST
Patent Text Reader

Abstract

Described herein are methods of expressing a polypeptide of interest, e.g., a therapeutic polypeptide or protein, via genome editing to insert a therapeutic transgene into the locus encoding transferrin via a nuclease-based method (e.g., CRISPR / Cas9), for expression of the therapeutic polypeptide as a fusion to transferrin. Uses of such an approach to modify the transferrin locus of the cell of a subject to thereby produce the polypeptide of interest (either as a transferrin fusion or first expressed as a transferrin fusion and then cleaved to be in a form lacking the transferrin domain) in the subject are also described, such as for the prevention or treatment of a disease or condition associated with a deficiency in the therapeutic polypeptide or protein. Such diseases or conditions may for example include a metabolic disease, such as a lysosomal disease. In embodiments, the disease or condition is Hurler syndrome or Pompe disease.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TRANSFERRIN FUSIONS AND USES THEREOF

[0002] CROSS REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Application Serial No. 63 / 641 ,239, filed on May 1 , 2024, which is incorporated herein by reference in its entirety.

[0004] REFERENCE TO SEQUENCE LISTING

[0005] This application contains a Sequence Listing in computer readable form entitled “G11229-00485.xml”, created April 28, 2025 and having a size of about 244,000 bytes. The computer readable form is incorporated herein by reference.

[0006] FIELD OF THE DISCLOSURE

[0007] The present disclosure generally relates to expression of polypeptides and uses thereof, and more specifically relates to expression of polypeptides prepared as transferrin fusion polypeptides via targeted modification of the transferrin locus in a cell.

[0008] BACKGROUND OF THE DISCLOSURE

[0009] Numerous genetic disorders exist, with genetic mutations resulting in the loss of such function or activity and in turn result in disease. For example, inborn errors of metabolism (lEMs) are a group of genetic disorders caused by disruptions in a specific enzymatic reaction within a metabolic pathway that induce pathology through toxic metabolite accumulation or deficiencies in downstream metabolites1. Among lEMs, lysosomal storage disorders (LSDs) form a diverse group of progressive and often severe diseases in which lysosomal functions are affected due to the loss of an enzyme that degrades a specific substrate within the lysosome. Current treatments for some LSDs include enzyme replacement therapy (ERT), in which patients receive frequent intravenous injections of recombinant enzymes. This approach relies on the unique pathway by which these enzymes reach the lysosome. Most soluble lysosomal enzymes acquire a mannose 6-phosphate (M6P) moiety during processing in the Golgi. M6P-modified proteins bind to the M6P receptor (M6PR), which transports most of the enzyme produced by a cell to its lysosomes. However, a small fraction (-10%) of the enzyme is secreted even though it contains M6P. Adjacent or distant cells can take up the secreted enzyme since the M6PR also appears on the surface of many cells, a process referred to as cross-correction. These treatments are lifelong, not curative, and fail to prevent longterm disease progression2. Moreover, these recombinant enzymes do not efficiently reach target tissues when injected intravenously and typically fail to cross the blood-brain barrier and reach the brain2’3.

[0010] There thus remains a need for improved therapeutic approaches, such as for genetic disorders.

[0011] SUMMARY OF THE DISCLOSURE

[0012] The present disclosure generally relates to expression of polypeptides and uses thereof, and more specifically relates to expression of polypeptides via transferrin fusions using targeted modification of the transferrin locus in a cell. In various aspects and embodiments, the present disclosure provides the following items:

[0013] 1 . A method for expressing a polypeptide of interest in a cell, the method comprising: preparing an expression construct, the expression construct encoding a fusion polypeptide comprising a first domain comprising transferrin or a fragment or derivative thereof (e.g., having transferrin activity) and a second domain comprising the polypeptide of interest, wherein preparing the expression construct comprises introducing, using a nuclease-based method, a first nucleic acid comprising a first nucleotide sequence encoding the polypeptide of interest into a target region within an endogenous transferrin gene of the cell; and allowing expression of the fusion polypeptide from the expression construct.

[0014] 2. The method of item 1 , wherein the second domain is C-terminal to the first domain.

[0015] 3. The method of item 1 or 2, wherein the fusion polypeptide comprises a linker region between the first and second domains.

[0016] 4. The method of any one of items 1-3, wherein the target region is within an intron of the transferrin gene.

[0017] 5. The method of item 4, wherein the target region is within an intron other than intron 1 of the transferrin gene.

[0018] 6. The method of item 4 or 5, wherein the target region is within intron 16 of the transferrin gene.

[0019] 7. The method of any one of items 4-6, wherein the first nucleic acid sequence further comprises, 5’ to the first nucleotide sequence, one or more exons of the transferrin gene 3’ to the target region.

[0020] 8. The method of item 7, wherein the target region is within intron 16 of the transferrin gene and wherein the one or more exons of the transferrin gene 3’ to the target region is exon 17 of the transferrin gene, and wherein the first nucleic acid comprises, in a 5’-3’ direction, exon 17 of the transferrin gene and the first nucleotide sequence.

[0021] 9. The method of any one of items 3-8, wherein the first nucleic acid further comprises a linker nucleotide sequence encoding the linker region, wherein the linker nucleotide sequence is located between the one or more exons of the transferrin gene 3’ to the target region and the first nucleotide sequence, such that the first nucleic acid comprises, in a 5’-3’ direction, (i) the one or more exons of the transferrin gene 3’ to the target region, (ii) the linker nucleotide sequence, and (iii) the first nucleotide sequence.

[0022] 10. The method of item 9, wherein the target region is within intron 16 of the transferrin gene, wherein the one or more exons of the transferrin gene 3’ to the target region is exon 17 of the transferrin gene and wherein the first nucleic acid comprises, in a 5’ -3’ direction, (i) exon 17 of the transferrin gene, (ii) the linker nucleotide sequence, and (iii) the first nucleotide sequence.

[0023] 11. The method of any one of items 1-10, wherein the first nucleotide sequence lacks a sequence encoding a signal sequence endogenous to the polypeptide of interest.

[0024] 12. The method of any one of items 1-11 , wherein the expressed fusion polypeptide is not cleaved between the first and second domains, thereby to produce the polypeptide of interest comprised within the fusion polypeptide. 13. The method of any one of items 1-11 , wherein the expressed fusion polypeptide is cleaved between the first and second domains, thereby to produce the polypeptide of interest in a form lacking the transferrin or fragment or derivative thereof (e.g., having transferrin activity).

[0025] 14. The method of any one of items 1-13, wherein the the method comprises providing the cell with (a) a CRISPR nuclease or a nucleic acid encoding the CRISPR nuclease, (b) one or more gRNAs comprising one or more guide sequences having one or more target sequences within the target region of the transferrin gene, or one or more nucleic acids encoding the one or more gRNAs, wherein the one or more target sequences are each contiguous to a protospacer adjacent motif recognized by the CRISPR nuclease, and (c) one or more donor nucleic acids comprising the first nucleic acid, wherein the one or more gRNAs direct the cleavage of the transferrin gene at the target region thereby to allow introduction of the first nucleic acid at the target region.

[0026] 15. The method of item 14, wherein the method comprises providing the cell with one or more vectors comprising (a) the nucleic acid encoding the CRISPR nuclease, (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids comprising the first nucleic acid.

[0027] 16. The method of item 15, wherein the method comprises providing the cell with a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and a further vector comprising (c) the one or more donor nucleic acids comprising the first nucleic acid.

[0028] 17. The method of any one of items 14-16, wherein the CRISPR nuclease is a Cas9 nuclease.

[0029] 18. The method of any one of items 14-17, wherein the first nucleic acid is introduced by homology-directed repair (HDR).

[0030] 19. The method of any one of items 1-15, wherein the polypeptide of interest has an activity which is deficient in the cell prior to expression of the polypeptide of interest.

[0031] 20. The method of any one of items 1-19, wherein the polypeptide of interest is a therapeutic protein.

[0032] 21 . The method of any one of items 1-20, wherein the polypeptide of interest is an enzyme.

[0033] 22. The method of item 21 , wherein the enzyme is a lysosomal enzyme.

[0034] 23. The method of item 21 or 22, wherein the enzyme is an alpha-L-iduronidase (IDUA), an acid alphaglucosidase (GAA), an iduronate 2-sulfatase, a heparan sulfamidase, an N-acetyl-alpha- glucosaminidase, a heparane-alpha-glucosaminide N-acetyltransferase, a glucosamine N-acetyl-6- sulfatase, or a betaglucuronidase.

[0035] 24. The method of any one of items 1-23, wherein the cell is a hepatic cell.

[0036] 25. The method of any one of items 6-24, wherein the target region is within SEQ ID NO: 20 (FIG. 9).

[0037] 26. The method of any one of items 1-25, wherein the method an in vitro method.

[0038] 27. The method of any one of items 1-25, wherein the method is an in vivo method and the cell is within a subject. 28. The method of item 27, wherein the polypeptide of interest is secreted from the cell.

[0039] 29. The method of item 27 or 28, wherein the polypeptide of interest is present in a tissue or body fluid of the subject following its expression.

[0040] 30. The method of item 29, wherein the polypeptide of interest is present in the serum of the subject following its expression.

[0041] 31 . The method of item 29, wherein the polypeptide of interest is present in the nervous system of the subject following its expression.

[0042] 32. One or more gRNAs as defined in any one of items 14-31 .

[0043] 33. One or more donor nucleic acids as defined in any one of items 14-32.

[0044] 34. An isolated nucleic acid comprising the expression construct as defined in any one of items 1-31 .

[0045] 35. An isolated polypeptide comprising the amino acid sequence of the fusion polypeptide as defined in any one of items 1-31.

[0046] 36. A vector comprising one or more nucleic acid sequences corresponding to the one or more gRNAs of item 32.

[0047] 37. The vector of item 36, further comprising a nucleic acid encoding a CRISPR nuclease.

[0048] 38. The vector of item 37, wherein the CRISPR nuclease is a Cas9 nuclease.

[0049] 39. A vector comprising the one more donor nucleic acids of item 33.

[0050] 40. A cell comprising the one or more gRNAs of item 32, the one or more donor nucleic acids of item 33, the expression construct as defined in any one of items 1-31 , the isolated polypeptide of item 35, and / or the vector of any one of items 36-39.

[0051] 41 . A composition comprising the one or more gRNAs of item 32, the one or more donor nucleic acids of item 33, the expression construct as defined in any one of items 1-31, the isolated polypeptide of item 35, the vector of any one of items 36-39, and / or the cell of item 40.

[0052] 42. The composition of item 41 , further comprising a pharmaceutically acceptable carrier.

[0053] 43. A method of preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest, comprising expressing the polypeptide interest in a cell of the subject according to the method of any one of items 1-31.

[0054] 44. A method of preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest, comprising administering to the subject an effective amount of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of items 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of item 33; or the composition of item 41 or 42.

[0055] 45. The method of item 44, comprising administering to the subject an effective amount of a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and a further vector comprising (c) the one or more donor nucleic acids.

[0056] 46. The method of any one of items 43-45, wherein the one or more vectors is / are a viral vector.

[0057] 47. The method of any one of items 43-46, wherein the disease or condition is a metabolic disease.

[0058] 48. The method of any one of items 43-47, wherein the disease or condition is a lysosomal disease.

[0059] 49. The method of any one of items 43-48, wherein the disease or condition is a mucopolysaccharidosis

[0060] (MPS).

[0061] 50. The method of item 49, wherein the disease or condition is a type I, II, III or VI MPS.

[0062] 51 . The method of any one of items 43-50, wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

[0063] 52. The method of any one of items 43-47, wherein the disease or condition is a glycogen storage disease.

[0064] 53. The method of any one of items 43-47 and 52, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

[0065] 54. One or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of items 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of item 33; or the composition of item 41 or 42, for use in preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

[0066] 55. The one or more vectors or composition for use of item 54, wherein the one or more vectors comprise (i) a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (ii) a further vector comprising (c) the one or more donor nucleic acids.

[0067] 56. The one or more vectors or composition for use of item 54 or 55, wherein the one or more vectors is / are a viral vector.

[0068] 57. The one or more vectors or composition for use of any one of items 54-56, wherein the disease or condition is a metabolic disease.

[0069] 58. The one or more vectors or composition for use of any one of items 54-57, wherein the disease or condition is a lysosomal disease.

[0070] 59. The one or more vectors or composition for use of any one of items 54-58, wherein the disease or condition is a mucopolysaccharidosis (MPS).

[0071] 60. The one or more vectors or composition for use of item 59, wherein the disease or condition is a type I, II, III or VI MPS.

[0072] 61 . The one or more vectors or composition for use of any one of items 54-60, wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

[0073] 62. The one or more vectors or composition for use of any one of items 54-57, wherein the disease or condition is a glycogen storage disease. 63. The one or more vectors or composition for use of any one of items 54-57 and 62, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

[0074] 64. Use of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of items 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of item 33; or the composition of item 41 or 42, for preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

[0075] 65. Use of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of items 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of item 33; or the composition of item 41 or 42, for the preparation of one or more medicaments for preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

[0076] 66. The use of item 64 or 65, wherein the one or more vectors comprise (i) a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (ii) a further vector comprising (c) the one or more donor nucleic acids.

[0077] 67. The use of any one of items 64 or 66, wherein the one or more vectors is / are a viral vector.

[0078] 68. The use of any one of items 64-67, wherein the disease or condition is a metabolic disease.

[0079] 69. The use of any one of items 64-68, wherein the disease or condition is a lysosomal disease.

[0080] 70. The use of any one of items 64-69, wherein the disease or condition is a mucopolysaccharidosis (MPS).

[0081] 71. The use of item 70, wherein the disease or condition is a type I, II, III or VI MPS.

[0082] 72. The use of any one of items 64-71 , wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

[0083] 73. The use of any one of items 64-69, wherein the disease or condition is a glycogen storage disease.

[0084] 74. The use of any one of items 64-69 and 73, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

[0085] Other objects, advantages, and features of the present disclosure will become more apparent upon reading of the following non-restrictive description of specific embodiments thereof, given by way of example only with reference to the accompanying drawings.

[0086] BRIEF DESCRIPTION OF DRAWINGS

[0087] In the appended drawings:

[0088] FIG. 1 : Schematic of the transferrin targeting strategy. A protein replacement strategy by insertion and fusion of a therapeutic transgene at the murine transferrin locus is shown. In this strategy, a donor template allowing for the integration of a cassette comprising a splicing acceptor sequence (SA), the coding sequence of the last exon of the murine transferrin locus, a linker, the coding sequence of a therapeutic transgene (CDS), two homology arms allowing targeting to the last intron of the murine Trf locus (HA-L and HA-R), a polyadenylation sequence (pA) and flanked by inverted terminal repeat sequences (ITRs) is delivered to the mouse liver. A Cas9 nuclease and a guide RNA targeting the intron create a double-strand DNA break facilitating the integration of the donor construct by homology-directed repair (HDR) or non-homologous end joining (NHEJ) of the full vector genome. Once the cassette is integrated, a fusion messenger RNA and protein will be produced. As transferrin is a highly expressed protein that bears a secretion signal peptide at its N-terminus, the fusion protein will be secreted by the liver.

[0089] FIGs. 2A-B: Long-term production and secretion of enzymatically active transferrin-IDUA fusion proteins by the liver. A) Neonatal (2 days old) ldua pups were injected into the retro-orbital sinus with saline or a combination of either 1E11 vector genomes (VGs) of recombinant adeno-associated viral vector (rAAV) allowing the expression of St1Cas9 and a single-guide RNA targeting the last intron of the Trf locus and 5E11 VGs of an rAAV donor vector for targeted integration at the transferrin locus as described in Figure 1 (transferrin) or 1E11 VGs of a nuclease rAAV vector targeting the first intron of the Alb locus and 5E11 VGs of a donor vector for targeted integration at the albumin locus as described by Sharma et al (albumin). Age-matched C57BL / 6J and Idua'- mice injected with saline were included as controls. The animals were sacrificed 6 months after injection. Western blot analysis was conducted with an antibody directed against human IDUA on serum, liver and heart samples. The IDUA proteins of the expected size are annotated (free IDUA or fused Trf-IDUA). Stain-free gel imaging is shown as the loading control for each blot. B) IDUA enzyme activity was determined for serum, liver, brain and heart samples of the animals described in A). Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles).

[0090] FIGs. 3A-B: Long-term production and secretion of enzymatically active transferrin-GAA fusion proteins by the liver. A) Neonatal (2 days old) Gaa4pups were injected into the retro-orbital sinus with a combination of 1E11 vector genomes (VGs) of recombinant adeno-associated viral vector (rAAV) allowing the expression of St1Cas9 and a single-guide RNA targeting the last intron of the Trf locus and 5E11 VGs of an rAAV donor vector for targeted integration at the transferrin locus as described in Figure 1. Age-matched C57BL / 6J and Gaa7' mice injected with saline were included as controls. The animals were sacrificed 8 months after injection. Western blot analysis was conducted with an antibody directed against human GAA on serum, liver, brain, heart, leg muscle and diaphragm samples from those animals. The GAA proteins of the expected size are annotated (processed lysosomal GAA or fused Trf-GAA. Stain-free gel imaging is shown as the loading control for each blot. B) GAA enzyme activity was determined for serum, liver, brain, heart, leg muscle and diaphragm samples of the animals described in 3A). Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles).

[0091] FIGs. 4A-C: Long-term treatment of Gaa'- animals with the transferrin strategy improves muscle function and corrects cardiomegaly. Muscle function from the animals described in Figure 3 was assessed at 6 months post-injection by rotarod (A) and grip strength (B) assays. For rotarod assays, the median latency to fall of 4 trials was determined for each animal. For grip strength assays, the peak force of each animal was measured in triplicate. (C) The heart-to-body weight ratio was determined for each animal post-sacrifice at 8 months post-injection. One male animal per group was sacrificed separately and excluded from analysis. Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles). The mean and standard error of the mean (SEM) are shown for each group.

[0092] FIG. 5: Screening for guide RNAs active at the transferrin locus. A schematic of the transferrin locus and the target region is shown. Mouse hepatoma Hepa1-6 cells (2 x 105per transfection) were transfected with 500ng of vectors encoding variants of a Cas9 nuclease from Streptococcus thermophiius (St1 Cas9) and a guide RNA for which the respective spacer sequences are shown, using an Amaxa 4D Nucleofector™ with the SF Cell Line kit and the EX-147 program. Cells were harvested 72h post-transfection and the editing activity was assessed using the TIDE assay. The results shown are from three independent experiments. In bold is the result of the guide RNA chosen for in vivo studies. Guide sequences in order from top to bottom correspond to SEQ ID NOs: 1-6.

[0093] FIGs. 6a-f: Purification and validation of an enzymatically active transferrin-IDUA fusion protein, (a) Recombinant transferrin-a-L-iduronidase (transferrin-IDUA), transferrin (b) and IDUA (c) were produced and secreted in 293F cells before purification. The various fractions of a purification procedure done with StrepTactin™ XT resin were loaded on a 8% polyacrylamide gel containing 0,5% (v / v) 2,2, 2-trich loroeth anol before migration and visualisation with a ChemiDoc™ MP imager. Initial, chosen fraction from the first purification step; FT, flow-through of the initial fraction after passage on the purification column; washes, wash fractions 1 and 5; elutions, elution fractions 1 to 3. (d) IDUA enzymatic activity of the proteins purified in (a), (b) and (c). The columns indicate the mean and the error bars indicate the standard deviation of three technical replicates, (e) Western blot to detect the recombinant transferrin-IDUA fusion protein. Liver homogenates from C57 / BI6J or Idua- / - mice were spiked with the indicated amounts of recombinant transferrin-IDUA fusion protein before loading on an acrylamide gel as described in (a), (b) and (c). The lower panel shows the total proteins as visualised on the gel with a ChemiDoc™ MP imager. The top panel shows the resulting blot with an antibody directed against human IDUA. (f) Same as (e) but for purified recombinant IDUA.

[0094] FIG. 7: Robust and widespread IDUA activity in Idua'- animals treated with the albumin and transferrin strategies and sacrificed at one month of age. Neonatal idua '- pups were injected with either saline, with 5 x 1011vector genomes (VGs) of a donor rAAV8 vector targeting an iDUA transgene either to the first intron of the albumin locus (albumin donor only) or the transferrin locus (transferrin donor only) or with a combination of the donor vector and 1 x 1011VGs of a nuclease rAAV8 vector driving the expression of St1 Cas9 and the corresponding sgRNA (albumin nuclease + donor, transferrin nuclease + donor). A group of age-matched C57BL / 6J animals were also included as controls. IDUA enzyme activity was determined for the serum, liver, brain and heart samples of these animals. Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles).

[0095] FIGs. 8A-C: Editing activity and detection of targeted integration of the IDUA donor constructs in Idua'- mice at 6 months of age. A) TIDE analysis was conducted on whole-liver genomic DNA from the animals described in Figure 2. Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles). A mouse injected with saline was used as the negative control for the TIDE assay. B) Targeting strategies and integration outcomes are respectively shown for the transferrin (Trf) and albumin {Alb) approaches. Also annotated are the primers used for “in-out” PCR (P1-P3). SA, splice acceptor sequence; IDUA CDS, coding sequence of the human a-L-iduronidase gene; pA, polyadenylation sequence; ITR, adeno-associated viral inverted terminal repeats. C) “In-out” PCR products, where one primer binds a sequence in the mouse genome located outside of the homology arms of the construct and a second primer binds a sequence inside the donor construct, are shown for each target. Annotated are the bands of expected size corresponding to integration of the donor construct by homology-directed repair (HDR) or direct integration of the rAAV donor vector by non-homologous end joining (NHEJ).

[0096] FIGs. 9A-B: Nuclease-driven integration at the transferrin locus allows for long-term production and secretion of enzymatically active transferrin-IDUA fusion proteins by the liver. A) Western blot analysis with an antibody directed against human IDUA on the serum (left) and liver samples (right) of C57 / BI6J mice, Iduad- mice injected with saline or Iduad- mice injected with both nuclease and donor vectors for the albumin or transferrin targeting strategy. The IDUA proteins of the expected size are annotated. Stain-free imaging after transfer is shown as the loading control for each membrane. B) IDUA enzyme activity was determined for serum, liver and brain samples of the animals described in Figure 8. Each symbol represents a different animal. Females (f; closed symbols), Males (m; open symbols).

[0097] FIG. 10: Robust and widespread GAA activity in Gaa'- animals treated with the transferrin targeting strategy and sacrificed at one month of age. Neonatal (2 days old) Gaa'- pups were injected into the retro-orbital sinus with a combination of 1E11 vector genomes (VGs) of recombinant adeno-associated viral vector (rAAV) allowing the expression of St1Cas9 and a single-guide RNA targeting either the last intron of the Trf locus and 5E11 VGs of an rAAV donor vector for targeted integration of a human GAA transgene, or with the donor vector alone, and sacrificed at 1 month of age. Age-matched groups of untreated Gaa4and C57BL / 6J animals were used as controls. GAA enzyme activity was determined for serum, liver, brain, heart, leg muscle (gastrocnemius) and diaphragm samples of the animals described above. Each symbol represents a different animal. Females (f; closed circles), Males (m; open circles)

[0098] FIGs. 11A-C: pAAV_mTrfJ16_hlDUA_donor sequence (SEQ ID NO: 17).

[0099] FIGs. 12A-D: pAAV_mTrfJ16_hGAA_donor sequence (SEQ ID NO: 18).

[0100] FIG. 13: Intron 16 sequence of human transferrin gene (SEQ ID NO: 20).

[0101] FIG. 14: cDNA sequence of human transferrin (SEQ ID NO: 21).

[0102] FIG. 15: Amino acid sequence of human transferrin (SEQ ID NO: 22). Positions 1-19 correspond to the signal sequence (SEQ ID NO: 23). Positions 20-698 are shown in bold and correspond to the mature protein sequence (SEQ ID NO: 24).

[0103] FIG. 16: Mouse transferrin intron 16 sequence (SEQ ID NO: 25).

[0104] FIG. 17: Mouse transferrin cDNA sequence (SEQ ID NO: 26). FIG. 18: Mouse transferrin amino acid sequence (SEQ ID NO: 27).

[0105] FIG. 19: Human IDUA amino acid sequence (SEQ ID NO: 28). Signal peptide / pre-peptide, absent from Trf fusion construct is shown in italics (SEQ ID NO: 29). IDUA precursor protein sequence found in Trf fusion construct is shown in bold (SEQ ID NO: 30).

[0106] FIG. 20: Human IDUA cDNA sequence (SEQ ID NO: 31). 5’ and 3’ UTR shown in lowercase. Natural IDUA sequence encoding the signal peptide / pre-peptide shown in uppercase italic. IDUA precursor coding sequence codon-optimized by DNA2.0 shown in uppercase bold (SEQ ID NO: 32).

[0107] FIG. 21 : Amino acid sequence of Trf-IDUA fusion (SEQ ID NO: 33) after successful integration. Mouse transferrin shown in normal uppercase. (GGGS)4 linker is underlined. IDUA precursor protein sequence (that does not contain the IDUA signal peptide / pre-peptide) is shown in bold.

[0108] FIGs. 22A-B: Trf-IDUA donor DNA sequence (integration cassette contained between the two AAV ITRs) (SEQ ID NO: 34). Homology arms to intron 16 of the mouse transferrin gene and restriction cloning sites are highlighted. Splice acceptor sequence shown in bold italic. Last exon of the mouse transferrin gene shown in normal uppercase. (GGGS)4 linker is underlined. IDUA precursor coding sequence (shown in bold) that does not contain the IDUA signal peptide / pre-peptide. This sequence was codon-optimized by DNA 2.0 Restriction cloning site and poly-A sequence is highlighted and underlined.

[0109] FIG. 23: Human GAA amino acid sequence (SEQ ID NO: 35). Signal peptide / pre-peptide absent from the Trf fusion construct is shown in italics (SEQ ID NO: 36). Pro-peptide absent from the Trf fusion construct shown in underlined italics (SEQ ID NO: 37). GAA precursor protein sequence found in Trf fusion construct shown in bold (SEQ ID NO: 38).

[0110] FIGs: 24A-B: Human GAA cDNA sequence (SEQ ID NO: 39). 5’ and 3’ UTR shown in lowercase. GAA sequence encoding the signal peptide / pre-peptide shown in uppercase italic. GAA sequence encoding the pro-peptide shown in underlined italic. GAA precursor coding sequence (bold; SEQ ID NO: 40) that does not contain the GAA signal peptide / pre-peptide and pro-peptide. This DNA sequence is not codon-optimized and corresponds to Genbank BC040431.1 :

[0111] FIG. 25: Trf-GAA amino acid sequence after successful integration (SEQ ID NO: 41). Mouse transferrin shown in normal uppercase. (GGGS)4 linker is underlined. GAA precursor protein sequence (that does not contain GAA signal peptide / pre-peptide or pro-peptide) is shown in bold.

[0112] FIGs. 26A-B: Trf-GAA donor DNA sequence (integration cassette contained between the two AAV ITRs) (SEQ ID NO: 42). Homology arms to intron 16 of the mouse transferrin gene and restriction cloning sites are highlighted. Last exon of the mouse transferrin gene in normal uppercase. (GGGS)4 linker is underlined. GAA precursor coding sequence that does not contain the GAA signal peptide / pre-peptide and pro-peptide is shown in bold. This DNA sequence corresponds to Genbank BC040431 .1 . Restriction cloning site and poly-A sequence is highlighted and underlined. DETAILED DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS

[0113] Described herein are methods of expressing a polypeptide of interest, e.g., a therapeutic polypeptide, via genome editing to insert a therapeutic transgene into the locus encoding transferrin, so that the therapeutic polypeptide is expressed as a fusion to transferrin.

[0114] Transferrin is the main transporter of iron in the blood, is highly expressed and secreted by the liver, has a long serum half-life and is known to be able to cross the blood-brain barrier through transcytosis4 5. Transferrin enters the cells via receptor-mediated endocytosis through the transferrin receptor which is ubiquitously expressed. In addition, fusion of therapeutic peptides to transferrin has also been employed as a strategy to increase their serum half-lives6.

[0115] The fusion strategy described herein can be used to produce functional fusion proteins and allow greater tissue penetration from the circulation than the therapeutic protein alone, thus providing a potent approach to correct lEMs such as lysosomal storage disorders. The studies described herein show that the targeting of therapeutic transgenes encoding lysosomal enzymes at the transferrin locus allows for high levels of active transferrin fusion proteins in serum and several tissues in animal models of Hurler syndrome and Pompe disease.

[0116] Accordingly, in an aspect, the present disclosure relates to a method for expressing a polypeptide of interest (e.g., a therapeutic polypeptide or protein) in a cell, the method comprising : providing or preparing an expression construct, the expression construct encoding a fusion polypeptide comprising a first domain comprising transferrin or a fragment or derivative thereof (e.g., having transferrin activity) and a second domain comprising the polypeptide of interest, wherein preparing the expression construct comprises introducing, using a nuclease-based method, a first nucleic acid comprising a first nucleotide sequence encoding the polypeptide of interest into a target region with an endogenous transferrin gene of the cell; and allowing expression of the fusion polypeptide from the expression construct.

[0117] The genomic DNA sequence of human transferrin is provided in SEQ ID NO: 19 (GenBank NG_013080.3), containing exons 1-17 and introns 1-16. The amino acid sequence of human transferrin (SEQ ID NO: 22) is shown in FIG. 11 , with positions 1-19 corresponding to the signal sequence (SEQ ID NO: 23) and positions 20-698 corresponding to the mature protein sequence (SEQ ID NO: 24).

[0118] In embodiments, the fragment or derivative of transferrin (e.g., having transferrin activity) comprises an amino acid sequence having at least 70, 75, 80, 85, 90, or 95% identity with the amino acid sequence of SEQ ID NO: 24.

[0119] In embodiments, transferrin activity comprises iron-binding activity and / or transferrin receptor binding activity.

[0120] In embodiments, the fusion polypeptide comprises a linker region between the first and second domains, i.e., between the transferrin or a fragment or derivative thereof and the polypeptide of interest. In embodiments, the linker region is at least 4, 5, 6, 7, 8, 9, 10, 12, 14, 18, 20, 25 or 30 amino acids in length. In embodiments the linker region is up to 40, 35, 30, 25 or 20 amino acids in length. In embodiments, the linker region is 5-25, in a further embodiment 10-20 amino acids in length. In an embodiment, the linker region comprises a repeating amino acid motif, such as a GGGS or GGGGS motif. In embodiments, the repeating amino acid motif is repeated 2, 3, 4, 5, or 6 times. In embodiments the linker is a cleavable linker, such as an in vivo cleavable linker.

[0121] In an embodiment, in the fusion polypeptide, the polypeptide of interest is fused to the C-terminal of transferrin, optionally via a linker region. In such a case, the expression construct comprises exons 1-17 of the transferrin gene 5’ to the nucleotide sequence encoding the polypeptide of interest. Such an expression construct may in embodiments be prepared via targeted insertion of a transgene at an intron of the transferrin gene. In the case of such insertion within intron 16, the transgene comprises exon 17, optionally a nucleotide sequence encoding the linker region, and the nucleotide sequence encoding the polypeptide of interest. As such, the endogenous exon 17 of the transferrin locus is replaced by an exon 17 inserted via the transgene.

[0122] In embodiments, the transgene may be inserted at a different target region of the transferrin gene, for example at an intron, and similarly the exons 3’ to the target region may be reintroduced via the transgene to generate a C- terminal fusion.

[0123] In embodiments, in cases where the polypeptide of interest in its endogenous form would be expressed with a signal sequence, it may be expressed in the fusion polypeptide without its endogenous signal sequence, with secretion rather being achieved via the transferrin signal sequence of the fusion polypeptide.

[0124] In embodiments, the polypeptide of interest may be retained in the transferrin fusion polypeptide in an active form, or may be cleaved from the transferrin fusion polypeptide to be expressed in a form lacking the transferrin domain.

[0125] As described herein, the transferrin portion of the fusion polypeptide may be the full mature form of transferrin or a fragment or derivative thereof having transferrin activity, for example the ability to bind iron and to bind to the transferrin receptor.

[0126] In embodiments, the nuclease-based method for inserting the transgene to prepare the expression construct for the fusion polypeptide within a cell is a CRISPR / Cas9 method, entailing the use of one or more gRNAs comprising one or more guide sequences having one or more target sequences within the target region of the transferrin gene, the one or more target sequences each being contiguous to a protospacer adjacent motif recognized by the CRISPR nuclease.

[0127] In embodiments, one or more donor or patch nucleic acids may be used, comprising the transgene to be inserted, comprising the nucleotide sequence encoding the polypeptide of interest, optionally also comprising an upstream linker-encoding nucleotide sequence as well optionally e.g., one or more exons of the transferrin gene, so that the fusion polypeptide is expressed as a C-terminal fusion to transferrin, optionally via a linker.

[0128] In embodiments, integration of the transgene may be via homology-directed repair (HDR) or non-homologous-end- joining (NHEJ).

[0129] In embodiments, the cell is a mammalian cell, such as a human cell. In an embodiment the cell is a hepatic cell. In embodiments, the methods described herein may be performed in vitro, ex vivo, in vivo (in a subject), or a combination thereof.

[0130] In embodiments, the polypeptide of interest (either as a transferrin fusion or first expressed as a transferrin fusion and then cleaved to be in a form lacking the transferrin domain) may, following its expression, be present in a tissue or body fluid of the subject. In an embodiment, the polypeptide of interest (either as a transferrin fusion or first expressed as a transferrin fusion and then cleaved to be in a form lacking the transferrin domain) may be present in the serum of the subject following its expression. In embodiments, the polypeptide of interest (either as a transferrin fusion or first expressed as a transferrin fusion and then cleaved to be in a form lacking the transferrin domain) may be present in the hepatic tissue of the subject. In embodiments, the polypeptide of interest (either as a transferrin fusion or first expressed as a transferrin fusion and then cleaved to be in a form lacking the transferrin domain) may be present in the nervous system of the subject.

[0131] The present disclosure also relates to one or more gRNAs and one or more donor nucleic acids described herein.

[0132] The present disclosure also relates to one or more vectors comprising nucleic acid sequence(s) corresponding to one or more components described herein, such as nucleic acid sequences encoding one or more gRNAs as described herein, one or more donor nucleic acids as described herein, an expression cassette as described herein, and / or a nucleic acid sequence encoding a CRISPR nuclease. In embodiments, such components may be present on separate vectors or any combination of two or more of such components may be present on the same vector. For example, such CRISPR / Cas9 methods may entail the use of one or more vectors encoding the gRNA(s), the Cas9 nuclease, and / or the donor or patch nucleic acids. In embodiments, such encoding sequences may be in separate vectors or in the same vector, with different combinations possible. For example, the nucleic acids encoding the gRNA(s) and the nucleic acid encoding the Cas9 nuclease may be comprised in the same vector, and the nucleic acid encoding the donor or patch nucleic acid may be comprised in a separate vector. In embodiments, the vectors may be viral vectors.

[0133] The present disclosure also relates to cells prepared by a method described herein, e.g., comprising an expression construct described herein.

[0134] The present disclosure also relates to cells comprising one or more components described herein, such as nucleic acid sequences encoding one or more gRNAs as described herein, one or more gRNAs as described herein, one or more donor nucleic acids as described herein, an expression cassette as described herein, a nucleic acid sequence encoding a CRISPR nuclease, a CRISPR nuclease, one or more vectors described herein, and / or one or more compositions as described herein.

[0135] The present disclosure also relates to a composition comprising one or more components described herein, such as nucleic acid sequences encoding one or more gRNAs as described herein, one or more gRNAs as described herein, one or more donor nucleic acids as described herein, an expression construct as described herein, a nucleic acid sequence encoding a CRISPR nuclease, a CRISPR nuclease, one or more vectors described herein, and / or one or more cells as described herein. The present disclosure also relates to various methods and uses of one or more components described herein, such as nucleic acid sequences encoding one or more gRNAs as described herein, one or more gRNAs as described herein, one or more donor nucleic acids as described herein, an expression construct or construct as described herein, a nucleic acid sequence encoding a CRISPR nuclease, a CRISPR nuclease, one or more vectors described herein, and / or one or more compositions as described herein. In embodiments, such methods or uses of the one or more components, including for expressing a polypeptide of interest in a cell, preventing or treating a disease, condition, or disorder in a subject, or the preparation of one or more medicaments for preventing or treating a disease, condition, or disorder in a subject.

[0136] The present disclosure also relates to one or more components described herein, such as nucleic acid sequences encoding one or more gRNAs as described herein, one or more gRNAs as described herein, one or more donor nucleic acids as described herein, an expression construct as described herein, a nucleic acid sequence encoding a CRISPR nuclease, a CRISPR nuclease, one or more vectors described herein, and / or one or more compositions as described herein, for various uses, such as for use in expressing a polypeptide of interest in a cell, preventing or treating a disease, condition, or disorder in a subject, or the preparation of one or more medicaments for treating a disease, condition, or disorder in a subject.

[0137] In an embodiment, the polypeptide of interest is a therapeutic protein, such as an enzyme, in further embodiments a lysosomal enzyme (e.g., an alpha-L-iduronidase (IDUA), an acid alpha-glucosidase (GAA), an iduronate 2- sulfatase, a heparan sulfamidase, an N-acetyl-alpha-glucosaminidase, a heparane-alpha-glucosaminide N- acetyltransferase, a glucosamine N-acetyl-6-sulfatase, or a betaglucuronidase).

[0138] In an embodiment, the disease, condition, or disorder is a condition in which the activity of a polypeptide is deficient and the disease, condition, or disorder can benefit from the expression of the polypeptide, such as a polypeptide of interest described herein.

[0139] In embodiments, the disease, condition, or disorder is a metabolic disease. In embodiments, the disease, condition, or disorder is a metabolic lysosomal disease. In embodiments, the disease, condition, or disorder is a mucopolysaccharidosis (MPS; e.g., a type I, II, III or VI MPS). In an embodiment, the disease, condition, or disorder is Hurler syndrome. In embodiments, the disease, condition, or disorder is a glycogen storage disease. In an embodiment, the disease, condition, or disorder is Pompe disease.

[0140] Definitions

[0141] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Generally, nomenclature used in connection with, and techniques of, cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are those well- known and commonly used in the art. The methods and techniques of the disclosure are generally performed according to conventional methods well-known in the art and as described in various general and more specific references that are cited and discussed throughout the specification unless otherwise indicated. See, e.g. : Sambrook J. and Russell D. Molecular Cloning: A Laboratory Manual, 3rded., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (2000); Ausubel F.M. et al., Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology, Wiley, John & Sons, Inc. (2002); Harlow E. and D. Lane, Using Antibodies: A Laboratory Manual; Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y. (1998); and Coligan J.E. et al., Short Protocols in Protein Science, Wiley, John & Sons, Inc. (2003). Any enzymatic reactions or purification techniques are performed according to manufacturer's specifications, as commonly accomplished in the art or as described herein. The nomenclature used in connection with, and the laboratory procedures and techniques of, analytical chemistry, synthetic organic chemistry, and medicinal and pharmaceutical chemistry described herein are those well-known and commonly used in the art.

[0142] The use of the terms "a" and "an" and "the" and similar referents in the context of describing the technology (especially in the context of the claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context.

[0143] The word “a” or “an” when used in conjunction with the term “comprising” in the claims and / or the specification may mean “one”, but it is also consistent with the meaning of “one or more”, “at least one”, and “one or more than one” unless the content clearly dictates otherwise. Similarly, the word “another'’ may mean at least a second or more unless the content clearly dictates otherwise.

[0144] As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “include” and “includes”) or “containing” (and any form of containing, such as “contain” and “contains”), are inclusive or open-ended and do not exclude additional, unrecited elements or process steps.

[0145] Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually recited herein. All subsets of values within the ranges are also incorporated into the specification as if they were individually recited herein. For example, for the range of 18-20, the numbers 18, 19, and 20 are explicitly contemplated, and for the range 6.0-7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are explicitly contemplated. The terms "such as" are used herein to mean, and is used interchangeably with, the phrase "such as but not limited to".

[0146] The use of any and all examples, or exemplary language (“e.g.”, "such as") provided herein, is intended merely to better illustrate the technology and does not pose a limitation on the scope of the claimed invention unless otherwise claimed.

[0147] Herein, the term "about" has its ordinary meaning. The term “about” is used to indicate that a value includes an inherent variation of error for the device or the method being employed to determine the value, or encompass values close to the recited values, for example within 10% of the recited values (or range of values). The terms "nucleic acid," "polynucleotide," and "oligonucleotide" are used interchangeably and refer to a deoxyribonucleotide or ribonucleotide polymer, in linear or circular conformation, and in either single- or doublestranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer. The terms can encompass known analogues of natural nucleotides, as well as nucleotides that are modified in the base, sugar and / or phosphate moieties (e.g., phosphorothioate backbones). In general, an analogue of a particular nucleotide has the same base-pairing specificity; i.e., an analogue of A will base-pair with T.

[0148] The terms "polypeptide," "peptide", and "protein" are used interchangeably to refer to a polymer of amino acid residues. The term also applies to amino acid polymers in which one or more amino acids are chemical analogues or modified derivatives of corresponding naturally-occurring amino acids.

[0149] As used herein, the term “non-conservative mutation” or “non-conservative substitution” in the context of polypeptides refers to a mutation in a polypeptide that changes an amino acid to a different amino acid with different biochemical properties {i.e., charge, hydrophobicity, and / or size). Although there are many ways to classify amino acids, they are often sorted into six main groups, on the basis of their structure and the general chemical characteristics of their R groups, (i) aliphatic (glycine, alanine, valine, leucine, and isoleucine); (ii) hydroxyl- or sulfur / selenium-containing (also known as polar amino acids; serine, cysteine, selenocysteine, threonine, and methionine); (iii) cyclic (proline); (iv) aromatic (phenylalanine, tyrosine, and tryptophan); (v) basic (histidine, lysine, and arginine), and (vi) acidic and their amide (aspartate, glutamate, asparagine, and glutamine). Thus, a non- conservative substitution includes one that changes an amino acid of one group with another amino acid of another group (e.g., an aliphatic amino acid for a basic, a cyclic, an aromatic, or a polar amino acid; a basic amino acid for an acidic amino acid; a negatively charged amino acid [aspartic acid or glutamic acid] for a positively charged amino acid [lysine, arginine, or histidine], etc.)

[0150] Conversely, a “conservative substitution” or “conservative mutation” in the context of polypeptides are mutations that change an amino acid to a different amino acid with similar biochemical properties (e.g. charge, hydrophobicity, and size). For example, a leucine and isoleucine are both aliphatic branched hydrophobic amino acids. Similarly, aspartic acid and glutamic acid are both small, negatively charged amino acids. Therefore, changing a leucine for an isoleucine (or vice versa) or changing an aspartic acid for a glutamic acid (or vice versa) are examples of conservative substitutions.

[0151] "Coding sequence" or "encoding nucleic acid" as used herein means the nucleic acids (RNA or DNA molecule) that comprise a nucleotide sequence which encodes e.g., a polypeptide or protein or a gRNA. The coding sequence can further include initiation and termination signals operably linked to regulatory elements including a promoter and polyadenylation signal capable of directing expression in the cells of an individual or mammal into which the nucleic acid is introduced. The coding sequence may be codon optimized.

[0152] "Complement" or "complementary" as used herein refers to Watson-Crick (e.g., A-T / U and C-G) or Hoogsteen base pairing between nucleotides or nucleotide analogs of nucleic acid molecules. "Complementarity" refers to a property shared between two nucleic acid sequences, such that when they are aligned antiparallel to each other, the nucleotide bases at each position will be complementary.

[0153] Sequence similarity

[0154] "Homology" and “homologous” refers to sequence similarity between two peptides or two nucleic acid molecules. Homology can be determined by comparing each position in the aligned sequences. A degree of homology between nucleic acid or between amino acid sequences is a function of the number of identical or matching nucleotides or amino acids at positions shared by the sequences. As the term is used herein, a nucleic acid sequence is "substantially homologous" to another sequence if the two sequences are substantially identical and the functional activity of the sequences is conserved (as used herein, the term “homologous” does not infer evolutionary relatedness, but rather refers to substantial sequence identity, and thus is interchangeable with the terms “identity” / “identical”). Two nucleic acid sequences are considered substantially identical if, when optimally aligned (with gaps permitted), they share at least about 50% sequence similarity or identity, or if the sequences share defined functional motifs. In alternative embodiments, sequence similarity in optimally aligned substantially identical sequences may be at least 60%, 70%, 75%, 80%, 85%, 90%, or 95%. For the sake of brevity, the units (e.g., 66, 67, ...81 , 82, ...91 , 92%, ...) have not systematically been recited but are considered, nevertheless, within the scope of the present invention.

[0155] Substantially complementary nucleic acids are nucleic acids in which the complement of one molecule is substantially identical to the other molecule. Two nucleic acid or protein sequences are considered substantially identical if, when optimally aligned, they share at least about 70% sequence identity. In alternative embodiments, sequence identity may for example be at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 98%, or at least 99%. Optimal alignment of sequences for comparisons of identity may be conducted using a variety of algorithms, such as the local homology algorithm of Smith and Waterman (Adv. Appl. Math. 2: 482-489, 1981), the homology alignment algorithm of Needleman and Wunsch (J. Mol. Biol. 48: 443-453, 1970), the search for similarity method of Pearson and Lipman (Proc. Natl. Acad. Sci. USA 85: 2444-2448, 1988), and the computerized implementations of these algorithms (such as GAP, BESTFIT, FASTA and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, Madison, Wl, U.S.A.). Sequence identity may also be determined using the BLAST algorithm, described by Altschul et al. (J. Mol. Biol. 215: 403-410, 1990 using the published default settings). Software for performing BLAST analysis may be available through the National Center for Biotechnology Information (through the internet at http: / / www.ncbi.nlm.nih.gov / ). The BLAST algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold. Initial neighborhood word hits act as seeds for initiating searches to find longer HSPs. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Extension of the word hits in each direction is halted when the following parameters are met: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. One measure of the statistical similarity between two sequences using the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. In alternative embodiments of the invention, nucleotide or amino acid sequences are considered substantially identical if the smallest sum probability in a comparison of the test sequences is less than about 1 , preferably less than about 0.1 , more preferably less than about 0.01 , and most preferably less than about 0.001.

[0156] An alternative indication that two nucleic acid sequences are substantially complementary is that the two sequences hybridize to each other under moderately stringent, or preferably, stringent conditions. Hybridization to filter-bound sequences under moderately stringent conditions may, for example, be performed in 0.5 M NaHPCU, 7% sodium dodecyl sulfate (SDS), 1 mM EDTA at 65°C, and washing in 0.2 x SSC / 0.1 % SDS at 42°C (Ausubel, F.M. eta / ., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 2010). Alternatively, hybridization to filter-bound sequences under stringent conditions may, for example, be performed in 0.5 M NaHPCU, 7% SDS, 1 mM EDTA at 65°C, and washing in 0.1 x SSC / 0.1% SDS at 68°C (Ausubel, 2010, supra). Hybridization conditions may be modified in accordance with known methods depending on the sequence of interest (Tijssen P, Hybridization with nucleic acid probes, Part II, Volume 24, 1st Edition, Part II. Probe labeling and hybridization techniques, Elsevier Science, 1993, 344 pages). Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point for the specific sequence at a defined ionic strength and pH.

[0157] "Binding" refers to a sequence-specific, non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid or between a sgRNA and a target polynucleotide or between a sgRNA and a CRISPR nuclease (e.g., Cas9, Cpf1). Not all components of a binding interaction need be sequence-specific (e.g., contacts with phosphate residues in a DNA backbone), as long as the interaction as a whole is sequence-specific. “Affinity” refers to the strength of binding: increased binding affinity being correlated with a lower Kd.

[0158] A "binding protein" is a protein that is able to bind non-covalently to another molecule. A binding protein can bind to, for example, a DNA molecule (a DNA-binding protein), an RNA molecule (an RNA-binding protein), and / or a protein molecule (a protein-binding protein). In the case of a protein-binding protein, it can bind to itself (to form homodimers, homotrimers, etc.) and / or it can bind to one or more molecules of a different protein or proteins. A binding protein can have more than one type of binding activity. For example, zinc finger proteins have DNA- binding, RNA-binding, and protein-binding activity.

[0159] As used herein, “a nuclease-based modification” refers to a modification in a polynucleotide e.g., an endogenous gene locus or genomic sequence) which involves the introduction of a cut (e.g., a double-stranded break in the polynucleotide) which ultimately will trigger a repair mechanism by the cell involving non-homologous-end-joining (NHEJ) or homologous recombination (HDR). The nuclease-based modification is made by site-specific nucleases targeting the polynucleotide of interest ( / .e., an endogenous gene locus or genomic sequence). Site-specific nucleases (engineered) are well known and include (but are not limited to) zinc finger nucleases, meganucleases, Mega-Tals, CRISPR nucleases, TALENs, etc.

[0160] "Recombination" refers to a process of exchange of genetic information between two polynucleotides. For the purposes of this disclosure, "homologous recombination” (HR) refers to the specialized form of such exchange that takes place, for example, during repair of double-strand breaks in cells via homology-directed repair (HDR) mechanisms. This process requires nucleotide sequence homology, uses a "donor" or “patch” nucleic acid as a template for repair of a "target" nucleic ( / .e., the one that experienced the double-strand break), and is variously known as "non-crossover gene conversion" or "short tract gene conversion" because it leads to the transfer of genetic information from the donor to the target. Without wishing to be bound by any particular theory, such transfer can involve mismatch correction of heteroduplex DNA that forms between the broken target and the donor, and / or "synthesis- dependent strand annealing", in which the donor is used to re-synthesize genetic information that will become part of the target, and / or related processes. Such specialized HR often results in an alteration of the sequence of the target molecule such that part or all of the sequence of the donor nucleic acid is incorporated into the target polynucleotide.

[0161] In the methods described herein, one or more targeted (site-specific) nucleases (e.g., sgRNA / CRISPR nuclease) create a double-stranded break in the target sequence (e.g., cellular chromatin) at a predetermined site. A "donor" polynucleotide, having homology to the nucleotide sequence in the region of the break, may be introduced into the cell if desired. The presence of the double-stranded break has been shown to facilitate integration of the donor sequence. The donor sequence may be physically integrated or, alternatively, the donor polynucleotide is used as a template for repair of the break via homologous recombination, resulting in the introduction of all or part of the nucleotide sequence as in the donor into the cellular chromatin. Thus, a first sequence in cellular chromatin can be altered and, in certain embodiments, can be converted into a sequence present in a donor polynucleotide. Thus, the use of the terms "replace" or "replacement" can be understood to represent replacement of one nucleotide sequence by another, ( / .e., replacement of a sequence in the informational sense), and does not necessarily require physical or chemical replacement of one polynucleotide by another. In any of the methods described herein, additional sgRNA / CRISPR nucleases, pair zinc-finger, meganucleases, Mega-Tals, and / or additional TALEN proteins can be used for additional double-stranded cleavage of additional target sites within the cell.

[0162] As used herein, the terms “donor” or “patch” nucleic acid are used interchangeably and refers to a nucleic acid that includes a fragment of the endogenous targeted gene of a cell (in some embodiments the entire targeted gene), but which includes desired modification(s) at specific nucleotides. The donor (patch) nucleic acid must be of sufficient size and similarity (e.g., in the right and left homology arms) to permit homologous recombination with the targeted gene. Preferably, the donor / patch nucleic acid is (or is flanked at the 5’ end and at the 3’ end by sequences) at least 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, or 98% identical to the endogenous targeted polynucleotide gene sequence. The patch nucleic acid may be provided for example as a ssODN, as a PCR product (amplicon) or within a vector. Preferably, the patch / donor nucleic acid will include modifications with respect to the endogenous gene which precludes it from being cut by a sgRNA once integrated in the genome of a cell and / or which facilitate the detection of the introduction of the patch nucleic acid by homologous recombination.

[0163] As used herein, a “target gene”, “target region” or “target polynucleotide” corresponds to the polynucleotide within a cell that will be modified by the introduction of the patch nucleic acid. It corresponds to an endogenous gene naturally present within a cell. One or both allele(s) of a target gene may be modified within a cell, in accordance with the present disclosure.

[0164] A “target polynucleotide” as used herein refers to any endogenous polynucleotide or nucleic acid present in the genome of a cell and encoding or not a known gene product. "Target gene" as used herein refers to any endogenous polynucleotide or nucleic acid present in the genome of a cell and encoding a known or putative gene product. The target gene or target polynucleotide further corresponds to the polynucleotide within a cell that will be modified by a nuclease of the present disclosure, alone or in combination with the introduction of one or more donor nucleic acid or patch nucleic acids.

[0165] "Promoter" as used herein means a synthetic or naturally-derived nucleic acid molecule which is capable of conferring, modulating, or controlling (e.g., activating, enhancing, and / or repressing) expression of a nucleic acid in a cell. A promoter may comprise one or more specific transcriptional regulatory sequences to further enhance or repress expression and / or to alter the spatial expression and / or temporal expression of same. A promoter may also comprise distal enhancer or repressor elements, which may be located as much as several thousand base pairs from the start site of transcription. A promoter may be derived from sources including viral, bacterial, fungal, plants, insects, and animals. A promoter may regulate the expression of a gene component constitutively or differentially with respect to cell, the tissue or organ in which expression occurs or, with respect to the developmental stage at which expression occurs, or in response to external stimuli such as physiological stresses, pathogens, metal ions, or inducing agents. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator-promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter, CMV IE promoter, U6 promoter, a liver-specific promoter (e.g., LP1 b; combining the human apolipoprotein E / C-l gene locus control region [ApoE-HCR] and a modified human a1 antitrypsin promoter [hAAT] coupled to an SV40 intron), human thyroxine binding globulin (TBG) promoter, CMV promoter, CAG promoter, CBH promoter, UbiC promoter, Ef1 a promoter, H1 promoter, and 7SK promoter, any of which may be used to express one or more gRNAs and / or a CRISPR nuclease in a cell. Sequences for the LP1 b and TBG promoters are provided in Table 8.

[0166] “Vector" as used herein means a nucleic acid sequence containing an origin of replication. A vector may be a viral vector, bacteriophage, bacterial artificial chromosome, or yeast artificial chromosome. A vector may be a DNA or RNA vector. A vector may be a self-replicating extrachromosomal vector, and preferably, is a DNA plasmid. For example, the vector may comprise nucleic acid sequence(s) that / which encode(s) a sgRNA, a donor (or patch) nucleic acid, and / or a CRISPR nuclease (e.g., Cas9 or Cpf1) of the present disclosure. A vector for expressing one or more sgRNA will comprise a “DNA” sequence of the sgRNA. In embodiments, a vector may comprise "regulatory" or "control" sequences, which refer to nucleic acid sequences necessary for the transcription and possibly translation of an operably linked coding sequence in a particular host cell. In addition to control sequences that govern transcription and translation, expression vectors may contain nucleic acid sequences that serve other functions as well.

[0167] In an embodiment, the vector further comprises a nucleic acid encoding a selectable marker or reporter protein. A selectable marker or reporter is defined herein to refer to a nucleic acid encoding a polypeptide that, when expressed, confers an identifiable characteristic (e.g., a detectable signal, resistance to a selective agent) to the cell permitting easy identification, isolation and / or selection of cells containing the selectable marker from cells without the selectable marker or reporter. Any selectable marker or reporter known to those of ordinary skill in the art is contemplated for inclusion as a selectable marker in the vector of the present disclosure. For example, the selectable marker may be a drug selection marker, an enzyme, or an immunologic marker. Examples of selectable markers or reporters include, but are not limited to, polypeptides conferring drug resistance (e.g., kanamycin / geneticin resistance), enzymes such as alkaline phosphatase and thymidine kinase, bioluminescent and fluorescent proteins such as luciferase, green fluorescent protein (GFP), yellow fluorescent protein (YFP), cyan fluorescent protein (CFP), blue fluorescent protein (BFP), citrine and red fluorescent protein from Discosoma sp. (dsRED), membrane bound proteins to which high affinity antibodies or ligands directed thereto exist or can be produced by conventional means, and fusion proteins comprising a membrane-bound protein appropriately fused to an antigen tag domain from, among others, hemagglutinin (HA) or Myc. The nucleic acid encoding the selectable marker or reporter protein may be under the control of the same promoter / enhancer as the nucleic acid of interest, or may be under the control of a distinct promoter / enhancer.

[0168] In embodiments, the vector may comprise additional elements, such as one or more origins of replication sites (often termed “on”), restriction endonuclease recognition sites (multiple cloning sites, MCS), and / or internal ribosome entry site (IRES) elements.

[0169] Nucleic acids encoding gRNAs, donor nucleic acids and CRISPR nucleases (e.g., Cas9) of the present disclosure may be delivered into cells using one or more various vectors such as viral vectors. Accordingly, preferably, the above-mentioned vector is a viral vector for introducing an sgRNA, donor nucleic acid and / or a nucleic acid encoding a nuclease of the present disclosure in a target cell. Non-limiting examples of viral vectors include retrovirus, lentivirus, herpesvirus, adenovirus, or adeno-associated virus (AAV), as well known in the art.

[0170] "Adeno-associated virus" or "AAV" as used interchangeably herein refers to a small virus belonging to the genus Dependoparvovirus of the Parvoviridae family that infects humans and some other primate species. AAV is not currently known to cause disease and, consequently, the virus may cause a very mild immune response.

[0171] In embodiments, the AAV vector preferably targets one or more cell types. Accordingly, the AAV vector may have enhanced cardiac, skeletal muscle, neuronal, liver, and / or pancreatic tissue (Langerhans cells) tropism. The AAV vector may be capable of delivering and expressing the at least one sgRNA, donor nucleic acid and / or nucleic acid encoding a nuclease of the present disclosure in the cell of a mammal. For example, the AAV vector may be an AAV-SASTG vector (Piacentino et a / ., Hum. Gene Ther. 23: 635-646, 2012). The AAV vector may deliver gRNAs, pegRNAs and nucleases to neurons, skeletal and cardiac muscle, and / or pancreas (Langerhans cells) in vivo. The AAV vector may be based on one or more of several capsid types, including AAV1 , AAV2, AAV5, AAV6, AAV8, and AAV9. The AAV vector may be based on AAV2 pseudotype with alternative muscle-tropic AAV capsids, such as AAV2 / 1 , AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, and AAV / SASTG vectors that efficiently transduce skeletal muscle or cardiac muscle by systemic and local delivery. In an embodiment, the AAV vector is a AAV-DJ vector. In an embodiment, the AAV vector is a AAV-DJ8 vector. In an embodiment, the AAV vector is a AAV2-DJ8 vector. In an embodiment, the AAV vector is a AAV-PHP.B vector. In an embodiment, the AAV vector is a AAV-PHP.B, AAV-9, or AAV-DJ8 (PHP.B: Deverman D.E. etal., Nat. Biotechnol. 34: 204-209, 2016; Jackson K.L. etal., Front. Mol. Neurosci. 9: 116, 2016; AAV DJ-8: www.cellbiolabs.com / news / aav-helper-free-expression-systems-aav-dj- aav-dj8, http: / / www.cellbiolabs.com / aav-expression-and-packaging; www.cellbiolabs.com / scaav-dj8-helper-free- complete-expression-systems; and AAV9: Saraiva J. et al., J. Control. Release 241 : 94-109, 2016; Inagaki K. et al., Mol. Ther. 14: 45-53, 2006).

[0172] Cells

[0173] In another aspect, the present disclosure provides a cell (host cell, engineered cell) comprising a nucleic acid, gRNA, expression construct, or vector / plasmid described herein. In an embodiment, the cell is a primary cell, for example a brain / neuronal cell, a peripheral blood cell (e.g., a B or T lymphocyte, a monocyte, or a NK cell), a cord blood cell, a bone marrow cell, a cardiac cell, an endothelial cell, an epidermal cell, an epithelial cell, a fibroblast, a hepatic cell, or a lung / pulmonary cell. In an embodiment, the cell is a bone marrow cell, peripheral blood cell or cord blood cell. In a further embodiment, the cell is an immune cell, such as a T cell (e.g., a CD3+T cell, a CD8+T cell, a CD4+T cell, a Regulatory T cell (Tregs, e.g., CD4+ / FOXP3+)), a B cell, or a NK cell.

[0174] In an embodiment, the cell is a stem cell. The term "stem cell" as used herein refers to a cell that has pluripotency which allows it to differentiate into a functional mature cell. It includes primitive hematopoietic cells, progenitor cells, as well as adult stem cells that are undifferentiated cells found in various tissue within the human body, which can renew themselves and give rise to specialized cell types and tissue from which the cells came (e.g., muscle stem cells, skin stem cells, brain or neural stem cells, mesenchymal stem cell, lung stem cells, liver stem cells).

[0175] In an embodiment, the cell is a primitive hematopoietic cell. As used herein, the term "primitive hematopoietic cell” is used to refers to cells having pluripotency which allows them to differentiate into functional mature blood cells of the myeloid and lymphoid lineages such as T cells, B cells, NK cells, granulocytes (e.g., promyelocytes, neutrophils, eosinophils, basophils), erythrocytes (e.g., reticulocytes, erythrocytes), thrombocytes (e.g., megakaryoblasts, platelet producing megakaryocytes, platelets), and monocytes (e.g., monocytes, macrophages), and that may or may not the ability to regenerate while maintaining their pluripotency (self-renewal). It encompasses "hematopoietic stem cells" or "HSCs", which are cells having both pluripotency which allows them to differentiate into functional mature cells such as granulocytes, erythrocytes, thrombocytes, and monocytes, and the ability to regenerate while maintaining their pluripotency (self-renewal), as well as pluripotent hematopoietic cells that do not have self-renewal capacity. It also encompasses embryonic stem cells (ESCs), which are pluripotent stem cells derived from the inner cell mass of a blastocyst, an early-stage pre-implantation embryo. In an embodiment, the population of cells comprises ESCs. In another embodiment, the population of cells comprises HSCs. HSCs may be obtained from the body or an organ of the body containing cells of hematopoietic origin. Such sources include unfractionated bone marrow (from femurs, hips, ribs, sternum, and other bones), umbilical cord blood, peripheral blood, liver, thymus, lymph, and spleen. All of the aforementioned crude or unfractionated blood products can be enriched for cells having HSC characteristics in ways known to those of skill in the art. HSCs are phenotypically identified by their small size, lack of lineage (lin) markers, low staining (side population) with vital dyes such as rhodamine 123 (rhodamineDULL, also called rho 0) or Hoechst 33342, and presence / absence of various antigenic markers on their surface many of which belongs to the cluster of differentiation series, such as CD34, CD38, CD90, CD133, CD105, CD45, and c-kit.

[0176] In an embodiment, the stem cell is an induced pluripotent stem cell (iPSC). The term iPSC refers to a pluripotent stem cell that can be generated directly from adult cells using appropriate factors to “reprogram” the cells.

[0177] In an embodiment, the cell is a mammalian cell, for example a human cell.

[0178] A nucleic acid, expression construct, or vector / plasmid described herein may be introduced into the cell using standard techniques for introducing nucleic acids into a cell, e.g., transfection, transduction, or transformation. In an embodiment, the vector is a viral vector, and the cell is transduced with the vector. As used herein, the term "transduction" refers to the stable transfer of genetic material from a viral particle (e.g., lentiviral) to a cell genome (e.g., hematopoietic cell genome). It also encompasses the introduction of non-integrating viral vectors into cells, which leads to the transient or episomal expression of the gene of interest present in the viral vector.

[0179] Viruses may be used to infect cells in vivo, ex vivo, or in vitro using techniques well known in the art. For example, when cells, for instance CD34+ cells or stem cells are transduced ex vivo, the vector particles may be incubated with the cells using a dose generally in the order of between 1 to 100 or 1 to 50 multiplicities of infection (MOI) which also corresponds to 1x105to 100 or 50 x 105transducing units of the viral vector per 105cells. This, of course, includes amount of vector corresponding to 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, and 50 MOI.

[0180] Prior to, during, and / or following transduction, the cells may be cultured in media suitable for the maintenance, growth, or proliferation of the cells. The culture conditions of the population of cells will vary depending on different factors, notably, the starting cell population. Suitable culture media and conditions are well known in the art. The culture may be carried out in natural medium, a semi-synthetic medium, or a synthetic medium in terms of composition, and may be a solid medium, a semisolid medium, or a liquid medium in terms of shape, and any nutrient medium used for cell culture, such as stem cell culture, which may be supplemented with one or more of growth factors. Such medium typically comprises sodium, potassium, calcium, magnesium, phosphorus, chlorine, amino acids, vitamins, cytokines, hormones, antibiotics, serum, fatty acids, saccharides, or the like. In the culture, other chemical components or biological components may be incorporated singly or in combination, as the case requires. Such components to be incorporated in the medium may be fetal calf serum, human serum, horse serum, insulin, transferrin, lactoferrin, cholesterol, ethanolamine, sodium selenite, monothioglycerol, 2-mercaptoethanol, bovine serum albumin, sodium pyruvate, polyethylene glycol, various vitamins, various amino acids, agar, agarose, collagen, methylcellulose, various cytokines, various growth factors, or the like. Examples of such basal medium appropriate for a method of expanding stem cells include, without limitation, StemSpan™ Serum-Free Expansion Medium (SFEM; StemCell Technologies®, Vancouver, Canada), StemSpan™ H3000-Defined Medium (StemCell Technologies®, Vancouver, Canada), CellGro™, SCGM (CellGenix™, Freiburg Germany), StemPro™-34 SFM (Invitrogen®), Dulbecco's Modified Eagle's Medium (DMEM), Ham's Nutrient Mixture H12 Mixture F12, McCoy's 5A medium, Eagle's Minimum Essential Medium (EMEM), MEM medium (alpha Modified Eagle's Minimum Essential Medium), RPMI 1640 medium, Isocove's Modified Dulbecco's Medium (IMDM), StemPro34™ (Invitrogen®), X-VIVO™ 10 (Cambrex®), X-VIVO™ 15 (Cambrex®), and Stemline™ II (Sigma-Aldrich®).

[0181] Following transduction, the transduced cells may be cultured under conditions suitable for their maintenance, growth and / or proliferation. In particular aspects, the transduced cells are cultured for about 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, or 14 days before transplantation.

[0182] Culture conditions for maintaining and / or expanding stem cells are well known in the art. Typically, the culturing conditions comprise the use of factors like cytokines and growth factors, generally known in the art for stem cell expansion. Such cytokines and growth factors can be biologies or small molecules and they include without limitation IL-1 , IL-3, IL-6, IL-11 , G-CSF, GM-CSF, SCF, FIT3-L, thrombopoietin (TPO), erythropoietin, and analogs thereof. As used herein, "analogs" include any structural variants of the cytokines and growth factors having the biological activity as the naturally occurring forms, including without limitation, variants with enhanced or decreased biological activity when compared to the naturally occurring forms or cytokine receptor agonists such as an agonist antibody against the TPO receptor (for example, VB22B scFv2 as detailed in patent publication WO 2007 / 145227, and the like). Cytokine and growth factor combinations are chosen to maintain / expand stem cells while limiting the production of terminally differentiated cells. In one specific embodiment, one or more cytokines and growth factors are selected from the group consisting of SCF, Flt3-L, and TPO.

[0183] Human IL-6 or interleukin-6, also known as B-cell stimulatory factor 2 has been described by (Kishimoto T., Ann. Rev. Immunol. 23: 1-21, 2005) and is commercially available. Human SCF or stem cell factor, also known as c-kit ligand, mast cell growth factor, or Steel factor has been described (Smith, M.A. et al., Acta Haematol., 105: 143- 150, 2001) and is commercially available. Flt3-L or FLT-3 Ligand, also referred as FL is a factor that binds to flt3- receptor. It has been described (Hannum C., Nature 368: 643-648, 1994) and is commercially available. TPO or thrombopoietin, also known as megakarayocyte growth factor (MGDF) or c-MpI ligand has been described (Kaushansky K., N. Engl. J. Med. 354: 2034-2045, 2006) and is commercially available.

[0184] The chemical components and biological components mentioned above may be used, not only by adding them to the medium, but also by immobilizing them onto the surface of the substrate or support used for the culture, specifically speaking, by dissolving a component to be used in an appropriate solvent, coating the substrate or support with the resulting solution and then washing away an excess of the component. Such a component to be used may be added to the substrate or support preliminarily coated with a substance which binds to the component.

[0185] Stem cells may be cultured in a culture vessel generally used for animal cell culture such as a Petri dish, a flask, a plastic bag, a Teflon™ bag, optionally after preliminary coating with an extracellular matrix or a cell adhesion molecule. The material for such a coating may be collagens I to XIX, fibronectin, vitronectin, laminins 1 to 12, nitrogen, tenascin, thrombospondin, von Willebrand factor, osteoponin, fibrinogen, various elastins, various proteoglycans, various cadherins, desmocolin, desmoglein, various integrins, E-selectin, P-selectin, L-selectin, immunoglobulin superfamily, Matrigel®, poly-D-lysine, poly-L-lysine, chitin, chitosan, Sepharose®, alginic acid gel, hydrogel, or a fragment thereof. Such a coating material may be a recombinant material having an artificially modified amino acid sequence. The stem cells may be cultured by using a bioreactor which can mechanically control the medium composition, pH and the like and obtain high density culture (Schwartz R.M. et al., Proc. Natl. Acad. Sci. U.S.A., 88: 6760-6764, 1991 ; Koller M.R. et al., Bone Marrow Transplant. 21 : 653-663, 1998; Koller M.R. et al., Blood 82: 378-384, 1993; Astori G. et al., Bone Marrow Transplant. 35: 1101-1106, 2005).

[0186] The cell population may then be washed to remove any component of the cell culture and resuspended in an appropriate cell suspension medium for short term use or in a long-term storage medium, for example a medium suitable for cryopreservation, for example DMEM with 40% FCS and 10% DMSO. Other methods for preparing frozen stocks for cultured cells also are available to those skilled in the art.

[0187] Compositions

[0188] In another aspect, the present disclosure provides a composition (e.g., a pharmaceutical composition) comprising a gRNA, donor nucleic acid and / or a CRISPR nuclease (e.g., Cas9), or nucleic acid(s) encoding same or vector(s) comprising such nucleic acid(s), or a cell, as described herein. In an embodiment, the composition further comprises one or more biologically or pharmaceutically acceptable carriers, excipients, and / or diluents.

[0189] As used herein, "pharmaceutically acceptable" (or "biologically acceptable") carriers, excipients, and / or diluents includes any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible, and which can be used pharmaceutically or in biological systems. Such materials are characterized by the absence of (or limited) toxic or adverse biological effects in vivo. It refers to those compounds, compositions, and / or dosage forms which are, within the scope of sound medical judgment, suitable for use in contact with the biological fluids and / or tissues and / or organs of a subject (e.g., human, animal) without excessive toxicity, irritation, allergic response, or other problem or complication, commensurate with a reasonable benefit / risk ratio.

[0190] When the excipient serves as a diluent, it can be a solid, semisolid, or liquid material, which acts as a vehicle, carrier, or medium for the active ingredient. Thus, the compositions can be in the form of tablets, pills, powders, lozenges, sachets, cachets, elixirs, suspensions, emulsions, solutions, syrups, aerosols (as a solid or in a liquid medium), ointments containing for example up to 10% by weight of the active compound, soft and hard gelatin capsules, suppositories, sterile injectable solutions, and sterile packaged powders (see Remington: The Science and Practice of Pharmacy by Alfonso R. Gennaro, 2003, 21thedition, Mack Publishing Company). In embodiments, the carrier may be suitable for intra-neural, parenteral, intravenous, intraperitoneal, intramuscular, subcutaneous, sublingual, or oral administration.

[0191] Some examples of suitable excipients include lactose, dextrose, sucrose, sorbitol, mannitol, starches, lecithin, phosphatidylcholine, gum acacia, calcium phosphate, alginates, tragacanth, gelatin, calcium silicate, microcrystalline cellulose, polyvinylpyrrolidone, cellulose, water, syrup, and methyl cellulose. The formulations can additionally include: lubricating agents such as talc, magnesium stearate, and mineral oil; wetting agents; emulsifying and suspending agents; preserving agents such as methyl- and propyl- hydroxybenzoates; sweetening agents; and flavoring agents. The compositions of the disclosure can be formulated to provide quick sustained or delayed release of the active ingredient after administration to the patient by employing procedures known in the art.

[0192] Pharmaceutical compositions suitable for use in the disclosure include compositions wherein the active ingredients are contained in an effective amount to achieve the intended purpose (e.g., preventing, treating, ameliorating, and / or inhibiting a disease or condition). The determination of an effective dose is well within the capability of those skilled in the art. For any compounds, the therapeutically effective dose can be estimated initially either in cell culture assays (e.g., cell lines) or in animal models, usually mice, rabbits, dogs, or pigs. The animal model may also be used to determine the appropriate concentration range and route of administration. Such information can then be used to determine useful doses and routes for administration in humans. An effective dose or amount refers to that amount of one or more active ingredient(s), which is sufficient for treating a specific disease or condition. Therapeutic efficacy and toxicity may be determined by standard pharmaceutical procedures in cell cultures or experimental animals, e.g., ED50 (the dose therapeutically effective in 50% of the population) and LD50 (the dose lethal to 50% of the population). The dose ratio between therapeutic and toxic effects is the therapeutic index, and it can be expressed as the ratio, LD50 / ED50. Pharmaceutical compositions, which exhibit large therapeutic indices, are preferred. The data obtained from cell culture assays and animal studies is used in formulating a range of dosage for human use. The dosage contained in such compositions is preferably within a range of circulating concentrations that include the ED50 with little or no toxicity. The dosage varies within this range depending upon the dosage form employed, sensitivity of the patient, and the route of administration. The exact dosage will be determined by the practitioner, in light of factors related to the subject that requires treatment. Dosage and administration are adjusted to provide sufficient levels of the active moiety or to maintain the desired effect. Factors, which may be taken into account, include the severity of the disease state, general health of the subject, age, weight, and gender of the subject, diet, time, and frequency of administration, drug combination(s), reaction sensitivities, and tolerance / response to therapy. Guidance as to particular dosages and methods of delivery is provided in the literature and generally available to practitioners in the art. In embodiments, dosages of an active ingredient of between about 0.01 and about 100 mg / kg body weight (in an embodiment, per day) may be used. In further embodiments, dosages of between about 0.5 and about 75 mg / kg body weight may be used. In further embodiments, dosages of between about 1 and about 50 mg / kg body weight may be used. In further embodiments, dosages of between about 10 and about 50 mg / kg body weight in further embodiments about 10, about 25 or about 50 mg / kg body weight, may be used.

[0193] The present disclosure further provides a kit or package comprising at least one container means having disposed therein at least one of an gRNA, donor nucleic acid, nuclease, nucleic acids, vector, cell, systems, combination, or composition as described herein. In an embodiment, the kit or package further comprises with instructions for use, such as for modification of a nucleotide sequence in a cell, or for the treatment of a disease, disorder, or condition.

[0194] Methods / uses

[0195] The present disclosure also relates to a method for inducing the expression of a nucleic acid of interest or gene of interest by a cell, the method comprising introducing the nucleic acid, expression construct, or vector described herein in the cell. The present disclosure also relates to a use of a nucleic acid, expression construct, or vector described herein for inducing the expression of a gene of interest by a cell. In an embodiment, the cell is a primary cell, for example a brain / neuronal cell, a peripheral blood cell (e.g., a B or T lymphocyte, a monocyte, or a NK cell), a cord blood cell, a bone marrow cell, a cardiac cell, an endothelial cell, an epidermal cell, an epithelial cell, a fibroblast, hepatic cell, or a lung / pulmonary cell. In an embodiment, the cell is a bone marrow cell, peripheral blood cell, or cord blood cell. In a further embodiment, the cell is an immune cell, such as a T cell (e.g., a CD3+T cell, a CD8+T cell, a CD4+T cell, a Regulatory T cell (Tregs, e.g., CD4+ / FOXP3+)), a B cell, or a NK cell.

[0196] The present disclosure also relates to a method for preventing or treating a disease, condition, or disorder in a subject, the method comprising administering a cell comprising a nucleic acid, expression construct, or vector described herein. The present disclosure also relates to the use of a nucleic acid, expression construct, or vector described herein for preventing or treating a disease, condition, or disorder in a subject. The present disclosure also relates to the use of a cell comprising a nucleic acid, expression construct, or vector described herein method for the manufacture of a medicament for preventing or treating a disease, condition, or disorder in a subject. In an embodiment, the disease, condition, or disorder is associated with an absence or deficiency of expression of a protein or the expression of a defective (e.g., mutated) protein, and the expression of an expression construct described herein encoding a functional (e.g., native) protein provides or restores a level of the absent or deficient expression or activity associated with the disease thereby to prevent, treat or ameliorate the disease, condition or disorder.

[0197] In embodiments, the disease or condition is associated with a deficiency of the polypeptide of interest, and / or is associated with a defective form of the polypeptide of interest.

[0198] In embodiments, the disease or condition is a metabolic disease. In embodiments, the disease or condition is a lysosomal disease. In embodiments, the disease or condition is a mucopolysaccharidosis (MPS), such as a type I, II, III or VI MPS. In embodiments, the disease or condition is an alpha-L-iduronidase (IDUA)-associated disease. In embodiments, the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L- iduronidase (IDUA). In embodiments, the disease or condition is a glycogen storage disease. In embodiments, the disease or condition is an acid alpha-glucosidase (GAA)-associated disease. In embodiments, the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA). In embodiments, the disease is a muscular disease, e.g., resulting in muscle damage / weakness or cardiomegaly (e.g., Pompe disease), and the treatment improves muscle function (e.g., muscle force, strength) and / or corrects cardiomegaly.

[0199] In embodiments, the polypeptide of interest is an alpha-L-iduronidase (IDUA) or a fragment or derivative of an IDUA having IDUA activity. In embodiments, the fragment or derivative of an IDUA having IDUA activity comprises an amino acid sequence having at least 70, 75, 80, 85, 90, or 95% identity with the amino acid sequence of SEQ ID NO: 32.

[0200] In embodiments, the polypeptide of interest is an acid alpha-glucosidase (GAA) or a fragment or derivative of a GAA having GAA activity. In embodiments, the fragment or derivative of a GAA having GAA activity comprises an amino acid sequence having at least 70, 75, 80, 85, 90, or 95% identity with the amino acid sequence of SEQ ID NO: 38.

[0201] As used herein, "treatment" (and grammatical variations thereof such as "treat" or "treating") refers to complete or partial amelioration or reduction of a disease or condition or disorder, or a symptom, adverse effect or outcome, or phenotype associated therewith. Desirable effects of treatment include, but are not limited to, preventing occurrence or recurrence of disease, alleviation of symptoms, diminishment of any direct or indirect pathological consequences of the disease, preventing metastasis, decreasing the rate of disease progression, amelioration or palliation of the disease state, and remission or improved prognosis. The terms do not imply complete curing of a disease or complete elimination of any symptom or effect(s) on all symptoms or outcomes.

[0202] In some embodiments, the cell therapy, is carried out by autologous transfer, in which the cells are isolated and / or otherwise prepared from the subject who is to receive the cell therapy, or from a sample derived from such a subject. Thus, in some aspects, the cells are derived from a subject, e.g., patient, in need of a treatment and the cells, following isolation and processing are administered to the same subject.

[0203] In some embodiments, the cell therapy, is carried out by allogeneic transfer, in which the cells are isolated and / or otherwise prepared from a subject other than a subject who is to receive or who ultimately receives the cell therapy, e.g., a first subject. In such embodiments, the cells then are administered to a different subject, e.g., a second subject, of the same species. In some embodiments, the first and second subjects are genetically identical. In some embodiments, the first and second subjects are genetically similar. In some embodiments, the second subject expresses the same HLA class or super type as the first subject. The cells can be administered by any suitable means. Dosing and administration may depend in part on whether the administration is brief or chronic. Various dosing schedules include, but are not limited to, single or multiple administrations over various time-points, bolus administration, and pulse infusion.

[0204] In certain embodiments, the cells, or individual populations of subtypes of cells, are administered to the subject at a range of about one million to about 100 billion cells and / or that amount of cells per kilogram of body weight, such as, e.g., 1 million to about 50 billion cells (e.g., about 5 million cells, about 25 million cells, about 500 million cells, about 1 billion cells, about 5 billion cells, about 20 billion cells, about 30 billion cells, about 40 billion cells, or a range defined by any two of the foregoing values), such as about 10 million to about 100 billion cells (e.g., about 20 million cells, about 30 million cells, about 40 million cells, about 60 million cells, about 70 million cells, about 80 million cells, about 90 million cells, about 10 billion cells, about 25 billion cells, about 50 billion cells, about 75 billion cells, about 90 billion cells, or a range defined by any two of the foregoing values), and in some cases about 100 million cells to about 50 billion cells (e.g., about 120 million cells, about 250 million cells, about 350 million cells, about 450 million cells, about 650 million cells, about 800 million cells, about 900 million cells, about 3 billion cells, about 30 billion cells, about 45 billion cells) or any value in between these ranges and / or per kilogram of body weight. Dosages may vary depending on attributes particular to the disease or disorder and / or patient and / or other treatments.

[0205] In some embodiments, the preventative or therapeutic methods or uses described herein are part of a combination treatment, such as simultaneously with or sequentially with, in any order, another therapeutic intervention, such as co-administered with one or more additional therapeutic agents or in connection with another therapeutic intervention, either simultaneously or sequentially in any order.

[0206] In some embodiments, the synthetic expression construct is used as a research tool, for example as reporter tool or in a commercial detection method (assay development). For example, the synthetic expression construct may be operably linked to a nucleic acid encoding a reporter protein, which may be used for the detection of the expression of a gene of interest in a specific cell type, e.g., to confirm that the gene of interest has been taken up by and is expressed by the cell. The term “reporter protein” refers to a protein that may be easily identified and measured such as fluorescent and luminescent proteins (e.g., GFP, YFP), as well as enzymes that are able to generate a detectable product from a substrate (e.g., luciferase). The synthetic expression construct may also be used for the cell-specific expression of a gene of interest in vitro, e.g., to assess the effect of the expression of the gene of interest in the targeted cells.

[0207] CRISPR system

[0208] The CRISPR technology is a system for genome editing, e.g., for modification of a nucleic acid sequence, and may also be used for example to modify the expression of a specific gene.

[0209] This system stems from findings in bacterial and archaea which have developed adaptive immune defenses termed clustered regularly interspaced short palindromic repeats (CRISPR) systems, which use crRNAs and Cas proteins to degrade complementary sequences present in invading viral and plasmid DNA. The original CRISPR systems comprised a CRISPR RNA (crRNA) and a trans-activating crRNA (tracrRNA), which form a hybrid (which guides a CRISPR nuclease, e.g. a Cas9).

[0210] Engineered CRISPR systems use for example a synthetically reconstituted "guide RNA" ("gRNA"), corresponding to a crRNA-tracrRNA fusion that obviates the need for RNase III and crRNA processing in general. The gRNA comprises a “gRNA guide sequence” or “gRNA target sequence” and an RNA sequence (Cas recognition sequence)”, which is necessary for CRISPR nuclease (e.g., Cas9) binding to the targeted gene. The gRNA guide sequence is the sequence that confers specificity. It hybridizes with (i.e., it is complementary to) the opposite strand of a target sequence {i.e., it corresponds to the RNA sequence of a DNA target sequence). Other CRISPR systems using different CRISPR nucleases have been developed and are known in the art (e.g., using the Cpf1 nuclease instead of a Cas9 nuclease).

[0211] Because the original Cas9 nuclease combined with a gRNA may produce off-target mutagenesis, one may alternatively use in accordance with the present disclosure a pair of specifically designed gRNAs in combination with a Cas9 nickase or in combination with a dCas9-Folkl nuclease to cut both strands of DNA.

[0212] Base editing: CRISPR / Cas technology has also been developed to permit base editing without inducing a DSB. For example, base editing may be performed using a guide RNA and a Cas9 nickase (cutting only one DNA strand) fused with a cytidine deaminase to chemically modify a cytidine into a thymine (Komor, A.C., et a / ., Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage, Nature 533: 420-424, 2016). Base editing may also be performed using a guide RNA and a Cas9 nickase fused with an adenosine deaminase to chemically modify an adenosine into an inosine, the equivalent of a guanine (Gaudelli, N.M., et a / ., Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551 : 464-471 , 2017).

[0213] Prime editing: Another use of CRISPR / Cas technology is prime editing, which uses an extended guide RNA (called a pegRNA) and a Cas9 nickase fused with a reverse transcriptase (the fusion protein is e.g., PE2) to replace any nucleotide by any other nucleotide7.

[0214] In embodiments, provided herein are CRISPR / nuclease-based engineered systems for use in modifying a target nucleic acid in cells. Introduction of DSBs can knockout a specific gene or allow modifying it by Homology Directed Repair (HDR), where one or more donor or patch nucleic acids comprising the desired modification(s) are provided to introduce the modification(s) by HDR. CRISPR / Cas9-induced DNA cleavage followed by Non-Homologous End Joining (NHEJ) repair has been used to generate loss-of-function alleles in protein-coding genes or to delete a very large DNA fragment. The CRISPR-based engineered systems of the present disclosure are designed to (i) target and cleave a gene of interest) to generate gene variants (e.g., creating insertions] and / or deletions, also referred to as INDELS).

[0215] Accordingly, in an aspect, the present disclosure involves the design and preparation of one or more gRNAs for inducing a DSB in a target gene of interest. In embodiments, the present disclosure also involves the design and preparation of one or more gRNAs for inducing a DSB in a target polynucleotide located at a different locus within the genome of target cells. The gRNAs and the nuclease are then used together to introduce the desired modification(s) (i.e., gene-editing events) by NHEJ or HDR within the genome of one or more target cells. When the desired modification(s) include specific point mutation(s) or insertions / deletion(s), one or more donor or patch nucleic acids comprising the desired modification(s) are provided to introduce the modification(s) by HDR. gRNAs In order to cut DNA at a specific site, CRISPR nucleases require the presence of a gRNA and a protospacer adjacent motif (PAM) on the targeted gene. The PAM immediately follows ( / .e., is adjacent to) the gRNA target sequence in the targeted polynucleotide gene sequence. The PAM is located at the 3’ end or 5’ end of the gRNA target sequence (depending on the CRISPR nuclease used) but is not included in the gRNA guide sequence. For example, the PAM for Cas9 CRISPR nucleases is located at the 3’ end of the gRNA target sequence on the target gene while the PAM for Cpf 1 nucleases is located at the 5’ end of the gRNA target sequence on the target gene. Different CRISPR nucleases also require a different PAM. Accordingly, selection of a specific polynucleotide gRNA target sequence is generally based on the CRISPR nuclease used. The PAM for the Streptococcus pyogenes Cas9 CRISPR system is 5 -NRG-3', where R is either A or G, and characterizes the specificity of this system in human cells. The PAM of Staphylococcus aureus Cas9 is NNGRR. The S. pyogenes Type II system naturally prefers to use an "NGG" sequence, where "N" can be any nucleotide, but also accepts other PAM sequences, such as "NAG" in engineered systems. Similarly, the Cas9 derived from Neisseria meningitidis (NmCas9) normally has a native PAM of NNNNGATT, but has activity across a variety of PAMs, including a highly degenerate NNNNGNNN PAM. The PAM for AsCpfl or LbCpfl CRISPR nuclease is TTTN. In an embodiment, the PAM for a Cas9 protein used in accordance with the present disclosure is a NGG trinucleotide-sequence (Cas9). In another embodiment, the PAM for a Cpf1 CRISPR nuclease used in accordance with the present disclosure is a TTTN nucleotide sequence. In a preferred embodiment, the St1Cas9 may be used, which corresponds to the PAM sequences NNAGAA and NNGGAA. In embodiments, different St1Cas9 PAM sequences may be used, for example, inferred consensus PAM sequences for St1Cas9 from strains CNRZ1066 and LMG13811 are NNACAA(W) and NNGCAA(A), respectively. Table 1 below provides a list of non-limiting examples of CRISPR / nuclease systems with their respective PAM sequences.

[0216] Table 1 : Non-exhaustive list of CRISPR-nuclease systems from different species (see Mohanraju P. et al., Science 353: aad5147, 2016; Shmakov S. et al., Mol. Cell 60: 385-397, 2015; and Zetsche B. et al., Cell 163: 759- 771 , 2015). Also included are engineered variants recognizing alternative PAM sequences (see Kleinstiver B.P. et al., Nat. Biotechnol. 33: 1293-1298, 2015 and Kleinstiver B.P. et al., Nature 523: 481-485, 2015).

[0217] As used herein, the expressions “guide RNA”, “gRNA”, “single guide RNA”, “sgRNA” are used interchangeably and refer to a polynucleotide sequence which works in combination with a CRISPR nuclease to hybridize with a target nucleic acid sequence and direct sequence-specific binding of a nucleic acid-targeting complex to the target nucleic acid sequence, which in turn can introduce a cut into DNA. The gRNA comprises a gRNA guide sequence and a “CRISPR nuclease recognition sequence”.

[0218] As used herein, the expression “gRNA guide sequence” refers to the corresponding RNA sequence of the “gRNA target sequence”. Therefore, it is the RNA sequence equivalent of the protospacer on the target polynucleotide gene sequence. It does not include the corresponding PAM sequence in the genomic DNA. It is the sequence that confers target specificity. The gRNA guide sequence is linked to a CRISPR nuclease recognition sequence which binds to the nuclease (e.g., Cas9 / Cpf1). The gRNA guide sequence recognizes and binds to the targeted gene of interest. It hybridizes with ( / .e., is complementary to) the opposite strand of a target gene sequence, which comprises the PAM ( / .e., it hybridizes with the DNA strand opposite to the PAM). As noted above, the “PAM” is the nucleic acid sequence, that immediately follows (is contiguous to) the target sequence or target polynucleotide but is not in the gRNA. Further, the embodiments and features described herein in respect of gRNAs also apply to pegRNAs, a type of gRNA used in prime editing.

[0219] A “CRISPR nuclease recognition sequence” as used herein refers broadly to one or more RNA sequences (or RNA motifs) required for the binding and / or activity (including activation) of the CRISPR nuclease on the target gene. Some CRISPR nucleases require longer RNA sequences than other to function. Also, some CRISPR nucleases require multiple RNA sequences (motifs) to function while others only require a single short RNA sequence / motif. For example, Cas9 proteins require a tracrRNA sequence in addition to a crRNA sequence to function while Cpf 1 only requires a crRNA sequence. Thus, unlike Cas9, which requires both crRNA sequence and a tracrRNA sequence (or a fusion or both crRNA and tracrRNA) to mediate interference, Cpf1 processes crRNA arrays independent of tracrRNA, and Cpf 1 -crRNA complexes alone cleave target DNA molecules, without the requirement for any additional RNA species (see Zetsche B. et a / ., Cell 163: 759-771 , 2015). The “CRISPR nuclease recognition sequence” included in the gRNA or pegRNA described herein is thus selected based on the specific CRISPR nuclease used. It includes direct repeat sequences and any other RNA sequence known to be necessary for the selected CRISPR nuclease binding and / or activity. Various RNA sequences which can be fused to an RNA guide sequence to enable proper functioning of CRISPR nucleases (referred to herein as CRISPR nuclease recognition sequence) are well known in the art and can be used in accordance with the present disclosure. The “CRISPR nuclease recognition sequence” may thus include a crRNA sequence only (e.g., for AsCpfl activity, such as the CRISPR nuclease recognition sequence UAAUUUCUAC UCUUGUAGAU (SEQ ID NO: 43)) or may include additional sequences (e.g., tracrRNA sequence necessary for Cas9 activity). Furthermore, in accordance with the present disclosure and as well known in the art, RNA motifs necessary for CRISPR nuclease binding and activity may be provided separately (e.g., [i] RNA guide sequence-crRNA CRISPR recognition sequence” [also known as crRNA] in one RNA molecule and [ii] a tracrRNA CRISPR recognition sequence on another, separate RNA molecule). Alternatively, all necessary RNA sequences (motifs) may be fused together in a single RNA guide. The CRISPR recognition sequence is preferably fused directly to the gRNA guide sequence (in 3’ [e.g., Cas9] or 5’ [Cpf1] depending on the CRISPR nuclease used) but may include a spacer sequence separating two RNA motifs. In embodiments, the CRISPR nuclease recognition sequence is a Cas9 recognition sequence having at least 65 nucleotides. In embodiments, the CRISPR nuclease recognition sequence is a Cas9 CRISPR nuclease recognition sequence having at least 85 nucleotides. In embodiments, the CRISPR nuclease recognition sequence is a Cpf1 recognition sequence (5’ direct repeat) having about 19 nucleotides. In an embodiment, the CRISPR nuclease recognition sequence is a St1Cas9 recognition sequence. The gRNA or pegRNA of the present disclosure may comprise any variant of the above noted sequences, provided that it allows for the proper functioning of the selected CRISPR nuclease (e.g., binding of the CRISPR nuclease protein to the gene of interest and / or target polynucleotide sequence[s]).

[0220] Together, the RNA guide sequence and CRISPR nuclease recognition sequence(s) provide both targeting specificity and scaffolding / binding ability for the CRISPR nuclease of the present disclosure. gRNAs and pegRNAs of the present disclosure do not exist in nature, i.e., is a non-naturally occurring nucleic acid(s).

[0221] A "target region", "target sequence" or "protospacer" in the context of gRNAs and CRISPR system of the present disclosure are used herein interchangeably and refers to the region of the target gene, which is targeted by the CRISPR / nuclease-based system, without the PAM. It refers to the sequence corresponding to the nucleotides that precede the PAM (i.e., in 5’ or 3’ of the PAM, depending on the CRISPR nuclease) in the genomic DNA. It is the sequence that is included into a gRNA expression construct (e.g., vector / plasmid / AAV). The CRISPR / nuclease- based system may include at least one {i.e., one or more) gRNAs, wherein each gRNA target different DNA sequences on the target gene. The target DNA sequences may be overlapping. The target sequence or protospacer is followed or preceded by a PAM sequence at an (3’ or 5’ depending on the CRISPR nuclease used) end of the protospacer. Generally, the target sequence is immediately adjacent {i.e., is contiguous) to the PAM sequence (it is located on the 5’ end of the PAM for SpC as9-l i ke nuclease and at the 3’ end for Cpf 1 -like nuclease). In embodiments, the gRNA of the present disclosure comprises a “gRNA guide sequence” or has a “gRNA target sequence” which corresponds to the target sequence on the gene of interest or target polynucleotide sequence that is followed or preceded by a PAM sequence (is adjacent to a PAM). The gRNA may comprise a "G" at the 5' end of its polynucleotide sequence. The presence of a “G” in 5’ is preferred when the gRNA is expressed under the control of the U6 promoter (Koo T. et al., Mol. Cells 38: 475-481, 2015). The CRISPR / nuclease system of the present disclosure may use gRNAs of varying lengths. The gRNA may comprise a gRNA guide sequence of at least 10 nucleotides (nts), at least 12 nts, at least 13 nts, at least 14 nts, at least 15 nts, at least 16 nts, at least 17 nts, at least 18 nts, at least 19 nts, at least 20 nts, at least 21 nts, at least 22 nts, at least 23 nts, at least 24 nts, at least 25 nts, at least 30 nts, or at least 35 nts of a target sequence of a gene of interest or target polynucleotide (such target sequence is followed or preceded by a PAM in the gene of interest or target polynucleotide but is not part of the gRNA). The length of the gRNA is selected based on the specific CRISPR nuclease used. In embodiments, the “gRNA guide sequence” or “gRNA target sequence” may be at least 17 nucleotides (17, 18, 19, 20, 21 , 22, or 23) long, preferably between 17 and 30 nts long, more preferably between 17-22 nucleotides long. In embodiments, the gRNA guide sequence is between 10-40, 10-30, 12-30, 15-30, 18-30, or 10-22 nucleotides long. In embodiments, the PAM sequence is "NGG", where "N" can be any nucleotide. In embodiments, the PAM sequence is "TTTN", where "N" can be any nucleotide. gRNAs may target any region of a target gene which is immediately adjacent (contiguous, adjoining, in 5’ or 3’) to a PAM (e.g., NGG / TTTN or CCN / NAAA for a PAM that would be located on the opposite strand) sequence. In embodiments, the gRNA of the present disclosure has a target sequence that is located in an exon (the gRNA guide sequence consists of the RNA sequence of the target [DNA] sequence which is located in an exon). In embodiments, the gRNA of the present disclosure has a target sequence that is located in an intron (the gRNA guide sequence consists of the RNA sequence of the target [DNA] sequence which is located in an intron). In embodiments, the gRNA may target any region (sequence) which is followed (or preceded, depending on the CRISPR nuclease used) by a PAM in the gene or target polynucleotide of interest.

[0222] Although a perfect match between the gRNA guide sequence and the DNA sequence on the targeted gene is preferred, a mismatch between a gRNA guide sequence and target sequence on the gene sequence of interest is also permitted, as long as it still allows hybridization of the gRNA with the complementary strand of the gRNA target polynucleotide sequence on the targeted gene. A seed sequence of between 8-12 consecutive nucleotides in the gRNA, which perfectly matches a corresponding portion of the gRNA target sequence is preferred for proper recognition of the target sequence. The remainder of the guide sequence may comprise one or more mismatches. In general, gRNA activity is inversely correlated with the number of mismatches. Preferably, the gRNA of the present disclosure comprises 7 mismatches, 6 mismatches, 5 mismatches, 4 mismatches, 3 mismatches, more preferably 2 mismatches, or less, and even more preferably no mismatch, with the corresponding gRNA target gene sequence (less the PAM). Preferably, the gRNA nucleic acid sequence is at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the gRNA target polynucleotide sequence in the gene of interest. Of course, the smaller the number of nucleotides in the gRNA guide sequence the smaller the number of mismatches tolerated. The binding affinity is thought to depend on the sum of matching gRNA-DNA combinations.

[0223] The number of gRNAs administered to or expressed in a target cell in accordance with the methods of the present disclosure may be at least 1 gRNA, at least 2 gRNAs, at least 3 gRNAs, at least 4 gRNAs, at least 5 gRNAs, at least 6 gRNAs, at least 7 gRNAs, at least 8 gRNAs, at least 9 gRNAs, at least 10 gRNAs, at least 11 gRNAs, at least 12 gRNAs, at least 13 gRNAs, at least 14 gRNAs, at least 15 gRNAs, at least 16 gRNAs, at least 17 gRNAs, or at least 18 gRNAs. The number of gRNAs administered to or expressed in a cell may be between at least 1 gRNA and 15 gRNAs, 1 gRNA and least 10 gRNAs, 1 gRNA and 8 gRNAs, 1 gRNA and 6 gRNAs, 1 gRNA and 4 gRNAs, 1 gRNA and gRNAs, 2 gRNA and 5 gRNAs, or 2 gRNAs and 3 gRNAs.

[0224] CRISPR nucleases

[0225] Recombinant dCas9-FoKI dimeric nucleases (RFNs) have been designed that can recognize extended sequences and edit endogenous genes with high efficiency in human cells. These nucleases comprise a dimerizationdependent wild type Fokl nuclease domain fused to a catalytically inactive Cas9 (dCas9) protein. Dimers of the fusion proteins mediate sequence specific DNA cleavage when bound to target sites composed of two half-sites (each bound to a dCas9 monomer domain, i.e., a Cas9 nuclease devoid of nuclease activity) with a spacer sequence between them. The dCas9-FoKI dimeric nucleases require dimerization for efficient genome editing activity and thus, use two gRNAs for introducing a cut into DNA.

[0226] The recombinant CRISPR nuclease that may be used in accordance with the present disclosure is i) derived from a naturally occurring Cas; and ii) has a nuclease (or nickase) activity to introduce a DSB (or two SSBs in the case of a nickase) in cellular DNA when in the presence of appropriate gRNA(s). Thus, as used herein, the term “CRISPR nuclease” refers to a recombinant protein which is derived from a naturally occurring Cas nuclease which has nuclease or nickase activity and which functions with the gRNAs of the present disclosure to introduce DSBs (or one or two SSBs) in the targets of interest. In an embodiment, the CRISPR nuclease is St1Cas9. In further embodiments, the CRISPR nuclease is SpCas9 or Cpf 1.

[0227] Exemplary CRISPR nucleases that may be used in accordance with the present disclosure are provided in Table 1 above.

[0228] CRISPR nucleases such as Cas9 / nucleases cut 3-4 bases upstream of the PAM sequence. CRISPR nucleases such as Cpf1 on the other hand, generate a 5’ overhang. The cut occurs 19 bases after the PAM on the targeted (+) strand and 23 bases on the opposite strand (Zetsche et al., 2015, PM ID 26422227). There can be some off- target DSBs using wildtype Cas9. The degree of off-target effects depends on a number of factors including: how closely homologous the off-target sites are compared to the on-target site, the specific site sequence, and the concentration of nuclease and guide RNA (gRNA). These considerations only matter if the PAM sequence is immediately adjacent to the nearly homologous target sites. The mere presence of additional PAM sequences should not be sufficient to generate off target DSBs; there needs to be extensive homology of the protospacer followed or preceded by PAM. In embodiments, CRISPR nuclease (Cas or other nuclease / nickase recombinant protein described herein) preferably comprises at least one Nuclear Localization Signal (NLS) to target the protein into the cell nucleus, and the vector further comprises one or more nucleotide sequences encoding the one or more NLSs. Accordingly, as used herein the expression “nuclear localization signal” or “NLS” refers to an amino acid sequence, which 'tags' a protein for import into the cell nucleus by nuclear transport. Typically, this signal consists of one or more short sequences of positively charged lysines or arginines exposed on the protein surface. Different nuclear localized proteins may share the same NLS. An NLS has the opposite function of a nuclear export signal, which targets proteins out of the nucleus. Classical NLSs can be further classified as either monopartite or bipartite. The first NLS to be discovered was the sequence PKKKRKV (SEQ ID NO: 44) in the SV40 Large T-antigen (a monopartite NLS). The NLS of nucleoplasmin, KR[PAATKKAGQA]KKKK (SEQ ID NO: 45), is the prototype of the ubiquitous bipartite signal: two clusters of basic amino acids, separated by a spacer of about 10 amino acids. The Cas9 protein exemplified herein is a Cas9 nuclease comprising one or more, preferably two, NLS sequences.

[0229] There are many other types of NLS, which are qualified as “non-classical”, such as the acidic M9 domain of hnRNP A1 , the sequence KIPIK in yeast transcription repressor Mata2, the complex signals of U snRNPs, as well as a recently identified class of NLSs known as PY-NLSs. Thus, any type of NLS (classical or non-classical) may be used in accordance with the present disclosure, as long as it targets the protein of interest into the nucleus of a target cell. In an embodiment, the NLS is derived from the simian virus 40 large T antigen. In an embodiment, the NLS of the recombinant protein of the present disclosure comprises or consists of the following amino acid sequence SPKKKRKVEAS (SEQ ID NO: 46). In an embodiment the NLS comprises or consists of the sequence KKKRKV (SEQ ID NO: 47). In an embodiment, the NLS comprises or consists of the sequence SPKKKRKVEASPKKKRKV (SEQ ID NO: 48). In another embodiment, the NLS comprises or consists of the sequence KKKRK (SEQ ID NO: 49). In another embodiment, the NLS comprises or consists of the sequence PKKKRKV (SEQ ID NO: 44).

[0230] In an embodiment, the CRISPR nuclease comprises a first NLS at its amino terminal end and a second NLS at its carboxy terminal end, and the vector comprises NLS-encoding nucleotide sequences flanking the CRISPR nuclease-encoding nucleotide sequence.

[0231] In embodiments, the CRISPR nuclease or Cas nickase may optionally advantageously be coupled to a protein transduction domain for entry of the protein into the target cells. Alternatively, the nucleic acid coding for the guide RNA or pegRNA and for the deaminase or reverse transcriptase may be delivered in targeted cells using various viral vectors, virus like particles (VLP), or exosomes.

[0232] Protein transduction domains (PTD) may be of various origins and allow intracellular delivery of a given therapeutic by facilitating the translocation of the protein / polypeptide into a cell membrane, organelle membrane, or vesicle membrane. PTD refers to a polypeptide, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates traversing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. A PTD attached to another molecule facilitates the molecule traversing a membrane, for example going from extracellular space to intracellular space, or cytosol to within an organelle including the mitochondria.

[0233] In an embodiment, a PTD is covalently linked to the amino terminus of a recombinant protein of the present disclosure. In another embodiment, a PTD is covalently linked to the carboxyl terminus of a recombinant protein of the present disclosure. Exemplary protein transduction domains include but are not limited to a minimal undecapeptide protein transduction domain (corresponding to residues 47-57 of HIV-1 TAT comprising YGRKKRRQRRR (SEQ ID NO: 50); a polyarginine sequence comprising a number of arginines sufficient to direct entry into a cell (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines); a VP22 domain; an Drosophila Antennapedia protein transduction domain; a truncated human calcitonin peptide; RRQRRTSKLMKR (SEQ ID NO: 51); Transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 52);

[0234] KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 53); and RQIKIWFQNRRMKWKK (SEQ ID NO: 54). Further exemplary PTDs include but are not limited to, KKRRQRRR (SEQ ID NO: 55), RKKRRQRRR (SEQ ID NO: 56); or an arginine homopolymer of from 3 arginine residues to 50 arginine residues. In an embodiment, the protein transduction domain is TAT or Pep-1. In an embodiment, the protein transduction domain is TAT and comprises the sequence SGYGRKKRRQRRRC (SEQ ID NO: 57). In another embodiment, the protein transduction domain is TAT and comprises the sequence YGRKKRRQRRR (SEQ ID NO: 58). In another embodiment, the protein transduction domain is TAT and comprises the sequence KKRRQRRR (SEQ ID NO: 53). In another embodiment, the protein transduction domain is Pep-1 and comprises the sequence KETWWETWWTEWSQPKKKRKV (SEQ ID NO: 59).

[0235] Other non-limiting examples of PTD include an endosomal escape peptide. Non-limiting examples of such endosomal escape peptides include DT, GALA, PEA, INF-7, LAH4, CM18, HGP, H5WYG, HA2, and EB1.

[0236] Optimization of codon degeneracy

[0237] Because CRISPR nuclease proteins are (or are derived from) proteins normally expressed in bacteria, it may be advantageous to modify their nucleic acid sequences for optimal expression in eukaryotic cells (e.g., mammalian cells) when designing and preparing CRISPR nuclease recombinant proteins. Similarly, donor or patch nucleic acids of the present disclosure used to introduce specific modifications in the target polynucleotide may use codon degeneracy (e.g., to introduce new restriction sites for enabling easier detection of the targeted modification).

[0238] Accordingly, the following codon chart (Table 3) may be used, in a site-directed mutagenic scheme, to produce nucleic acids encoding the same or slightly different amino acid sequences of a given nucleic acid:

[0239] Table 3: Codons encoding the same amino acid

[0240] In an embodiment, a nucleic acid, vector, or expression construct described herein comprises a nucleic acid of interest, and in embodiments, the methods, uses and products herein relate to the use of a nucleic acid of interest. In embodiments, the term "nucleic acid of interest" or “gene of interest” is used to refer to a nucleic acid that encodes a functional peptide or polypeptide (protein) of interest (native or modified peptides / proteins). In an embodiment, the functional peptide or polypeptide is a therapeutic peptide or polypeptide, i.e., a peptide or polypeptide that can be administered to a subject for the purpose of treating or preventing a disease. Any nucleic acid encoding a peptide or polypeptide of interest known to those of ordinary skill in the art is contemplated for inclusion in the synthetic expression construct. The peptide or polypeptide of interest may be an enzyme, a signaling molecule (e.g., kinase, phosphatase), a receptor, a growth factor (e.g., cytokines), a chemotactic protein (e.g., chemokines), a structural protein (cytoskeletal proteins), a transcription factor, a cell adhesion protein, an antibody or antigen-binding fragment thereof, etc. The peptide or polypeptide may be a naturally-occurring peptide or polypeptide, a fragment or variant thereof, chimeric versions thereof, etc.

[0241] EXAMPLES

[0242] The present disclosure is illustrated in further detail by the following non-limiting examples.

[0243] Example 1 : Materials and methods

[0244] Hepa1-6 cell culture and transfection

[0245] Hepa1-6 cells were obtained from the ATCC (CRL-1830) and maintained at 37°C under 5% CO2 in Dulbecco’s modified Eagle’s medium (DMEM, high glucose, GlutaMAX™ Supplement) supplemented with 10% fetal bovine serum and 1% penicillin-streptomycin. The cells were tested and found negative for mycoplasma contamination. Hepa1-6 cells (2x105cells / transfection) were transfected with 500ng of a vector expressing CAG-driven St1Cas9 and a corresponding single guide RNA targeting the transferrin locus (see Table 3) using a Lonza 4D Nucleofector™ with a SF nucleofection kit (Lonza) and the EX-147 program. Cells were incubated for 5 minutes at room temperature after nucleofection before resuspension in fresh culture medium. Apart from these modifications, general manufacturer’s recommendations were followed. Cells were harvested 72h post-nucleofection.

[0246] 293F cell culture, transfection and protein production

[0247] 293F cells were purchased from Gibco and maintained at 37°C under 5% CO2 with 120 RPM orbital shaking in Freestyle™ 293 Expression Medium (Gibco). Recombinant transferrin-IDUA, transferrin and IDUA were produced as described elsewhere8. Briefly, 2 x 105293F cells were transfected with 250ng of eSpCas9(1.1)_No_FLAG_AAVS1_T2 (addgene #79888) and 750ng of a donor vector allowing integration within the AAVS1 locus and expression of one of the 3 proteins from the CAG promoter and based off addgene #164077 using Lipofectamine™ 3000 (Invitrogen). Puromycin selection (0.5pg / ml) was carried out for 10 days. The pools of puromycin-selected cells were expanded in fresh Freestyle™ 293 Expression Medium and cultured for 10 days before the culture medium was harvested. Protein purification from the supernatant was first carried out with Ni Sepharose™ 6 Fast Flow (Cytiva) followed by purification with Strep-Tactin®XT 4Flow® high capacity resin (I BA Lifesciences) according to the manufacturers’ instructions.

[0248] Genome editing vectors

[0249] Single guide RNAs were designed with CRISPOR9and manual sequence inspection and their sequences are provided in Table 3. Guide RNAs were cloned into mammalian expression vectors described elsewhere10(addgene #110626, #136651 and #136655 for guide RNAs with NNAGAA, NNACAA or NNGAAA protospacer- adjacent motifs (PAMs), respectively). The plasmid vector used to produce the nuclease rAAV8 vector targeting transferrin is based on addgene #110624 but expresses a variant of St1Cas9 that recognizes an NNACAA PAM. The plasmid vector used to produce the nuclease rAAV8 vector targeting albumin intron 1 is also based on addgene #110624, but expresses a variant of St1Cas9 that recognizes an NNGAAA PAM. The plasmid vectors used to produce the IDUA and GAA transferrin donor vectors were cloned using the backbone of the albumin intron 1 donor plasmid from Sharma et aD which was also directly used to produce the albumin intron 1 IDUA donor rAAV8. The IDUA sequence in both transferrin and albumin donor vectors is codon-optimised, whereas the GAA cDNA used in the transferrin-IDUA donor vector is not codon-optimised and corresponds to Genbank #BC040431 .1 (Horizon Discovery).

[0250] Table 3: guide RNAs used in the studies described herein

[0251] Adeno-associated virus production

[0252] The rAAV8s were produced by the Canadian Neurophotonics Platform’s Viral Vector Core (The Molecular Tools Platform) using the triple plasmid transfection method, as described12. Briefly, HEK293T17 (ATCC CRL-11268) cells were transfected using polyethylenimine with the helper plasmid pxx-68029, the rep / cap hybrid plasmid pAAV2 / 8 (addgene #112864), and the rAAV vectors described above. After 24 hours, the medium was replaced with medium without FBS, and the cells were harvested 24 hours later. We purified rAAV particles from the cell extracts using freeze / thaw lysis followed by a discontinuous iodixanol gradient. Viruses were resuspended in phosphate-buffered saline containing 320 mM NaCI, 5% D-sorbitol, and 0.001 % pluronic acid (F- 68), aliquoted, and stored at -80°C. The rAAVs were titrated by digital droplet PCR. Vector yields were 1 x1013- 3x1013VG / mL. The purity of viral preparations was determined by sodium dodecyl sulfate (SDS)-polyacrylamide gel electrophoresis on a 4-15% Mini-PROTEAN™ TGX™ Stain-Free Gel (Bio-Rad) in Tris-glycine-SDS buffer.

[0253] Animal experiments lduaW392Xmice (strain 017681)13were purchased from The Jackson Laboratory and backcrossed in a C57 / BI6J background. C57BL / 6J mice were also purchased from The Jackson Laboratory. Gaa6neomice (strain 004154)14 were also purchased from The Jackson Laboratory and backcrossed in a C57 / BI6J background. All mice were group-housed and fed a standard chow diet (Harlan #2018SX) with free access to food and water. Mice were exposed to a 12: 12-h dark-light cycle and kept at an ambient temperature of 23±1 °C. Animals were cared for and handled according to the Canadian Guide forthe Care and Use of Laboratory Animals. The Laval University Animal Care and Use Committee approved the procedures.

[0254] Neonatal (2-day-old) pups were injected intravenously in the retro-orbital sinus7with saline or the indicated amount of rAAV8, adjusted to an injection volume of 20 piL per eye with saline. All animals were weaned at 21 days old. Urine samples were collected before sacrifice by applying gentle pressure to the urinary bladder. Animals were perfused under anesthesia with 30ml of saline before sacrifice at predetermined time points. All collected solid tissues were frozen on dry ice.

[0255] TIDE assays

[0256] Genomic DNA was extracted from 2x105Hepa1-6 cells with QuickExtract™ DNA extraction solution (EpiCentre) following manufacturer’s recommendations or 30 mg of mouse liver using a EZ-10 Spin Column Animal Genomic DNA Miniprep Kit (Bio Basic), per the manufacturers’ recommendations. Loci were amplified by polymerase chain reaction (PCR) using the primers listed in Table 4. TIDE analysis was performed using a significance threshold value for decomposition of P < 0.00115.

[0257] Table 4: PCR primers used in the studies described herein.

[0258] In-out PCRs

[0259] Primers used in this study are provided in Table 4. PCR amplifications were performed with 30 cycles of amplification with Phusion™ polymerase. To detect targeted integration, in-out PCRs were performed with a primer binding outside the homology region and a primer binding inside the integration cassette. Genomic DNA from saline-treated animals was used as a control for all PCRs. l / l estem blot analyses

[0260] Western blot on mouse serum was performed on samples diluted in Laemmli buffer. For liver samples, 50mg of liver were homogenized with a FastPrep-24™ homogenizer (MP Biochemicals) in 1 ml of ice-cold phosphate- buffered saline (PBS) containing 0,1% Triton™ X-100 before centrifugation at 10,000g for 20mins at 4°C. Halt™ Protease and Phosphatase Inhibitor Cocktail (Thermo Scientific) was added to the clarified homogenates. Liver homogenate was next diluted in Laemmli buffer. All western blot samples were boiled for 10 min before loading on a 4-15% Mini-PROTEAN™ TGX™ Stain-Free Gel (Bio-Rad). Transfer was performed on 0,2pim nitrocellulose membranes. Total proteins on the membranes after transfer were visualized by stain-free imaging with a ChemiDoc™ MP imager (Bio-Rad). Membranes were next blocked with PBS containing 0,1% Tween-20 (PBST) and 5% milk. For IDUA detection, the following antibodies were used: anti-IDUA (1 / 3000, AF4119, Novus Biologicals), anti-sheep IgG HRP-conjugated (1 / 5000, HAF016, Novus Biologicals). For GAA detection, the following antibodies were used: anti-GAA (1 / 3000, ab137068, abeam), anti-rabbit IgG HRP-conjugated (1 / 5000, #7074, Cell Signaling Technology).

[0261] IDUA enzymatic assays

[0262] IDUA enzyme activity was determined with a fluorometric assay using 4-Methylumbelliferyl alpha-L-iduronide as previously described16. Tissues were lysed in PBS containing 0,1 % Triton X-100 as described above using a FastPrep-24 homogenizer. Briefly, 25piL of recombinant protein, serum or tissue lysates pre-diluted in PBS containing 0,1 % Triton X-100 were added to 25piL of a mix containing 360piM of 4-Methylumbelliferyl alpha-L- iduronide (Glycosynth, #44076) and 0.4M sodium formate buffer (pH 3,5). Tissue and serum samples were incubated 30 min at 37°C before adding 200piL of 200mM glycine-carbonate buffer (pH 10,4) to stop the reaction. For the assays with recombinant proteins, samples were incubated 10 min at 37°C. 4-Methylumbelliferone (4-MU, Sigma #M1381) was used to make the standard curve. Fluorescence was measured using a SpectraMax™ i3 plate reader. Protein concentrations were determined using the Pierce Protein Assay Reagent (Thermo Fisher Scientific). Data are shown as nmol of 4-MU released per hour, per mL of serum or mg of protein for assays with tissue and serum samples, or nmol of 4-MU released per minute per pig of protein for assays with recombinant protein.

[0263] GAA enzymatic assays

[0264] GAA enzyme activity was determined with a fluorometric assay using 4-Methylumbelliferyl alpha-L-iduronide as previously described17. Briefly, 25piL of or tissue lysates pre-diluted in PBS containing 0,1 % Triton™ X-100 were added to 25piL of a mix containing 3mM of 4-Methylumbelliferyl alpha-D-glucopyranoside (Glycosynth, #44051) and 0.2M sodium acetate buffer (pH 4,3). Samples were incubated 1h at 37°C before adding 200piL of 200mM glycine-carbonate buffer (pH 10,4) to stop the reaction. 4-Methylumbelliferone (4-MU, Sigma #M 1381 ) was used to make the standard curve. Fluorescence was measured using a SpectraMax™ i3 plate reader. Protein concentrations were determined using the Pierce Protein Assay Reagent (Thermo Fisher Scientific). Data are shown as nmol of 4-MU released per hour, per mL of serum or mg of protein.

[0265] Glycosaminoglycan quantification in urine and tissue samples

[0266] Urinary glycosaminoglycan (GAGs) levels were quantified by LC-MS / MS as previously described18. Urine samples were pooled before analysis. For determination of liver and brain GAG levels, 100mg of tissue were homogenized in 1 ml of ice-cold methanol with a FastPrep-24 homogenizer before a final dilution at 25mg tissue / mL. Tissue glycosaminoglycans were quantified by UPLC-MS / MS as others have described19.

[0267] Rotarod and grip strength assays

[0268] Muscle function was assessed with rotarod and grip strength assays the Gaa / - animals 6 months after treatment and age-matched C57BL / 6J animals. For the Rotarod tests, a ROTA-ROD / RS (Panlab Harvard Apparatus) was programmed to accelerate at 0,12 RPM / min and latency to fall of 4 trials was measured. Animals were rested for 10mins between trials. Forelimb grip strength was measured in triplicate with a Grip Strength Meter (Columbus Instruments).

[0269] Example 2: Selection of the integration site within the transferrin locus and screening for active singleguide RNAs

[0270] Transferrin is a monomeric ~80kDa glycoprotein comprised of an N-terminal and a C-terminal lobe each containing an iron-binding site5’20. It has been shown in vitro that addition of a short tag at either terminus of transferrin does not impair binding to the transferrin receptors21. Others have also shown that it is possible to fuse a-L-iduronidase to the C-terminus of transferrin and express a protein with detectable activity in vitro and in vivo in a mouse model of mucopolysaccharidosis type I (MPS I aka Hurler syndrome)4. We chose the last intron of the mouse transferrin gene as the integration site for the therapeutic transgene to generate a C-terminal fusion product, as shown in Figure 1.

[0271] Figure 1 shows a protein replacement strategy by insertion and fusion of a therapeutic transgene at the murine transferrin locus. This strategy should result in high circulating levels of therapeutic fusion proteins secreted by the liver and able to reach tissues expressing the transferrin receptors. This is notably the case of the endothelial cells composing the blood-brain barrier, which means that the fusion protein could potentially cross from the circulation to the brain parenchyma and have a therapeutic effect there. Of note, endocytosis of the transferrin-LSD enzymes could occur via the M6PR and / or the transferrin receptors with a potential synergy between these two pathways for tissue uptake.

[0272] We then screened several single-guide RNAs (sgRNAs) targeting a Cas9 nuclease from Streptococcus thermophiius (St1 Cas9)10to this region by transient plasmid transfection in the Hepa1-6 mouse hepatoma cell line as shown in Figure 5.

[0273] Example 3: Purification and validation of an enzymatically active transferrin-IDUA fusion protein As a proof of concept, we next produced and purified recombinant mouse transferrin, alpha-L-lduronidase (IDUA), and a transferrin-IDUA fusion in 293F cells (Figure 6)8 22. IDUA is the missing lysosomal hydrolase involved in Hurler syndrome, a disease characterized by the body-wide accumulation of glycosaminoglycans (GAGs), mostly heparan and dermatan sulfate with skeletal, cardiac and neurological involvement30.

[0274] We found that the protein obtained with the transferrin-IDUA expression construct was of the same molecular weight as the expected gene product following integration of the donor within the transferrin locus as shown in Figure 1, with the purified transferrin and IDUA proteins also corresponding to their reported molecular weights (Figures 6A-C). We determined that both IDUA and the transferrin-IDUA fusion showed enzymatic activity against the fluorogenic substrate 4-methylumbelliferyl-a-L-iduronide (Figure 6D)16. Both purified proteins were also detectable by western blot when spiked in Idua-'- mouse liver extracts and were approximately of the expected molecular weight (Figures 6E-F).

[0275] Example 4: Targeted integration of the IDUA donor constructs in neonatal Idua '- mice

[0276] Once we determined that recombinant transferrin-IDUA fusion showed in vitro enzymatic activity in our system, we set out to test our strategy in vivo in iduaW392Xmice (referred to here as Idua' mice), an animal model of the lysosomal storage disease mucopolysaccharidosis type I, also known as Hurler syndrome13. These mice show complete lack of IDUA enzyme activity coupled with progressive accumulation of dermatan sulfate and heparan sulfate, two glycosaminoglycans seen in Hurler patients, in most organs13.

[0277] We produced two recombinant adeno-associated viral vectors serotype 8 (rAAV8), one containing an expression construct for St1Cas9 and the selected guide RNA targeting the transferrin locus and the other built as in Figure 1 and designed to generate mature human IDUA fused at the C-terminus of transferrin. As a control, we targeted human IDUA to the first intron of albumin in order to produce the native enzyme at high levels3’11’23. This specific strategy results in native enzymes being secreted by the liver into circulation. We have previously identified a sgRNA targeting St1Cas9 to the same target site within the first intron of the mouse albumin locus as the zinc finger nuclease (ZFN) used in these earlier studies2’10. We thus produced an rAAV8 human IDUA donor construct identical to the one used by Sharma et ai. and a second rAAV8 vector containing an expression construct for St1 Cas9 and the guide RNA targeting the albumin locus.

[0278] We then injected neonatal (2 days old) Idua ' pups with a combination of both nuclease and donor vectors for each strategy and sacrificed them at 6 months of age. A group of ldua / - animals were injected with saline as controls, and age-matched C57 / BI6J animals were used as wild-type controls. We also injected Idua' animals with only the donor vector to validate that efficient transgene integration and subsequent protein production was achieved through nuclease-mediated targeting. We detected robust supraphysiological IDUA enzymatic activity in the serum, liver, brain and heart of the animals injected with both nuclease and donor vectors for both targeting strategies (Figure 7). We also observed that animals injected with only the donor vectors consistently showed enzymatic activity orders of magnitude below those of animals injected with both nuclease and donor vectors, confirming that nuclease-driven integration is necessary for efficient protein production in this system.

[0279] Example 5: Nuclease-driven integration at the transferrin locus allows for long-term production and secretion of enzymatically active transferrin-IDUA fusion proteins by the liver

[0280] We next injected Iduar'- neonates either with both nuclease and donor vectors for the albumin and transferrin targeting approaches and sacrificed the animals at 6 months of age to assess the long-term therapeutic effects of both strategies. Western blot analysis on liver, serum and heart samples provided insights into the production, secretion, uptake and processing of the therapeutic proteins. We also conducted IDUA enzymatic activity assays on the serum, liver, brain and heart of treated animals. Saline-treated lduaW392Xmice and C57BL / 6J of a similar age were used as controls (Figure 2). Western blot analysis showed that the expected transferrin-IDUA fusion product was indeed secreted at high steady-state levels into circulation, whereas the IDUA protein produced through the albumin strategy was only detected in the liver although at higher levels than for the transferrin strategy (Figure 2A). Moreover, while only the full length transferrin-IDUA fusion was detected in the serum, a mix of full- length transferrin-IDUA and cleaved IDUA protein was found in the liver. Of note, the IDUA polypeptide is naturally cleaved and glycosylated during its maturation on its way to the lysosome3’11’23. The IDUA banding pattern observed in the liver in the albumin targeting strategy and the cleaved protein for the transferrin strategy are identical, suggesting proper processing of the transferrin-IDUA fusion. This also suggests that the fusion protein is efficiently secreted by hepatocytes into circulation and is processed into the mature form upon re-entry and routing to the lysosome. Alternatively, a small fraction of transferrin-IDUA may get directly routed to the lysosome before secretion in hepatocytes while the majority of the fusion protein is secreted.

[0281] We next found that the levels of IDUA enzyme activity in the serum were ~2 log higher in the transferrin-targeted animals while activity in the liver was higher in the albumin-targeted group (Figure 2B). We observed low but detectable levels of activity in the brain for all transferrin-targeted animals, suggesting that the transferrin-IDUA fusion protein can cross the blood-brain barrier. We also found high levels of enzymatic activity in the heart of all the animals treated with the transferrin approach.

[0282] Next, DNA was extracted from whole liver pieces. We first confirmed editing at each locus as measured by the TIDE assay for each respective targeting strategy (Figure 8A).15. We then assessed the integration of the donor construct for each strategy. Both transferrin and albumin targeting strategies produced targeted integration of the IDUA transgene by homology-directed repair (HDR) or non-homologous end joining (NHEJ) that was detectable by “in-out” PCR (Figure 8B).

[0283] Example 6: Long-term treatment with the transferrin strategy allows complete normalization of glycosaminoglycan levels in the urine, liver and brain of treated animals

[0284] Next, we sought to determine if the high IDUA activity levels detected resulted in the durable reduction or elimination of the lysosomal substrates typical of Hunter syndrome in those older Iduar'- animals. Glycosaminoglycan (GAGs) levels in the urine, liver and brain were determined for each strategy by ultraperformance (UP)LC-MSZMS (Table 5)19.

[0285] Table 5: Long-term treatment of Idua'- animals with the transferrin strategy allows complete normalization of glycosaminoglycan levels. Tissue and urine samples from animals described in Figure 4 were analyzed by ultra-performance (UP)LC-MS / MS to determine levels of dermatan sulfate (DS) and heparan sulfate (HS). Mean values are indicated for each group.

[0286] Dermatan and heparan sulfate levels showed complete normalization in the urine, liver and brain of the animals treated with the transferrin approach, suggesting systemic tissue uptake of the transferrin-IDUA fusion protein from the circulation and successful entry into the lysosome. Partial normalization was also observed in the brain of the mice treated with the albumin approach, an observation similar to what others have shown when high circulating levels of the native enzyme are reached in neonatal mice2’23.

[0287] Collectively, these data show that the transferrin targeting strategy can provide significant long-term biochemical relief in a mouse model of Hurler syndrome.

[0288] Example 7: Efficient expression and muscle distribution of a transferrin fusion protein in a mouse model of Pompe disease

[0289] We then adapted our transferrin targeting approach in GaaBneo(referred here as Gaaz) mice, an animal model of Pompe disease. These animals show complete lack of GAA activity coupled with progressive glycogen accumulation in multiple tissues, cardiomegaly, reduced mobility and progressive muscle weakness14We first designed an rAAV8 donor vector to allow expression of the mature human GAA protein fused at the C-terminus of transferrin, similar to our transferrin-IDUA construct. We next injected neonatal (2 days old) Gaa '- pups with a combination of both nuclease and donor vectors or with the donor vector alone and sacrificed them at 1 month of age. Age-matched groups of untreated Gaa4and C57 / BI6J animals were used as controls. Here, we included animals injected only with the donor vector to show that St1Cas9-driven targeted integration at the transferrin locus was required for efficient integration and protein production. We performed western blot analysis as well as GAA enzymatic activity assays17on serum and liver samples. As Pompe disease predominantly affects the central nervous system as well as skeletal and cardiac muscles, we also conducted GAA activity assays on brain, heart, leg muscle and diaphragm samples (Figure 10)24’17’25’26. We detected GAA activity in all tested tissues irrespective of sex and observed high levels of serum GAA enzyme activity. We also saw levels of enzymatic activity superior to those observed in C57 / BL6J animals in cardiac and skeletal muscles. This observation is important, as native GAA is known to be inefficiently secreted and to present low uptake in skeletal muscle27’24 17’26. Furthermore, we also observed low but detectable levels of GAA activity in the brain, suggesting that fusion with transferrin allowed a fraction of the secreted enzyme to cross the blood-brain barrier. Animals injected with the donor vector alone showed minimal GAA activity.

[0290] We also performed western blots directed against GAA on serum and liver samples. Similar to what we observed in Iduad- animals, we were able to detect the transferrin-GAA protein in the serum and a mix of fused and processed GAA protein in the liver. The processed protein in the liver showed a similar banding pattern to native GAA, which is cleaved at both its N- and C-terminus28.

[0291] Of note, no GAA protein was observed in western blots samples from animals injected with the donor vector alone, and GAA enzyme levels were similar to those of untreated animals for all tested tissues. This confirms that nuclease-driven integration is necessary for efficient protein production in this system.

[0292] Example 8: Targeting at the transferrin locus produces sustained transgene expression and phenotype correction in a mouse model of Pompe disease

[0293] We finally injected a group of Gaa7neonates with both donor and nuclease vectors for the transferrin targeting strategy, conducted functional assays at 6 months of age and sacrificed all animals at 8 months of age. Age- matched saline-injected Gaa7' and C57BL / 6J animals were used as controls.

[0294] We first performed western blots directed against GAA on serum, liver, heart, leg muscle and diaphragm samples and enzymatic activity assays on serum, liver, brain, heart, leg muscle and diaphragm samples (Figure 4). Similar to what we observed in Idua-'- animals, we were able to detect the transferrin-GAA protein in the serum, processed GAA protein in the liver and a mix of fused and processed protein in the cardiac and skeletal muscles (Figure 4A). The processed protein showed a banding pattern identical to native GAA, in which the precursor protein is cleaved at both it’s N- and C-termini upon entry in the lysosome39. We also found that while activity levels in the brain were comparable to those of saline-treated animals, sustained enzyme levels were achieved for all other tested tissues and were generally similar to those of C57BL / 6 J animals for the tested muscles (Figure 4B).

[0295] Collectively, our data show that targeting the transferrin locus in the liver to generate transferrin fusion proteins is a promising platform approach for the systemic treatment of a number of disorders, such as lysosomal storage disorders.

[0296] Table 6: Sequences described herein

[0297] Although the present invention has been described hereinabove by way of specific embodiments thereof, it can be modified, without departing from the spirit and nature of the subject invention as defined in the appended claims. In the claims, the word "comprising" is used as an open-ended term, substantially equivalent to the phrase "including, but not limited to". The singular forms "a", "an" and "the" include corresponding plural references unless the context clearly dictates otherwise.

[0298] REFERENCES

[0299] 1 . Lanpher, B., Brunetti-Pierri, N. & Lee, B. Inborn errors of metabolism: the flux from Mendelian to complex diseases. Nat Rev Genet 7, 449-460 (2006).

[0300] 2. Platt, F.M., d’Azzo, A., Davidson, B.L., Neufeld, E.F. & Tifft, C.J. Lysosomal storage diseases. Nature Reviews Disease Primers 4, 27 (2018).

[0301] 3. Ou, L. et al. ZFN-Mediated In Vivo Genome Editing Corrects Murine Hurler Syndrome. Molecular therapy 27, 178-187 (2019).

[0302] 4. Osborn, M.J., McElmurry, R.T., Peacock, B., Tolar, J. & Blazar, B.R. Targeting of the CNS in MPS-IH Using a Nonviral Transferrin-alpha-l-iduronidase Fusion Gene Product. Molecular Therapy 16, 1459-1466 (2008).

[0303] 5. Gomme, P.T., McCann, K.B. & Bertolini, J. Transferrin: structure, function and potential therapeutic actions. Drug Discovery Today 10, 267-273 (2005).

[0304] 6. Kim, B.-J. et al. Transferrin Fusion Technology: A Novel Approach to Prolonging Biological Half-Life of Insulinotropic Peptides. Journal of Pharmacology and Experimental Therapeutics 334, 682-692 (2010).

[0305] 7. Yardeni, T., Eckhaus, M., Morris, H.D., Huizing, M. & Hoogstraten-Miller, S. Retro-orbital injections in mice. Lab Anim (NY) 40, 155-160 (2011).

[0306] 8. Rivest, J.F., Goupil, C. & Doyon, Y., Vol. 2021 (protocols. io; 2021).

[0307] 9. Concordet, J.P. & Haeussler, M. CRISPOR: intuitive guide selection for CRISPR / Cas9 genome editing experiments and screens. Nucleic acids research 46, W242-w245 (2018).

[0308] 10. Agudelo, D. et al. Versatile and robust genome editing with Streptococcus thermophilus CRISPR1-Cas9. Genome research 30, 107-117 (2020).

[0309] 11. Sharma, R. et al. In vivo genome editing of the albumin locus as a platform for protein replacement therapy. Blood 126, 1777-1784 (2015).

[0310] 12. Gray, S.J. et al. Production of recombinant adeno-associated viral vectors and use in in vitro and in vivo administration. Curr Protoc Neurosci Chapter 4, Unit 4 17 (2011).

[0311] 13. Wang, D. et al. Characterization of an MPS l-H knock-in mouse that carries a nonsense mutation analogous to the human IDUA-W402X mutation. Molecular Genetics and Metabolism 99, 62-71 (2010).

[0312] 14. Raben, N. et al. Targeted Disruption of the Acid Alpha-Glucosidase Gene in Mice Causes an Illness with Critical Features of Both Infantile and Adult Human Glycogen Storage Disease Type II * Journal of Biological Chemistry 273, 19086-19092 (1998).

[0313] 15. Brinkman, E.K., Chen, T., Amendola, M. & van Steensel, B. Easy quantitative assessment of genome editing by sequence trace decomposition. Nucleic acids research 42, e168-e168 (2014).

[0314] 16. Ou, L, Herzog, T.L., Wilmot, C.M. & Whitley, C.B. Standardization of a-L-iduronidase enzyme assay with Michaelis-Menten kinetics. Molecular Genetics and Metabolism 111, 113-115 (2014). 17. Meena, N.K., Randazzo, D., Raben, N. & Puertollano, R. AAV-mediated delivery of secreted acid a- glucosidase with enhanced uptake corrects neuromuscular pathology in Pompe mice. JCI Insight 8 (2023).

[0315] 18. Auray-Blais, C. et al. Efficient analysis of urinary glycosaminoglycans by LC-MS / MS in mucopolysaccharidoses type I, II and VI. Molecular Genetics and Metabolism 102, 49-56 (2011).

[0316] 19. Menkovic, I., Lavoie, P., Boutin, M. & Auray-Blais, C. Distribution of heparan sulfate and dermatan sulfate in mucopolysaccharidosis type II mouse tissues pre- and post-enzyme-replacement therapy determined by UPLC-MS / MS. Bioanalysis 11, 727-740 (2019).

[0317] 20. Zak, O. & Aisen, P. A poly-His tag method for obtaining the C-terminal lobe of human transferrin. Protein Expression and Purification 28, 120-124 (2003).

[0318] 21. Mason, A.B. et al. Differential Effect of a His Tag at the N- and C-Termini: Functional Studies with Recombinant Human Serum Transferrin. Biochemistry 41 , 9448-9454 (2002).

[0319] 22. Dalvai, M. et al. A Scalable Genome-Editing-Based Approach for Mapping Multiprotein Complexes in Human Cells. Cell Reports 13, 621-633 (2015).

[0320] 23. Ou, L. et al. A Highly Efficacious PS Gene Editing System Corrects Metabolic and Neurological Complications of Mucopolysaccharidosis Type I. Molecular Therapy 28, 1442-1454 (2020).

[0321] 24. Baik, A.D. et al. Cell type-selective targeted delivery of a recombinant lysosomal enzyme for enzyme therapies. Moleculartherapy : the journal of the American Society of Gene Therapy2S, 3512-3524 (2021).

[0322] 25. Kishnani, P.S., Sun, B. & Koeberl, D.D. Gene therapy for glycogen storage diseases. Human Molecular Genetics 28, R31-R41 (2019).

[0323] 26. Puzzo, F. et al. Rescue of Pompe disease in mice by AAV-mediated liver delivery of secretable acid a- glucosidase. Science translational medicine 9, eaam6375 (2017).

[0324] 27. Kishnani, P.S. & Koeberl, D.D. Liver depot gene therapy for Pompe disease. Annals of Translational Medicine 7, 288 (2019).

[0325] 28. Moreland, R.J. et al. Species-specific differences in the processing of acid a-glucosidase are due to the amino acid identity at position 201. Gene 491, 25-30 (2012).

[0326] 29. Gray, S.J. et al. Production of recombinant adeno-associated viral vectors and use in in vitro and in vivo administration. Curr Protoc Neurosci. 2011 October ; CHAPTER: Unit 4.17, pp. 1-36.

[0327] 30. Bie, H. et al. Insights into mucopolysaccharidosis I from the structure and action of a-L-iduronidase. Nature Chemical Biology 9, 739-745 (2013).

Claims

CLAIMS:1 . A method for expressing a polypeptide of interest in a cell, the method comprising: preparing an expression construct, the expression construct encoding a fusion polypeptide comprising a first domain comprising transferrin or a fragment or derivative thereof having transferrin activity and a second domain comprising the polypeptide of interest, wherein preparing the expression construct comprises introducing, using a nuclease-based method, a first nucleic acid comprising a first nucleotide sequence encoding the polypeptide of interest into a target region within an endogenous transferrin gene of the cell; and allowing expression of the fusion polypeptide from the expression construct.

2. The method of claim 1 , wherein the second domain is C-terminal to the first domain.

3. The method of claim 1 or 2, wherein the fusion polypeptide comprises a linker region between the first and second domains.

4. The method of any one of claims 1-3, wherein the target region is within an intron of the transferrin gene.

5. The method of claim 4, wherein the target region is within an intron other than intron 1 of the transferrin gene.

6. The method of claim 4 or 5, wherein the target region is within intron 16 of the transferrin gene.

7. The method of any one of claims 4-6, wherein the first nucleic acid sequence further comprises, 5’ to the first nucleotide sequence, one or more exons of the transferrin gene 3’ to the target region.

8. The method of claim 7, wherein the target region is within intron 16 of the transferrin gene and wherein the one or more exons of the transferrin gene 3’ to the target region is exon 17 of the transferrin gene, and wherein the first nucleic acid comprises, in a 5’-3’ direction, exon 17 of the transferrin gene and the first nucleotide sequence.

9. The method of any one of claims 3-8, wherein the first nucleic acid further comprises a linker nucleotide sequence encoding the linker region, wherein the linker nucleotide sequence is located between the one or more exons of the transferrin gene 3’ to the target region and the first nucleotide sequence, such that the first nucleic acid comprises, in a 5’-3’ direction, (i) the one or more exons of the transferrin gene 3’ to the target region, (ii) the linker nucleotide sequence, and (iii) the first nucleotide sequence.

10. The method of claim 9, wherein the target region is within intron 16 of the transferrin gene, wherein the one or more exons of the transferrin gene 3’ to the target region is exon 17 of the transferrin gene and wherein the first nucleic acid comprises, in a 5’ -3’ direction, (i) exon 17 of the transferrin gene, (ii) the linker nucleotide sequence, and (iii) the first nucleotide sequence.11 . The method of any one of claims 1-10, wherein the first nucleotide sequence lacks a sequence encoding a signal sequence endogenous to the polypeptide of interest.

12. The method of any one of claims 1-11, wherein the expressed fusion polypeptide is not cleaved between the first and second domains, thereby to produce the polypeptide of interest comprised within the fusion polypeptide.

13. The method of any one of claims 1-11, wherein the expressed fusion polypeptide is cleaved between the first and second domains, thereby to produce the polypeptide of interest in a form lacking the transferrin or fragment or derivative thereof having transferrin activity.

14. The method of any one of claims 1-13, wherein the the method comprises providing the cell with (a) a CRISPR nuclease or a nucleic acid encoding the CRISPR nuclease, (b) one or more gRNAs comprising one or more guide sequences having one or more target sequences within the target region of the transferrin gene, or one or more nucleic acids encoding the one or more gRNAs, wherein the one or more target sequences are each contiguous to a protospacer adjacent motif recognized by the CRISPR nuclease, and (c) one or more donor nucleic acids comprising the first nucleic acid, wherein the one or more gRNAs direct the cleavage of the transferrin gene at the target region thereby to allow introduction of the first nucleic acid at the target region.

15. The method of claim 14, wherein the method comprises providing the cell with one or more vectors comprising (a) the nucleic acid encoding the CRISPR nuclease, (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids comprising the first nucleic acid.

16. The method of claim 15, wherein the method comprises providing the cell with a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and a further vector comprising (c) the one or more donor nucleic acids comprising the first nucleic acid.

17. The method of any one of claims 14-16, wherein the CRISPR nuclease is a Cas9 nuclease.

18. The method of any one of claims 14-17, wherein the first nucleic acid is introduced by homology- directed repair (HDR).

19. The method of any one of claims 1-15, wherein the polypeptide of interest has an activity which is deficient in the cell prior to expression of the polypeptide of interest.

20. The method of any one of claims 1-19, wherein the polypeptide of interest is a therapeutic protein.21 . The method of any one of claims 1-20, wherein the polypeptide of interest is an enzyme.

22. The method of claim 21, wherein the enzyme is a lysosomal enzyme.

23. The method of claim 21 or 22, wherein the enzyme is an alpha-L-iduronidase (IDUA), an acid alphaglucosidase (GAA), an iduronate 2-sulfatase, a heparan sulfamidase, an N-acetyl-alpha- glucosaminidase, a heparane-alpha-glucosaminide N-acetyltransferase, a glucosamine N-acetyl-6- sulfatase, or a betaglucuronidase.

24. The method of any one of claims 1-23, wherein the cell is a hepatic cell.

25. The method of any one of claims 6-24, wherein the target region is within SEQ ID NO: 20 (FIG. 9).

26. The method of any one of claims 1-25, wherein the method an in vitro method.

27. The method of any one of claims 1-25, wherein the method is an in vivo method and the cell is within a subject.

28. The method of claim 27, wherein the polypeptide of interest is secreted from the cell.

29. The method of claim 27 or 28, wherein the polypeptide of interest is present in a tissue or body fluid of the subject following its expression.

30. The method of claim 29, wherein the polypeptide of interest is present in the serum of the subject following its expression.31 . The method of claim 29, wherein the polypeptide of interest is present in the nervous system of the subject following its expression.

32. One or more gRNAs as defined in any one of claims 14-31.

33. One or more donor nucleic acids as defined in any one of claims 14-32.

34. An isolated nucleic acid comprising the expression construct as defined in any one of claims 1-31 .

35. An isolated polypeptide comprising the amino acid sequence of the fusion polypeptide as defined in any one of claims 1-31.

36. A vector comprising one or more nucleic acid sequences corresponding to the one or more gRNAs of claim 32.

37. The vector of claim 36, further comprising a nucleic acid encoding a CRISPR nuclease.

38. The vector of claim 37, wherein the CRISPR nuclease is a Cas9 nuclease.

39. A vector comprising the one more donor nucleic acids of claim 33.

40. A cell comprising the one or more gRNAs of claim 32, the one or more donor nucleic acids of claim 33, the expression construct as defined in any one of claims 1-31 , the isolated polypeptide of claim 35, and / or the vector of any one of claims 36-39.41 . A composition comprising the one or more gRNAs of claim 32, the one or more donor nucleic acids of claim 33, the expression construct as defined in any one of claims 1-31 , the isolated polypeptide of claim 35, the vector of any one of claims 36-39, and / or the cell of claim 40.

42. The composition of claim 41 , further comprising a pharmaceutically acceptable carrier.

43. A method of preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest, comprising expressing the polypeptide interest in a cell of the subject according to the method of any one of claims 1-31.

44. A method of preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest, comprising administering to the subject an effective amount of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of claims 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of claim 33; or the composition of claim 41 or 42.

45. The method of claim 44, comprising administering to the subject an effective amount of a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and a further vector comprising (c) the one or more donor nucleic acids.

46. The method of any one of claims 43-45, wherein the one or more vectors is / are a viral vector.

47. The method of any one of claims 43-46, wherein the disease or condition is a metabolic disease.

48. The method of any one of claims 43-47, wherein the disease or condition is a lysosomal disease.

49. The method of any one of claims 43-48, wherein the disease or condition is a mucopolysaccharidosis(MPS).

50. The method of claim 49, wherein the disease or condition is a type 1, 11, 111 or VI MPS.51 . The method of any one of claims 43-50, wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

52. The method of any one of claims 43-47, wherein the disease or condition is a glycogen storage disease.

53. The method of any one of claims 43-47 and 52, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

54. One or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of claims 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of claim 33; or the composition of claim 41 or 42, for use in preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

55. The one or more vectors or composition for use of claim 54, wherein the one or more vectors comprise (i) a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (ii) a further vector comprising (c) the one or more donor nucleic acids.

56. The one or more vectors or composition for use of claim 54 or 55, wherein the one or more vectors is / are a viral vector.

57. The one or more vectors or composition for use of any one of claims 54-56, wherein the disease or condition is a metabolic disease.

58. The one or more vectors or composition for use of any one of claims 54-57, wherein the disease or condition is a lysosomal disease.

59. The one or more vectors or composition for use of any one of claims 54-58, wherein the disease or condition is a mucopolysaccharidosis (MPS).

60. The one or more vectors or composition for use of claim 59, wherein the disease or condition is a type I, II, III or VI MPS.61 . The one or more vectors or composition for use of any one of claims 54-60, wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

62. The one or more vectors or composition for use of any one of claims 54-57, wherein the disease or condition is a glycogen storage disease.

63. The one or more vectors or composition for use of any one of claims 54-57 and 62, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

64. Use of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of claims 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of claim 33; or the composition of claim 41 or 42, for preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

65. Use of one or more vectors comprising (a) a nucleic acid encoding a CRISPR nuclease, (b) one or more nucleic acid sequences corresponding to the one or more gRNAs as defined in any one of claims 14-31 for expressing the one or more gRNAs, and (c) the one or more donor nucleic acids of claim 33; or the composition of claim 41 or 42, for the preparation of one or more medicaments for preventing or treating a disease or condition in a subject that can benefit from the expression of a polypeptide of interest.

66. The use of claim 64 or 65, wherein the one or more vectors comprise (i) a vector comprising (a) the nucleic acid encoding the CRISPR nuclease and (b) the one or more nucleic acid sequences corresponding to the one or more gRNAs for expressing the one or more gRNAs, and (ii) a further vector comprising (c) the one or more donor nucleic acids.

67. The use of any one of claims 64 or 66, wherein the one or more vectors is / are a viral vector.

68. The use of any one of claims 64-67, wherein the disease or condition is a metabolic disease.

69. The use of any one of claims 64-68, wherein the disease or condition is a lysosomal disease.

70. The use of any one of claims 64-69, wherein the disease or condition is a mucopolysaccharidosis(MPS).71 . The use of claim 70, wherein the disease or condition is a type I, II, III or VI MPS.

72. The use of any one of claims 64-71 , wherein the disease or condition is Hurler syndrome and the polypeptide of interest is an alpha-L-iduronidase (IDUA).

73. The use of any one of claims 64-69, wherein the disease or condition is a glycogen storage disease.

74. The use of any one of claims 64-69 and 73, wherein the disease or condition is Pompe disease and the polypeptide of interest is an acid alpha-glucosidase (GAA).

Citation Information

Patent Citations

  • Compositions and methods for gene editing by targeting transferrin

    WO2019140330A1

  • Genetically-modified cells comprising a modified transferrin gene

    WO2020146807A1