Compositions and methods for the treatment of abnormal hemoglobin disorders

The CRISPR system targets the WIZ gene in hematopoietic stem cells to increase fetal hemoglobin and decrease β-globin expression, effectively treating abnormal hemoglobin disorders by enhancing erythrocyte differentiation and function.

JP7846007B2Active Publication Date: 2026-04-14NOVARTIS AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NOVARTIS AG
Filing Date
2020-12-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current treatments for abnormal hemoglobin disorders such as sickle cell anemia and β-thalassemia are inadequate in effectively increasing fetal hemoglobin production and reducing disease-causing β-globin expression in erythrocytes.

Method used

The use of a CRISPR system to target and modify the WIZ gene in hematopoietic stem cell progenitor cells, leading to increased fetal hemoglobin expression and decreased β-globin expression, with the modified cells being capable of differentiation and engraftment in organisms, and treated ex vivo to promote expansion and proliferation.

Benefits of technology

The CRISPR-modified cells demonstrate significant upregulation of fetal hemoglobin and reduction in sickle β-globin, resulting in decreased sickle cell count and increased normal red blood cell count in erythroid offspring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846007000203
    Figure 0007846007000203
  • Figure 0007846007000204
    Figure 0007846007000204
  • Figure 0007846007000205
    Figure 0007846007000205
Patent Text Reader

Abstract

The present invention relates to compositions and methods for treating hemoglobinopathies.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Sequence List This application includes a sequence listing submitted electronically in ASCII format, which is incorporated herein by reference in its entirety. The ASCII copy, created on 16 December 2020, is named PAT058744-WO-PCT_SL.txt and has a size of 1,094,185 bytes.

[0002] Claim of Profit This application claims priority to U.S. Provisional Patent Applications No. 62 / 950,025 and No. 62 / 950,048, filed on 18 December 2019, the disclosures of which are incorporated herein by reference in their entirety. [Background technology]

[0003] CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) evolved in bacteria as an adaptive immune system to defend against viral attacks. Upon exposure to a virus, a short segment of viral DNA is incorporated into the CRISPR locus of the bacterial genome. RNA is transcribed from a portion of the CRISPR locus containing the viral sequence. This RNA, containing a sequence complementary to the viral genome, mediates the targeting of the Cas9 protein to a sequence in the viral genome. The Cas9 protein cleaves the viral target, thereby silencing it.

[0004] In recent years, the CRISPR / Cas system has been applied to genome editing in eukaryotic cells. By introducing site-directed single-strand breaks (SSBs) or double-strand breaks (DSBs), it becomes possible to modify target sequences using, for example, non-homologous end joining (NHEJ) or homologous recombination repair (HDR). [Overview of the Initiative] [Means for solving the problem]

[0005] While not bound by theory, the present invention is in part based on unexpected findings regarding the relationship between WIZ gene expression / protein activity and hemoglobin F (HbF) production. As demonstrated in the examples and figures, knockdown or knockout of the WIZ gene or WIZ protein in cells (by various modalities / compositions described herein) significantly increased HbF induction in such cells, thereby treating HbF-related conditions and disorders (e.g., abnormal hemoglobin disorders, such as sickle cell anemia and β-thalassemia). The present invention is also, in part, based on the discovery that modifying cells (e.g., hematopoietic stem cell progenitor cells (HSPCs)) with the WIZ gene, for example, using a CRISPR system, for example, the Cas9 CRISPR system, as described herein, can increase fetal hemoglobin (HbF) expression and / or decrease β-globin expression (e.g., β-globin gene with disease-causing mutations) in the offspring of the modified cells, for example, erythrocyte offspring, and that modified cells (e.g., modified HSPCs) can be used to treat abnormal hemoglobin disorders, such as sickle cell disease and β-thalassemia. In one embodiment, surprisingly, this specification shows that when a gene editing system, such as the CRISPR system, as described herein, is introduced into cells targeting the WIZ gene (e.g., HSPCs), modified HSPCs (e.g., HSPCs containing one or more indels, as described herein) are produced that efficiently engraft in organisms, persist for long periods in the engrafted organisms, and are capable of differentiation, including differentiation into erythrocytes accompanied by increased fetal hemoglobin expression. In addition, these modified HSPCs can be ex vivo cultured under conditions that promote their expansion and proliferation while maintaining their stem cell properties, for example, in the presence of a stem cell proliferation agent (e.g., as described herein).For example, when a gene editing system as described herein, such as the CRISPR system, is introduced into HPSCs derived from patients with sickle cell disease, the modified cells and their offspring (e.g., erythroid offspring) unexpectedly show not only upregulation of fetal hemoglobin, but also a significant decrease in sickle β-globin, as well as a significant decrease in sickle cell count and an increase in normal red blood cell count compared to the unmodified cell population.

[0006] Accordingly, in one embodiment, the present invention provides a CRISPR system comprising one or more gRNA molecules as described herein, for example, one (e.g., a Cas CRISPR system, e.g., a Cas9 CRISPR system, e.g., a Streptococcus pyogenes (S. pyogenes) Cas9 CRISPR system). Any of the gRNA molecules described herein may be used in such a system, as well as in the methods and cells described herein.

[0007] In one embodiment, the present invention provides a gRNA molecule comprising tracr and crRNA, wherein the crRNA comprises a targeting domain complementary to the target sequence of a WIZ gene (e.g., the human WIZ gene). In embodiments, the WIZ gene comprises the genomic nucleic acid sequence, -strand, hg38, or a fragment or variant thereof at Chr19:15419978-15451624.

[0008] In embodiments, the targeting domain comprises, for example, one of sequence numbers 1 to 3106 (see, for example, Tables 1 to 3). In embodiments, the gRNA molecule comprises a targeting domain comprising (for example, one of) a fragment of any of the above sequences.

[0009] In any of the embodiments and models described herein, the gRNA molecule may further have the regions and / or properties described herein. In an embodiment, the gRNA molecule comprises a fragment of any of the targeting domains described herein. In an embodiment, the targeting domain comprises, for example, 17, 18, 19, or 20 consecutive nucleic acids of any one of the described targeting domain sequences. In an embodiment, the 17, 18, 19, or 20 consecutive nucleic acids of any one of the described targeting domain sequences are 17, 18, 19, or 20 consecutive nucleic acids located at the 3' end of the described targeting domain sequence. In another embodiment, the 17, 18, 19, or 20 consecutive nucleic acids of any one of the described targeting domain sequences are 17, 18, 19, or 20 consecutive nucleic acids located at the 5' end of the described targeting domain sequence. In other embodiments, the 17, 18, 19, or 20 consecutive nucleic acids of any one of the described targeting domain sequences do not include either the 5' nucleic acid or the 3' nucleic acid of the described targeting domain sequence. In embodiments, the targeting domain consists of the described targeting domain sequence.

[0010] In some embodiments, including those described herein, a portion of crRNA and a portion of tracr hybridize to form a flagpole containing SEQ ID NO: 3110 or 3111. In some embodiments, including those described herein, the flagpole further includes a first flagpole extension located 3' to the crRNA portion of the flagpole, wherein the first flagpole extension contains SEQ ID NO: 3112. In some embodiments, including those described herein, the flagpole further includes a second flagpole extension located 3' to the crRNA portion of the flagpole, and, if present, a first flagpole extension, wherein the second flagpole extension contains SEQ ID NO: 3113.

[0011] In some embodiments, including those described herein, tracr comprises SEQ ID NO: 3152 or SEQ ID NO: 3153. In some embodiments, including those described herein, tracr comprises SEQ ID NO: 3160 and optionally further comprises an additional 1, 2, 3, 4, 5, 6, or 7 uracil (U) nucleotides at its 3' end. In some embodiments, including those described herein, crRNA comprises a [targeting domain]-: a) SEQ ID NO: 3110; b) SEQ ID NO: 3111; c) SEQ ID NO: 3127; d) SEQ ID NO: 3128; e) SEQ ID NO: 3129; f) SEQ ID NO: 3130; or g) SEQ ID NO: 3154 from 5' to 3'.

[0012] In some embodiments, including those described above, tracr has a 5' to 3' position of a) SEQ ID NO: 3115; b) SEQ ID NO: 3116; c) SEQ ID NO: 3131; d) SEQ ID NO: 3132; e) SEQ ID NO: 3152; f) SEQ ID NO: 3153; g) SEQ ID NO: 232; h) SEQ ID NO: 3155; i) (SEQ ID NO: 3156; j) SEQ ID NO: 3157; k) at least 1, 2, 3, 4, 5, 6, or 7 uracil (U) nucleotides, for example, 1, 2, 3, 4, 5, 6, or 7 uracil (U) nucleotides Any of a) to j) above, further comprising at the 3' end; l) any of a) to k) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides, e.g., 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 3' end; or m) any of a) to l) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides, e.g., 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 5' end (e.g., the 5' side).

[0013] In some embodiments, including those described herein, the targeting domain and tracr are located on separate nucleic acid molecules. In some embodiments, including those described herein, the targeting domain and tracr are located on separate nucleic acid molecules, and the nucleic acid molecule containing the targeting domain optionally includes SEQ ID NO: 3129 located immediately 3' to the targeting domain, and the nucleic acid molecule containing tracr includes SEQ ID NO: 3152, for example. In some embodiments, including those described herein, the crRNA portion of the flagpole includes SEQ ID NO: 3129 or SEQ ID NO: 3130. In some embodiments, including those described herein, tracr includes SEQ ID NO: 3115 or 3116, and optionally, if a first flagpole extension is present, a first tracr extension located 5' to SEQ ID NO: 3115 or 3116, the first tracr extension including SEQ ID NO: 3117.

[0014] In some embodiments, including those described above, the targeting domain and tracr are located on a single nucleic acid molecule, for example, where tracr is located on the 3' side of the targeting domain. In some embodiments, the gRNA molecule includes a loop located on the 3' side of the targeting domain and the 5' side of tracr. In some embodiments, the loop includes SEQ ID NO: 3114. In some embodiments, including those described above, the gRNA molecule includes any of (a) to (e) above, with the 5' to 3' end further comprising [targeting domain]-: (a) SEQ ID NO: 3123; (b) SEQ ID NO: 3124; (c) SEQ ID NO: 3125; (d) SEQ ID NO: 3126; (e) SEQ ID NO: 3159; or (f) 1, 2, 3, 4, 5, 6, or 7 uracil (U) nucleotides at the 3' end.

[0015] In some embodiments, including those described above, the targeting domain and tracr are arranged on a single nucleic acid molecule, where the nucleic acid molecule comprises, for example, the targeting domain and SEQ ID NO: 3159, optionally positioned immediately 3' to the targeting domain.

[0016] In some embodiments, including those described in any of the above-mentioned aspects and embodiments, one or more nucleic acid molecules containing a gRNA molecule are, a) One or more phosphorothioate modifications, for example, three, at the 3' end of the one or more nucleic acid molecules; b) One or more phosphorothioate modifications, for example, three, at the 5' end of the one or more nucleic acid molecules; c) One or more 2'-O-methyl modifications at the 3' end of the one or more nucleic acid molecules; for example, three 2'-O-methyl modifications; d) One or more 2'-O-methyl modifications at the 5' end of the one or more nucleic acid molecules; for example, three 2'-O-methyl modifications; e) 2'O-methyl modification at the fourth, third, and second 3' residues from the terminal of each of the one or more nucleic acid molecules; f) 2'O-methyl modification at the fourth, third, and second 5' residues from the terminal of each of the one or more nucleic acid molecules; or f) Any combination of these Includes.

[0017] In some embodiments, including those described above, the present invention provides a gRNA molecule, wherein, When a CRISPR system containing this gRNA molecule (e.g., RNP as described herein) is introduced into a cell, an indel is formed at or near a target sequence complementary to the targeting domain of the gRNA molecule.

[0018] In some embodiments, including those described above, the present invention provides a gRNA molecule in which, when a CRISPR system (e.g., RNP as described herein) containing this gRNA molecule is introduced into a population of cells, an indel is formed in at or near a target sequence complementary to the targeting domain of the gRNA molecule in at least about 15%, e.g., at least about 17%, e.g., at least about 20%, e.g., at least about 30%, e.g., at least about 40%, e.g., at least about 50%, e.g., at least about 55%, e.g., at least about 60%, e.g., at least about 70%, e.g., at least about 75% of the cells in the population. In some embodiments, including those described above, the indel contains at least one nucleotide from the WIZ gene region. In an embodiment, at least about 15% of the cells in the population contain an indel containing at least one nucleotide from the WIZ gene region. In an embodiment, the indel is measured by next-generation sequencing (NGS).

[0019] In some embodiments, including those described above, the present invention provides a gRNA molecule in which, when a CRISPR system (e.g., an RNP as described herein) containing this gRNA molecule is introduced into cells, the expression of fetal hemoglobin in the cells or their offspring, e.g., their erythroid offspring, e.g., their erythrocyte offspring, increases. In embodiments, when a CRISPR system (e.g., an RNP as described herein) containing a gRNA molecule is introduced into a population of cells, the proportion of F cells in the population or its offspring, e.g., its erythroid offspring, e.g., its erythrocyte offspring, increases by at least about 15%, e.g., at least about 17%, e.g., at least about 20%, e.g., at least about 25%, e.g., at least about 30%, e.g., at least about 35%, e.g., at least about 40%, compared to the proportion of F cells in a population of cells or its offspring in which the gRNA molecule has not been introduced, e.g., its erythroid offspring, e.g., its erythrocyte at least about 40%. In the embodiment, the cells or their offspring, for example, their erythrocyte offspring, for example, their erythrocyte offspring, produce at least about 6 picograms per cell (e.g., at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms, or about 9 to about 10 picograms) of fetal hemoglobin.

[0020] In some embodiments, including those described above, the present invention provides a gRNA molecule in which, when a CRISPR system (e.g., RNP as described herein) containing this gRNA molecule is introduced into cells, no off-target indels are formed in the cells, such as those detectable by next-generation sequencing and / or nucleotide insertion assays, for example, and no off-target indels are formed outside the WIZ gene region (e.g., within the range of a gene, e.g., the coding region of a gene).

[0021] In some embodiments, including those described above, the present invention provides a gRNA molecule in which, when a CRISPR system (e.g., an RNP as described herein) containing this gRNA molecule is introduced into a population of cells, not more than about 5%, e.g., not more than about 1%, e.g., not more than about 0.1%, e.g., not more than about 0.01% of the cell population, in which off-target indels are detected, e.g., outside the WIZ gene (e.g., within the range of a gene, e.g., within the coding region of a gene), as detectable by, for example, next-generation sequencing and / or nucleotide insertion assays.

[0022] In some embodiments, including any of the aforementioned aspects and embodiments, the cells are mammalian, primate, or human cells (or a population of cells including them), for example, human cells; for example, the cells are HSPCs (or a population of cells including them), for example, the HSPCs are CD34+; for example, the HSPCs are CD34+CD90+. In some embodiments, the cells are self to the patient receiving the cells. In other embodiments, the cells are allogeneic to the patient receiving the cells.

[0023] In some embodiments, the gRNA molecules, genome editing systems (e.g., CRISPR systems), and / or methods described herein relate to cells, for example, as described herein, that include or result in one or more of the following characteristics: (a) At least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the cell populations described herein contain an indel in a genomic DNA sequence complementary to the targeting domain of the gRNA molecule described herein, or in the vicinity thereof; (b) The cells described herein (e.g., a population of cells) have the ability to differentiate into differentiated cells of the erythrocyte lineage (e.g., erythrocytes), wherein the differentiated cells exhibit an increased fetal hemoglobin level compared to, for example, unmodified cells (e.g., a population of cells); (c) A population of cells described herein has the ability to differentiate into a population of differentiated cells, for example, a population of cells of the erythrocyte lineage (for example, a population of erythrocytes), wherein the population of differentiated cells has an increased proportion of F cells (for example, at least about 15%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher proportion of F cells) compared to, for example, a population of unmodified cells; (d) The cells described herein (e.g., a population of cells) have the ability to differentiate into differentiated cells, e.g., cells of the erythrocyte lineage (e.g., erythrocytes), wherein the differentiated cells (e.g., a population of differentiated cells) produce at least about 6 picograms per cell (e.g., at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms, or about 9 to about 10 picograms) of fetal hemoglobin; (e) No off-target indels are formed in the cells described herein, for example, as detectable by next-generation sequencing and / or nucleotide insertion assays, for example, no off-target indels are formed outside the WIZ gene region (for example, within the range of a gene, e.g., the coding region of a gene); (f) Of the cell population described herein, not more than about 5%, for example not more than about 1%, for example not more than about 0.1%, for example not more than about 0.01%, in which cells have off-target indels detected, for example, outside the WIZ gene region (for example, within the range of a gene, for example, the coding region of a gene), as detectable by next-generation sequencing and / or nucleotide insertion assays; (g) The cells or their offspring described herein are detectable in a patient who receives them, at a time beyond 16 weeks, 20 weeks, or 24 weeks after transplantation, by detecting an indel (optionally, this indel is a large deletion indel) in or near a genomic DNA sequence complementary to the targeting domain of any of the gRNA molecules of SEQ ID NOs. 1 to SEQ ID NOs. 3106, for example, detectable in bone marrow or peripheral blood; (h) A population of cells described herein has the ability to differentiate into a population of differentiated cells, for example, a population of cells of the erythrocyte lineage (for example, a population of erythrocytes), wherein the population of differentiated cells includes, for example, a reduced proportion of sickle cells compared to a population of unmodified cells (for example, a proportion of sickle cells that is at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%); and / or (i) The cells or populations of cells described herein have the ability to differentiate into a population of differentiated cells, for example, a population of erythrocyte-derived cells (for example, a population of erythrocytes), wherein the population of differentiated cells includes cells that produce sickle hemoglobin (HbS) at a reduced level (for example, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, or at least about 90%) compared to, for example, a population of unmodified cells (population).

[0024] In one embodiment, the present invention is 1) One or more gRNA molecules (including the first gRNA molecule) of any of the gRNA embodiments and examples described herein, and, for example, Cas9 molecules described herein; 2) Nucleic acids comprising one or more gRNA molecules (including the first gRNA molecule) of any of the gRNA embodiments and examples described herein, for example, and a nucleotide sequence encoding a Cas9 molecule as described herein; 3) Nucleic acids comprising one or more nucleotide sequences each encoding one of the gRNA molecules (including the first gRNA molecule) of any of the gRNA embodiments and examples described herein, and Cas9 molecules as described herein; 4) nucleic acids comprising one or more nucleotide sequences each encoding one of the gRNA molecules (including the first gRNA molecule) of any of the gRNA embodiments and embodiments described herein, for example, and nucleic acids encoding the Cas9 molecule described herein; or 5) Any of the above 1) to 4), and template nucleic acid; or 6) A nucleic acid comprising any of the above 1) to 4) and a nucleotide sequence encoding a template nucleic acid. The present invention provides a composition containing the following:

[0025] In one embodiment, the present invention provides a composition comprising a first gRNA molecule, for example, one of the gRNA embodiments and embodiments described herein, and further comprising, for example, a Cas9 molecule, for example, wherein the Cas9 molecule is active or inactivated Streptococcus pyogenes (s.pyogenes) Cas9, for example, wherein the Cas9 molecule comprises SEQ ID NO: 3133. In one embodiment, the Cas9 molecule comprises, for example, (a) SEQ ID NO: 3161; (b) SEQ ID NO: 3162; (c) SEQ ID NO: 3163; (d) SEQ ID NO: 3164; (e) SEQ ID NO: 3165; (f) SEQ ID NO: 3166; (g) SEQ ID NO: 3167; (h) SEQ ID NO: 3168; (i) SEQ ID NO: 3169; (j) SEQ ID NO: 3170; (k) SEQ ID NO: 3171 or (l) SEQ ID NO: 3172.

[0026] In some embodiments, including those described in any of the aforementioned compositions and embodiments, the first gRNA molecule and the Cas9 molecule are present in a ribonucleoprotein complex (RNP).

[0027] In some embodiments, including those in any of the embodiments and forms of the compositions described above, the present invention provides a composition further comprising a second gRNA molecule; a second gRNA molecule and a third gRNA molecule; or a second gRNA molecule, optionally a third gRNA molecule, and optionally a fourth gRNA molecule, wherein the second gRNA molecule, optionally a third gRNA molecule, and optionally a fourth gRNA molecule are gRNA molecules described herein, for example, gRNA molecules in any of the embodiments and forms of the gRNA molecules described above, wherein each gRNA molecule in the composition is complementary to a different target sequence. In embodiments, two or more of the first gRNA molecule, the second gRNA molecule, optionally a third gRNA molecule, and optionally a fourth gRNA molecule are complementary to target sequences within the same gene or region. In this embodiment, the first gRNA molecule, the second gRNA molecule, an optional third gRNA molecule, and an optional fourth gRNA molecule are complementary to target sequences that are only 6000 nucleotides or less, 5000 nucleotides or less, 500 or less, 400 nucleotides or less, 300 or less, 200 nucleotides or less, 100 nucleotides or less, 90 nucleotides or less, 80 nucleotides or less, 70 nucleotides or less, 60 nucleotides or less, 50 nucleotides or less, 40 nucleotides or less, 30 nucleotides or less, 20 nucleotides or less, or 10 nucleotides or less. In some embodiments, including those described in any of the embodiments and models of the composition described above, the composition comprises a first gRNA molecule and a second gRNA molecule (for example, consisting of a first gRNA molecule and a second gRNA molecule), wherein the first gRNA molecule and the second gRNA molecule are (a) independently selected from and complementary to different target sequences; (b) independently selected from the gRNA molecules of Table 1 and complementary to different target sequences; (c) independently selected from the gRNA molecules of Table 2 and complementary to different target sequences; or (d) independently selected from the gRNA molecules of Table 3 and complementary to different target sequences; or (f) independently selected from the gRNA molecules of any of the embodiments and models described above and complementary to different target sequences.

[0028] In some embodiments, including those described in any of the aforementioned forms and embodiments of the composition, the composition comprises a first gRNA molecule and a second gRNA molecule, wherein a) The first gRNA molecule is complementary to a target sequence containing at least one nucleotide (e.g., 20 consecutive nucleotides) within the range of Chr19:15419978-15451624, -chain, and hg38; b) The second gRNA molecule is complementary to a target sequence containing at least one nucleotide (e.g., 20 consecutive nucleotides) within the range of Chr19:15419978-15451624, -chain, and hg38.

[0029] In one embodiment, with respect to the gRNA molecular component of the composition, the composition comprises a first gRNA molecule and a second gRNA molecule.

[0030] In some embodiments, including those described in any of the embodiments and configurations of the compositions described herein, each of the gRNA molecules is present in a ribonucleoprotein complex (RNP) together with, for example, the Cas9 molecule described herein.

[0031] In some embodiments, including those described in any of the embodiments and models of the composition described above, the composition comprises a template nucleic acid, wherein the template nucleic acid comprises nucleotides corresponding to nucleotides in or near the target sequence of a first gRNA molecule. In embodiments, the template nucleic acid comprises a nucleic acid encoding a human WIZ gene or a fragment thereof.

[0032] In some embodiments, including those described in any of the aforementioned aspects and embodiments of the composition, the composition is formulated in a medium suitable for electroporation.

[0033] In some embodiments, including those described in any of the embodiments and models of the compositions described herein, each of the gRNA molecules of the composition is present in RNP together with the Cas9 molecule described herein, where each of the RNPs is at a concentration of less than about 10 μM, e.g., less than about 3 μM, e.g., less than about 1 μM, e.g., less than about 0.5 μM, e.g., less than about 0.3 μM, e.g., less than 0.1 μM. In embodiments, the RNP is at a concentration of about 1 μM. In embodiments, the RNP is at a concentration of about 2 μM. In embodiments, the concentrations are, for example, the concentrations of RNP in a composition containing cells as described herein, and optionally, the composition containing cells and RNP is suitable for electroporation.

[0034] In one embodiment, the present invention provides a nucleic acid sequence encoding one or more gRNA molecules, for example, of any of the gRNA molecule embodiments and models described herein. In an embodiment, the nucleic acid comprises a promoter operably ligated to the sequence encoding one or more gRNA molecules, for example, the promoter being a promoter recognized by RNA polymerase II or RNA polymerase III, or for example, the promoter being a U6 promoter or an HI promoter.

[0035] In some embodiments, including those in any of the nucleic acid embodiments and models described above, the nucleic acid further encodes a Cas9 molecule, for example, comprising any of the following: (a) SEQ ID NO: 3133, (a) SEQ ID NO: 3161; (b) SEQ ID NO: 3162; (c) SEQ ID NO: 3163; (d) SEQ ID NO: 3164; (e) SEQ ID NO: 3165; (f) SEQ ID NO: 3166; (g) SEQ ID NO: 3167; (h) SEQ ID NO: 3168; (i) SEQ ID NO: 3169; (j) SEQ ID NO: 3170; (k) SEQ ID NO: 3171 or (l) SEQ ID NO: 3172. In embodiments, the nucleic acid comprises a promoter operably linked to the sequence encoding the Cas9 molecule, for example, an EF-1 promoter, a CMV IE gene promoter, an EF-1α promoter, a ubiquitin C promoter, or a phosphoglycerate kinase (PGK) promoter.

[0036] Some embodiments provided herein include vectors comprising nucleic acids of any of the nucleic acid embodiments and embodiments described above. In embodiments, the vector is selected from the group consisting of lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmides, and RNA vectors.

[0037] In one aspect provided herein, a method for modifying a cell (e.g., a population of cells) in or near a target sequence within the cell (e.g., modifying the structure (e.g., sequence) of a nucleic acid), wherein the cell (e.g., a population of cells) 1) One or more gRNA molecules described herein (e.g., any of the embodiments and examples of the gRNA molecules described herein) and, for example, a Cas9 molecule described herein; 2) Nucleic acids comprising one or more gRNA molecules (e.g., any of the gRNA molecule embodiments and examples described herein) and a nucleotide (nucelotide) sequence encoding, for example, a Cas9 molecule described herein; 3) Nucleic acids comprising one or more nucleotide sequences each encoding one gRNA molecule as described herein (for example, any of the embodiments and models of the gRNA molecule described herein), and Cas9 molecules as described herein; 4) Nucleic acids comprising one or more nucleotide sequences each encoding one gRNA molecule as described herein (for example, any of the gRNA molecule embodiments and examples described herein), and nucleic acids comprising a nucleotide (nucelotide) sequence encoding, for example, a Cas9 molecule as described herein; 5) Any of the above 1) to 4), and template nucleic acid; 6) A nucleic acid comprising any of the above 1) to 4) and a nucleotide sequence encoding a template nucleic acid; 7) Compositions described herein, for example, any of the embodiments and models of the compositions described above; or 8) Vectors described herein, for example, any of the vectors described above in any of the embodiments and models of the vectors described herein. This includes methods that involve bringing it into contact with (for example, introducing it into).

[0038] In some embodiments, including those described in any of the methods and embodiments described above, a gRNA molecule or nucleic acid encoding a gRNA molecule and a Cas9 molecule or nucleic acid encoding a Cas9 molecule are formulated into a single composition. In other embodiments, a gRNA molecule or nucleic acid encoding a gRNA molecule and a Cas9 molecule or nucleic acid encoding a Cas9 molecule are formulated into two or more compositions. In some embodiments, the two or more compositions are delivered simultaneously or sequentially.

[0039] In some embodiments of the methods described herein, including those in any of the aforementioned aspects and embodiments, the cells are animal cells, e.g., mammalian cells, primate cells, or human cells, e.g., the cells are hematopoietic stem cell / progenitor cells (HSPCs) (e.g., a population of HSPCs), e.g., the cells are CD34+ cells, e.g., the cells are CD34+CD90+ cells. In embodiments of the methods described herein, the cells are placed in a composition comprising a population of cells enriched with respect to CD34+ cells. In embodiments of the methods described herein, the cells (e.g., a population of cells) are isolated from bone marrow, mobilized peripheral blood, or umbilical cord blood. In embodiments of the methods described herein, the cells are autologous or allogeneic to the patient receiving the cells, e.g., autologous.

[0040] In some embodiments of the methods described herein, including those in any of the embodiments and models of the methods described above, a) the modification results in an indel in or near a genomic DNA sequence complementary to the targeting domain of one or more gRNA molecules; or b) the modification results in a deletion in the WIZ gene region that is complementary to the targeting domain of one or more gRNA molecules (e.g., at least 90% complementary to the gRNA targeting domain, e.g., fully complementary to the gRNA targeting domain), e.g., substantially the entirety of such sequence. In embodiments of the methods, the indel is an insertion or deletion of less than about 40 nucleotides, e.g., less than 30 nucleotides, e.g., less than 20 nucleotides, e.g., less than 10 nucleotides, e.g., a single nucleotide deletion.

[0041] In some embodiments of the methods described herein, including those in any of the embodiments and models of the methods described above, the Method yields a population of cells, for example, containing indels, in which at least about 15% of the population, e.g., at least about 17%, e.g., at least about 20%, e.g., at least about 30%, e.g., at least about 40%, e.g., at least about 50%, e.g., at least about 55%, e.g., at least about 60%, e.g., at least about 70%, e.g., at least about 75%, has been modified.

[0042] In some embodiments of the methods described herein, including those in any of the aforementioned aspects and embodiments, the modification results in cells (e.g., a population of cells) having the ability to differentiate into differentiated cells of the erythrocyte lineage (e.g., erythrocytes), wherein these differentiated cells exhibit increased fetal hemoglobin levels compared to, for example, unmodified cells (e.g., a population of cells).

[0043] In some embodiments of the methods described herein, including those in any of the embodiments and models of the methods described above, the modification results in a population of cells having the ability to differentiate into a population of differentiated cells, for example, a population of erythrocyte-derived cells (e.g., a population of erythrocytes), wherein the population of differentiated cells has an increased proportion of F cells (e.g., at least about 15%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher proportion of F cells) compared to, for example, the population of unmodified cells.

[0044] In some embodiments of the methods described herein, including those in any of the aforementioned aspects and embodiments, the modification results in cells having the ability to differentiate into differentiated cells, for example, cells of the erythrocyte lineage (e.g., erythrocytes), where the differentiated cells produce at least about 6 picograms per cell (e.g., at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms, or about 9 to about 10 picograms) of fetal hemoglobin.

[0045] In one embodiment, the present invention provides cells modified by any of the methods described herein, for example, by any of the embodiments and models of the methods described above.

[0046] In one embodiment, the present invention provides cells that can be obtained by any of the methods described herein, for example, by any of the embodiments and models of the methods described above.

[0047] In one embodiment, the present invention provides a cell comprising, for example, a first gRNA molecule of any of the aforementioned embodiments or forms of the gRNA molecule described herein, or, for example, a composition of any of the aforementioned embodiments or forms of the composition described herein, or, for example, a nucleic acid of any of the aforementioned embodiments or forms of the nucleic acid described herein, or, for example, a vector of any of the aforementioned embodiments or forms of the vector described herein.

[0048] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and models, the cell further comprises a Cas9 molecule, for example, one of the Cas9 molecules described herein, e.g., SEQ ID NO: 3133, (a) SEQ ID NO: 3161; (b) SEQ ID NO: 3162; (c) SEQ ID NO: 3163; (d) SEQ ID NO: 3164; (e) SEQ ID NO: 3165; (f) SEQ ID NO: 3166; (g) SEQ ID NO: 3167; (h) SEQ ID NO: 3168; (i) SEQ ID NO: 3169; (j) SEQ ID NO: 3170; (k) SEQ ID NO: 3171 or (l) SEQ ID NO: 3172.

[0049] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and forms, the cell contains, has contained, or will contain, a second gRNA molecule or nucleic acid encoding the gRNA molecule, for example, one of the aforementioned gRNA molecule embodiments or forms described herein, wherein the first gRNA molecule and the second gRNA molecule contain non-identical targeting domains.

[0050] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and models, the expression of fetal hemoglobin is increased in the cells or their offspring (e.g., their erythroid offspring, e.g., their erythroid cell offspring) compared with cells of the same cell type that are not modified to include gRNA molecules or their offspring.

[0051] In some embodiments of the cells described herein, including those in any of the aforementioned cell forms and embodiments, the cells have the ability to differentiate into differentiated cells, such as cells of the erythrocyte lineage (e.g., erythrocytes), wherein the differentiated cells exhibit increased levels of fetal hemoglobin compared to cells of the same type that are not modified to include, for example, gRNA molecules.

[0052] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and models, differentiated cells (e.g., cells of the erythrocyte lineage, e.g., erythrocytes) produce at least about 6 picograms (e.g., at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms, or about 9 to about 10 picograms) of fetal hemoglobin compared to differentiated cells of the same type that are not modified to include, for example, gRNA molecules.

[0053] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and models, the cells are treated with a stem cell growth agent, for example, a)(1r,4r)-N 1 -(2-benzyl-7-(2-methyl-2H-tetrazole-5-yl)-9H-pyrimido[4,5-b]indole-4-yl)cyclohexane-1,4-diamine; b)4-(3-piperidine-1-ylpropylamino)-9H-pyrimido[4,5-b]indole-7-carboxylate methyl; c)4-(2-(2-(benzo[b]thiophen-3-yl)-9-isopropyl-9H-purine-6-ylamino)ethyl)phenol; d)(S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purine-9-yl)propan-l-ol; or e) combinations of these (e.g., (1r,4r)-N 1 -(2-benzyl-7-(2-methyl-2H-tetrazole-5-yl)-9H-pyrimido[4,5-b]indole-4-yl)cyclohexane-1,4-diamine is contacted with a stem cell growth agent selected from (S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purine-9-yl)propan-l-ol), for example, by ex vivo contact. In embodiments, the stem cell growth agent is (S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purine-9-yl)propan-l-ol.

[0054] In some embodiments of the cells described herein, including those in any of the aforementioned cell embodiments and embodiments, the cell contains: a) an indel in or near a genomic DNA sequence complementary to the targeting domain of a gRNA molecule described herein, for example, one of the gRNA molecules described herein or in any of the embodiments or embodiments; or b) a deletion in the WIZ gene region that is complementary to the targeting domain of a gRNA molecule described herein, for example (e.g., at least 90% complementary to the gRNA targeting domain, e.g., fully complementary to the gRNA targeting domain), e.g., substantially all of that sequence. In some embodiments, the indel is an insertion or deletion of less than about 40 nucleotides, e.g., less than 30 nucleotides, e.g., less than 20 nucleotides, e.g., less than 10 nucleotides, e.g., the indel is a single nucleotide deletion.

[0055] In some embodiments of the cells described herein, including those in any of the aforementioned cell forms and embodiments, the cells are animal cells, for example, mammalian cells, primate cells, or human cells. In some embodiments, the cells are hematopoietic stem cell / progenitor cells (HSPCs) (e.g., a population of HSPCs), for example, the cells are CD34+ cells, for example, the cells are CD34+CD90+ cells. In embodiments, the cells (e.g., a population of cells) are isolated from bone marrow, mobilized peripheral blood, or umbilical cord blood. In embodiments, the cells are autologous to the patient receiving the cells. In embodiments, the cells are allogeneic to the patient receiving the cells.

[0056] In one embodiment, the present invention provides a population of cells described herein, for example, a population of cells comprising any of the aforementioned embodiments and forms of cells described herein. In one embodiment, the present invention provides a population of cells in which at least about 50%, for example, at least about 60%, for example, at least about 70%, for example, at least about 80%, for example, at least about 90% (for example, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%) of the population are cells described herein, for example, any of the aforementioned embodiments and forms of cells. In one embodiment, a population of cells (e.g., cells of a population of cells) has the ability to differentiate into a population of differentiated cells, for example, a population of cells of the erythrocyte lineage (e.g., a population of erythrocytes), where the population of differentiated cells has an increased proportion of F cells (e.g., at least about 15%, at least about 17%, at least about 20%, at least about 25%, at least about 30%, or at least about 40% higher proportion of F cells) compared to, for example, a population of unmodified cells of the same type. In one embodiment, the F cells of the population of differentiated cells produce at least about 6 picograms per cell on average (e.g., at least about 7 picograms, at least about 8 picograms, at least about 9 picograms, at least about 10 picograms, or about 8 to about 9 picograms, or about 9 to about 10 picograms) of fetal hemoglobin.

[0057] In some embodiments, including those in any of the aforementioned cell population configurations and embodiments, the present invention provides a cell population comprising: 1) a patient's body weight of 1 kg receiving at least 1 e6 CD34+ cells / cells; 2) a patient's body weight of 1 kg receiving at least 2 e6 CD34+ cells / cells; 3) a patient's body weight of 1 kg receiving at least 3 e6 CD34+ cells / cells; 4) a patient's body weight of 1 kg receiving at least 4 e6 CD34+ cells / cells; or 5) a patient's body weight of 1 kg receiving 2 e6 to 10 e6 CD34+ cells / cells. In embodiments, at least about 40%, for example, at least about 50% (e.g., at least about 60%, at least about 70%, at least about 80%, or at least about 90%) of the cells in the population are CD34+ cells. In embodiments, at least about 5%, for example, at least about 10%, for example, at least about 15%, for example, at least about 20%, for example, at least about 30% of the cells in the population are CD34+CD90+ cells. In embodiments, the cell population is derived from umbilical cord blood, peripheral blood (e.g., mobilized peripheral blood), or bone marrow, for example, derived from bone marrow. In embodiments, the cell population includes, for example, mammalian cells, such as human cells, for example. In embodiments, the cell population is self to the patient to whom it is administered. In other embodiments, the cell population is allogeneic to the patient to whom it is administered.

[0058] In one embodiment, the present invention provides a composition comprising cells as described herein, for example, cells of any of the aforementioned embodiments and forms of cells, or a population of cells as described herein, for example, a population of cells of any of the aforementioned embodiments and forms of the population of cells. In one embodiment, the composition comprises a pharmaceutically acceptable medium, for example, a pharmaceutically acceptable medium suitable for cryopreservation.

[0059] In one embodiment, the present invention provides a method for treating an abnormal hemoglobin disorder, comprising administering to a patient any of the cells described herein, for example, any of the aforementioned embodiments and forms of cells, any of the populations of cells described herein, for example, any of the aforementioned embodiments and forms of the populations of cells, or any of the compositions described herein, for example, any of the aforementioned embodiments and forms of compositions.

[0060] In one embodiment, the present invention provides a method for increasing fetal hemoglobin expression in mammals, comprising administering to a patient cells described herein, for example, any of the aforementioned embodiments and forms of cells, a population of cells described herein, for example, any of the aforementioned embodiments and forms of the population of cells, or a composition described herein, for example, any of the aforementioned embodiments and forms of the composition. In one embodiment, the abnormal hemoglobin disorder is β-thalassemia. In one embodiment, the abnormal hemoglobin disorder is sickle cell anemia.

[0061] In one embodiment, the present invention is (a) A step of providing cells (e.g., a population of cells) (e.g., HSPCs (e.g., a population of HSPCs)); (b) the step of ex vivo culturing the cells (e.g., the population of cells) in a cell culture medium containing a stem cell growth agent; and (c) a first gRNA molecule, for example, as described herein, e.g., a first gRNA molecule of any of the embodiments and forms of the gRNA molecule described herein; a nucleic acid molecule encoding the first gRNA molecule; a composition, for example, as described herein, e.g., a composition of any of the embodiments and forms of the composition described herein; or a vector, for example, a vector of any of the embodiments and forms described herein, a step of introducing the vector into the cells. A method is provided for preparing cells (e.g., a population of cells), including the following: In an embodiment of this method, after introducing step (c), the cells (e.g., a population of cells) have the ability to differentiate into differentiated cells (e.g., a population of differentiated cells), for example, cells of the erythrocyte lineage (e.g., a population of erythrocyte lineage cells), for example, erythrocytes (e.g., a population of erythrocytes), wherein the differentiated cells (e.g., a population of differentiated cells) produce increased fetal hemoglobin compared to, for example, the same cells not subjected to step (c). In this embodiment of the method, the stem cell proliferation agent is a) (1r,4r)-N1-(2-benzyl-7-(2-methyl-2H-tetrazole-5-yl)-9H-pyrimido[4,5-b]indole-4-yl)cyclohexane-1,4-diamine; b) 4-(3-piperidine-1-ylpropylamino)-9H-pyrimido[4,5-b]indole-7-carboxylate methyl; c) 4-(2-(2-(benzo[b]thiophen-3-yl)-9-isopropyl-9H-purine-6-ylamino)ethyl)phenol; d) (S)-2-(6-(2-(1H-indole (S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purin-9-yl)propan-l-ol; or e) a combination thereof (for example, a combination of (1r,4r)-N1-(2-benzyl-7-(2-methyl-2H-tetrazole-5-yl)-9H-pyrimido[4,5-b]indole-4-yl)cyclohexane-1,4-diamine and (S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purin-9-yl)propan-l-ol). In embodiments, the stem cell growth agent is (S)-2-(6-(2-(1H-indole-3-yl)ethylamino)-2-(5-fluoropyridine-3-yl)-9H-purin-9-yl)propan-l-ol. In one embodiment, the cell culture medium comprises thrombopoietin (Tpo), Flt3 ligand (Flt-3L), and human stem cell factor (SCF). In another embodiment, the cell culture medium further comprises human interleukin-6 (IL-6).In one embodiment, the cell culture medium contains thrombopoietin (Tpo), Flt3 ligand (Flt-3L), and human stem cell factor (SCF) at concentrations ranging from approximately 10 ng / mL to approximately 1000 ng / mL, for example, approximately 50 ng / mL each. In another embodiment, the cell culture medium contains human interleukin-6 (IL-6) at concentrations ranging from approximately 10 ng / mL to approximately 1000 ng / mL, for example, approximately 50 ng / mL each. In another embodiment, the cell culture medium contains a stem cell proliferation agent at a concentration ranging from approximately 1 nM to approximately 1 mM, for example, approximately 1 uM to approximately 100 nM, for example, approximately 500 nM to approximately 750 nM. In another embodiment, the cell culture medium contains a stem cell proliferation agent at a concentration of approximately 500 nM, for example, 500 nM. In this embodiment, the cell culture medium contains a stem cell proliferation agent at a concentration of approximately 750 nM, for example, 750 nM.

[0062] In embodiments of a method for preparing cells (e.g., a population of cells), the culture in step (b) includes a culture period prior to the introduction of step (c), for example, the culture period prior to the introduction of step (c) is at least 12 hours, for example, over a period of about 1 to about 12 days, for example, over a period of about 1 to about 6 days, for example, over a period of about 1 to about 3 days, for example, over a period of about 1 to about 2 days, for example, over a period of about 2 days. In embodiments of a method for preparing cells (e.g., a population of cells), including any of the aforementioned aspects and embodiments of the Method, the culture in step (b) includes a culture period after the introduction of step (c), for example, the culture period after the introduction of step (c) is at least 12 hours, for example, over a period of about 1 to about 12 days, for example, over a period of about 1 to about 6 days, for example, over a period of about 2 to about 4 days, for example, over a period of about 2 days, or over a period of about 3 days, or over a period of about 4 days. In embodiments of the method for preparing cells (e.g., a population of cells), including any of the aforementioned aspects and embodiments of the present method, the population of cells proliferates at least four times, for example, at least five times, or for example, at least ten times, compared to cells that have not been cultured, for example, according to step (b).

[0063] In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the present method, the introduction of step (c) includes electroporation. In embodiments, the electroporation includes 1 to 5 pulses, for example, 1 pulse, where each pulse has a pulse voltage in the range of 700 to 2000 volts and a pulse duration in the range of 10 ms to 100 ms. In embodiments, the electroporation includes 1 pulse, for example, consisting of 1 pulse. In embodiments, the voltage of the pulse (or 2 or more pulses) is in the range of 1500 to 1900 volts, for example, 1700 volts. In embodiments, the pulse duration of 1 or 2 or more pulses is in the range of 10 ms to 40 ms, for example, 20 ms.

[0064] In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the Method, the cells (e.g., a population of cells) provided in step (a) are human cells (e.g., a population of human cells). In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the Method, the cells (e.g., a population of cells) provided in step (a) are isolated from bone marrow, peripheral blood (e.g., mobilized peripheral blood), or umbilical cord blood. In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the Method, the cells (e.g., a population of cells) provided in step (a) are isolated from bone marrow, for example, from the bone marrow of a patient suffering from an abnormal hemoglobin disorder.

[0065] In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the present method, the population of cells provided in step (a) is enriched with respect to CD34+ cells.

[0066] In embodiments of the method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the present method, the cells (e.g., a population of cells) are cryopreserved after step (c).

[0067] In embodiments of a method for preparing cells (e.g., a population of cells), including those described in any of the aforementioned aspects and embodiments of the Method, after introducing step (c), the cells (e.g., a population of cells) include: a) an indel in or near a genomic DNA sequence complementary to the targeting domain of the first gRNA molecule; or b) a deletion in the WIZ gene region that is complementary to the targeting domain of the first gRNA molecule (e.g., at least 90% complementary to the gRNA targeting domain, e.g., fully complementary to the gRNA targeting domain), e.g., substantially all of that sequence.

[0068] In embodiments of a method for preparing cells (e.g., a population of cells), including any of the aforementioned aspects and embodiments of the Method, after introducing step (c), at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% of the cell population contains an indel in a genomic DNA sequence complementary to the targeting domain of the first gRNA molecule or in its vicinity.

[0069] In one embodiment, the present invention provides cells (e.g., a population of cells) that can be obtained by a method for preparing cells (e.g., a population of cells) described herein, for example, in any of the embodiments and models of the method for preparing the aforementioned cells.

[0070] In one embodiment, the present invention provides a method for treating an abnormal hemoglobin disorder in a human patient, comprising the step of administering to the human patient a composition comprising cells described herein, for example, cells of any of the aforementioned embodiments and forms of cells; or a population of cells described herein, for example, a population of cells of any of the aforementioned embodiments and forms of cell populations. In one embodiment, the abnormal hemoglobin disorder is β-thalassemia. In one embodiment, the abnormal hemoglobin disorder is sickle cell disease.

[0071] In one embodiment, the present invention provides a method for increasing fetal hemoglobin expression in a human patient, comprising the step of administering to the human patient a composition comprising cells described herein, for example, cells of any of the aforementioned embodiments and forms of cells; or a population of cells described herein, for example, a population of cells of any of the aforementioned embodiments and forms of cell populations. In one embodiment, the human patient has β-thalassemia. In one embodiment, the human patient has sickle cell disease.

[0072] In an embodiment of a method for treating abnormal hemoglobin disorders or increasing fetal hemoglobin expression, a human patient is administered a composition containing at least about 1e6 cells per kg of the patient's body weight (e.g., cells as described herein), for example, at least about 1e6 CD34+ cells per kg of the patient's body weight (e.g., cells as described herein). In an embodiment of a method for treating abnormal hemoglobin disorders or increasing fetal hemoglobin expression, a human patient is administered a composition containing at least about 2e6 cells per kg of the patient's body weight (e.g., cells as described herein), for example, at least about 2e6 CD34+ cells per kg of the patient's body weight (e.g., cells as described herein). In an embodiment of a method for treating abnormal hemoglobin disorders or a method for increasing fetal hemoglobin expression, a human patient is administered a composition containing, for example, about 2e6 cells per kg of the human patient's body weight (e.g., cells as described herein), or about 2e6 CD34+ cells per kg of the human patient's body weight (e.g., cells as described herein). In an embodiment of a method for treating abnormal hemoglobin disorders or a method for increasing fetal hemoglobin expression, a human patient is administered a composition containing at least about 3e6 cells per kg of the human patient's body weight (e.g., cells as described herein), or at least about 3e6 CD34+ cells per kg of the human patient's body weight (e.g., cells as described herein). In an embodiment of a method for treating abnormal hemoglobin disorders or a method for increasing fetal hemoglobin expression, a human patient is administered a composition containing, for example, about 3e6 cells per kg of the human patient's body weight (e.g., cells as described herein), or about 3e6 CD34+ cells per kg of the human patient's body weight (e.g., cells as described herein). In an embodiment of a method for treating abnormal hemoglobin disorders or a method for increasing fetal hemoglobin expression, a human patient is administered a composition containing, for example, about 2e6 to about 10e6 cells per kg of the human patient's body weight (e.g., cells as described herein), or about 2e6 to about 10e6 CD34+ cells per kg of the human patient's body weight (e.g., cells as described herein).

[0073] Furthermore, this specification also provides a method for treating abnormal hemoglobin disorders by administering to a patient a composition that reduces WIZ gene expression and / or WIZ protein activity, or a composition that reduces WIZ gene expression and / or WIZ protein activity as described herein. In one embodiment, the composition that reduces WIZ gene expression and / or WIZ protein activity includes small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof. In one embodiment, the abnormal hemoglobin disorder is β-thalassemia or sickle cell disease.

[0074] Furthermore, this specification also provides a method for increasing mammalian fetal hemoglobin expression by administering to a patient a cell or population of cells described herein or a composition containing such cells or population of cells, or a composition that reduces WIZ gene expression and / or WIZ protein activity. In some embodiments, the composition that reduces WIZ gene expression and / or WIZ protein activity includes small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof.

[0075] In some embodiments, the present invention provides gRNA molecules described herein, e.g., gRNA molecules of any of the embodiments and embodiments of the gRNA molecules described herein; compositions described herein, e.g., compositions of any of the embodiments and embodiments of the compositions described herein; nucleic acids described herein, e.g., nucleic acids of any of the embodiments and embodiments of the nucleic acids described herein; vectors described herein, e.g., vectors of any of the embodiments and embodiments of the vectors described herein; cells described herein, e.g., cells of any of the embodiments and embodiments of the cells described herein; or populations of cells described herein, e.g., populations of cells of any of the embodiments and embodiments of the populations of cells described herein; or embodiments and embodiments of compositions that reduce WIZ gene expression and / or WIZ protein activity. In some embodiments, compositions that reduce WIZ gene expression and / or WIZ protein activity include small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof.

[0076] In some embodiments, the present invention provides gRNA molecules described herein, e.g., gRNA molecules of any of the aforementioned embodiments and embodiments of gRNA molecules; compositions described herein, e.g., compositions of any of the aforementioned embodiments and embodiments of compositions; nucleic acids described herein, e.g., nucleic acids of any of the aforementioned embodiments and embodiments of nucleic acids; vectors described herein, e.g., vectors of any of the aforementioned embodiments and embodiments of vectors; cells described herein, e.g., cells of any of the aforementioned embodiments and embodiments of cells; or populations of cells described herein, e.g., populations of cells of any of the aforementioned embodiments and embodiments of cell populations; or embodiments of compositions that reduce WIZ gene expression and / or WIZ protein activity. In some embodiments, compositions that reduce WIZ gene expression and / or WIZ protein activity include small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof.

[0077] In some embodiments, the present invention provides gRNA molecules described herein, e.g., gRNA molecules of any of the embodiments and embodiments of the gRNA molecules described herein; compositions described herein, e.g., compositions of any of the embodiments and embodiments of the compositions described herein; nucleic acids described herein, e.g., nucleic acids of any of the embodiments and embodiments of the nucleic acids described herein; vectors described herein, e.g., vectors of any of the embodiments and embodiments of the vectors described herein; cells described herein, e.g., cells of any of the embodiments and embodiments of the cells described herein; or populations of cells described herein, e.g., populations of cells of any of the embodiments and embodiments of the populations of cells described herein; or embodiments and embodiments of compositions that reduce WIZ gene expression and / or WIZ protein activity. In some embodiments, compositions that reduce WIZ gene expression and / or WIZ protein activity include small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof.

[0078] In one embodiment, the present invention provides gRNA molecules described herein, e.g., gRNA molecules of any of the aforementioned embodiments and embodiments of gRNA molecules; compositions described herein, e.g., compositions of any of the aforementioned embodiments and embodiments of compositions; nucleic acids described herein, e.g., nucleic acids of any of the aforementioned embodiments and embodiments of nucleic acids; vectors described herein, e.g., vectors of any of the aforementioned embodiments and embodiments of vectors; cells described herein, e.g., cells of any of the aforementioned embodiments and embodiments of cells; or populations of cells described herein, e.g., populations of cells of any of the aforementioned embodiments and embodiments of populations of cells; or embodiments of compositions that reduce WIZ gene expression and / or WIZ protein activity, wherein the disease is an abnormal hemoglobin disorder, e.g., β-thalassemia or sickle cell disease. In one embodiment, the composition that reduces WIZ gene expression and / or WIZ protein activity comprises small molecule compounds (e.g., WIZ degraders), siRNA, shRNA, antisense oligonucleotides (ASOs), miRNAs, anti-microRNA oligonucleotides (AMOs), or any combination thereof. [Brief explanation of the drawing]

[0079] [Figure 1A] A volcano plot of genes with differential expression in WIZ KO cells compared to scrambled gRNA controls. Each point represents a gene. The HBG1 / 2 genes are upregulated differently by WIZ_6 and WIZ_18 gRNAs, which target the WIZ gene. [Figure 1B] Frequency of HbF+ cells resulting from shRNA-mediated WIZ deficiency in human mobilized peripheral blood CD34+ erythroid cells. [Figure 1C] Frequency of HbF+ cells resulting from CRISPR / Cas9-mediated WIZ deficiency in human mobilized peripheral blood CD34+ erythroid cells. [Modes for carrying out the invention]

[0080] Abbreviation ACN Acetonitrile Acetic acid (ACOH) AMO anti-microRNA oligonucleotides aq. aqueous solution ASO Antisense Oligonucleotides Boc2O di-tert-butyl dicarbonate br Wide line BSA (Bovine Serum Albumin) Cas9 CRISPR-related protein 9 CRISPR clustering creates a regular arrangement of short palindromic repeats. crRNA CRISPR RNA d double line DCE 1,2-Dichloroethane DCM Dichloromethane dd double line double line ddd Double line double line double line ddq quadruple line double line double line ddt Triple line double line double line DIPEA N,N-diisopropylethylamine DIPEA (DIEA) Diisopropylethylamine DMA N,N-dimethylacetamide DMAP 4-dimethylaminopyridine DME 1,2-Dimethoxyethane DMEM Dulbecco's modified Eagle medium DMF (N,N-dimethylformamide) DMSO (Dimethyl Sulfoxide) DMSO (Dimethyl Sulfoxide) dq quadruple double line dt Triple line double line dtbbpy 4,4'-di-tert-butyl-2,2'-dipyridyl dtd Double line Triple line Double line DTT (Dithiothreitol) EC50 50% effective concentration EDTA (Ethylenediaminetetraacetic acid) eGFP (Highly Sensitive Green Fluorescent Protein) ELSD Evaporative Light Scattering Detector Et2O Diethyl ether Et3N triethylamine HCl ethyl acetate EtOH Ethanol FACS fluorescence-activated cell sorting FBS Fetal Bovine Serum FITC Fluorescein Flt3L Fms-related tyrosine kinase 3 ligand, Flt3L g grams g / min (grams per minute) h or hr hours HbF (fetal hemoglobin) HCl (hydrogen chloride) HEPES (4-(2-hydroxyethyl)-1-piperazineethanesulfonic acid) hept septet HPLC (High-Performance Liquid Chromatography) HRMS high-resolution mass spectrometry IC50 50% inhibitory concentration IMDM (Iskov Modified Dulbecco's Medium) IPA (iPrOH) Isopropyl alcohol Ir[(dF(CF3)ppy)2dtbbpy]PF6[4,4'-bis(1,1-dimethylethyl)-2,2'-bipyridine-N1,N1']bis[3,5-difluoro-2-[5-(trifluoromethyl)-2-pyridinyl-N]phenyl-C]iridium(III) hexafluorophosphate KCl Potassium Chloride LC-MS (Liquid Chromatography-Mass Spectrometry) m multiplet M molar concentration MeCN acetonitrile MeOH methanol mg milligrams MHz (megahertz) min mL (milliliter) mmol millimol mPB mobilized peripheral blood MS mass spectrometry MsCl methanesulfonyl chloride (CH3SO2Cl) MsOH Methanesulfonic acid (CH3SO3H) Sodium sulfate (Na2SO4) NaBH(OAc)3 Sodium borotriacetoxyhydride NaHCO3 (sodium bicarbonate) NMR nuclear magnetic resonance on overnight PBS (phosphate-buffered saline) Complex of PdCl2(dppf)·DCM [1,1'-bis(diphenylphosphin)ferrocene]dichloropalladium(II) with dichloromethane q quadruple line qd Double line quadruple line quint quintet quintd (double line, quintuple line) rbf round-bottom flask rhEPO Recombinant Human Erythropoietin rhIL-3 Recombinant Human Interleukin-3 rhIL-6 Recombinant Human Interleukin-6 rhSCF Recombinant Human Stem Cell Factor rhTPO recombinant human thrombopoietin RNP (Ribonucleoprotein) Rt retention time rt or rt room temperature s single line SEM 2-(trimethylsilyl)ethoxymethyl shRNA (low-molecular-weight hairpin RNA) t triple line td Double line triple line tdd Double line double line triple line TEA(NEt3) Triethylamine TFA (Trifluoroacetic Acid) TfOH trifluic acid THF (Tetrahydrofuran) TLC (Thin-Layer Chromatography) TMP 2,2,6,6-tetramethylpiperidine tracrRNA (trans-activated crRNA) Ts Tosil tt Mie Line Mie Line ttd Double line triple line triple line μW or uW microwave UPLC Ultra-High-Speed ​​Liquid Chromatography WIZ: Zinc finger-containing proteins arranged at wide intervals.

[0081] definition The terms “CRISPR system,” “Cas system,” or “CRISPR / Cas system” refer to a set of molecules comprising a gRNA molecule and an RNA-guided nuclease or other effector molecule that together are necessary and sufficient to induce and achieve nucleic acid modification by an RNA-guided nuclease or other effector molecule at a target sequence. In one embodiment, the CRISPR system comprises a gRNA and a Cas protein, such as the Cas9 protein. Such a system comprising Cas9 or a modified Cas9 molecule is referred to herein as the “Cas9 system” or “CRISPR / Cas9 system.” In one example, the gRNA molecule and the Cas molecule may complex together to form a ribonucleoprotein (RNP) complex.

[0082] The terms “guide RNA,” “guide RNA molecule,” “gRNA molecule,” or “gRNA” are used synonymously and refer to a set of nucleic acid molecules that facilitate the specific guidance of an RNA-guided nuclease or other effector molecule (typically complexed with a gRNA molecule) to a target sequence. In some embodiments, such guidance is achieved by hybridizing a portion of the gRNA to DNA (e.g., via the gRNA targeting domain) and by binding a portion of the gRNA molecule to an RNA-guided nuclease or other effector molecule (e.g., at least via the gRNA tracr). In some embodiments, the gRNA molecule consists of a single, consecutive polynucleotide molecule and is referred to herein as “single guide RNA” or “sgRNA,” etc. In other embodiments, the gRNA molecule consists of multiple, usually two, polynucleotide molecules that themselves are typically capable of associating through hybridization and is referred herein as “dual guide RNA” or “dgRNA,” etc. The gRNA molecule is described in more detail below, but generally includes a targeting domain and a tracr. In one embodiment, the targeting domain and tracr are located on a single polynucleotide. In another embodiment, the targeting domain and tracr are located on separate polynucleotides.

[0083] The term "targeting domain," when used in relation to gRNA, refers to a part of a gRNA molecule that recognizes, for example, a target sequence within a cellular nucleic acid, for example, a target sequence within a gene, for example, and is complementary to it.

[0084] The term "crRNA," when used in relation to gRNA molecules, refers to a portion of a gRNA molecule that includes a targeting domain and a region that interacts with tracr to form a flagpole region.

[0085] The term "target sequence" refers to a nucleic acid sequence that is complementary, for example, perfectly complementary, to the gRNA targeting domain. In embodiments, the target sequence is located on genomic DNA. In some embodiments, the target sequence is adjacent (either on the same strand or complementary strand of DNA) to a protospacer-adjacent motif (PAM) sequence recognized by a protein having nuclease activity or other effector activity, such as a PAM sequence recognized by Cas9. In embodiments, the target sequence is intragenetic or intralocusal, affecting the expression of globin genes, such as intragenetic or intralocusal, affecting the expression of β-globin or fetal hemoglobin (HbF). In embodiments, the target sequence is intragenetic or intralocusal, affecting the expression of WIZ gene regions.

[0086] When the term "flagpole" is used herein in relation to gRNA molecules, it refers to the portion of gRNA that binds to or hybridizes with crRNA and tracr.

[0087] The term "tracr," when used herein in relation to a gRNA molecule, refers to a portion of gRNA that binds to a nuclease or other effector molecule. In embodiments, tracr includes a nucleic acid sequence that specifically binds to Cas9. In embodiments, tracr includes a nucleic acid sequence that forms part of a flagpole.

[0088] The term "Cas9" or "Cas9 molecule" refers to the enzyme in the bacterial type II CRISPR / Cas system involved in DNA cleavage. Cas9 includes the wild-type protein as well as its functional and non-functional mutants. In embodiments, Cas9 is the Cas9 of Streptococcus pyogenes.

[0089] The term "complementary," when used in relation to nucleic acids, refers to the pairing of bases A with T or U, and G with C. The term complementary refers to nucleic acid molecules that are perfectly complementary, i.e., A pairs with T or U and G pairs with C throughout the entire reference sequence, as well as molecules that are at least 80%, 85%, 90%, 95%, or 99% complementary.

[0090] "Template nucleic acid" refers to the nucleic acid that is inserted into the modification site by the CRISPR system donor sequence for gene repair (insertion) at the cleavage site when used in relation to homologous recombination repair or homologous recombination.

[0091] When the term “indel” is used herein, it refers to a nucleic acid comprising one or more nucleotide insertions, one or more nucleotide deletions, or a combination of nucleotide insertions and deletions compared to a reference nucleic acid, resulting from a composition comprising a gRNA molecule, for example, after exposure to a CRISPR system. Indels can be determined by sequencing the nucleic acid after exposure to a composition comprising a gRNA molecule, for example, by NGS. With respect to the location of an indel, an indel is said to be “at or near” the reference site (e.g., the site complementary to the targeting domain of the gRNA molecule) if it comprises at least one insertion or deletion within a range of about 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from the reference site, or overlaps with part or all of the said reference site (e.g., overlaps with a site complementary to the targeting domain of the gRNA molecule, for example, as described herein, or comprises at least one insertion or deletion within a range of 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide from there). In embodiments, the indel is a large deletion containing nucleic acid, for example, more than about 1 kb, more than about 2 kb, more than about 3 kb, more than about 4 kb, more than about 5 kb, more than about 6 kb, or more than about 10 kb. In embodiments, the 5' end, 3' end, or both of the 5' and 3' ends of the large deletion are located in or near the target sequence of the gRNA molecule described herein. In embodiments, the large deletion contains about 4.9 kb of DNA located within the WIZ gene region, for example, between the target sequences of the gRNA molecule described herein.

[0092] When this term is used herein, “indel pattern” refers to a set of indels that occur after exposure to a composition containing a gRNA molecule. In one embodiment, the indel pattern consists of the top three indels based on frequency of occurrence. In one embodiment, the indel pattern consists of the top five indels based on frequency of occurrence. In one embodiment, the indel pattern consists of indels present at a frequency of more than about 1% of all sequencing reads. In one embodiment, the indel pattern consists of indels present at a frequency of more than about 5% of all sequencing reads. In one embodiment, the indel pattern consists of indels present at a frequency of more than about 10% of the total number of indel sequencing reads (i.e., reads that do not consist of unmodified reference nucleic acid sequences). In one embodiment, the indel pattern includes any three of the top five most frequently observed indels. The indel pattern may be determined, for example, by sequencing cells of a cell population exposed to a gRNA molecule, for example, by the method described herein.

[0093] When used herein, “off-target indel” refers to an indel located outside or near the target sequence of the targeting domain of a gRNA molecule. Such a site may contain, for example, one, two, three, four, five or more mismatched nucleotides compared to the sequence of the targeting domain of the gRNA. In exemplary embodiments, such sites are detected using in silico prediction-based targeted sequencing of off-target sites or by insertion methods known in the art. With respect to the gRNAs described herein, an example of an off-target indel is an indel formed in a sequence outside the WIZ gene region. In exemplary embodiments, the off-target indel is formed within the gene sequence, for example, the gene coding sequence.

[0094] The terms "a" and "an" refer to one or more (i.e., at least one) of the grammatical referents of the articles. For example, "a element" means one or more elements.

[0095] The term "and / or" means either "and" or "or" unless otherwise specified.

[0096] When referring to measurable values ​​such as quantity or duration, the term “about” is intended to include variations of ±20%, or in some cases ±10%, or in some cases ±5%, or in some cases ±1%, or in some cases ±0.1% from the specified value (as such variations are appropriate for carrying out the methods of this disclosure).

[0097] The terms “antigen” or “Ag” refer to a molecule that elicits an immune response. This immune response may include either antibody production or activation of specific immune-qualified cells, or both. Those skilled in the art will understand that any macromolecule, including virtually any protein or peptide, can be an antigen. Furthermore, antigens may originate from recombinant DNA or genomic DNA. Those skilled in the art will understand that any DNA containing a nucleotide sequence or partial nucleotide sequence encoding a protein that elicits an immune response will encode an “antigen” as the term is used herein. Furthermore, those skilled in the art will understand that an antigen does not have to be encoded by a full-length nucleotide sequence of a gene. It is readily apparent that the present invention includes, but is not limited to, the use of partial nucleotide sequences of two or more genes, and that these nucleotide sequences are sequenced in various combinations to encode a polypeptide that elicits a desired immune response. Furthermore, those skilled in the art will understand that an antigen does not have to be encoded by a “gene” at all. It is readily apparent that antigens may be synthesized, obtained from biological samples, or may be macromolecules other than polypeptides. Such biological samples may include, but are not limited to, tissue samples, cells, or bodily fluids containing other biological components.

[0098] The term "self-derived" refers to any material that originates from the same individual into which it will later be reintroduced.

[0099] The term "homogeneous" refers to any material originating from different animals of the same species as the individual into which it is introduced. Two or more individuals are said to be homogeneous if they do not have identical genes at one or more loci. In some embodiments, homogeneous material from individuals of the same species may be genetically distinct enough to interact as antigens.

[0100] The term "heterogeneous" refers to a graft derived from an animal of a different species.

[0101] The phrase "derived from" indicates a relationship between a first molecule and a second molecule when used herein. This generally refers to a structural similarity between the first molecule and the second molecule and does not imply or include any limitation of the process or source to the first molecule derived from the second molecule.

[0102] The term "coding" refers to the inherent property of a specific nucleotide sequence in a polynucleotide, such as a gene, cDNA, or mRNA, to serve as a template for the synthesis of any defined nucleotide sequence (e.g., rRNA, tRNA, and mRNA) or a defined amino acid sequence, and other polymers and macromolecules in biological processes having the biological properties derived therefrom. Thus, a gene, cDNA, or RNA codes for a protein if the transcription and translation of the mRNA corresponding to that gene produces a protein in a cell or other biological system. Both the coding strand, whose nucleotide sequence is identical to the mRNA sequence and is usually provided in a sequence listing, and the non-coding strand, which is used as a transcription template for the gene or cDNA, may be said to code for a protein or other product of that gene or cDNA.

[0103] Unless otherwise specified, "nucleotide sequences encoding an amino acid sequence" includes all nucleotide sequences that encode the same amino acid sequence, including degenerate versions of each other. The phrase "nucleotide sequence encoding a protein or RNA" may also include introns, insofar as the nucleotide sequence encoding that protein may contain one or more introns in any version.

[0104] The terms “effective dose” and “therapeutic dose” are used synonymously herein and refer to an amount of any compound, formulation, material, or composition as described herein that is effective in achieving a specific biological outcome, such as a reduction or inhibition of enzyme or protein activity, or improving symptoms, alleviating a condition, slowing or delaying the progression of a disease, or preventing a disease. In one embodiment, the term “therapeutic dose” refers to an amount of any compound of the Disclosure that, upon administration to a subject, is effective in at least partially alleviating, preventing, and / or improving a condition, disorder, or disease characterized by (i) WIZ-mediated, (ii) WIZ activity, or (iii) WIZ (normal or abnormal) activity; (2) reducing or inhibiting WIZ activity; or (3) reducing or inhibiting the expression level of the WIZ gene and / or protein. In another embodiment, the term “therapeutic dose” means an amount of the compound of the Disclosure that, upon administration to cells, or tissues, or noncellular biological materials, or culture media, is effective in at least partially reducing or inhibiting the activity of WIZ; or effective in at least partially reducing or inhibiting the expression levels of the WIZ gene and / or protein.

[0105] As used herein, the terms “inhibit,” “inhibit,” or “to inhibit” mean the reduction or suppression of a given pathological condition, symptom, disorder, or disease, or a significant decrease in the baseline activity of a biological activity or process, or a significant decrease in the baseline expression level of the gene and / or protein of interest.

[0106] The term "endogenous" refers to any material that originates from or is produced within an organism, cell, tissue, or system.

[0107] The term "exogenous" refers to any material introduced from or produced outside of an organism, cell, tissue, or system.

[0108] The term "expression" refers to the transcription and / or translation of a specific nucleotide sequence driven by a promoter.

[0109] The term "transfer vector" refers to a composition containing isolated nucleic acid that can be used for the delivery of isolated nucleic acid into a cell. Many vectors are known in the art, including, but are not limited to, linear polynucleotides, polynucleotides associated with ionic or amphiphilic compounds, plasmids, and viruses. Therefore, the term "transfer vector" includes self-replicating plasmids or viruses. The term should also be interpreted to further include non-plasmid and non-viral compounds that facilitate the transfer of nucleic acids into cells, such as polylysine compounds and liposomes. Examples of viral transfer vectors include, but are not limited to, adenovirus vectors, adeno-associated virus vectors, retroviral vectors, and lentiviral vectors.

[0110] The term "expression vector" refers to a vector containing recombinant polynucleotides that include an expression regulatory sequence operably ligated to the nucleotide sequence to be expressed. An expression vector contains sufficient cis-acting elements for expression. Other expression elements can be supplied by the host cell or to an in vitro expression system. Expression vectors include all known in the art, including cosmids, plasmids (e.g., naked or liposome-containing) and viruses (e.g., lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses) incorporating recombinant polynucleotides.

[0111] The terms "homologous" or "identical" refer to the sequence identity of subunits between two polymer molecules, between two nucleic acid molecules such as two DNA molecules or two RNA molecules, or between two polypeptide molecules. When subunit positions in both molecules are occupied by the same monomeric subunit, for example, when each position in two DNA molecules is occupied by adenine, they are homologous or identical at that position. Homologousity between two sequences is a direct function of the number of matching or homologous positions. For example, if half of the positions in two sequences (e.g., five positions in a polymer of 10 subunits) are homologous, then those two sequences are 50% homologous. If 90% of the positions (e.g., nine out of ten) are matching or homologous, then those two sequences are 90% homologous.

[0112] The term "isolated" means that something has been modified or removed from its natural state. For example, nucleic acids or peptides that are naturally present in living animals are not "isolated," but the same nucleic acids or peptides that have been partially or completely separated from the coexisting material in their natural state are "isolated." Isolated nucleic acids or proteins can exist in a substantially purified form or in a non-natural environment, such as in a host cell.

[0113] The terms "operably linked" or "transcriptional regulation" refer to a functional link between a regulatory sequence and a heterologous nucleic acid sequence that results in the expression of the latter. For example, when a first nucleic acid sequence is functionally related to a second nucleic acid sequence, the first nucleic acid sequence is operably linked to the second nucleic acid sequence. For example, when a promoter affects the transcription or expression of a coding sequence, the promoter is operably linked to the coding sequence. The operably linked DNA sequences may be adjacent to each other; for example, if two protein coding regions need to be joined together, they may be within the same reading frame.

[0114] The term "parenteral" administration of immunogenic compositions includes, for example, subcutaneous (sc), intravenous (iv), intramuscular (im), or intrasternal injection, intratumoral, or infusion techniques.

[0115] The terms “nucleic acid” or “polynucleotide” refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in either single-stranded or double-stranded forms, and their polymers. Unless specifically limited, the term includes nucleic acids, including known analogues of natural nucleotides, that have similar binding properties to a reference nucleic acid and are metabolized in the same way as naturally occurring nucleotides. Unless otherwise indicated, a specific nucleic acid sequence also implicitly includes its conservedly modified variants (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as sequences explicitly indicated. Specifically, degenerate codon substitution can be achieved by creating sequences in which the third position of one or more (or all) codons is substituted with a mixed base and / or a deoxyinosine residue (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)).

[0116] The terms “peptide,” “polypeptide,” and “protein” are used synonymously and refer to compounds composed of amino acid residues covalently linked by peptide bonds. A protein or peptide must contain at least two amino acids, and there is no limit to the maximum number of amino acids that may be included in a protein sequence or peptide sequence. A polypeptide includes any peptide or protein containing two or more amino acids linked together by peptide bonds. As used herein, this term refers to both short chains, also commonly called peptides, oligopeptides, and oligomers in the art, and longer chains of many types, generally referred to in the art as proteins. Polypeptides particularly include, for example, biologically active fragments, substantially homologous polypeptides, oligopeptides, homodimers, heterodimers, polypeptide variants, modified polypeptides, derivatives, analogs, and fusion proteins. Polypeptides include native peptides, recombinant peptides, or combinations thereof.

[0117] The term "promoter" refers to a DNA sequence recognized by a cellular synthetic mechanism, or an introduced synthetic mechanism, that is necessary to initiate the specific transcription of a polynucleotide sequence.

[0118] The term "promoter / regulatory sequence" refers to a nucleic acid sequence necessary for the expression of a gene product that is operably ligated to that promoter / regulatory sequence. In some cases, this sequence may be a core promoter sequence, while in other cases, it may also include enhancer sequences and other regulatory elements necessary for the expression of the gene product. A promoter / regulatory sequence may, for example, cause tissue-specific expression of a gene product.

[0119] The term "constitutive" promoter refers to a nucleotide sequence that, when operably linked to a polynucleotide encoding or designating a gene product, causes the cell to produce the gene product under most or all physiological conditions.

[0120] The term "inducible" promoter refers to a nucleotide sequence that, when operably linked to a polynucleotide encoding or designating a gene product, causes the production of a gene product in a cell only if a substantially equivalent inducer is present in the cell.

[0121] The term "tissue-specific" promoter refers to a nucleotide sequence that, when operably linked to a polynucleotide encoding or designated by a gene, causes the production of a gene product in a cell only if the cell is substantially the tissue type corresponding to the promoter.

[0122] As used herein, “modulator” or “degrader” means, for example, a compound of the Disclosure that effectively modulates, reduces, or diminishes the level of a specific protein (e.g., WIZ), or degrades a specific protein (e.g., WIZ). The amount of specific protein (e.g., WIZ) that is degraded can be measured by comparing the amount of specific protein (e.g., WIZ) remaining after treatment with the compound of the Disclosure with the initial amount or level of specific protein (e.g., WIZ) present when measured before treatment with the compound of the Disclosure.

[0123] As used herein, “selective modifier,” “selective degrader,” or “selective compound” means a compound of the Disclosure that effectively modulates, reduces, or degrades the level of a specific protein (e.g., WIZ) to a greater extent than any other protein, or degrades the specific protein (e.g., WIZ). “Selective modifier,” “selective degrader,” or “selective compound” can be identified, for example, by comparing the compound’s ability to modulate, reduce, or degrade the level of a specific protein (e.g., WIZ) or degrade it with its ability to modulate, reduce, or degrade the level of other proteins. In some embodiments, selectivity is defined as the EC of the compound. 50 or IC 50It can be identified by measuring [the relevant factor]. Degradation may also be achieved through the intervention of an E3 ligase, such as an E3-ligase complex containing the protein cereblon.

[0124] When used herein in relation to messenger RNA (mRNA), the 5' cap (also referred to as the RNA cap, RNA7-methylguanosine cap, or RNA m7G cap) is a modified guanine nucleotide added to the "pre" or 5' end of eukaryotic messenger RNA immediately after transcription initiation. The 5' cap consists of a terminal group linked to the first transcription nucleotide. Its presence is important for ribosome recognition and protection from RNases. Capping is linked to transcription and occurs synchronously, with each influencing the other. Immediately after transcription initiation, a cap synthesis complex associated with RNA polymerase binds to the 5' end of the synthesized mRNA. This enzyme complex catalyzes the chemical reactions necessary for mRNA capping. Synthesis proceeds as a multi-step biochemical reaction. The capping portion can be modified to regulate the function of the mRNA, such as its stability or translation efficiency.

[0125] As used herein, “in vitro transcription RNA” or “IVT RNA” refers to RNA synthesized in vitro, preferably mRNA. Generally, in vitro transcription RNA is prepared from an in vitro transcription vector. An in vitro transcription vector contains a template used to prepare in vitro transcription RNA.

[0126] As used herein, "poly(A)" refers to a series of adenosines added to mRNA by polyadenylation. In a preferred embodiment of a construct for transient expression, poly(A) is 50 to 5000 (SEQ ID NO: 3118), preferably more than 64, more preferably more than 100, and most preferably more than 300 or 400. The poly(A) sequence can be chemically or enzymatically modified to modulate mRNA function such as localization, stability, or translation efficiency.

[0127] As used herein, “polyadenylation” refers to the covalent bonding of a polyadenylyle moiety or a modified variant thereof to a messenger RNA molecule. In eukaryotes, most messenger RNA (mRNA) molecules are polyadenylated at their 3' end. The 3' poly(A) tail is a long sequence of adenine nucleotides (often several hundred) added to the mRNA precursor by the action of the enzyme, polyadenylate polymerase. In higher eukaryotes, the poly(A) tail is added to a transcript containing a specific sequence, the polyadenylation signal. The poly(A) tail and the protein bound to it help protect mRNA from degradation by exonucleases. Polyadenylation is also important for transcription termination, export of mRNA from the nucleus, and translation. Polyadenylation occurs in the nucleus immediately after transcription from DNA to RNA, but can also occur later in the cytoplasm. After transcription is complete, the mRNA strand is cleaved by the action of endonuclease complexes associated with RNA polymerase. The cleavage site is usually characterized by the presence of the nucleotide sequence AAUAAA near the cleavage site. After the mRNA is cleaved, an adenosine residue is added to the free 3' end of the cleavage site.

[0128] As used herein, “transient” refers to the expression of a non-integrated transgene over a period of several hours, several days, or several weeks, the duration of which is shorter than the duration of expression of the gene when integrated into the host cell genome or contained within a stable plasmid replicon.

[0129] As used herein, the terms “to treat,” “treatment,” and “treating” mean a reduction or improvement in the progression, severity, and / or duration of a disorder, e.g.,

[0130] As used herein, the terms “prevent,” “prevent,” or “prevention” of any disease or disorder mean the preventive treatment of a disease or disorder; or delaying the onset or progression of a disease or disorder.

[0131] As used herein, “HbF-dependent disease or disorder” means any disease or disorder that is directly or indirectly affected by the regulation of HbF protein levels. Preferred examples of such diseases or disorders are abnormal hemoglobin disorders, such as sickle cell disease or thalassemia (e.g., β-thalassemia).

[0132] As used herein, an object “needs treatment” if such treatment would benefit the object biologically, medically, or in terms of quality of life.

[0133] The term "signaling pathway" refers to the biochemical relationships between various signaling molecules that play a role in the transmission of signals from one part of a cell to another. The term "cell surface receptor" includes molecules and complexes of molecules that have the ability to receive signals and transmit them across the cell membrane.

[0134] The term “subject” is intended to include living organisms (e.g., mammals, humans) from which an immune response can be elicited. Preferably, the term “subject” refers to primates (e.g., humans, males or females), dogs, rabbits, guinea pigs, pigs, rats, and mice. In certain embodiments, the subject is a primate. In yet other embodiments, the subject is a human.

[0135] The term "substantially purified" refers to cells that essentially contain no other cell types. Substantially purified cells also refer to cells that are isolated from the other cell types to which they normally bind in their natural state. In some examples, a population of substantially purified cells refers to a homogeneous population of cells. In other examples, the term simply refers to cells that are isolated from the cells to which they naturally bind in their natural state. In some embodiments, these cells are cultured in vitro. In other embodiments, these cells are not cultured in vitro.

[0136] The term "therapeutic," as used herein, means treatment. A therapeutic effect is achieved by reducing, suppressing, relieving, or eradicating a disease condition.

[0137] The term "prevention," as used herein, means the preventive or protective treatment of a disease or disease condition.

[0138] The terms "transfected," "transformed," or "transduced" refer to the process of introducing or transferring exogenous nucleic acids and / or proteins into host cells. A "transfected," "transformed," or "transduced" cell is one that has been transfected, transformed, or transduced with exogenous nucleic acids and / or proteins. Cells include primary target cells and their offspring.

[0139] The term "specifically binds" refers to a molecule that recognizes and binds to a binding partner (e.g., a protein or nucleic acid) present in the sample, but does not substantially recognize or bind to other molecules in the sample.

[0140] The term "biologically equivalent" refers to an amount of a drug other than the reference compound that is required to produce an effect equivalent to that produced by the reference dose or amount of the reference compound.

[0141] As used herein, "refractory" refers to a disease that does not respond to treatment, such as a hemoglobin disorder. In some embodiments, a refractory hemoglobin disorder may be resistant to treatment before or at the start of treatment. In other embodiments, a refractory hemoglobin disorder may become resistant during treatment. A refractory hemoglobin disorder is also referred to as a resistant hemoglobin disorder.

[0142] When used herein, "recurrent" refers to the recurrence of a disease (e.g., a disorder) or signs and symptoms of a disease, such as a disordered hemoglobinosis, after a period of improvement, for example, after prior treatment with a certain therapy, such as a disordered hemoglobinosis therapy.

[0143] Scope: Throughout this disclosure, various aspects of the invention may be presented in the form of scopes. It should be understood that descriptions in the form of scopes are merely for convenience and for brevity, and should not be interpreted as restricting the scope of the invention inflexibly. Accordingly, descriptions of scopes should be considered to specifically disclose any possible partial scopes and the individual numbers within those scopes. For example, a description of a scope such as 1-6 should be considered to specifically disclose partial scopes such as 1-3, 1-4, 1-5, 2-4, 2-6, 3-6, and the individual numbers within those scopes, e.g., 1, 2, 2.7, 3, 4, 5, 5.3, and 6. Another example is a scope such as 95-99% identity, which includes those having 95%, 96%, 97%, 98%, or 99% identity, and partial scopes such as 96-99%, 96-98%, 96-97%, 97-99%, 97-98%, and 98-99% identity. This applies regardless of the scope.

[0144] The term "WIZ" refers to the gene encoding the Widely-Interspaced Zinc Finger-Containing Protein, or its variants or homologs that maintain its transcriptional activity (e.g., at least 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of WIZ's activity), along with all introns and exons, and its regulatory regions such as promoters and enhancers. This gene encodes the zinc finger protein. WIZ is also known as zinc finger protein 803, ZNF803, the widely-interspaced zinc finger motif, or WIZ zinc finger. This term encompasses all isoforms and splice variants of WIZ. The human gene encoding WIZ is mapped to chromosome 19: 15,419,980–15,449,951 (according to Ensembl). For human and mouse amino acid and nucleic acid sequences, public databases such as GenBank, UniProt, and Swiss-Prot can be consulted. For the human WIZ genome sequence, NC_000019.10 can be found in GenBank. The WIZ gene refers to this genomic location, including all introns and exons. There are several known isotypes of WIZ. In some embodiments, mutants or homologs have at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a contiguous segment of 50, 100, 150, or 200 amino acids) compared to the naturally occurring WIZ protein. Exemplary WIZ transcription mutants and their genomic coordinates are shown in Table 4.

[0145] [Table 1]

[0146] [Table 2]

[0147] [Table 3]

[0148] In an embodiment, exemplary WIZ transfer variants and their nucleotide sequences are shown in Table 5 below.

[0149] [Table 4]

[0150] [Table 5]

[0151] The peptide sequence of human WIZ isoform 1 is as follows: [Chemical formula] Accession number 3173 (UniProt accession number O95785-1)

[0152] The sequences of other WIZ protein isoforms are provided below: Isoform 2: UniProt O95785-2 Isoform 3: UniProt O95785-3 Isoform 4: UniProt O95785-4.

[0153] Alternatively, the isoforms of the WIZ protein have the amino acid sequences of NCBI reference sequences NP_067064.2, NP_001317324.2, NP_001358518.1, NP_001358532.2, XP_005260064.1, XP_005260062.1, XP_005260063.1, XP_005260065.1, XP_005260068.1, XP_006722891.1, XP_005260067.1, XP_011526465.1, or XP_024307397.1.

[0154] When used herein, human WIZ proteins also include proteins having at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with the WIZ isoforms disclosed herein over their full length, wherein such proteins still possess at least one of the functions of WIZ.

[0155] The term "complementary," when used in relation to nucleic acids, refers to the pairing of bases A with T or U, and G with C. The term complementary refers to nucleic acid molecules that are perfectly complementary, i.e., A pairs with T or U and G pairs with C throughout the entire reference sequence, as well as molecules that are at least 80%, 85%, 90%, 95%, or 99% complementary.

[0156] The terms “hematopoietic stem cells / progenitor cells” or “HSPC” are used synonymously and refer to a cell population that includes both hematopoietic stem cells ("HSCs") and hematopoietic progenitor cells ("HPCs"). Such cells are characterized, for example, as CD34+. In exemplary embodiments, HSPCs are isolated from bone marrow. In other exemplary embodiments, HSPCs are isolated from peripheral blood. In other exemplary embodiments, HSPCs are isolated from umbilical cord blood. In one embodiment, HSPCs are characterized as CD34+ / CD38- / CD90+ / CD45RA-. In an embodiment, HSPCs are characterized as CD34+ / CD90+ / CD49f+ cells. In an embodiment, HSPCs are characterized as CD34+ cells. In an embodiment, HSPCs are characterized as CD34+ / CD90+ cells. In an embodiment, HSPCs are characterized as CD34+ / CD90+ / CD45RA- cells.

[0157] When used herein, “stem cell growth agent” refers to a compound that causes cells, e.g., HSPCs, HSCs, and / or HPCs, to proliferate at a faster rate than the same cell type without the agent, for example, by increasing their number. In one exemplary embodiment, the stem cell growth agent is an antagonist of the aryl hydrocarbon receptor pathway. Further examples of stem cell growth agents are provided below. In embodiments, proliferation, e.g., increase in number, is achieved ex vivo.

[0158] "Engraftment" or "to engraft" refers to the uptake of a population of cells or tissues, such as HSPCs, into the body of a recipient, such as a mammal or human subject. In one example, engraftment includes the growth, proliferation, and / or differentiation of the engrafted cells in the recipient. In one example, engraftment of HSPCs includes the differentiation and growth of the HSPCs into erythroid cells within the body of the recipient.

[0159] The term “hematopoietic progenitor cell” (HPC), as used herein, refers to primitive hematopoietic cells with limited self-renewal capacity and the potential for multi-lineage differentiation (e.g., bone marrow, lymphoid), mono-lineage differentiation (e.g., bone marrow or lymphoid), or cell-type-limited differentiation (e.g., erythroid progenitor cells), depending on their position in the hematopoietic hierarchy (Doulatov et al., Cell Stem Cell 2012).

[0160] When used herein, “hematopoietic stem cells” (HSCs) refer to immature hematopoietic cells that have the ability to self-replicate and differentiate into more mature hematopoietic cells, including granulocytes (e.g., promyelocytes, neutrophils, eosinophils, basophils), erythrocytes (e.g., reticulocytes, erythrocytes), thrombus cells (e.g., megakaryoblasts, platelet-producing megakaryocytes, platelets), and monocytes (e.g., monocytes, macrophages). Throughout this specification, HSCs are described synonymously with stem cells. It is known in the art that such cells may or may not include CD34+ cells. CD34+ cells are immature cells that express the CD34 cell surface marker. CD34+ cells are thought to constitute a subpopulation of cells having the stem cell characteristics defined above. It is well known in the art that HSCs are pluripotent cells that can give rise to primitive progenitor cells (e.g., pluripotent progenitor cells) and / or progenitor cells committed to specific hematopoietic lineages (e.g., lymphocyte progenitor cells). Stem cells committed to a specific hematopoietic lineage may be from T cell lineages, B cell lineages, dendritic cell lineages, Langerhans cell lineages, and / or lymphoid tissue-specific macrophage cell lineages. In addition, HSCs also refer to long-term HSCs (LT-HSCs) and short-term HSCs (ST-HSCs). ST-HSCs are more active and proliferative than LT-HSCs. However, LT-HSCs have unlimited self-renewal (i.e., they survive throughout adulthood), while ST-HSCs have limited self-renewal (i.e., they survive for only a limited period). Any of these HSCs can be used in any of the methods described herein. ST-HSCs are useful by choice because they are proliferative and therefore the number of HSCs and their offspring increases rapidly. Hematopoietic stem cells are obtained by choice from blood products. Blood products include preparations obtained from the body or organs of the body that contain cells of hematopoietic origin. Such sources include unfractionated bone marrow, umbilical cord, peripheral blood (e.g., mobilized peripheral blood, e.g., mobilized with mobilizing agents such as G-CSF or Plerixafor® (AMD3100), or a combination of G-CSF and Plerixafor® (AMD3100)), liver, thymus, lymph, and spleen.All of the aforementioned crude or unfractionated blood products can be enriched in ways known to those skilled in the art with respect to cells having hematopoietic stem cell characteristics. In one embodiment, HSCs are characterized as CD34+ / CD38- / CD90+ / CD45RA-. In another embodiment, HSCs are characterized as CD34+ / CD90+ / CD49f+ cells. In yet another embodiment, HSCs are characterized as CD34+ cells. In yet another embodiment, HSCs are characterized as CD34+ / CD90+ cells. In yet another embodiment, HSCs are characterized as CD34+ / CD90+ / CD45RA- cells.

[0161] In relation to cells, "proliferation" or "proliferating" refers to an increase in the number of one or more characteristic cell types from an initial population of cells, which may or may not be the same. The initial cells used for proliferation do not have to be the same as the cells produced by proliferation.

[0162] "Cell population" refers to cells of a eukaryotic mammal, preferably human, isolated from a biological source, such as a blood product or tissue, and derived from two or more cells.

[0163] When used in relation to cell populations, "enriched" refers to a population of cells selected based on the presence of one or more markers, such as CD34+.

[0164] The term "CD34+ cells" refers to cells that express the CD34 marker on their surface. CD34+ cells can be detected and counted, for example, using flow cytometry and fluorescently labeled anti-CD34 antibodies.

[0165] "CD34+ cell enrichment" means that the cell population has been selected based on the presence of the CD34 marker. Therefore, the percentage of CD34+ cells in the cell population after the selection method is higher than the percentage of CD34+ cells in the initial cell population before the CD34 marker selection step. For example, CD34+ cells may account for at least 50%, 60%, 70%, 80%, or at least 90% of the cells in the CD34+ cell enriched cell population.

[0166] The terms “F cell” and “F- cell” refer to cells that contain and / or produce (e.g., express) fetal hemoglobin, usually red blood cells (e.g., erythrocytes). For example, an F- cell is a cell that contains or produces a detectable level of fetal hemoglobin. For example, an F- cell is a cell that contains or produces at least 5 picograms of fetal hemoglobin. In another example, an F- cell is a cell that contains or produces at least 6 picograms of fetal hemoglobin. In another example, an F- cell is a cell that contains or produces at least 7 picograms of fetal hemoglobin. In another example, an F- cell is a cell that contains or produces at least 8 picograms of fetal hemoglobin. In another example, an F- cell is a cell that contains or produces at least 9 picograms of fetal hemoglobin. In another example, an F- cell is a cell that contains or produces at least 10 picograms of fetal hemoglobin. Fetal hemoglobin levels can be measured using the assays described herein or by other methods known in the art, such as flow cytometry, high-performance liquid chromatography, mass spectrometry, or enzyme-linked immunosorbent assays using anti-fetal hemoglobin detection reagents.

[0167] An "inhibitor" is, for example, an siRNA (e.g., shRNA, miRNA, snoRNA), gRNA, compound, or small molecule that inhibits cellular function (e.g., replication) by binding, partially or completely blocking stimulation, reducing, preventing, or delaying activation, or inactivating, desensitizing, or downregulating signal transduction, gene expression, or enzyme activity required for protein activity. A "WIZ inhibitor" refers to a substance that results in a detectably lower level of WIZ gene or WIZ protein expression or a lower level of WIZ protein activity when compared to the level in the absence of such a substance. In some embodiments, the WIZ inhibitor is a small molecule compound (e.g., a small molecule compound that can target WIZ for degradation, also known as a "WIZ degrader"). In some embodiments, the WIZ inhibitor is an anti-WIZ shRNA. In some embodiments, the WIZ inhibitor is an anti-WIZ siRNA. In some embodiments, the WIZ inhibitor is an anti-WIZ ASO. In some embodiments, the WIZ inhibitor is an anti-WIZ AMO. In some embodiments, the WIZ inhibitor is an anti-WIZ antisense nucleic acid. In some embodiments, the WIZ inhibitor is a composition or cell or population of cells described herein (including those containing the gRNA molecules described herein).

[0168] When referred to herein, “antisense nucleic acid” is a nucleic acid (e.g., a DNA or RNA molecule) complementary to at least a portion of a specific target nucleic acid (e.g., protein-translatable mRNA) that has the ability to reduce the transcription of the target nucleic acid (e.g., from DNA to mRNA) or reduce the translation of the target nucleic acid (e.g., mRNA) or to modify the splicing of the transcript (e.g., single-stranded morpholino oligo). See, for example, Weintraub, Scientific American, 262:40 (1990). Typically, synthetic antisense nucleic acids (e.g., oligonucleotides) are generally 15 to 25 nucleotides long. Thus, antisense nucleic acids have the ability to hybridize with the target nucleic acid (e.g., target mRNA) (e.g., selective hybridization with it). In embodiments, the antisense nucleic acid hybridizes to the target nucleic acid sequence (e.g., mRNA) under stringent hybridization conditions. In embodiments, the antisense nucleic acid hybridizes to the target nucleic acid (e.g., mRNA) under moderately stringent hybridization conditions. The antisense nucleic acid may include naturally occurring nucleotides or modified nucleotides, such as phosphorothioates, methylphosphonates, and anomeric sugar phosphates, and skeletal modified nucleotides.

[0169] Within a cell, antisense nucleic acids hybridize to the corresponding mRNA, forming a double-stranded molecule. Since the cell does not translate the double-stranded mRNA, the antisense nucleic acid interferes with mRNA translation. Inhibition of in vitro gene translation using antisense methods is well known in the art (Marcus-Sakura, Anal. Biochem., 172:289, (1988)). Furthermore, antisense molecules that directly bind to DNA may be used. Antisense nucleic acids can be single-stranded or double-stranded nucleic acids. Non-limiting examples of antisense nucleic acids include siRNA (including nucleotide analogs, derivatives or precursors thereof), small hairpin RNA (shRNA), microRNA (miRNA), saRNA (small activated RNA), and small nucleolar RNA (snoRNA), or certain derivatives or precursors thereof.

[0170] "siRNA" refers to a nucleic acid that forms double-stranded RNA, which, when present in the same cell as the gene or target gene (e.g., when expressed), has the ability to reduce or inhibit the expression of that gene or target gene. siRNA is typically about 5 to 100 nucleotides long, more typically about 10 to 50 nucleotides long, more typically about 15 to 30 nucleotides long, most typically about 20 to 30 nucleotides long, or about 20 to 25 or 24 to 29 nucleotides long, for example, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. siRNA molecules and methods for their preparation are described, for example, in Bass, 2001, Nature, 411, 428-429; Elbashir et al., 2001, Nature, 411, 494-498; International Publication No. 00 / 44895; International Publication No. 01 / 36646; International Publication No. 99 / 32619; International Publication No. 00 / 01846; International Publication No. 01 / 29058; International Publication No. 99 / 07409; and International Publication No. 00 / 44914. DNA molecules that transcribe dsRNA or siRNA (e.g., as hairpin double helix) also provide RNAi. DNA molecules for transcribing dsRNA are described in U.S. Patent No. 6,573,099, U.S. Patent Publication Nos. 2002 / 0160393 and 2003 / 0027783, and Tuschl and Borkhardt, Molecular Interventions, 2:158 (2002).

[0171] Of the double-stranded RNA of an siRNA, the strand that is at least partially complementary to at least a portion of a specific target nucleic acid (e.g., a target nucleic acid sequence), such as an mRNA molecule (e.g., a target mRNA molecule), is called the antisense (or guide strand); and the other strand is called the sense (or passenger strand). The passenger strand is degraded, and the guide strand is incorporated into the RNA-induced silencing complex (RISC).

[0172] Small hairpin RNA (shRNA / hairpin vector) is an artificial RNA molecule with a tight hairpin turn that can be used to silence target gene expression via RNA interference (RNAi).

[0173] Antisense oligonucleotides (ASOs) are single strands of DNA or RNA complementary to a sequence of choice. In the case of antisense RNA, they interfere with the protein translation of certain messenger RNA strands by binding to them in a process called hybridization. Antisense oligonucleotides can be used to target specific complementary (coding or non-coding) RNAs. Once binding occurs, this hybrid can be degraded by the enzyme RNase H.

[0174] Anti-miRNA oligonucleotides (also known as AMOs) refer to synthetically designed molecules (e.g., oligonucleotides) used to neutralize the function of cellular microRNAs (miRNAs) for a desired response.

[0175] The term “miRNA” is used in its simple, common sense to refer to a small non-coding RNA molecule that has the ability to regulate gene expression after transcription. In one embodiment, a miRNA is a nucleic acid that has substantial or complete identity with a target gene. In an embodiment, a miRNA inhibits gene expression by interacting with complementary cellular mRNA, thereby interfering with the expression of that complementary mRNA. Typically, a miRNA is at least about 15–50 nucleotides long (e.g., each complementary sequence of miRNA is 15–50 nucleotides long, and miRNA is about 15–50 base pairs long). In other embodiments, the length is 20–30 base nucleotides, preferably about 20–25 or about 24–29 nucleotides long, for example, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long.

[0176] "Nucleic acid" refers to a deoxyribonucleotide or ribonucleotide and its polymers or complements in single-stranded, double-stranded, or multi-stranded form. The term "polynucleotide" or "oligonuceltodie" refers to a linear sequence of nucleotides. The term "nucleotide" typically refers to a single unit, i.e., monomer, of a polynucleotide. A nucleotide may be a ribonucleotide, a deoxyribonucleotide, or a modified version thereof. Examples of polynucleotides as intended herein include single-stranded and double-stranded DNA, single-stranded and double-stranded RNA (including siRNA), and hybrid molecules having mixtures of single-stranded and double-stranded DNA and RNA. Nucleic acids may be linear or branched. For example, a nucleic acid may be a linear nucleotide chain, or it may be branched, for example, such that the nucleic acid contains one or more arms or branches of nucleotides. Optionally, branched nucleic acids repeatedly branch to form higher-order structures such as dendrimers.

[0177] These terms also encompass known nucleotide analogs or nucleic acids containing modified skeleton residues or bonds, which are synthetic, naturally occurring, and non-naturally occurring, possessing similar binding properties to the reference nucleic acid and being metabolized in a similar manner to the reference nucleotide. Examples of such analogs include, without limitation, phosphoramidates, phosphorodiamidates, phosphorothioates (also known as phosphothioates), phosphorodithioates, phosphonocarboxylic acids, phosphonocarboxylates, phosphonoacetic acids, phosphonoformic acids, methylphosphonates, boronphosphonates, or phosphate diester derivatives containing O-methylphosphoamidite bonds (see Eckstein, Oligonucleotides and Analogues: A Practical Approach, Oxford University Press), as well as peptide nucleic acid skeletons and bonds. Other nucleic acid analogs include those having a positive skeleton; a nonionic skeleton, modified sugars, and a non-ribose skeleton (e.g., phosphorodiamidate morpholino oligo or locked nucleic acid (LNA)), including those described in U.S. Patent Nos. 5,235,033 and 5,034,506, and Chapters 6 and 7, ASC Symposium Series 580, Carbohydrate Modifications in Antisense Research, Sanghui & Cook, eds. Nucleic acids containing one or more carbocyclic sugars also fall within one definition of nucleic acid. Modification of the ribose-phosphate skeleton may be performed for various reasons, for example, to increase the stability and half-life of such molecules in a physiological environment or as probes on biochips. Mixtures of naturally occurring nucleic acids and analogs may be made; or mixtures of different nucleic acid analogs, as well as mixtures of naturally occurring nucleic acids and analogs, may be made. In embodiments, the internucleotide bonds in DNA are phosphate diesters, phosphate diester derivatives, or a combination of both.

[0178] Unless otherwise specified, all genome or chromosome coordinates are based on hg38.

[0179] The gRNA molecules, compositions, and methods described herein relate to genome editing in eukaryotic cells using the CRISPR / Cas9 system. More specifically, the gRNA molecules, compositions, and methods described herein are useful for regulating globin levels, for example, in regulating the expression and production of globin genes and proteins. These gRNA molecules, compositions, and methods may be useful in the treatment of hemoglobin disorders.

[0180] I.gRNA molecule gRNA molecules may have several domains, as will be described in more detail below; however, gRNA molecules typically include at least a crRNA domain (including a targeting domain) and a tracr. The gRNA molecules of the present invention, used as a component of the CRISPR system, are useful for modifying (e.g., modifying sequences) DNA at or near a target site. Such modifications include, for example, deletions and / or insertions that result in reduced or absent expression of functional products of genes containing the target site. These uses and further uses will be described in more detail below.

[0181] In one embodiment, a single molecule, i.e., sgRNA, preferably comprising crRNA (containing a targeting domain complementary to the target sequence and a region forming part of the flagpole (i.e., the crRNA flagpole region))); a loop; and tracr (containing a domain complementary to the crRNA flagpole region and a domain for further binding of a nuclease or other effector molecule, e.g., a Cas molecule, e.g., a Cas9 molecule), may take the following format (from 5' to 3'): [Targeting domain]-[crRNA flagpole region]-[Optional first flagpole extension]-[Loop]-[Optional first tracr extension]-[tracr flagpole region]-[tracr nuclease binding domain].

[0182] In one embodiment, the tracr nuclease-binding domain binds to a Cas protein, such as the Cas9 protein.

[0183] In one embodiment, a bimodal molecule, i.e., dgRNA, comprises two polynucleotides; a first polynucleotide, preferably from 5' to 3', containing a crRNA (a targeting domain complementary to the target sequence and a region forming part of the flagpole); and a second polynucleotide, preferably from 5' to 3', containing a tracr (a domain complementary to the crRNA flagpole region and a domain for further binding of a nuclease or other effector molecule, e.g., a Cas molecule, e.g., a Cas9 molecule), and can take the following format (from 5' to 3'): Polynucleotide 1 (crRNA): [Targeting domain] - [crRNA flagpole region] - [Optional first flagpole extension region] - [Optional second flagpole extension region] Polynucleotide 2 (tracr): [Optional first tracr extension region] - [tracr flagpole region] - [tracr nuclease-binding domain].

[0184] In one embodiment, the tracr nuclease-binding domain binds to a Cas protein, such as the Cas9 protein.

[0185] In some embodiments, the targeting domain comprises or consists of a targeting domain sequence described herein, for example, the targeting domains listed in Tables 1 to 3, or 17, 18, 19, or 20 (preferably 20) consecutive nucleotides of the targeting domain sequence listed in Tables 1 to 3.

[0186] In some embodiments, the flagpole, for example, the crRNA flagpole region, contains GUUUUAGAGCUA (SEQ ID NO: 3110) from 5' to 3'.

[0187] In some embodiments, the flagpole, for example, the crRNA flagpole region, contains GUUUAAGAGCUA (SEQ ID NO: 3111) from 5' to 3'.

[0188] In some embodiments, the loop includes GAAA (sequence number 3114) from 5' to 3'.

[0189] In some embodiments, tracr is used for gRNA molecules that contain UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 3115) from 5' to 3', and preferably SEQ ID NO: 3110.

[0190] In some embodiments, tracr is used for gRNA molecules containing UAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 3116) from 5' to 3', preferably SEQ ID NO: 3111.

[0191] In some embodiments, gRNA may also contain additional U nucleic acids at its 3' end. For example, gRNA may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 additional U nucleic acids (SEQ ID NO: 3177) at its 3' end. In one embodiment, gRNA contains 4 additional U nucleic acids at its 3' end. In the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain additional U nucleic acids at its 3' end. For example, in the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 additional U nucleic acids (SEQ ID NO: 3177) at its 3' end. In one embodiment, in the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain 4 additional U nucleic acids at its 3' end. In embodiments of dgRNA, only the polynucleotide containing tracr contains additional U nucleic acids, for example, four U nucleic acids. In embodiments of dgRNA, only the polynucleotide containing the targeting domain contains additional U nucleic acids. In embodiments of dgRNA, both the polynucleotide containing the targeting domain and the polynucleotide containing tracr contain additional U nucleic acids, for example, four U nucleic acids.

[0192] In some embodiments, gRNA may also contain additional A nucleic acids at its 3' end. For example, gRNA may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 additional A nucleic acids (SEQ ID NO: 3178) at its 3' end. In one embodiment, gRNA contains 4 additional A nucleic acids at its 3' end. In the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain additional A nucleic acids at its 3' end. For example, in the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 additional A nucleic acids (SEQ ID NO: 3178) at its 3' end. In one embodiment, in the case of dgRNA, one or more of the polynucleotides of dgRNA (e.g., polynucleotides containing a targeting domain and polynucleotides containing tracr) may contain 4 additional A nucleic acids at its 3' end. In embodiments of dgRNA, only the polynucleotide containing tracr contains additional A nucleic acids, for example, four A nucleic acids. In embodiments of dgRNA, only the polynucleotide containing the targeting domain contains additional A nucleic acids. In embodiments of dgRNA, both the polynucleotide containing the targeting domain and the polynucleotide containing tracr contain additional U nucleic acids, for example, four A nucleic acids.

[0193] In this embodiment, one or more polynucleotides of the gRNA molecule may contain a cap at their 5' end.

[0194] In one embodiment, a single molecule, i.e., sgRNA, preferably from 5' to 3', includes crRNA (containing a targeting domain complementary to the target sequence; crRNA flagpole region; first flagpole extension; loop; first tracr extension (containing a domain complementary to at least a portion of the first flagpole extension); and tracr (containing a domain complementary to the crRNA flagpole region and a domain for further binding of a Cas9 molecule). In some embodiments, the targeting domain includes a targeting domain sequence described herein, for example, the targeting domains listed in Tables 1 to 3, or 17, 18, 19, or 20 (preferably 20) consecutive nucleotides of a targeting domain sequence listed in Tables 1 to 3, for example, 17, 18, 19, or 20 (preferably 20) consecutive nucleotides on the 3' side of a targeting domain sequence listed in Tables 1 to 3, or a targeting domain comprising such a sequence.

[0195] In embodiments including a first flagpole extension and / or a first tracr extension, the flagpole, loop, and tracr sequences may be as described above. Generally, any first flagpole extension and first tracr extension can be used, provided they are complementary. In embodiments, the first flagpole extension and the first tracr extension consist of 3, 4, 5, 6, 7, 8, 9, 10 or more complementary nucleotides.

[0196] In some embodiments, the first flagpole extension includes UGCUG (SEQ ID NO: 3112) from 5' to 3'. In some embodiments, the first flagpole extension consists of SEQ ID NO: 3112.

[0197] In some embodiments, the first tracr extension includes CAGCA (SEQ ID NO: 3117) from 5' to 3'. In some embodiments, the first tracr extension consists of SEQ ID NO: 3117.

[0198] In some embodiments, the dgRNA comprises two nucleic acid molecules. In some embodiments, the dgRNA comprises a first nucleic acid preferably comprising, from 5' to 3', a targeting domain complementary to the target sequence; a crRNA flagpole region; optionally a first flagpole extension; and optionally a second flagpole extension; and preferably, from 5' to 3', an optional first tracr extension; and a second nucleic acid (which may be referred to herein as tracr) comprising tracr (which comprises a domain complementary to the crRNA flagpole region and a domain for further binding of Cas, e.g., a Cas9 molecule), and at least a domain for binding of a Cas molecule, e.g., a Cas9 molecule). The second nucleic acid may also include an additional U nucleic acid at its 3' end (e.g., on the 3' side of tracr). For example, tracr may contain an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 U nucleic acids (SEQ ID NO: 3177) at its 3' end (e.g., on the 3' side of tracr). The second nucleic acid may, in addition or instead, contain an additional A nucleic acid at its 3' end (e.g., on the 3' side of tracr). For example, tracr may contain an additional 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 A nucleic acids (SEQ ID NO: 3178) at its 3' end (e.g., on the 3' side of tracr). In some embodiments, the targeting domain includes a targeting domain sequence described herein, e.g., the targeting domains listed in Tables 1 to 3, or a targeting domain comprising 17, 18, 19, or 20 (preferably 20) consecutive nucleotides of a targeting domain sequence listed in Tables 1 to 3.

[0199] In embodiments involving dgRNA, the crRNA flagpole region, an optional first flagpole extension, an optional first tracr extension, and the tracr sequence may be as described above.

[0200] In some embodiments, an optional second flagpole extension includes UUUUG (SEQ ID NO: 3113) from 5' to 3'.

[0201] In embodiments, the 1, 2, 3, 4, or 5 nucleotides on the 3' end of a gRNA molecule (and, in the case of a dgRNA molecule, the polynucleotide containing the targeting domain and / or the polynucleotide containing tracr), the 1, 2, 3, 4, or 5 nucleotides on the 5' end, or the 1, 2, 3, 4, or 5 nucleotides on both the 3' and 5' ends are modified nucleic acids, as will be explained in more detail in Section XIII below.

[0202] These domains are briefly discussed below: 1) Targeting domain: For guidance on selecting targeting domains, see, for example, Fu Y el al. NAT BIOTECHNOL 2014 (doi:10.1038 / nbt.2808) and Sternberg SH el al. NATURE 2014 (doi:10.1038 / naturel3011).

[0203] The targeting domain contains a nucleotide sequence that is complementary to the target sequence on the target nucleic acid, for example, at least 80, 85, 90, 95, or 99% complementary, or even perfectly complementary. The targeting domain is part of the RNA molecule and therefore will contain the base uracil (U), while any DNA encoding the gRNA molecule will contain the base thymine (T). Although we do not wish to be constrained by theory, the complementarity of the targeting domain with the target sequence is thought to contribute to the specificity of the interaction between the gRNA molecule / Cas9 molecule complex and the target nucleic acid. In the pairing of the targeting domain and the target sequence, it is understood that the uracil base in the targeting domain can pair with the adenine base in the target sequence.

[0204] In one embodiment, the targeting domain is 5 to 50, e.g., 10 to 40, e.g., 10 to 30, e.g., 15 to 30, e.g., 15 to 25 nucleotides long. In one embodiment, the targeting domain is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In one embodiment, the targeting domain is 16 nucleotides long. In one embodiment, the targeting domain is 17 nucleotides long. In one embodiment, the targeting domain is 18 nucleotides long. In one embodiment, the targeting domain is 19 nucleotides long. In one embodiment, the targeting domain is 20 nucleotides long. In one embodiment, the aforementioned 16, 17, 18, 19, or 20 nucleotides include 5'-16, 17, 18, 19, or 20 nucleotides from the targeting domains listed in Tables 1 to 3. In the embodiment, the aforementioned 16, 17, 18, 19, or 20 nucleotides include 3'-16, 17, 18, 19, or 20 nucleotides from the targeting domains listed in Tables 1 to 3.

[0205] Although not constrained by theory, 8, 9, 10, 11, or 12 nucleic acids located at the 3' end of the targeting domain are considered important for targeting the target sequence and can therefore be referred to as the “core” region of the targeting domain. In some embodiments, the core domain is perfectly complementary to the target sequence.

[0206] A chain of target nucleic acids to which the targeting domain is complementary is referred to herein as the target sequence. In some embodiments, the target sequence is located on a chromosome and is, for example, a target within a gene. In some embodiments, the target sequence is located within an exon of a gene. In some embodiments, the target sequence is located within an intron of a gene. In some embodiments, the target sequence contains or is adjacent to a binding site for a regulatory element of the gene of interest, such as a promoter or transcription factor binding site (for example, within 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or 1000 nucleic acids). Some or all of the nucleotides of the domain may have modifications, for example, modifications listed in Section XIII herein.

[0207] 2) crRNA flagpole region: The flagpole comprises portions from both crRNA and tracr. The crRNA flagpole region is complementary to a portion of tracr, and in some embodiments, it has sufficient complementarity with a portion of tracr to form a double-stranded region under at least some physiological conditions, e.g., under normal physiological conditions. In some embodiments, the crRNA flagpole region is 5 to 30 nucleotides long. In some embodiments, the crRNA flagpole region is 5 to 25 nucleotides long. The crRNA flagpole region may share homology with or derive from a naturally occurring portion of a repeat sequence from a bacterial CRISPR array. In some embodiments, this has at least 50% homology with the crRNA flagpole regions disclosed herein, e.g., the crRNA flagpole regions of Streptococcus pyogenes or S. thermophilus.

[0208] In one embodiment, the flagpole, for example, the crRNA flagpole region, includes sequence number 3110. In one embodiment, the flagpole, for example, the crRNA flagpole region, includes a sequence having at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homology with sequence number 3110. In one embodiment, the flagpole, for example, the crRNA flagpole region, includes at least 5, 6, 7, 8, 9, 10, or 11 nucleotides of sequence number 3110. In one embodiment, the flagpole, for example, the crRNA flagpole region, includes sequence number 3111. In one embodiment, the flagpole includes a sequence having at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, or 99% homology with sequence number 3111. In one embodiment, the flagpole, for example, the crRNA flagpole region, includes at least 5, 6, 7, 8, 9, 10, or 11 nucleotides of SEQ ID NO: 3111.

[0209] Some or all of the nucleotides of the domain may have modifications, such as those described in Section XIII of this specification.

[0210] 3) First flagpole extension section When a tracr containing a first tracr elongation region is used, the crRNA may contain a first flagpole elongation region. In general, any first flagpole elongation region and a first tracr elongation region can be used, provided they are complementary. In embodiments, the first flagpole elongation region and the first tracr elongation region consist of 3, 4, 5, 6, 7, 8, 9, 10 or more complementary nucleotides.

[0211] The first flagpole extension may contain nucleotides complementary to the nucleotides of the first tracr extension, for example, 80%, 85%, 90%, 95%, or 99%, or even perfectly complementary nucleotides. In some embodiments, the nucleotides of the first flagpole extension that hybridize with the complementary nucleotides of the first tracr extension are continuous. In some embodiments, the nucleotides of the first flagpole extension that hybridize with the complementary nucleotides of the first tracr extension are discontinuous and include, for example, two or more hybridization regions separated by nucleotides that do not base-pair with the nucleotides of the first tracr extension. In some embodiments, the first flagpole extension contains at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments, the first flagpole extension includes UGCUG (SEQ ID NO: 3112) from 5' to 3'. In some embodiments, the first flagpole extension consists of SEQ ID NO: 3112. In some embodiments, the first flagpole extension includes a nucleic acid that has at least 80%, 85%, 90%, 95%, or 99% homology to SEQ ID NO: 3112.

[0212] Some or all of the nucleotides in the first tracr extension may have modifications, such as those listed in Section XIII of this specification.

[0213] 3) Loop The loop serves to link the crRNA flagpole region of the sgRNA (or optionally the first flagpole extension, if present) to the tracr (or optionally the first tracr extension, if present). The loop may covalently or noncovalently link the crRNA flagpole region and the tracr. In some embodiments, this linkage is covalent. In some embodiments, the loop covalently couples the crRNA flagpole region and the tracr. In some embodiments, the loop covalently couples the first flagpole extension to the first tracr extension. In some embodiments, the loop is or includes a covalent bond interposed between the crRNA flagpole region and the domain of the tracr that hybridizes to the crRNA flagpole region. Typically, the loop contains one or more nucleotides, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.

[0214] In dgRNA molecules, the two molecules can associate through hybridization between at least a portion of crRNA (e.g., the crRNA flagpole region) and at least a portion of tracr (e.g., a tracr domain complementary to the crRNA flagpole region).

[0215] A variety of loops are suitable for use in sgRNA. The loops may be covalent or short, such as one or a few nucleotides long, for example, 1, 2, 3, 4, or 5 nucleotides long. In some embodiments, the loops are 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, or 25 nucleotides long or longer. In some embodiments, the loops are 2-50, 2-40, 2-30, 2-20, 2-10, or 2-5 nucleotides long. In some embodiments, the loops share homology with or derive homology from naturally occurring sequences. In some embodiments, the loops have at least 50% homology with the loops disclosed herein. In some embodiments, the loops include SEQ ID NO: 3114.

[0216] Some or all of the nucleotides of the domain may have modifications, such as those described in Section XIII of this specification.

[0217] 4) Second flagpole extension section In some embodiments, the dgRNA may include an additional sequence, referred to herein as a second flagpole extension, at the 3' end of the crRNA flagpole region, or, if present, the first flagpole extension. In some embodiments, the second flagpole extension is 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, or 2-4 nucleotides long. In some embodiments, the second flagpole extension is 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides long or longer. In some embodiments, the second flagpole extension includes sequence number 3113.

[0218] 5) Tracr: tracr is a nucleic acid sequence required for the binding of nucleases, such as Cas9. Although not theoretically constrained, it is thought that each Cas9 species is associated with a specific tracr sequence. Tracr sequences are utilized in both sgRNA and dgRNA systems. In some embodiments, tracr includes or is derived from sequences from *Streptococcus pyogenes* tracr. In some embodiments, tracr has a portion that hybridizes to the flagpole portion of crRNA, for example, a portion that is complementary to the crRNA flagpole region to form a double-stranded region under at least some physiological conditions (sometimes referred to herein as the tracr flagpole region or the tracr domain complementary to the crRNA flagpole region). In some embodiments, the tracr domain that hybridizes with the crRNA flagpole region includes at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides that hybridize with the complementary nucleotide of the crRNA flagpole region. In some embodiments, the tracr nucleotides that hybridize with the complementary nucleotide of the crRNA flagpole region are contiguous. In some embodiments, the tracr nucleotides that hybridize with the complementary nucleotide of the crRNA flagpole region are discontinuous and include, for example, two or more hybridization regions separated by nucleotides that do not base-pair with the nucleotides of the crRNA flagpole region. In some embodiments, a portion of the tracr that hybridizes with the crRNA flagpole region includes UAGCAAGUUAAAA (SEQ ID NO: 3119) from 5' to 3'. In some embodiments, a portion of the tracr that hybridizes to the crRNA flagpole region includes UAGCAAGUUUAAA (SEQ ID NO: 3120) from 5' to 3'. In embodiments, the sequence that hybridizes to the crRNA flagpole region is located on the 5' side of the tracr, which further binds a nuclease, such as a Cas molecule, such as a Cas9 molecule.

[0219] tracr further includes a domain that further binds to a nuclease, such as a Cas molecule, such as a Cas9 molecule. Although not constrained by theory, it is thought that different species of Cas9 bind to different tracr sequences. In some embodiments, tracr includes a sequence that binds to a Streptococcus pyogenes (S. pyogenes) Cas9 molecule. In some embodiments, tracr includes a sequence that binds to a Cas9 molecule disclosed herein. In some embodiments, the domain that further binds the Cas9 molecule includes UAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 3121) from 5' to 3'. In some embodiments, the domain that further binds the Cas9 molecule includes UAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 3122) from 5' to 3'.

[0220] In some embodiments, tracr includes sequence number 3115. In some embodiments, tracr includes sequence number 3116.

[0221] Some or all of the nucleotides in the tracr may have modifications, for example, the modifications described in Section XIII of this specification. In embodiments, a gRNA (e.g., the tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, contains an inverted debase residue at the 5' end, the 3' end, or both the 5' and 3' ends of the gRNA. In embodiments, a gRNA (e.g., the tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, contains one or more phosphorothioate bonds between residues at the 5' end of the polynucleotide, for example, a phosphorothioate bond between the first two 5' residues, between each of the first three 5' residues, between each of the first four 5' residues, or between each of the first five 5' residues. In embodiments, the gRNA or gRNA component may, in lieu of or in addition to, include one or more phosphorothioate bonds between the 3' terminal residues of the polynucleotide, for example, between the first two 3' residues, between each of the first three 3' residues, between each of the first four 3' residues, or between each of the first five 3' residues. In one embodiment, the gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNA or gRNA components described above, includes phosphorothioate bonds between each of the first four 5' residues (e.g., including three phosphorothioate bonds at the 5' end, e.g., consisting of such) and includes phosphorothioate bonds between each of the first four 3' residues (e.g., including three phosphorothioate bonds at the 3' end, e.g., consisting of such). In one embodiment, any of the phosphorothioate modifications described above is combined with an inverted debasing residue at the 5' end, 3' end, or both the 5' and 3' ends of the polynucleotide. In such embodiments, the inverted debasing nucleotide may be linked to the 5' and / or 3' nucleotide by a phosphate bond or a phosphorothioate bond.In an embodiment, the gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, contains one or more nucleotides with 2'O-methyl modifications. In an embodiment, each of the first one, two, three or more 5' residues contains a 2'O-methyl modification. In an embodiment, each of the first one, two, three or more 3' residues contains a 2'O-methyl modification. In an embodiment, the fourth, third, and second 3' residues from the terminal contain a 2'O-methyl modification. In an embodiment, each of the first one, two, three or more 5' residues contains a 2'O-methyl modification, and each of the first one, two, three or more 3' residues contains a 2'O-methyl modification. In one embodiment, each of the first three 5' residues contains a 2'O-methyl modification, and each of the first three 3' residues contains a 2'O-methyl modification. In the embodiment, each of the first three 5' residues contains a 2'O-methyl modification, and the fourth, third, and second 3' residues from the terminal contain a 2'O-methyl modification. In the embodiment, any of the 2'O-methyl modifications, for example, as described above, may be combined with one or more phosphorothioate modifications, for example, as described above, and / or one or more inverse debasing modifications, for example, as described above. In one embodiment, a gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, comprises, for example, a phosphorothioate bond between each of the first four 5' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a phosphorothioate bond between each of the first four 3' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a 2'O-methyl modification on each of the first three 5' residues, and a 2'O-methyl modification on each of the first three 3' residues.In one embodiment, a gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNA or gRNA components described above, comprises, for example, a phosphorothioate bond between each of the first four 5' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a phosphorothioate bond between each of the first four 3' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a 2'O-methyl modification on each of the first three 5' residues, and a 2'O-methyl modification on each of the fourth, third, and second 3' residues from the terminal, for example, comprising the same.

[0222] In one embodiment, a gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, comprises, for example, a phosphorothioate bond between each of the first four 5' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a phosphorothioate bond between each of the first four 3' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide), a 2'O-methyl modification on each of the first three 5' residues, a 2'O-methyl modification on each of the first three 3' residues, and additional inverted debase residues at each of the 5' and 3' ends.

[0223] In one embodiment, a gRNA (e.g., tracr and / or crRNA of sgRNA or dgRNA), for example, any of the gRNAs or gRNA components described above, comprises, for example, a phosphorothioate bond between each of the first four 5' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide, e.g., consisting thereof), a phosphorothioate bond between each of the first four 3' residues (e.g., comprising three phosphorothioate bonds at the 5' end of a polynucleotide, e.g., consisting thereof), a 2'O-methyl modification on each of the first three 5' residues, and a 2'O-methyl modification on each of the fourth, third, and second 3' residues from the terminal, and an additional inverted debase residue on each of the 5' and 3' ends, e.g., consisting thereof.

[0224] In one embodiment, the gRNA is a dgRNA and includes, for example, the following: crRNA: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAU*mG*mC*mU(Sequence ID 3179) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends); and tracr: AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU(Sequence ID 3152) (Optionally, it has an inverted debase residue at the 5' and / or 3' end.)

[0225] In one embodiment, the gRNA is a dgRNA and includes, for example, the following: crRNA: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAU*mG*mC*mU(Sequence ID 3179) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends); and tracr: mA*mA*mC*AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU*mU*mU*mU(Sequence ID 3174) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends).

[0226] In one embodiment, the gRNA is a dgRNA and includes, for example, the following: crRNA: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAUGCUGUU*mU*mU*mG(Sequence ID 3180) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends); and tracr: AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU(Sequence ID 3152) (Optionally, it has an inverted debase residue at the 5' and / or 3' end.)

[0227] In one embodiment, the gRNA is a dgRNA and includes, for example, the following: crRNA: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAUGCUGUU*mU*mU*mG(Sequence ID 3180) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends); and tracr: mA*mA*mC*AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU*mU*mU*mU(Sequence ID 3174) (In the formula, m represents a base with a 2'O-methyl modification, and * represents a phosphorothioate bond) (Optionally, it has an inverted debase residue at the 5' and / or 3' end).

[0228] In one embodiment, the gRNA is a dgRNA and includes, for example, the following: crRNA: NNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAUGCUGUUUUG(Sequence ID 3181) (wherein N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends); and tracr: mA*mA*mC*AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU*mU*mU*mU(Sequence ID 3174) (In the formula, m represents a base with a 2'O-methyl modification, and * represents a phosphorothioate bond) (Optionally, it has an inverted debase residue at the 5' and / or 3' end).

[0229] In one embodiment, the gRNA is an sgRNA and includes, for example, the following: NNNNNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU(Sequence ID 3182) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends).

[0230] In one embodiment, the gRNA is an sgRNA and includes, for example, the following: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCU*mU*mU*mU(Sequence ID 3183) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends).

[0231] In one embodiment, the gRNA is an sgRNA and includes, for example, the following: mN*mN*mN*NNNNNNNNNNNNNNNNNGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCmU*mU*mU*U(Sequence ID 3184) (wherein m represents a base having a 2'O-methyl modification, * represents a phosphorothioate bond, and N represents a residue of the targeting domain as described herein, for example) (optionally having inverted debase residues at the 5' and / or 3' ends).

[0232] 6) First Tracr extension If the gRNA contains a first flagpole elongation region, the tracr may contain a first tracr elongation region. The first tracr elongation region may contain nucleotides complementary to the nucleotides of the first flagpole elongation region, for example, 80%, 85%, 90%, 95%, or 99%, or even perfectly complementary nucleotides. In some embodiments, the first tracr elongation nucleotides that hybridize with the complementary nucleotides of the first flagpole elongation region are continuous. In some embodiments, the first tracr elongation nucleotides that hybridize with the complementary nucleotides of the first flagpole elongation region are discontinuous and include, for example, two or more hybridization regions separated by nucleotides that do not base-pair with the nucleotides of the first flagpole elongation region. In some embodiments, the first tracr extension comprises at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides. In some embodiments, the first tracr extension comprises SEQ ID NO: 3117. In some embodiments, the first tracr extension comprises a nucleic acid that has at least 80%, 85%, 90%, 95%, or 99% homology to SEQ ID NO: 3117.

[0233] Some or all of the nucleotides in the first tracr extension may have modifications, such as those listed in Section XIII of this specification.

[0234] In some embodiments, the sgRNA may include the following located 5' to 3', positioned on the 3' side of the targeting domain: a) GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3123); b) GUUUAAGAGCUAGAAAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3124); c) GUUUUAGAGCUAUGCUGGAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3125); d) GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3126); e) any of a) to d) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 uracil(U) nucleotides at the 3' end, for example, 1, 2, 3, 4, 5, 6, or 7 uracil(U) nucleotides; f) any of the above a) to d) further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 3' end, for example, 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides; or g) Any of a) to f) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 5' end (e.g., the 5' end, e.g., the 5' side of the targeting domain), e.g., 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides. In embodiments, any of a) to g) above is located immediately to the 3' side of the targeting domain.

[0235] In one embodiment, the sgRNA of the present invention has a [targeting domain] from 5' to 3'. GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU(Sequence ID 3159) For example, it includes, or consists of.

[0236] In one embodiment, the sgRNA of the present invention has a [targeting domain] from 5' to 3'. GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU(Sequence ID 3155) For example, it includes, or consists of.

[0237] In some mechanisms, dgRNA may include the following: crRNA containing the following, preferably located immediately to the 3' side of the targeting domain, from 5' to 3': a) GUUUUAGAGCUA (Sequence ID 3110); b) GUUUAAGAGCUA (Sequence ID 3111); c)GUUUUAGAGCUAUGCUG(Sequence ID 3127); d)GUUUAAGAGCUAUGCUG(Sequence ID 3128); e)GUUUUAGAGCUAUGCUGUUUUG(Sequence ID 3129); f)GUUUAAGAGCUAUGCUGUUUUG(Sequence ID 3130); or g)GUUUUAGAGCUAUGCU(Sequence ID 3154): And the tracr from 5' to 3' includes the following: a) UAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3115); b) UAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3116); c) CAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (Sequence ID 3131); d) CAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC(Sequence ID 3200); e) AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU(Sequence ID 3152); f) AACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUUU(Sequence ID 3153); g) AACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (Sequence ID 3160) h) GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU(Sequence ID 3155); i) AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (Sequence ID 3156); j) GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU(Sequence ID 3157); k) Any of a) to j) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 uracil(U) nucleotides at the 3' end, for example, 1, 2, 3, 4, 5, 6, or 7 uracil(U) nucleotides; l) any of the above a) to j) further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 3' end, for example, 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides; or m) Any of a) to l) above, further comprising at least 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides at the 5' end (e.g., at the 5' side end), e.g., 1, 2, 3, 4, 5, 6, or 7 adenine(A) nucleotides.

[0238] In one embodiment, the sequence of k) above includes the 3' sequence UUUUUU, for example, when the U6 promoter is used for transcription. In one embodiment, the sequence of k) above includes the 3' sequence UUUU, for example, when the HI promoter is used for transcription. In one embodiment, the sequence of k) above includes a number of 3'U that may vary depending on the termination signal of the pol-III promoter used, for example. In one embodiment, the sequence of k) above includes a variable 3' sequence derived from the DNA template, for example, when the T7 promoter is used. In one embodiment, the sequence of k) above includes a variable 3' sequence derived from the DNA template, for example, when the RNA molecule is created using in vitro transcription. In one embodiment, the sequence of k) above includes a variable 3' sequence derived from the DNA template, for example, when transcription is driven using the pol-II promoter.

[0239] In one embodiment, the crRNA comprises, for example, a sequence including a targeting domain and sequence number 3129 located 3' to the targeting domain (for example, located immediately 3' to the targeting domain), and a tracr including, for example, sequence number 3152.

[0240] In one embodiment, the crRNA comprises, for example, a sequence including a targeting domain and sequence number 3130 located 3' to the targeting domain (for example, located immediately 3' to the targeting domain), and a tracr including, for example, sequence number 3153 (sequence number 3153).

[0241] In one embodiment, the crRNA comprises, for example, a sequence including a targeting domain and GUUUUAGAGCUAUGCU (SEQ ID NO: 3154), located 3' to the targeting domain (for example, located immediately 3' to the targeting domain), and a tracr including, for example, GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUU (SEQ ID NO: 3155).

[0242] In one embodiment, the crRNA comprises, for example, a sequence including a targeting domain and GUUUUAGAGCUAUGCU (SEQ ID NO: 3154), located 3' to the targeting domain (for example, located immediately 3' to the targeting domain), and a tracr including, for example, AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 3156).

[0243] In one embodiment, the crRNA comprises, for example, a sequence including a targeting domain and GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 3129), which is located 3' to the targeting domain (for example, immediately 3' to the targeting domain), and a tracr including, for example, GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUU (SEQ ID NO: 3157).

[0244] II. gRNA targeting domain for the WIZ gene Tables 1 to 3 (at the end of this document) provide targeting domains for the WIZ gene region for the gRNA molecule of the present invention and for use in various embodiments of the present invention, for example, in the modification of the expression of globin genes, such as fetal hemoglobin genes or hemoglobin β genes.

[0245] III. Methods for designing gRNA Methods for designing gRNAs, including methods for selecting, designing, and validating target sequences, are described herein. Exemplary targeting domains are also provided herein. The targeting domains considered herein can be incorporated into the gRNAs described herein.

[0246] For information on target sequence selection and validation methods, as well as off-target analysis, please refer to, for example, Mali el al., 2013 SCIENCE 339(6121):823-826; Hsu et al, 2013 NAT BIOTECHNOL, 31(9):827-32; Fu et al, 2014 NAT BIOTECHNOL, doi:10.1038 / nbt.2808. PubMed PM ID:24463574; Heigwer et al, 2014 NAT METHODS ll(2):122-3. doi:10.1038 / nmeth.2812. PubMed PMID:24481216; Bae el al, 2014 BIOINFORMATICS PubMed PMID:24463181; Xiao A el al, 2014 BIOINFORMATICS It is listed in PubMed PMID:24389662.

[0247] For example, a software tool can be used to optimize gRNA selection within the user's target sequence range, minimizing overall off-target activity across the genome. Off-target activity may be other than cleavage. For each possible gRNA selection, the tool can identify all off-target sequences across the genome (e.g., those preceding either NAG PAM or NGG PAM) containing mismatched base pairs up to a specific number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) using, for example, Streptococcus pyogenes Cas9. The cleavage efficiency at each off-target sequence can be predicted, for example, using an experimentally derived weighting scheme. Each possible gRNA is then ranked according to its overall predicted off-target cleavage. Higher-ranked gRNAs correspond to those most likely to have maximum on-target cleavage and minimum off-target cleavage. Other functions, such as automated reagent design for CRISPR construction, primer design for on-target Surveyor assays, and primer design for high-throughput detection and quantification of off-target cleavage by next-generation sequencing, may also be included in the tool. Candidate gRNA molecules can be determined by methods known in the art or as described herein.

[0248] While software algorithms can be used to generate an initial list of potential gRNA molecules, cleavage efficiency and specificity do not necessarily reflect the predicted values. Therefore, gRNA molecules typically need to be screened in specific cell lines, such as primary human cell lines (e.g., human HSPCs, e.g., human CD34+ cells), to determine, for example, cleavage efficiency, indel formation, cleavage specificity, and desired phenotypic changes. These properties can be assayed by the methods described herein.

[0249] IV.Cas molecule Cas9 molecule In preferred embodiments, the Cas molecule is a Cas9 molecule. Various species of Cas9 molecules can be used in the methods and compositions described herein. While the *Streptococcus pyogenes* Cas9 molecule is the subject of much of the disclosure herein, Cas9 molecules of other species of Cas9 proteins listed herein, Cas9 molecules derived therefrom, or Cas9 molecules based thereon can also be used. In other words, other Cas9 molecules, such as *S. thermophilus*, *Staphylococcus aureus*, and / or *Neisseria meningitidis* Cas9 molecules, can be used in the systems, methods, and compositions described herein. Further Cas9 species include Acidovorax avenae, Actinobacillus pleuropneumoniae, Actinobacillus succinogenes, Actinobacillus suis, Actinomyces sp., Alicycliphilus denitrificans, Aminomonas paucivorans, Bacillus cereus, Bacillus smithii, Bacillus thuringiensis, and Bacteroides species. (sp.), Blastopirellula marina, Bradyrhizobium sp.), Brevibacillus laterosporus, Campylobacter coli, Campylobacter jejuni, Campylobacter lad, Candidatus Puniceispirillum, Clostridium cellulolyticum, Clostridium perfringens, Corynebacterium accolens, Corynebacterium diphtheria, Corynebacterium matruchotii, Dinoroseobacter shibae), Eubacterium dolichum, gamma proteobacterium, Gluconacetobacter diazotrophicus, Haemophilus parainfluenzae, Haemophilus sputorum, Helicobacter canadensis, Helicobacter cinaedi, Helicobacter mustelae, Ilyobacter polytropus, Kingella kingae, Lactobacillus crispatus, Listeria ivanovii, Listeria monocytogenes (Listeria (monocytogenes), Listeriaceae bacteria, Methylocystis sp.), Methylosinus trichosporium, Mobiluncus mulieris, Neisseria bacilliformis, Neisseria cinerea, Neisseria flavescens, Neisseria lactamica, Neisseria sp., Neisseria wadsworthii, Nitrosomonas sp., Parvibaculum lavamentivorans, Pasteurella multocida, Phascolarctobacterium succinate succinatutens), Ralstonia syzygii, Rhodopseudomonas palustris, Rhodovulum sp., Simonsiella muelleri, Sphingomonas sp., Sporolactobacillus vineae, Staphylococcus lugdunensis, Streptococcus sp., Subdoligranulum sp., Tistrella mobilis, Treponema sp., or Verminephrobacter eisenie eiseniae is one example.

[0250] When the term Cas9 is used herein, it refers to a molecule that interacts with a gRNA molecule (e.g., the domain sequence of tracr) and, in cooperation with the gRNA molecule, can localize to (e.g., target or hom to) a site containing a target sequence and a PAM sequence.

[0251] In some embodiments, the Cas9 molecule has the ability to cleave target nucleic acid molecules, which may be referred to herein as an active Cas9 molecule. In some embodiments, the active Cas9 molecule comprises one or more of the following activities: nickase activity, i.e., the ability to cleave a single strand of a nucleic acid molecule, e.g., a non-complementary or complementary strand; double-strand nuclease activity, i.e., the ability to cleave both strands of a double-stranded nucleic acid to create a double-strand break, which in some embodiments is the presence of two nickase activities; endonuclease activity; exonuclease activity; and helicase activity, i.e., the ability to unwind the helical structure of a double-stranded nucleic acid.

[0252] In one embodiment, an enzymatically active Cas9 molecule cleaves both DNA strands, resulting in a double-strand break. In another embodiment, the Cas9 molecule cleaves only one strand, for example, the strand to which the gRNA hybridizes, or the strand complementary to the strand to which the gRNA hybridizes. In another embodiment, the active Cas9 molecule includes cleavage activity associated with an HNH-like domain. In yet another embodiment, the active Cas9 molecule includes cleavage activity associated with an N-terminal RuvC-like domain. In yet another embodiment, the active Cas9 molecule includes cleavage activity associated with an HNH-like domain and cleavage activity associated with an N-terminal RuvC-like domain. In yet another embodiment, the active Cas9 molecule includes an active or cleavage-competent HNH-like domain and an inactive or cleavage-incompetent N-terminal RuvC-like domain. In yet another embodiment, the active Cas9 molecule includes an inactive or cleavage-incompetent HNH-like domain and an active or cleavage-competent N-terminal RuvC-like domain.

[0253] In some embodiments, the ability of an active Cas9 molecule to interact with and cleave a target nucleic acid is PAM sequence-dependent. The PAM sequence is a sequence in the target nucleic acid. In some embodiments, cleavage of the target nucleic acid occurs upstream of the PAM sequence. Active Cas9 molecules of different bacterial species can recognize different sequence motifs (e.g., PAM sequences). In some embodiments, the active Cas9 molecule of Streptococcus pyogenes recognizes the sequence motif NGG and leads to cleavage of the target nucleic acid sequence 1 to 10, e.g., 3 to 5 base pairs upstream from that sequence. See, for example, Mali el ai, SCIENCE 2013;339(6121):823-826. In some embodiments, the active Cas9 molecule of S. thermophilus recognizes the sequence motifs NGGNG and NNAGAAW (W=A or T) and leads to cleavage of the core target nucleic acid sequence 1 to 10, e.g., 3 to 5 base pairs upstream from these sequences. See, for example, Horvath et al., SCIENCE 2010;327(5962):167-170 and Deveau et al, J BACTERIOL 2008;190(4):1390-1400. In one embodiment, the active Cas9 molecule of S. mutans recognizes a sequence motif NGG or NAAR (RA or G), leading to the cleavage of a core target nucleic acid sequence 1 to 10, e.g., 3 to 5 base pairs upstream from this sequence. See, for example, Deveau et al., J BACTERIOL 2008;190(4):1390-1400.

[0254] In one embodiment, the active Cas9 molecule of Staphylococcus aureus (S. aureus) recognizes the sequence motif NNGRR (R=A or G), leading to the cleavage of a target nucleic acid sequence 1 to 10 base pairs upstream, e.g., 3 to 5 base pairs. See, for example, Ran F. et al., NATURE, vol. 520, 2015, pp. 186-191. In another embodiment, the active Cas9 molecule of Neisseria meningitidis (N. meningitidis) recognizes the sequence motif NNNNGATT, leading to the cleavage of a target nucleic acid sequence 1 to 10 base pairs upstream, e.g., 3 to 5 base pairs. See, for example, Hou et al., PNAS EARLY EDITION 2013, 1-6. The ability of the Cas9 molecule to recognize PAM sequences can be determined using, for example, the transformation assay described in Jinek et al., SCIENCE 2012, 337:816.

[0255] Some Cas9 molecules have the ability to interact with gRNA molecules and, together with the gRNA molecule, home to (e.g., be targeted or localized to) the core target domain, but they do not have the ability to cleave the target nucleic acid, or to cleave it in an efficient ratio. Cas9 molecules that have no cleavage activity or substantially no cleavage activity may be referred to herein as inactive Cas9 (enzymatically inactive Cas9), dead Cas9, or dCas9 molecules. For example, inactive Cas9 molecules may lack cleavage activity or have substantially low cleavage activity, such as less than 20, 10, 5, 1, or 0.1% of a reference Cas9 molecule when measured by the assays described herein.

[0256] Exemplary naturally occurring Cas9 molecules are described in Chylinski et al, RNA Biology 2013;10:5,727-737. Such Cas9 molecules include the cluster 1 bacterial family, cluster 2 bacterial family, cluster 3 bacterial family, cluster 4 bacterial family, cluster 5 bacterial family, cluster 6 bacterial family, cluster 7 bacterial family, cluster 8 bacterial family, cluster 9 bacterial family, cluster 10 bacterial family, cluster 11 bacterial family, cluster 12 bacterial family, cluster 13 bacterial family, cluster 14 bacterial family, cluster 1 bacterial family, cluster 16 bacterial family, cluster 17 bacterial family, cluster 18 bacterial family, cluster 19 bacterial family, cluster 20 bacterial family, cluster 21 bacterial family, cluster 22 bacterial family, cluster 23 bacterial family, cluster 24 bacterial family, cluster 25 bacterial family, cluster 26 bacterial family, cluster 27 bacterial family, cluster 28 bacterial family, cluster 29 bacterial family, cluster 30 bacterial family, and cluster 3 1 bacterial family, cluster 32 bacterial family, cluster 33 bacterial family, cluster 34 bacterial family, cluster 35 bacterial family, cluster 36 bacterial family, cluster 37 bacterial family, cluster 38 bacterial family, cluster 39 bacterial family, cluster 40 bacterial family, cluster 41 bacterial family, cluster 42 bacterial family, cluster 43 bacterial family, cluster 44 bacterial family, cluster 45 bacterial family, cluster 46 bacterial family, cluster 47 bacterial family, cluster 48 bacterial family, cluster 49 bacterial family, cluster 50 bacterial family, cluster 51 bacterial family, cluster 52 bacterial family, cluster 53 bacterial family, cluster 54 bacterial family, cluster 55 bacterial family, cluster 56 bacterial family, cluster 57 bacterial family, cluster 58 bacterial family, cluster 59 bacterial family, cluster 60 bacterial family, cluster 61 bacterial family,This includes Cas9 molecules from the following bacterial families: Cluster 62, Cluster 63, Cluster 64, Cluster 65, Cluster 66, Cluster 67, Cluster 68, Cluster 69, Cluster 70, Cluster 71, Cluster 72, Cluster 73, Cluster 74, Cluster 75, Cluster 76, Cluster 77, or Cluster 78.

[0257] Examples of naturally occurring Cas9 molecules include those from the Cluster 1 bacterial family. Examples include Streptococcus pyogenes (e.g., strains SF370, MGAS 10270, MGAS 10750, MGAS2096, MGAS315, MGAS5005, MGAS6180, MGAS9429, NZ131, and SSI-1), S. thermophilus (e.g., strain LMD-9), S. pseudoporcinus (e.g., strain SPIN 20026), S. mutans (e.g., strains UA 159, NN2025), S. macacae (e.g., strain NCTC1 1558), and S. gallolyticus (e.g., strains UCN34, ATCC). S. equinus (e.g., strains ATCC 9812, MGCS 124), S. dysgalactiae (e.g., strain GGS 124), S. bovis (e.g., strain ATCC 700338), S. anginosus (e.g., strain F0211), S. agalactiae (e.g., strains NEM316, A909), Listeria monocytogenes (e.g., strain F6854), Listeria innocua (L. innocua, e.g., strain Clip l 262), Enterococcus italicus (e.g., strain DSM) Examples include the Cas9 molecules of Enterococcus faecium (e.g., strains 1,231,408) (15952). Further exemplary Cas9 molecules are those of Neisseria meningitidis (Hou et al. PNAS Early Edition 2013, 1-6) and Staphylococcus aureus (S. aureus).

[0258] In one embodiment, the Cas9 molecule, for example, an active Cas9 molecule or an inactive Cas9 molecule, is any Cas9 molecular sequence described herein or a naturally occurring Cas9 molecular sequence, for example, one listed herein or one of the following: Chylinski et al., RNA Biology 2013, 10:5, 'I2'I-T, 1, Hou et al. PNAS Early Edition It contains an amino acid sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology with the Cas9 molecules of the species described in 2013,1-6, and differs from it by 1%, 2%, 5%, 10%, 15%, 20%, 30%, or 40% or less of amino acid residues, or differs from it by only 1, 2, 5, 10, or 20 amino acids and up to 100, 80, 70, 60, 50, 40, or 30 amino acids, or is identical to it.

[0259] In one embodiment, the Cas9 molecule comprises an amino acid sequence having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology with Streptococcus pyogenes Cas9 (NCBI reference sequence: WP_010922251.1; SEQ ID NO: 3133); and the number of different amino acid residues when compared to it does not exceed 1%, 2%, 5%, 10%, 15%, 20%, 30%, or 40%; and differs from it by 1, 2, 5, 10, or 20 amino acids or 100, 80, 70, 60, 50, 40, or 30 amino acids or less; or is identical to it.

[0260] In one embodiment, the Cas9 molecule is a S. pyogenes Cas9 variant of SEQ ID NO: 3133, comprising one or more mutations introducing an uncharged or nonpolar amino acid, such as alanine, to a positively charged amino acid (e.g., lysine, arginine, or histidine) at the aforementioned position. In one embodiment, the mutation is a mutation to one or more positively charged amino acids in the nt groove of Cas9. In one embodiment, the Cas9 molecule is a S. pyogenes Cas9 variant of SEQ ID NO: 3133, comprising a mutation at position 855 of SEQ ID NO: 3133, for example, a mutation to an uncharged amino acid, such as alanine, at position 855 of SEQ ID NO: 3133. In one embodiment, the Cas9 molecule has a mutation to an uncharged amino acid, such as alanine, only at position 855 of SEQ ID NO: 3133, compared to SEQ ID NO: 3133. In embodiments, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133, comprising mutations at position 810, position 1003, and / or position 1060 of SEQ ID NO: 3133, for example, mutations to alanine at positions 810, 1003, and / or 1060 of SEQ ID NO: 3133. In embodiments, the Cas9 molecule has mutations only at positions 810, 1003, and 1060 of SEQ ID NO: 3133 compared to SEQ ID NO: 3133, for example, where each mutation is a mutation to an uncharged amino acid, for example, alanine. In embodiments, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133, comprising mutations at position 848, position 1003, and / or position 1060 of SEQ ID NO: 3133, for example, mutations to alanine at positions 848, 1003, and / or 1060 of SEQ ID NO: 3133. In embodiments, the Cas9 molecule has mutations only at positions 848, 1003, and 1060 of SEQ ID NO: 3133 compared to SEQ ID NO: 3133, for example, where each mutation is a mutation to an uncharged amino acid, for example, alanine.In this embodiment, the Cas9 molecule is the Cas9 molecule as described in Slaymaker et al., Science Express (available online as of December 1, 2015, at Science DOI: 10.1126 / science.aad5227).

[0261] In an embodiment, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133 containing one or more mutations. In an embodiment, the Cas9 variant contains a mutation at position 80 of SEQ ID NO: 3133, for example, containing leucine at position 80 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the C80L mutation). In an embodiment, the Cas9 variant contains a mutation at position 574 of SEQ ID NO: 3133, for example, containing glutamic acid at position 574 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the C574E mutation). In an embodiment, the Cas9 variant contains mutations at position 80 and position 574 of SEQ ID NO: 3133, for example, containing leucine at position 80 of SEQ ID NO: 3133 and glutamic acid at position 574 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the C80L mutation and the C574E mutation). Although not constrained by theory, such mutations are thought to improve the solubility of the Cas9 molecule.

[0262] In an embodiment, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133 containing one or more mutations. In an embodiment, the Cas9 variant contains a mutation at position 147 of SEQ ID NO: 3133, for example, containing tyrosine at position 147 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the D147Y mutation). In an embodiment, the Cas9 variant contains a mutation at position 411 of SEQ ID NO: 3133, for example, containing threonine at position 411 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the P411T mutation). In embodiments, Cas9 mutants include mutations at position 147 and position 411 of SEQ ID NO: 3133, for example, tyrosine at position 147 and threonine at position 411 (i.e., including, for example, SEQ ID NO: 3133 having the D147Y mutation and the P411T mutation). Although not theoretically constrained, such mutations are thought to improve the targeting efficiency of the Cas9 molecule, for example, in yeast.

[0263] In embodiments, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 mutant of SEQ ID NO: 3133 containing one or more mutations. In embodiments, the Cas9 mutant contains a mutation at position 1135 of SEQ ID NO: 3133, for example, containing glutamic acid at position 1135 of SEQ ID NO: 3133 (i.e., containing, for example, SEQ ID NO: 3133 having the D1135E mutation). Although not theoretically constrained, such mutations are thought to improve the selectivity of the Cas9 molecule for the NGG PAM sequence compared to the NAG PAM sequence.

[0264] In one embodiment, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133, which includes one or more mutations introducing an uncharged or nonpolar amino acid, such as alanine, at specific positions. In another embodiment, the Cas9 molecule is a Streptococcus pyogenes (S. pyogenes) Cas9 variant of SEQ ID NO: 3133, which includes mutations at positions 497, 661, 695, and / or 926 of SEQ ID NO: 3133, for example, mutations to alanine at positions 497, 661, 695, and / or 926 of SEQ ID NO: 3133. In another embodiment, the Cas9 molecule has mutations only at positions 497, 661, 695, and 926 of SEQ ID NO: 3133 compared to SEQ ID NO: 3133, for example, where each mutation is a mutation to an uncharged amino acid, such as alanine. Although not constrained by theory, such mutations are thought to reduce off-target cleavage by the Cas9 molecule.

[0265] It will be understood that the mutations to the Cas9 molecule described herein may be combined, and that the Cas9 molecule may be tested in the assay described herein in combination with any of the fusions or other modifications described herein.

[0266] Various types of Cas molecules can be used in carrying out the invention disclosed herein. In some embodiments, Cas molecules of type II Cas systems are used. In other embodiments, Cas molecules of other Cas systems are used. For example, type I or type III Cas molecules may be used. Exemplary Cas molecules (and Cas systems) are described in Haft et al., PLoS COMPUTATIONAL BIOLOGY 2005, 1(6):e60 and Makarova et al., NATURE REVIEW MICROBIOLOGY 201 1, 9:467-477 (the contents of both references are incorporated herein by reference in their entirety).

[0267] In one embodiment, the Cas9 molecule comprises one or more of the following activities: nickase activity; double-strand cleavage activity (e.g., endonuclease and / or exonuclease activity); helicase activity; or the ability to localize to a target nucleic acid in conjunction with a gRNA molecule.

[0268] Modified Cas9 molecule Naturally occurring Cas9 molecules possess several properties, including nickase activity, nuclease activity (e.g., endonuclease and / or exonuclease activity); helicase activity; the ability to functionally associate with gRNA molecules; and the ability to target (or localize to) sites on nucleic acids (e.g., PAM recognition and specificity). In some embodiments, the Cas9 molecule may possess all or some of these properties. In a typical embodiment, the Cas9 molecule interacts with gRNA molecules and has the ability to cooperate with gRNA molecules to localize to sites on nucleic acids. Other activities, such as PAM specificity, cleavage activity, or helicase activity, may vary more broadly in the Cas9 molecule.

[0269] Cas9 molecules possessing desired properties can be prepared in several ways, for example, by modifying a parent Cas9 molecule, such as a naturally occurring Cas9 molecule, to provide a modified Cas9 molecule with the desired properties. For example, one or more mutations or differences can be introduced compared to the parent Cas9 molecule. Such mutations and differences include substitutions (e.g., conservative substitutions or substitutions of non-essential amino acids); insertions; or deletions. In one embodiment, a Cas9 molecule may contain one or more mutations or differences compared to a reference Cas9 molecule, for example, 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, or 50 or more mutations, or fewer than 200, 100, or 80 mutations.

[0270] In some embodiments, one or more mutations have no substantial effect on Cas9 activity, e.g., the Cas9 activity described herein. In some embodiments, one or more mutations have a substantial effect on Cas9 activity, e.g., the Cas9 activity described herein. In some embodiments, exemplary activities include one or more of PAM specificity, cleavage activity, and helicase activity. One or more mutations may be located, for example, in one or more RuvC-like domains, e.g., an N-terminal RuvC-like domain; an HNH-like domain; or in the region outside the RuvC-like and HNH-like domains. In some embodiments, one or more mutations are located in the N-terminal RuvC-like domain. In some embodiments, one or more mutations are located in the HNH-like domain. In some embodiments, mutations are located in both the N-terminal RuvC-like domain and the HNH-like domain.

[0271] Whether a detailed sequence, such as a substitution, may affect one or more activities, such as targeting activity or cleavage activity, can be determined or predicted, for example, by determining whether the mutation is conserved, or by the methods described in Section III. In one embodiment, a “non-essential” amino acid residue is a residue that, when used in association with the Cas9 molecule, can be modified from the wild-type sequence of a Cas9 molecule, such as a naturally occurring Cas9 molecule, such as an active Cas9 molecule, without invalidating or, more preferably substantially altering, the Cas9 activity (e.g., cleavage activity), whereas a change in an “essential” amino acid residue results in a substantial loss of activity (e.g., cleavage activity).

[0272] Cas9 molecules with or without modified PAM recognition Naturally occurring Cas9 molecules can recognize specific PAM sequences, such as those found in Streptococcus pyogenes, S. thermophilus, S. mutans, Staphylococcus aureus, and Neisseria meningitidis, as described above.

[0273] In one embodiment, the Cas9 molecule has the same PAM specificity as the naturally occurring Cas9 molecule. In another embodiment, the Cas9 molecule has PAM specificity unrelated to the naturally occurring Cas9 molecule, or unrelated to the naturally occurring Cas9 molecule with which it has the closest sequence homology. For example, by modifying the naturally occurring Cas9 molecule, it is possible to modify, for example, PAM recognition, modify the PAM sequence recognized by the Cas9 molecule to reduce off-target sites and / or improve specificity, or eliminate the PAM recognition requirement. In one embodiment, by modifying the Cas9 molecule, it is possible to reduce off-target sites and increase specificity by, for example, increasing the length of the PAM recognition sequence and / or improving Cas9 specificity to a high level of identity. In one embodiment, the length of the PAM recognition sequence is at least 4, 5, 6, 7, 8, 9, 10, or 15 amino acids. Cas9 molecules that recognize different PAM sequences and / or have reduced off-target activity can be created using directional evolution. Exemplary methods and systems that can be used for directional evolution of Cas9 molecules are described, for example, by Esvelt el al, Nature 2011, 472(7344):499-503. Candidate Cas9 molecules can be evaluated, for example, by the methods described herein.

[0274] Non-cleaving and modified cleaving Cas9 molecules In some embodiments, the Cas9 molecule has cleavage properties that differ from naturally occurring Cas9 molecules, for example, from naturally occurring Cas9 molecules with the closest homology. For example, the Cas9 molecule may differ from naturally occurring Cas9 molecules, such as the Cas9 molecule of Streptococcus pyogenes, in the following ways: for example, its ability to modulate, for example, decrease or increase, double-strand break cleavage (endonuclease and / or exonuclease activity) compared to naturally occurring Cas9 molecules (for example, the Cas9 molecule of Streptococcus pyogenes); for example, its ability to modulate, for example, decrease or increase, single-strand break cleavage (nickase activity) of nucleic acids, such as the non-complementary or complementary strand of a nucleic acid molecule, compared to naturally occurring Cas9 molecules (for example, the Cas9 molecule of Streptococcus pyogenes); or it may have the ability to cleave nucleic acid molecules, such as double-stranded or single-stranded nucleic acid molecules.

[0275] Modified truncated active Cas9 molecule In one embodiment, the active Cas9 molecule includes one or more of the following activities: cleavage activity related to the N-terminal RuvC-like domain; cleavage activity related to the HNH-like domain; and cleavage activity related to the HNH domain and cleavage activity related to the N-terminal RuvC-like domain.

[0276] In one embodiment, the Cas9 molecule is a Cas9 nickase that, for example, cleaves only a single strand of DNA. In one embodiment, the Cas9 nickase includes a mutation at position 10 and / or at position 840 of SEQ ID NO: 3133, for example, the D10A and / or H840A mutations for SEQ ID NO: 3133.

[0277] Non-cleavable inactive Cas9 molecule In some embodiments, the modified Cas9 molecule is an inactive Cas9 molecule that does not cleave nucleic acid molecules (neither double-stranded nor single-stranded nucleic acid molecules) or cleaves nucleic acid molecules with significantly lower efficiency, for example, less than 20, 10, 5, 1, or 0.1% of the cleavage activity of a reference Cas9 molecule when measured by the assay described herein. The reference Cas9 molecule may be a naturally occurring unmodified Cas9 molecule, such as a Cas9 molecule from Streptococcus pyogenes, S. thermophilus, Staphylococcus aureus, or Neisseria meningitidis. In some embodiments, the reference Cas9 molecule is a naturally occurring Cas9 molecule with the closest sequence identity or homology. In some embodiments, the inactive Cas9 molecule lacks substantial cleavage activity associated with the N-terminal RuvC-like domain and cleavage activity associated with the HNH-like domain.

[0278] In one embodiment, the Cas9 molecule is dCas9. (Tsai et al. (2014), Nat. Biotech. 32:569-577.)

[0279] A catalytically inactive Cas9 molecule may be fused with a transcriptional repressor. The inactive Cas9 fusion protein complexes with gRNA and localizes to the DNA sequence specified by the gRNA's targeting domain, but unlike active Cas9, it does not cleave the target DNA. By fusing an effector domain, such as a transcriptional repressor domain, to the inactive Cas9, it becomes possible to recruit that effector to any DNA site specified by the gRNA. Site-specific targeting of the Cas9 fusion protein to the promoter region of a gene can block or affect polymerase binding to that promoter region, such as Cas9 fusion with transcription factors (e.g., transcription activators) and / or binding of transcriptional enhancers to nucleic acids, thereby increasing or inhibiting transcriptional activation. Alternatively, site-specific targeting of a Cas9 fusion with a transcriptional repressor to the promoter region of a gene can be used to reduce transcriptional activation.

[0280] Transcriptional repressors or domains that can be fused with an inactive Cas9 molecule may include Kruppel-associated boxes (KRAB or SKD), Mad mSIN3 interaction domains (SID), or ERF repressor domains (ERD).

[0281] In another embodiment, the inactive Cas9 molecule may be fused with a chromatin-modifying protein. For example, the inactive Cas9 molecule may be fused with heterochromatin protein 1 (HPl), histone lysine methyltransferases (e.g., SUV39H1, SUV39H2, G9A, ESET / SETDB l, Pr-SET7 / 8, SUV4-20H1, RIZ1), histone lysine demethylases (e.g., LSD1 / BHC110, SpLsdl / Sw,l / Safl 10, Su(var)3-3, JMJD2A / JHDM3A, JMJD2B, JMJD2C / GASC1, JMJD2D, Rphl, JARID 1A / RBP2, JARIDIB / PLU-I, JAR1D It may also be fused with 1C / SMCX, JARID1D / SMCY, Lid, Jhn2, Jmj2), histone lysine deacetylases (e.g., HDAC1, HDAC2, HDAC3, HDAC8, Rpd3, Hosl, Cir6, HDAC4, HDAC5, HDAC7, HDAC9, Hdal, Cir3, SIRT1, SIRT2, Sir2, Hstl, Hst2, Hst3, Hst4, HDAC11) and DNA methylases (DNMT1, DNMT2a / DMNT3b, MET1). Inactive Cas9-chromatin modification molecule fusion proteins can be used to modify the chromatin state and reduce the expression of target genes.

[0282] A heterologous sequence (e.g., a transcriptional repressor domain) may be fused to the N-terminus or C-terminus of an inactive Cas9 protein. In an alternative embodiment, the heterologous sequence (e.g., a transcriptional repressor domain) may be fused to the inner portion (i.e., a portion other than the N-terminus or C-terminus) of an inactive Cas9 protein.

[0283] The ability of a Cas9 molecule / gRNA molecule complex to bind to and cleave a target nucleic acid can be determined, for example, by the methods described in Section III of this specification. The activity of a Cas9 molecule, for example, active Cas9 or inactive Cas9, either alone or in complex with a gRNA molecule, can also be determined by methods well known in the art, including gene expression assays and chromatin-based assays, such as chromatin immunoprecipitation (ChiP) and chromatin in vivo assays (CiA).

[0284] Other Cas9 molecular fusions In embodiments, the Cas9 molecule, for example, Cas9 of Streptococcus pyogenes, may further contain one or more amino acid sequences that confer further activity.

[0285] In some embodiments, the Cas9 molecule may contain one or more nuclear localization sequences (NLSs), e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the Cas9 molecule contains at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the amino terminus, or at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs at or near the carboxyl terminus, or a combination thereof (e.g., one or more NLSs at the amino terminus and one or more NLSs at the carboxyl terminus). When two or more NLSs are present, each may be selected independently of the others, and thus a single NLS may be present in two or more copies, and / or in combination with one or more other NLSs present in one or more copies. In some embodiments, an NLS is considered to be near the N-terminus or C-terminus when the nearest amino acid of the NLS is within approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 amino acids, or more, from the N-terminus or C-terminus along the polypeptide chain. Typically, an NLS consists of one or more short sequences of positively charged lysine or arginine exposed on the protein surface, although other types of NLSs are known.Non-limiting examples of NLS include: NLS of the SV40 virus large T antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 3134); NLS of nucleoplasmin (e.g., nucleoplasmin vipertite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 3135); c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 3136) or RQRRNELKRSP (SEQ ID NO: 3137); and hRNPA1 M9 having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 3138). NLS; Importin-α IBB domain sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 3139); Myoma T protein sequences VSRKRPRP (SEQ ID NO: 3140) and PPKKARED (SEQ ID NO: 3141); Human p53 sequence PQPKKKPL (SEQ ID NO: 3142); Mouse c-ab1 Examples of NLS sequences include the sequence SALIKKKKKMAP (SEQ ID NO: 3143) of influenza IV; the sequences DRLRR (SEQ ID NO: 3144) and PKQKKRK (SEQ ID NO: 3145) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 3146) of hepatitis virus δ antigen; the sequence REKKKFLKRR (SEQ ID NO: 3147) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 3148) of human poly(ADP-ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 3149) of the steroid hormone receptor (human) glucocorticoid. Other suitable NLS sequences are known in the art (e.g., Sorokin, Biochemistry (Moscow) (2007) 72:13, 1439-1457; Lange J Biol Chem. (2007) 282:8, 5101-5).

[0286] In one embodiment, the Cas9 molecule, for example, the Streptococcus pyogenes (S. pyogenes) Cas9 molecule, includes, for example, an NLS sequence of SV40 located at the N-terminus of the Cas9 molecule. In another embodiment, the Cas9 molecule, for example, the Streptococcus pyogenes (S. pyogenes) Cas9 molecule, includes an NLS sequence of SV40 located at the N-terminus of the Cas9 molecule and an NLS sequence of SV40 located at the C-terminus of the Cas9 molecule. In yet another embodiment, the Cas9 molecule, for example, the Streptococcus pyogenes (S. pyogenes) Cas9 molecule, includes an NLS sequence of SV40 located at the N-terminus of the Cas9 molecule and an NLS sequence of nucleoplasmin located at the C-terminus of the Cas9 molecule. In any of the embodiments described above, the molecule may further include, for example, a tag at the N-terminus or C-terminus, such as a His tag, such as a His(6) tag (SEQ ID NO: 3175) or a His(8) tag (SEQ ID NO: 3176).

[0287] In some embodiments, the Cas9 molecule may include one or more amino acid sequences, such as a tag, that enable the Cas9 molecule to be specifically recognized. In one embodiment, the tag is a histidine tag, for example, a histidine tag containing at least 3, 4, 5, 6, 7, 8, 9, 10 or more histidine amino acids. In one embodiment, the histidine tag is a His6 tag (6 histidines) (SEQ ID NO: 3175). In another embodiment, the histidine tag is a His8 tag (8 histidines) (SEQ ID NO: 3176). In one embodiment, the histidine tag may be separated from one or more other parts of the Cas9 molecule by a linker. In one embodiment, the linker is a GGS. An example of such a fusion is the Cas9 molecule iProt106520.

[0288] In some embodiments, the Cas9 molecule may include one or more amino acid sequences recognized by a protease (e.g., including a protease cleavage site). In embodiments, the cleavage site is a tobacco etch virus (TEV) cleavage site, for example, including the sequence ENLYFQG (SEQ ID NO: 3158). In some embodiments, the protease cleavage site, e.g., the TEV cleavage site, is positioned between a tag, e.g., a His tag, e.g., a His6 tag (SEQ ID NO: 3175) or a His8 tag (SEQ ID NO: 3176) and the rest of the Cas9 molecule. Although not theoretically constrained, such introduction is thought to allow the tag to be used, for example, in the purification of the Cas9 molecule and then cleaved, so that the tag does not interfere with the function of the Cas9 molecule.

[0289] In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) comprises an N-terminal NLS and a C-terminal NLS (e.g., an NLS-Cas9-NLS from N-terminus to C-terminus), where each NLS is SV40 NLS (PKKKRKV (SEQ ID NO: 3134)). In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) comprises an N-terminal NLS, a C-terminal NLS, and a C-terminal His6 tag (SEQ ID NO: 3175) (e.g., an NLS-Cas9-NLS-His tag from N-terminus to C-terminus), where each NLS is SV40 NLS (PKKKRKV (SEQ ID NO: 3134)). In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) comprises an N-terminal His tag (e.g., a His6 tag (SEQ ID NO: 3175)), an N-terminal NLS, and a C-terminal NLS (e.g., from N-terminus to C-terminus, comprising His tag-NLS-Cas9-NLS), for example, where each NLS is SV40 NLS (PKKKRKV (SEQ ID NO: 3134)). In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) comprises an N-terminal NLS and a C-terminal His tag (e.g., a His6 tag (SEQ ID NO: 3175)) (e.g., from N-terminus to C-terminus, comprising His tag-Cas9-NLS), for example, where the NLS is SV40 NLS (PKKKRKV (SEQ ID NO: 3134)). In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) comprises an N-terminal NLS and a C-terminal His tag (e.g., a His6 tag (SEQ ID NO: 3175)) (e.g., an NLS-Cas9-His tag from the N-terminus to the C-terminus), for example, where the NLS is SV40 NLS (PKKKRKV (SEQ ID NO: 3134)).In embodiments, the Cas9 molecule (e.g., the Cas9 molecule as described herein) includes an N-terminal His tag (e.g., His8 tag (SEQ ID NO: 3176)), an N-terminal cleavage domain (e.g., a tobacco etch virus (TEV) cleavage domain (e.g., including sequence ENLYFQG (SEQ ID NO: 3158))), an N-terminal NLS (e.g., SV40 NLS; SEQ ID NO: 3134), and a C-terminal NLS (e.g., SV40 NLS; SEQ ID NO: 3134) (e.g., from N-terminus to C-terminus, including His tag-TEV-NLS-Cas9-NLS). In any of the embodiments described above, Cas9 has the sequence of SEQ ID NO: 3133. Alternatively, in any of the embodiments described above, Cas9 has the sequence of a Cas9 variant of SEQ ID NO: 3133, as described herein. In any of the embodiments described above, the Cas9 molecule includes a linker, e.g., a GGS linker, between the His tag and another part of the molecule. The amino acid sequences of the exemplary Cas9 molecules described above are provided below.

[0290] iProt105026 (also known as iProt106154, iProt106331, iProt106545, and PID426303, prepared by protein preparation) (SEQ ID NO: 3161): [ka] iProt106518 (Sequence ID 3162): [ka] iProt106519 (Sequence ID 3163): [ka] iProt106520 (Sequence ID 3164): [ka] iProt106521 (Sequence ID 3165): [ka] iProt106522 (Sequence ID 3166): [ka] iProt106658 (Sequence ID 3167): [ka] iProt106745 (Sequence ID 3168): [ka] iProt106746 (Sequence ID 3169): [ka] iProt106747 (Sequence ID 3170): [ka] iProt106884 (Sequence ID 3171): [ka] iProt20109496 (Sequence ID 3172) [ka]

[0291] Nucleic acid encoding the Cas9 molecule Nucleic acids encoding Cas9 molecules, such as active Cas9 molecules or inactive Cas9 molecules, are provided herein.

[0292] Exemplary nucleic acids encoding the Cas9 molecule are described in Cong et al, SCIENCE 2013, 399(6121):819-823; Wang et al, CELL 2013, 153(4):910-918; Mali et al., SCIENCE 2013, 399(6121):823-826; and Jinek et al, SCIENCE 2012, 337(6096):816-821.

[0293] In one embodiment, the nucleic acid encoding the Cas9 molecule may be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule may be chemically modified, as described, for example, in Section XIII. In one embodiment, the Cas9 mRNA has one or more, for example, all of the following properties: it is capped, polyadenylated, and substituted with 5-methylcytidine and / or pseudouridine.

[0294] In addition, or instead, the synthetic nucleic acid sequence may be codon-optimized, for example, by replacing at least one rare or low-frequency codon with a high-frequency codon. For example, the synthetic nucleic acid can lead to the synthesis of optimized messenger mRNA, which is optimized for expression in a mammalian expression system, for example, as described herein.

[0295] Below are exemplary codon-optimized nucleic acid sequences encoding the Cas9 molecule of Streptococcus pyogenes. [ka] [ka] [ka]

[0296] Below are exemplary codon-optimized nucleic acid sequences encoding the Cas9 molecule, including sequence number 3172: [ka] [ka]

[0297] It is understood that if the above Cas9 sequence is fused with a peptide or polypeptide at its C-terminus (for example, an inactive Cas9 fused with a transcriptional repressor at its C-terminus), the stop codon will be removed.

[0298] Furthermore, this specification also provides nucleic acids, vectors, and cells for producing the Cas9 molecule, for example, the Cas9 molecule described herein. Recombinant production of polypeptide molecules can be achieved using techniques known to those skilled in the art. This specification describes molecules and methods for the recombinant production of polypeptide molecules, such as the Cas9 molecule, for example, as described herein. When used in connection with this specification, “recombinant” molecule and production include any polypeptide prepared, expressed, created, or isolated by recombinant means (for example, the Cas9 molecule, as described herein, for example), such as polypeptides isolated from transgenic or transchromosomal animals (e.g., mice) for the nucleic acid encoding the molecule of interest, hybridomas prepared therefrom, molecules isolated from, for example, transfectomas from host cells transformed to express the molecule, molecules isolated from a recombinant combinatorial library, and molecules prepared, expressed, created, or isolated by any other means involving splicing from all or part of the gene encoding the molecule (or part thereof) to another DNA sequence. Recombinant production may originate from host cells, for example, host cells containing the molecules described herein, for example, the Cas9 molecule, for example, the nucleic acid encoding the Cas9 molecule described herein.

[0299] This specification provides nucleic acid molecules that encode molecules (e.g., Cas9 molecules and / or gRNA molecules), such as those described herein. Specifically, this specification provides nucleic acid molecules that encode any one of SEQ ID NOs. 3161 to 3172, or a fragment of any of SEQ ID NOs. 3161 to 3172, or a sequence that encodes a polypeptide having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence homology with any of SEQ ID NOs. 3161 to 3172.

[0300] This specification provides a vector, for example, as described herein, comprising any of the nucleic acid molecules described above. In embodiments, the nucleic acid molecule is operably linked to a promoter, for example, a promoter that is activatable in the host cell into which the vector is introduced.

[0301] This specification provides host cells comprising one or more nucleic acid molecules and / or vectors as described herein. In embodiments, the host cell is a prokaryotic host cell. In embodiments, the host cell is a eukaryotic host cell. In embodiments, the host cell is a yeast or Escherichia coli (e. coli) cell. In embodiments, the host cell is a mammalian cell, such as a human cell. Such host cells can be used for the production of recombinant molecules as described herein, such as Cas9 or gRNA molecules as described herein, for example.

[0302] Other Cas molecules Any Cas9 variant or class II CRISPR endonuclease may be used in any of the compositions and methods described herein.

[0303] The term "Cas9 mutant" refers to a protein that has at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity across the entire sequence or a functional portion of the sequence (e.g., a sequence of 50, 100, 150, or 200 consecutive amino acids) compared to the wild-type Cas9 protein, and has one or more mutations that increase its binding specificity to PAM compared to the wild-type Cas9 protein. Exemplary Cas9 mutants are listed in Table 6 below.

[0304] [Table 6]

[0305] When referred to herein, "Cpf1" or "Cpf1 protein" or "Cas12a" includes either a recombinant or naturally occurring form of Cpf1 (CxxC finger protein 1) endonuclease, or a variant or homolog thereof that maintains Cpf1 endonuclease enzyme activity (e.g., activity within 50%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% compared to Cpf1). In some embodiments, the variant or homolog has at least 90%, 95%, 96%, 97%, 98%, 99%, or 100% amino acid sequence identity over the entire sequence or a portion of the sequence (e.g., a contiguous portion of 50, 100, 150, or 200 amino acids) compared to the naturally occurring Cpf1 protein. In this embodiment, the Cpf1 protein is substantially identical to the protein identified by UniProt reference number Q9P0U4, or to a variant or homolog having substantial identity therewith.

[0306] The term "Class II CRISPR endonuclease" refers to an endonuclease that has endonuclease activity similar to Cas9 and is involved in the Class II CRISPR system. An exemplary Class II CRISPR system is the Type II CRISPR locus of Streptococcus pyogenes SF370, which contains a cluster of four genes, Cas9, Cas1, Cas2, and Csn1, as well as two non-coding RNA elements, tracrRNA, and a characteristic series of repeat sequences (direct repeats) interspersed with short, consecutive non-repetitive sequences (spacers, each approximately 30 bp). In this system, targeted DNA double-strand breaks (DSBs) can be produced in four sequential steps. First, two non-coding RNAs, a pre-crRNA array and tracrRNA, can be transcribed from the CRISPR locus. Secondly, the tracrRNA may hybridize to the direct repeats of the precrRNA, which are then processed to form a mature crRNA containing individual spacer sequences. Thirdly, the mature crRNA:tracrRNA complex can lead Cas9 to a DNA target consisting of a protospacer and a corresponding PAM by heterodouble helix formation between the crRNA spacer region and the protospacer DNA. Finally, Cas9 can mediate cleavage of the target DNA upstream of the PAM, creating a double-strand break (DSB) within the protospacer.

[0307] V. Functional analysis of candidate molecules Candidate Cas9 molecules, candidate gRNA molecules, and candidate Cas9 / gRNA molecule complexes can be evaluated by methods known in the art or as described herein. For example, an exemplary method for evaluating the endonuclease activity of a Cas9 molecule is described, for example, in Jinek el al., SCIENCE 2012;337(6096):816-821.

[0308] VI. Template nucleic acid (for nucleic acid introduction) The terms “template nucleic acid” or “donor template,” as used herein, refer to a nucleic acid that is modified, for example, cleaved, and inserted into or near a target sequence by the CRISPR system of the present invention. In some embodiments, a nucleic acid sequence in or near the target sequence is modified to have, typically, part or all of, a sequence of the template nucleic acid at or near one or more cleavage sites. In some embodiments, the template nucleic acid is single-stranded. In alternative embodiments, the template nucleic acid is double-stranded. In some embodiments, the template nucleic acid is DNA, for example, double-stranded DNA. In alternative embodiments, the template nucleic acid is single-stranded DNA.

[0309] In some embodiments, the template nucleic acid comprises a sequence encoding a globin protein, such as β-globin, and includes, for example, a β-globin gene. In some embodiments, the β-globin encoded by the nucleic acid comprises one or more mutations, such as an anti-sickle mutation. In some embodiments, the β-globin encoded by the nucleic acid comprises the mutation T87Q. In some embodiments, the β-globin encoded by the nucleic acid comprises the mutation G16D. In some embodiments, the β-globin encoded by the nucleic acid comprises the mutation E22A. In some embodiments, the β-globin gene comprises the mutations G16D, E22A, and T87Q. In some embodiments, the template nucleic acid further comprises one or more regulatory elements, such as a promoter (e.g., a human β-globin promoter), a 3' enhancer, and / or at least a portion of a globin locus regulatory region (e.g., one or more DNase I hypersensitive sites (e.g., HS2, HS3, and / or HS4 of the human globin locus)).

[0310] In other embodiments, the template nucleic acid includes a sequence encoding gamma globin, for example, a gamma globin gene. In embodiments, the template nucleic acid includes a sequence encoding two or more copies of gamma globin protein, for example, two or more, for example, two gamma globin gene sequences. In embodiments, the template nucleic acid further includes one or more regulatory elements, for example, a promoter and / or enhancer.

[0311] In some embodiments, the template nucleic acid modifies the structure of the target site by participating in homologous recombination repair events. In some embodiments, the template nucleic acid modifies the sequence of the target site. In some embodiments, the template nucleic acid results in the incorporation of modified or non-naturally occurring bases into the target nucleic acid.

[0312] Mutations in the genes or pathways described herein can be corrected using one of the methods discussed herein. In one embodiment, mutations in the genes or pathways described herein are corrected by homologous recombination repair (HDR) using a template nucleic acid. In another embodiment, mutations in the genes or pathways described herein are corrected by homologous recombination (HR) using a template nucleic acid. In yet another embodiment, mutations in the genes or pathways described herein are corrected by non-homologous end joining (NHEJ) repair using a template nucleic acid. In yet another embodiment, a nucleic acid encoding the molecule of interest can be inserted into or near a site modified by the CRISPR system of the present invention. In one embodiment, the template nucleic acid includes, for example, one or more promoters and / or enhancers operably linked to a nucleic acid sequence encoding the molecule of interest, as described herein.

[0313] HDR or HR repair and template nucleic acids As described herein, target sequences can be modified and genomic mutations corrected (e.g., repaired or edited) using nuclease-induced homologous recombination repair (HDR) or homologous recombination (HR). While not intended to be constrained by theory, target sequence modification is thought to occur by repair based on a donor template or template nucleic acid. For example, a donor template or template nucleic acid results in target sequence modification. Plasmid donors or linear double-stranded templates are intended to be used as templates for homologous recombination. Furthermore, single-stranded donor templates are intended to be used as templates for target sequence modification by alternative homologous recombination repair methods (e.g., single-stranded annealing) between the target sequence and the donor template. Target sequence modification achieved by a donor template may depend on cleavage by the Cas9 molecule. Cas9 cleavage may include double-strand breaks, single single-strand breaks, or two single-strand breaks.

[0314] In one embodiment, the mutation can be corrected by either a single double-strand break or two single-strand breaks. In one embodiment, the mutation can be corrected by providing a CRISPR / Cas9 system and template that produces (1) one double-strand break, (2) two single-strand breaks, (3) two double-strand breaks resulting in breaks on each side of the target sequence, (4) one double-strand break and two single-strand breaks resulting in a double-strand break and two single-strand breaks on each side of the target sequence, (5) four single-strand breaks resulting in a pair of single-strand breaks on each side of the target sequence, or (6) one single-strand break.

[0315] Correction mediated by double-strand breaks In one embodiment, double-strand breaks are achieved by a Cas9 molecule having cleavage activity associated with an HNH-like domain and cleavage activity associated with a RuvC-like domain, such as the N-terminal RuvC-like domain, e.g., wild-type Cas9. Such embodiments require only a single gRNA.

[0316] Correction mediated by single-strand breaks In other embodiments, two single-strand breaks, or nicks, are achieved by a Cas9 molecule having nickase activity, for example, cleavage activity related to an HNH-like domain or cleavage activity related to an N-terminal RuvC-like domain. Such embodiments require two gRNAs, one for each single-strand break configuration. In one embodiment, the nickase-active Cas9 molecule cleaves the strand with which the gRNA hybridizes, but not the strand complementary to the strand with which the gRNA hybridizes. In another embodiment, the nickase-active Cas9 molecule does not cleave the strand with which the gRNA hybridizes, but rather cleaves the strand complementary to the strand with which the gRNA hybridizes.

[0317] In one embodiment, the nickase is a Cas9 molecule having HNH activity, for example, a Cas9 molecule with inactivated RuvC activity, for example, a mutation at D10, for example, a Cas9 molecule with a D10A mutation. D10A inactivates RuvC. Therefore, the Cas9 nickase has HNH activity (only) and cleaves the gRNA hybridizing strand (for example, a complementary strand that does not have NGG PAM). In another embodiment, a Cas9 molecule having H840, for example, an H840A mutation, can be used as the nickase. H840A inactivates HNH. Therefore, the Cas9 nickase has RuvC activity (only) and cleaves the non-complementary strand (for example, a strand that has NGG PAM and whose sequence is identical to that of the gRNA).

[0318] In embodiments where two single-stranded nicks are positioned using nickase and two gRNAs, one nick is on the + strand of the target nucleic acid and the other nick is on the - strand. The PAM faces outward. The gRNAs can be selected so that they are separated by only about 0-50, 0-100, or 0-200 nucleotides. In some embodiments, there is no overlap between the targeting domains and complementary target sequences of the two gRNAs. In some embodiments, the gRNAs do not overlap and are separated by about 50, 100, or 200 nucleotides. In some embodiments, using two gRNAs may increase specificity, for example, by reducing off-target binding (Ran el al., CELL 2013).

[0319] In one embodiment, HDR can be induced using a single nick. This specification intends to demonstrate that a single nick can be used to increase the HDR, HR, or NHEJ rate at a given cut site.

[0320] Arrangement of double-strand or single-strand breaks relative to the target location. The double-strand or single-strand break on one side of the strand must be close enough to the target site for the correction to occur. In some embodiments, the distance is 50, 100, 200, 300, 350, or 400 nucleotides or less. Although we do not wish to be constrained by theory, it is considered that the break point must be close enough to the target site so that it is within the region where exonuclease-mediated removal occurs between terminal resections. As the donor sequence can only be used to correct sequences within the terminal resection region, if the distance between the target site and the break point is too great, the terminal resection may not contain the mutation and therefore may not be corrected.

[0321] In an embodiment in which a gRNA (single molecule (or chimeric) or modular gRNA) and Cas9 nuclease induce double-strand breaks for the purpose of inducing HDR-mediated or HR-mediated modification, the cleavage site is located 0 to 200 bp away from the target site (e.g., 0 to 175, 0 to 150, 0 to 125, 0 to 100, 0 to 75, 0 to 50, 0 to 25, 25 to 200, 25 to 175, 25 to 150, 25 to 125, 25 to 100, 25 to 75, 25 to 50, 50 to 200, 50 to 175, 50 to 150, 50 to 125, 50 to 100, 50 to 75, 75 to 200, 75 to 175, 75 to 150, 75 to 125, 75 to 100 bp). In one embodiment, the cutting site is located 0 to 100 bp away from the target location (e.g., 0 to 75, 0 to 50, 0 to 25, 25 to 100, 25 to 75, 25 to 50, 50 to 100, 50 to 75, or 75 to 100 bp).

[0322] In an embodiment in which two gRNAs (independently, single-molecule (or chimeric) or modular gRNAs) complexed with Cas9 nickase induce two single-strand breaks for the purpose of inducing HDR-mediated modification, the closer nick is located 0-200 bp from the target site (e.g., 0-175, 0-150, 0-125, 0-100, 0-75, 0-50, 0-25, 25-200, 25-175, 25-150, 25-125, 25-100, 25-75, 25-50, 50-200, 50-175, 50-150, 50-125, 50-100, 50-75, 75-20) The two nicks are separated by 0, 75-175, 75-150, 75-125, and 75-100 bp, and ideally the two nicks are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 30-55, 30-50, 30-45, 30-40, 30-35, 35-55, 35-50, 35-45, 35-40, 40-55, 40-50, 40-45 bp) and separated by 100 bp or less (e.g., separated by 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 bp or less). In one embodiment, the cutting site is located 0 to 100 bp away from the target location (e.g., 0 to 75, 0 to 50, 0 to 25, 25 to 100, 25 to 75, 25 to 50, 50 to 100, 50 to 75, or 75 to 100 bp).

[0323] In one embodiment, two gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place double-strand breaks on either side of the target site. In an alternative embodiment, three gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place a double-strand break (i.e., one gRNA complexes with Cas9 nuclease) and two single-strand breaks or paired single-strand breaks (i.e., two gRNAs complex with Cas9 nickase) on either side of the target site (for example, the first gRNA is used to target the upstream (i.e., 5' side) of the target site, and the second gRNA is used to target the downstream (i.e., 3' side) of the target site). In another embodiment, four gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to produce two pairs of single-strand breaks on either side of the target site (i.e., two pairs of gRNAs complex with Cas9 nickase) (e.g., the first gRNA targets the upstream (i.e., 5' side) of the target site, and the second gRNA targets the downstream (i.e., 3' side) of the target site). The closer of the two single-strand nicks in the double-strand break or pair may ideally be within 0 to 500 bp from the target site (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target site). When nickase is used, the two nicks in a pair are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 50-55, 45-55, 40-55, 35-55, 30-55, 30-50, 35-50, 40-50, 45-50, 35-45, or 40-45 bp) and separated by no more than 100 bp (e.g., 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp or less).

[0324] In one embodiment, two gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place double-strand breaks on either side of the target site. In an alternative embodiment, three gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place double-strand breaks (i.e., one gRNA complexes with Cas9 nuclease) and two single-strand breaks or paired single-strand breaks (i.e., two gRNAs complex with Cas9 nickase) at two target sequences (for example, the first gRNA is used to target the upstream (i.e., 5' side) target sequence of the insertion site, and the second gRNA is used to target the downstream (i.e., 3' side) target sequence). In another embodiment, four gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to produce two pairs of single-strand breaks on either side of the insertion site (i.e., two pairs of gRNAs complex with Cas9 nickase) (e.g., the first gRNA targets the upstream (i.e., 5' side) target sequence described herein, and the second gRNA targets the downstream (i.e., 3' side) target sequence described herein). The closer of the two single-strand nicks in the double-strand break or pair is ideally within 0 to 500 bp from the target site (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target site). When nickase is used, the two nicks in a pair are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 50-55, 45-55, 40-55, 35-55, 30-55, 30-50, 35-50, 40-50, 45-50, 35-45, or 40-45 bp) and separated by no more than 100 bp (e.g., 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp or less).

[0325] Length of homology arm Homology arms must extend at least to the region where terminal resection may occur, so that, for example, a resectioned single-stranded overhang can find a complementary region within the donor template. The total length may be limited by parameters such as plasmid size or viral packaging limits. In some embodiments, homology arms do not extend to repeating elements, such as ALU repeats or LINE repeats. A template may have two homology arms of the same or different lengths.

[0326] Exemplary homology arm lengths include at least 25, 50, 100, 250, 500, 750, or 1000 nucleotides.

[0327] As used herein, a target site refers to a region on a target nucleic acid (e.g., a chromosome) that is modified by a Cas9 molecule-dependent process. For example, a target site may be a cleavage of the target nucleic acid by a modified Cas9 molecule and a modification, e.g., correction, of the target site by a template nucleic acid. In some embodiments, a target site may be a region between two nucleotides on the target nucleic acid to which one or more nucleotides are added, e.g., between adjacent nucleotides. A target site may include one or more nucleotides that are modified, e.g., corrected, by a template nucleic acid. In some embodiments, a target site is within the range of a target sequence (e.g., a sequence to which gRNA binds). In some embodiments, a target site is upstream or downstream of a target sequence (e.g., a sequence to which gRNA binds).

[0328] Typically, a template sequence undergoes cleavage-mediated or catalytic recombination with a target sequence. In one embodiment, the template nucleic acid includes a sequence corresponding to a site on the target sequence that is cleaved by a Cas9-mediated cleavage event. In another embodiment, the template nucleic acid includes a sequence corresponding to both a first site on the target sequence that is cleaved by a first Cas9-mediated event and a second site on the target sequence that is cleaved by a second Cas9-mediated event.

[0329] In one embodiment, the template nucleic acid may include sequences that cause alteration of the coding sequence of the translation sequence, for example, substitutions of one amino acid in a protein product with another amino acid, such as converting a mutant allele to a wild-type allele, converting a wild-type allele to a mutant allele, and / or introducing a stop codon, insertion of an amino acid residue, deletion of an amino acid residue, or a nonsense mutation.

[0330] In other embodiments, the template nucleic acid may include sequences that result in modifications of non-coding sequences, such as modifications of exons or 5' or 3' untranslated or non-transcribed regions. Such modifications include modifications of control elements, such as promoters, enhancers, and cis- or trans-acting control elements.

[0331] The template nucleic acid may contain sequences that, when incorporated, produce the following: Decreased activity of the positive control element; Increased activity of the positive control element; Decreased activity of the negative control element; Increased activity of the negative regulatory element; Decreased gene expression; Increased gene expression; Increased resistance to injury or disease; Increased resistance to viral invasion; Correction of mutations or modification of undesirable amino acid residues; Conferring, increasing, losing, or decreasing the biological properties of a gene product, for example, increasing the enzymatic activity of an enzyme, or increasing the ability of a gene product to interact with another molecule.

[0332] The template nucleic acid may contain sequences that produce the following: A change of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides or more in the target sequence.

[0333] In one embodiment, the template nucleic acid has a nucleotide length of 20±10, 30±10, 40±10, 50±10, 60±10, 70±10, 80±10, 90±10, 100±10, 110±10, 120±10, 130±10, 140±10, 150±10, 160±10, 170±10, 180±10, 190±10, 200±10, 210±10, 220±10, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-2000, 2000-3000, or greater than 3000.

[0334] The template nucleic acid contains the following components: [5' homology arm]-[insertion sequence]-[3' homology arm].

[0335] Homology arms induce recombination of the chromosome, thereby allowing undesirable elements, such as mutations or signatures, to be replaced by substitutional sequences. In one embodiment, the homology arm is adjacent to the most distal cleavage site.

[0336] In one embodiment, the 3' end of the 5' homology arm is located adjacent to the 5' end of the substitution sequence. In one embodiment, the 5' homology arm may extend at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 180, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides toward the 5' end of the substitution sequence.

[0337] In one embodiment, the 5' end of the 3' homology arm is located adjacent to the 3' end of the substitution sequence. In one embodiment, the 3' homology arm may extend at least 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 120, 150, 180, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, or 2000 nucleotides toward the 3' end of the substitution sequence.

[0338] In this specification, it is intended that the inclusion of certain sequence repeat elements, such as Alu repeats and LINE elements, can be avoided by shortening one or both homology arms. For example, the 5' homology arm can be shortened to avoid sequence repeat elements. In other embodiments, the 3' homology arm can be shortened to avoid sequence repeat elements. In some embodiments, both the 5' and 3' homology arms can be shortened to avoid the inclusion of certain sequence repeat elements.

[0339] This specification envisions the design of template nucleic acids for mutation correction to be used as single-stranded oligonucleotides (ssODNs). When using ssODNs, the 5' and 3' homology arms can be up to approximately 200 base pairs (bp) in length, for example, in the range of at least 25, 50, 75, 100, 125, 150, 175, or 200 bp. For ssODNs, longer homology arms are also envisioned to allow for continued improvement in oligonucleotide synthesis.

[0340] NHEJ methods for gene targeting As described herein, gene-specific knockout can be targeted using nuclease-inducible non-homologous end joining (NHEJ). Nuclease-inducible NHEJ can also be used to remove (e.g., delete) the sequence of the target gene.

[0341] While we do not wish to be constrained by theory, in certain embodiments, genome modifications associated with the methods described herein are thought to rely on the error-prone nature of nuclease-induced NHEJs and the NHEJ repair pathway. NHEJs repair DNA double-strand breaks by ligating both ends together. However, generally, the original sequence is restored only if the two compatible ends are completely ligated in the same way they were formed by the double-strand break. Since DNA ends at double-strand breaks are frequently subjected to enzymatic processing, nucleotide additions or deletions occur on one or both strands before those ends can be rejoined. As a result, insertion and / or deletion (indel) mutations are present in the DNA sequence of the NHEJ repair site. Two-thirds of these mutations alter the leading frame and may therefore produce non-functional proteins. In addition, mutations that maintain the leading frame but insert or delete large amounts of sequences may disrupt protein function. This is locus-dependent because mutations in critical functional domains are more likely to be unacceptable than mutations in less critical regions of a protein.

[0342] Indel mutations produced by NHEJ are inherently unpredictable. However, certain indel sequences are preferred at a given cleavage site and appear frequently in the population. Deletion lengths can vary widely, most commonly in the range of 1–50 bp, but can easily reach over 100–200 bp. Insertions tend to be shorter and often involve short duplications of sequences directly surrounding the cleavage site. However, larger insertions are possible, in which case the inserted sequence often extends to other regions of the genome or even to plasmid DNA present in the cell.

[0343] Because NHEJ is a mutagenic process, small sequence motifs can be deleted using NHEJ, as long as the creation of a specific final sequence is not required. When a double-strand break is targeted near a short target sequence, the deletion mutation caused by NHEJ repair often extends to an undesirable nucleotide, thus removing it. For deletions of larger DNA segments, introducing two double-strand breaks, one on each side of the sequence, can cause an NHEJ between the ends, potentially removing the entire intervening sequence. Both of these methods can be used for deletions of specific DNA sequences. However, the error-prone nature of NHEJ can still lead to indel mutations at the repair site.

[0344] The methods and compositions described herein may use both double-strand break Cas9 molecules and single-strand or nickase Cas9 molecules to create NHEJ-mediated indels. A gene of interest can be knocked out (i.e., its expression can be eliminated) using an NHEJ-mediated indel that targets a gene, for example, the coding region of the gene of interest, for example, the early coding region. For example, the early coding region of the gene of interest includes a sequence immediately after the transcription start site, within the range of the first exon of the coding sequence, or within 500 bp (e.g., less than 500, 450, 400, 350, 300, 250, 200, 150, 100, or 50 bp) from the transcription start site.

[0345] Arrangement of double-strand or single-strand breaks relative to the target location In embodiments in which a gRNA and Cas9 nuclease create double-strand breaks for the purpose of inducing NHEJ-mediated indels, the gRNA, e.g., a single molecule (or chimeric) or modular gRNA molecule, is configured to place a single double-strand break very close to a nucleotide at the target site. In some embodiments, the cleavage site is located 0 to 500 bp away from the target site (e.g., less than 500, 400, 300, 200, 100, 50, 40, 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 bp away from the target site).

[0346] In embodiments in which two gRNAs complex with Cas9 nickase to induce two single-strand breaks for the purpose of inducing NHEJ-mediated indels, the two gRNAs, for example, independently, monomolecular (or chimeric) or modular gRNAs, are configured to place two single-strand breaks to result in NHEJ repair of nucleotides at target sites. In one embodiment, the gRNAs are configured to essentially mimic double-strand breaks by placing breaks at the same site on different strands or within a few nucleotides of each other. In one embodiment, the closer nick is located 0 to 30 bp away from the target site (e.g., less than 30, 25, 20, 1, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 bp away from the target site), and the two nicks are located within 25 to 55 bp of each other (e.g., 25 to 50, 25 to 45, 25 to 40, 25 to 35, 25 to 30, 50 to 55, 45 to 55, 40 to 55, 35 to 55, 30 to 50, 35 to 50, 40 to 50, 45 to 50, 35 to 45, or 40 to 45 bp) and are located less than or equal to 100 bp away from each other (e.g., less than or equal to 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp). In one embodiment, the gRNA is configured to place a single-strand break on either side of the nucleotide at the target site.

[0347] In the methods and compositions described herein, both double-strand break Cas9 molecules and single-strand or nickase Cas9 molecules may be used to create breaks on both sides of the target site. Double-strand or paired single-strand breaks may be created on both sides of the target site to remove the nucleic acid sequence between the two cuts (e.g., a region between the two breaks is deleted). In one embodiment, two gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place double-strand breaks on both sides of the target site (e.g., the first gRNA is used to target the upstream (i.e., 5' side) of the mutation in the gene or pathway described herein, and the second gRNA is used to target the downstream (i.e., 3' side) of the mutation in the gene or pathway described herein). In an alternative embodiment, three gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to place a double-strand break on either side of the target site (i.e., one gRNA complexes with Cas9 nuclease) and two single-strand breaks or paired single-strand breaks (i.e., two gRNAs complex with Cas9 nickase) (for example, using the first gRNA to target the upstream (i.e., 5' side) of the mutation in the gene or pathway described herein, and using the second gRNA to target the downstream (i.e., 3' side) of the mutation in the gene or pathway described herein). In another embodiment, four gRNAs, for example, independently, single (or chimeric) or modular gRNAs, are configured to induce two pairs of single-strand breaks on either side of the target site (i.e., two pairs of gRNAs complex with Cas9 nickase) (e.g., the first gRNA is used to target the upstream (i.e., 5' side) of the mutation in the gene or pathway described herein, and the second gRNA is used to target the downstream (i.e., 3' side) of the mutation in the gene or pathway described herein). The closer of the two single-strand nicks in one or more double-strand breaks or pairs may ideally be within 0 to 500 bp from the target site (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target site).When nickase is used, the two nicks in a pair are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 50-55, 45-55, 40-55, 35-55, 30-55, 30-50, 35-50, 40-50, 45-50, 35-45, or 40-45 bp) and separated by no more than 100 bp (e.g., 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp or less).

[0348] In other embodiments, insertion of template nucleic acids can be mediated by microhomology end joining (MMEJ). See, for example, Saksuma et al., "MMEJ-assisted gene knock-in using TALENs and CRISPR-Cas9 with the PITCh systems," Nature Protocols 11, 118-133 (2016) doi:10.1038 / nprot.2015.140, published online December 17, 2015 (this content is incorporated by reference in its entirety).

[0349] VII. Systems containing two or more gRNA molecules While not intended to be constrained by theory, this specification has shown that targeting two very close target sequences on a continuous nucleic acid (e.g., by two gRNA / Cas9 molecule complexes, each inducing a single-strand or double-strand break in or near its respective target sequence) induces excision (e.g., deletion) of nucleic acid sequences (or at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% of nucleic acid sequences) located between the two target sequences. In some embodiments, this disclosure provides the use of two or more gRNA molecules containing targeting domains that target very close target sequences on a continuous nucleic acid, such as a chromosome, such as a gene or locus, including its introns, exons, and regulatory elements. This use is obtained, for example, by introducing two or more gRNA molecules together with one or more Cas9 molecules (or two or more gRNA molecules and / or nucleic acids encoding one or more Cas9 molecules) into a cell.

[0350] In some embodiments, the target sequences of two or more gRNA molecules are located at a distance of 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10,000, 11,000, 12,000, 13,000, 14,000, or 15,000 nucleotides or more, and at a distance of 25,000 nucleotides or less, on the consecutive nucleic acid. In embodiments, the target sequences are located at a distance of approximately 4000 to 6000 nucleotides. In some embodiments, the target sequences are located at a distance of approximately 4000 nucleotides. In one embodiment, the target sequence is located approximately 5000 nucleotides away from the target sequence, or approximately 4000 nucleotides away. In another embodiment, the target sequence is located approximately 6000 nucleotides away.

[0351] In some embodiments, multiple gRNA molecules each target a sequence within the same gene or locus. In other embodiments, multiple gRNA molecules each target a sequence within two or more different genes or loci.

[0352] In some embodiments, the present invention provides compositions and cells comprising a plurality of gRNA molecules of the present invention, for example, two or more, for example, two gRNA molecules, wherein the plurality of gRNA molecules target sequences separated by less than 15,000, less than 14,000, less than 13,000, less than 12,000, less than 11,000, less than 10,000, less than 9,000, less than 8,000, less than 7,000, less than 6,000, less than 5,000, less than 4,000, less than 3,000, less than 2,000, less than 1,000, less than 900, less than 800, less than 700, less than 600, less than 50, less than 40, or less than 30 nucleotides. In some embodiments, these target sequences lie on the same strand of a double-stranded nucleic acid. In one embodiment, these target sequences are located on different strands of a double-stranded nucleic acid.

[0353] In one embodiment, the present invention provides a method for cleaving (e.g., deleting) a nucleic acid located between two gRNA binding sites that are separated by less than 25,000, less than 20,000, less than 15,000, less than 14,000, less than 13,000, less than 12,000, less than 11,000, less than 10,000, less than 9,000, less than 8,000, less than 7,000, less than 6,000, less than 5,000, less than 4,000, less than 3,000, less than 2,000, less than 1,000, less than 900, less than 800, less than 700, less than 600, less than 500, less than 400, less than 30, less than 200, less than 100, less than 90, less than 80, less than 70, less than 60, less than 50, less than 40, or less than 30 nucleotides on the same or different strands of a double-stranded nucleic acid. In one embodiment, this method results in a deletion of more than 50%, more than 60%, more than 70%, more than 80%, more than 85%, more than 86%, more than 87%, more than 88%, more than 89%, more than 90%, more than 91%, more than 92%, more than 93%, more than 94%, more than 95%, more than 96%, more than 97%, more than 98%, more than 99%, or 100% of the nucleotides located between the PAM sites associated with each gRNA binding site. In an embodiment, this deletion further includes one or more nucleotides located within the range of one or more PAM sites associated with each gRNA binding site. In an embodiment, this deletion also includes one or more nucleotides located outside the region between the PAM sites associated with each gRNA binding site.

[0354] In one embodiment, two or more gRNA molecules include targeting domains that target a target sequence adjacent to a gene regulatory element, such as a promoter binding site, enhancer region, or repressor region, thereby causing upregulation or downregulation of the target gene by cleaving an intervening sequence (or a portion thereof). In other embodiments, two or more gRNA molecules include targeting domains that target a sequence adjacent to a gene, so that cleavage of an intervening sequence (or a portion thereof) results in deletion of the target gene.

[0355] In one embodiment, two or more gRNA molecules each contain a targeting domain, for example, comprising a targeting domain sequence from Table 1, Table 2, or Table 3. In another embodiment, two or more gRNA molecules each contain a targeting domain, for example, comprising a targeting domain of a gRNA molecule that results in an upregulation of at least 15% of the number of F cells in a population of erythrocytes differentiated exovivoally by gRNA-edited HSPCs (e.g., on day 7 post-editing) by the method described herein. In one embodiment, two or more gRNA molecules each contain a targeting domain complementary to a sequence in the same gene or region, for example, the WIZ gene region. In another embodiment, two or more gRNA molecules each contain a targeting domain complementary to sequences in different genes or regions, for example, one in the WIZ intron region and the other in the WIZ exon region.

[0356] In one embodiment, two or more gRNA molecules include targeting domains that target target sequences adjacent to gene regulatory elements, such as promoter binding sites, enhancer regions, or repressor regions, so that cleavage of an intervening sequence (or a portion of an intervening sequence) results in upregulation or downregulation of the target gene. In another embodiment, two or more gRNA molecules include targeting domains that target target sequences adjacent to a gene, so that cleavage of an intervening sequence (or a portion of an intervening sequence) results in deletion of the target gene. For example, two or more gRNA molecules include targeting domains that target target sequences adjacent to the WIZ gene, so that the WIZ gene is excised.

[0357] In one embodiment, two or more gRNA molecules include, for example, a targeting domain selected from Table 1.

[0358] In one embodiment, two or more gRNA molecules include, for example, a targeting domain comprising, a targeting domain sequence listed in Table 2. In another embodiment, two or more gRNA molecules include, for example, a targeting domain comprising, a targeting domain sequence of a gRNA listed in Table 3.

[0359] VIII. Characteristics of gRNA Furthermore, surprisingly, this specification states that a single gRNA molecule may have target sequences at two or more loci (e.g., loci with high sequence homology), and that such loci may be on the same chromosome, for example, less than approximately 15,000 nucleotides, less than approximately 14,000 nucleotides, less than approximately 13,000 nucleotides, less than approximately 12,000 nucleotides, less than approximately 11,000 nucleotides, less than approximately 10,000 nucleotides, less than approximately 9,000 nucleotides, less than approximately 8,000 nucleotides, and less than approximately 7,000 nucleotides. When located within the range of less than ocide, less than approximately 6,000 nucleotides, less than approximately 5,000 nucleotides, less than approximately 4,000 nucleotides, or less than approximately 3,000 nucleotides (for example, between approximately 4,000 and approximately 6,000 nucleotides apart), such gRNA molecules have been shown to cause cleavage of intervening sequences (or portions thereof), thereby yielding beneficial effects, such as upregulation of fetal hemoglobin in erythroid cells differentiated from modified HSPCs (as described herein). Accordingly, in some embodiments, the present invention provides gRNA molecules having target sequences (for example, as described in Tables 1 to 3) at two loci, for example, having target sequences at loci on the same chromosome, for example, in the WIZ intron region and the WIZ exon region. Although not constrained by theory, such gRNAs may cause genomic cleavage at two or more locations (for example, at target sequences in each of the two regions), and subsequent repair may result in deletion of intervening nucleic acid sequences. In this case as well, although not constrained by theory, the deletion of the aforementioned intervening sequence may have a desired effect on the expression or function of one or more proteins.

[0360] While not constrained by theory, it is conceivable that some indel patterns may be more advantageous than others. For example, indels primarily containing insertions and / or deletions resulting in "frameshift mutations" (e.g., 1 or 2 base pair insertions or deletions, or any insertion or deletion where n / 3 (where n = the number of nucleotides in the insertion or deletion) is not an integer) may be beneficial in reducing or eliminating the expression of functional proteins. Similarly, indels primarily containing "large deletions" (e.g., deletions greater than 10, 11, 12, 13, 14, 15, 20, 25, or 30 nucleotides, including sequences located between the first and second binding sites of a gRNA, as described herein, for example) may also be beneficial in removing critically important regulatory sequences, such as promoter binding sites, or in altering the structure or function of a locus that may similarly affect the expression of functional proteins. Surprisingly, the indel patterns induced by a given gRNA / CRISPR system are consistently reproduced for given cell types, gRNAs, and CRISPR systems as described herein; however, introducing a gRNA / CRISPR system into a given cell does not necessarily result in the formation of any single indel structure.

[0361] Accordingly, the present invention provides gRNA molecules that produce beneficial indel patterns or structures, for example, gRNA molecules having indel patterns or structures mainly composed of large deletions. Such gRNA molecules may be selected by evaluating the indel patterns or structures produced by candidate gRNA molecules by NGS in test cells (e.g., HEK293 cells) or target cells, for example, HSPC cells, as described herein. As shown in the examples, gRNA molecules have been discovered that, when introduced into a desired cell population, result in a cell population in which cells with large deletions at or near the target sequence of the gRNA are predominant. In some cases, the rate of formation of large deletion indels is as high as 75%, 80%, 85%, 90%, or even higher. Accordingly, the present invention provides a population of cells in which at least about 40% (e.g., at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%) cells have a large deletion, for example, in or near the target site of the gRNA molecule described herein. The present invention also provides a population of cells in which at least about 50% (e.g., at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%) cells have a large deletion, for example, in or near the target site of the gRNA molecule described herein.

[0362] Accordingly, the present invention provides a method for selecting gRNA molecules to be used in the therapeutic method of the present invention, comprising: 1) providing a plurality of gRNA molecules for a target of interest; 2) evaluating the indel pattern or structure produced by the use of the gRNA molecules; 3) selecting gRNA molecules that form an indel pattern or structure mainly composed of frameshift mutations, large deletions, or combinations thereof; and 4) using the selected gRNAs in the method of the present invention.

[0363] Accordingly, the present invention provides a method for selecting gRNA molecules for use in a therapeutic method of the present invention, comprising: 1) providing a plurality of gRNA molecules to a target of interest having, for example, target sequences at two or more positions; 2) evaluating the indel pattern or structure produced by the use of the gRNA molecules; 3) selecting gRNA molecules that form sequence cleavages including nucleic acid sequences located between two target sequences in, for example, at least about 25% of cells in a population of cells exposed to the gRNA molecules; and 4) using the selected gRNA molecules in a method of the present invention.

[0364] The present invention further provides a method for modifying cells and modified cells, wherein a specific indel pattern is consistently produced in a given gRNA / CRISPR system. The indel pattern can be determined using the method of the Examples, including the top five most frequently occurring indels observed in the gRNA / CRISPR system described herein, as disclosed, for example, in the Examples. As shown in the examples, a population of cells is created in which a large proportion of the cells contain one of the top five indels (for example, a population of cells in which more than 30%, more than 40%, more than 50%, more than 60%, or more cells in the population contain one of the top five indels). Thus, the present invention provides cells containing any one of the top five indels observed in a given gRNA / CRISPR system, e.g., HSPCs (as described herein). Furthermore, the present invention provides a population of cells in which a high percentage of cells contain one of the top five indels described herein for a given gRNA / CRISPR system, e.g., a population of HSPCs (as described herein), when evaluated, for example, by NGS. Indel pattern analysis and When used in relation to the above, “high percentage” means that at least about 50% of the population (e.g., at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99%) of the cells contain one of the top five indels described herein for a given gRNA / CRISPR system. In other embodiments, a population of cells is such that at least about 25% (e.g., about 25% to about 60%, e.g., about 25% to about 50%, e.g., about 25% to about 40%, e.g., about 25% to about 35%) of the cells have one of the top five indels described herein for a given gRNA / CRISPR system.

[0365] Furthermore, it has been discovered that certain gRNA molecules do not form indels at off-target sequences within the genome range of the target cell type (e.g., off-target sequences outside the WIZ gene region), or that the formation of indels at off-target sites (e.g., off-target sequences outside the WIZ region) occurs at an extremely low frequency (e.g., less than 5% of cells in the population) compared to the frequency of indel formation at target sites. Therefore, the present invention provides a gRNA molecule and a CRISPR system that do not exhibit off-target indel formation in the target cell type, or that exhibit off-target indel formation at a frequency of less than 5%, for example, indels at any off-target site outside the WIZ gene region at a frequency of less than 5%. In embodiments, the present invention provides a gRNA molecule and a CRISPR system that do not exhibit any off-target indel formation in the target cell type. Accordingly, the present invention further provides cells, for example, populations of cells, for example, populations of HSPCs, for example, as described herein, which include indels (e.g., frameshift indels, or any of the top five indels generated by a given gRNA / CRISPR system, for example, as described herein) at or near the target site of the gRNA molecule described herein, but do not include indels at any off-target sites of the gRNA molecule, for example, any off-target sites outside the WIZ gene region.In other embodiments, the present invention further provides a population of cells, e.g., HSPCs, as described herein, in which at least 20%, e.g., at least 30%, e.g., at least 40%, e.g., at least 50%, e.g., at least 60%, e.g., at least 70%, e.g., at least 75%, of cells having indels at or near the target site of the gRNA molecule described herein (e.g., frameshift indels, or any of the top five indels generated by a given gRNA / CRISPR system, e.g., as described herein), but less than 5%, e.g., less than 4%, less than 3%, less than 2%, or less than 1%, of cells having indels at any off-target site of the gRNA molecule, e.g., cells having indels at any off-target site outside the WIZ gene region. In other embodiments, the present invention further provides a population of cells, e.g., HSPCs, as described herein, in which at least 20%, e.g., at least 30%, e.g., at least 40%, e.g., at least 50%, e.g., at least 60%, e.g., at least 70%, e.g., at least 75%, e.g., at least 80%, e.g., at least 90%, e.g., at least 95%, are cells having indels within the WIZ gene region (e.g., in a sequence at least 90% homologous to the target sequence of the gRNA, or in its vicinity), but less than 5%, e.g., less than 4%, less than 3%, less than 2%, or less than 1%, are cells having indels in any off-target site outside the WIZ gene region or in its vicinity. In embodiments, the off-target indels are formed within the range of the gene sequence, e.g., within the range of the gene coding sequence. In embodiments, off-target indels are not formed within the range of the gene sequence, e.g., within the range of the gene coding sequence, in the target cells, e.g., as described herein.

[0366] IX. Delivery / Construction The components, such as a Cas9 molecule or a gRNA molecule, or both, can be delivered, formulated, or administered in various forms. In non-limiting examples, the gRNA molecule and the Cas9 molecule can be formulated (in one or more compositions) and delivered or administered directly to cells in which a genome editing event is desired. Alternatively, one or more components, such as nucleic acids encoding a Cas9 molecule or a gRNA molecule, or both, can be formulated (in one or more compositions) and delivered or administered. In one embodiment, the gRNA molecule is provided as DNA encoding the gRNA molecule, and the Cas9 molecule is provided as DNA encoding the Cas9 molecule. In one embodiment, the gRNA molecule and the Cas9 molecule are encoded on separate nucleic acid molecules. In one embodiment, the gRNA molecule and the Cas9 molecule are encoded on the same nucleic acid molecule. In one embodiment, the gRNA molecule is provided as RNA, and the Cas9 molecule is provided as DNA encoding the Cas9 molecule. In one embodiment, the gRNA molecule is provided with one or more modifications, such as those described herein. In one embodiment, the gRNA molecule is provided as RNA, and the Cas9 molecule is provided as mRNA encoding the Cas9 molecule. In one embodiment, the gRNA molecule is provided as RNA, and the Cas9 molecule is provided as a protein. In one embodiment, the gRNA and Cas9 molecules are provided as a ribonucleoprotein complex (RNP). In one embodiment, the gRNA molecule is provided as DNA encoding the gRNA molecule, and the Cas9 molecule is provided as a protein.

[0367] Delivery, for example, of RNP (to HSPC cells as described herein), can be achieved, for example, by electroporation (as known in the art) or by other means of making the cell membrane permeable to nucleic acid and / or polypeptide molecules. In embodiments, a CRISPR system, for example, of RNP as described herein, is delivered by electroporation using a 4D-Nucleofector (Lonza), for example, using the 4D-Nucleofector (Lonza) program CM-137. In embodiments, a CRISPR system, for example, of RNP as described herein, is delivered by electroporation using voltages of approximately 800 volts to approximately 2000 volts, for example, approximately 1000 volts to approximately 1800 volts, for example, approximately 1200 volts to approximately 1800 volts, for example, approximately 1400 volts to approximately 1800 volts, for example, approximately 1600 volts to approximately 1800 volts, for example, approximately 1700 volts. In embodiments, the pulse width / pulse length is approximately 10 ms to approximately 50 ms, for example, approximately 10 ms to approximately 40 ms, for example, approximately 10 ms to approximately 30 ms, for example, approximately 15 ms to approximately 25 ms, for example, approximately 20 ms, for example, 20 ms. In embodiments, one, two, three, four, five, or more pulses are used, for example, two, for example, one pulse. In one embodiment, a CRISPR system, for example, RNP as described herein, is delivered by electroporation using a single pulse with a voltage of approximately 1700 volts (e.g., 1700 volts) and a pulse width of approximately 20 ms (e.g., 20 ms). In embodiments, electroporation is achieved using a Neon electroporator.Further techniques for achieving membrane permeability are known in the art and include, for example, cell squeezing (as described in, for example, International Publication No. 2015 / 023982 and International Publication No. 2013 / 059343, the contents of which are incorporated herein by reference in their entirety), nanoneedles (as described in, for example, Chiappini et al., Nat. Mat., 14; 532-39, or U.S. Patent Application Publication No. 2014 / 0295558, the contents of which are incorporated herein by reference in their entirety), and nanostraws (as described in, for example, Xie, ACS Nano, 7(5); 4351-58, the contents of which are incorporated herein by reference in their entirety).

[0368] When a component is encoded and delivered via DNA, the DNA typically includes a regulatory region, such as a promoter, to induce expression. Useful promoters for the Cas9 molecular sequence include the CMV, EF-1α, MSCV, PGK, and CAG regulatory promoters. Useful promoters for gRNA include the H1, EF-1a, and U6 promoters. Similar or different intensities of promoters can be selected to modulate component expression. The Cas9 molecular sequence may include a nuclear localization signal (NLS), such as the SV40 NLS. In some embodiments, the promoter for the Cas9 molecule or the gRNA molecule may be independently inducible, tissue-specific, or cell-specific.

[0369] DNA-based delivery of Cas9 molecules and / or gRNA molecules DNA encoding the Cas9 molecule and / or gRNA molecule can be administered to a subject or delivered to cells by methods known in the art or as described herein. For example, Cas9-coding DNA and / or gRNA-coding DNA can be delivered, for example, by a vector (e.g., a viral or non-viral vector), by a non-vector-based method (e.g., using naked DNA or a DNA complex), or a combination thereof.

[0370] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by a vector (e.g., a viral vector / virus, plasmid, minicircle, or nanoplasmid).

[0371] The vector may contain sequences encoding Cas9 molecules and / or gRNA molecules. The vector may also contain sequences encoding signal peptides (e.g., for nuclear, nucleolar, or mitochondrial translocation) fused to the Cas9 molecular sequence, for example. For example, the vector may contain one or more nuclear translocation sequences (e.g., derived from SV40) fused to the Cas9 molecular sequence.

[0372] The vector may include one or more regulatory elements, such as a promoter, enhancer, intron, polyadenylation signal, Kozak consensus sequence, intrasequence ribosome entry site (IRES), 2A sequence, and splice acceptor or donor. In some embodiments, the promoter is recognized by RNA polymerase II (e.g., CMV promoter). In other embodiments, the promoter is recognized by RNA polymerase III (e.g., U6 promoter). In some embodiments, the promoter is a regulatory promoter (e.g., inductive promoter). In other embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is a tissue-specific promoter. In some embodiments, the promoter is a viral promoter. In other embodiments, the promoter is a non-viral promoter.

[0373] In some embodiments, the vector or delivery medium is a minicircle. In some embodiments, the vector or delivery medium is a nanoplasmid.

[0374] In some embodiments, the vector or delivery medium is a viral vector (e.g., for the creation of recombinant viruses). In some embodiments, the virus is a DNA virus (e.g., a dsDNA or ssDNA virus). In other embodiments, the virus is an RNA virus (e.g., an ssRNA virus).

[0375] Examples of viral vectors / viruses include, for example, retroviruses, lentiviruses, adenoviruses, adeno-associated viruses (AAVs), vaccinia viruses, poxviruses, and herpes simplex viruses. Viral vector technology is well known in the art and is described, for example, in Sambrook et al., 2012, MOLECULAR CLONING: A LABORATORY MANUAL, volumes 1-4, Cold Spring Harbor Press, NY), as well as in other virology and molecular biology manuals.

[0376] In some embodiments, the virus infects dividing cells. In other embodiments, the virus infects non-dividing cells. In some embodiments, the virus infects both dividing and non-dividing cells. In some embodiments, the virus can be integrated into the host genome. In some embodiments, the virus is engineered to induce immunosuppression, for example, in humans. In some embodiments, the virus is replication competent. In other embodiments, the virus is replication deficient, for example, in which one or more coding regions of genes necessary for further virion replication and / or packaging are replaced by or deleted from other genes. In some embodiments, the virus induces transient expression of Cas9 molecules and / or gRNA molecules. In other embodiments, the virus induces persistent expression of Cas9 molecules and / or gRNA molecules, for example, for at least one week, two weeks, one month, two months, three months, six months, nine months, one year, two years, or permanently. The packaging capacity of a virus can vary, for example, from at least about 4kb to at least about 30kb, and may be at least about 5kb, 10kb, 15kb, 20kb, 25kb, 30kb, 35kb, 40kb, 45kb, or 50kb.

[0377] In some embodiments, Cas9 and / or gRNA coding DNA is delivered by a recombinant retrovirus. In some embodiments, the retrovirus (e.g., Moloney mouse leukemia virus) includes, for example, a reverse transcriptase that enables integration into the host genome. In some embodiments, the retrovirus is replication competent. In other embodiments, the retrovirus is replication deficient, for example, in which one or more coding regions of genes necessary for further virion replication and packaging are replaced by or deleted from other genes.

[0378] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by a recombinant lentivirus. For example, the lentivirus is replication-deficient and, for example, lacks one or more genes necessary for viral replication.

[0379] In some embodiments, Cas9 and / or gRNA-coding DNA is delivered by recombinant adenovirus. In some embodiments, the adenovirus is engineered to induce immunosuppression in humans.

[0380] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by recombinant AAV. In some embodiments, the AAV can integrate its genome into the genome of a host cell, for example, a target cell as described herein. In some embodiments, the AAV is a self-complementary adeno-associated virus (scAAV), for example, an scAAV that packages both strands together to form a double-stranded DNA when annealed as a whole. AAV serotypes that can be used in the methods of this disclosure include, for example, AAV1, AAV2, modified AAV2 (e.g., modified in Y444F, Y500F, Y730F and / or S662V), AAV3, modified AAV3 (e.g., modified in Y705F, Y731F and / or T492V), AAV4, AAV5, AAV6, modified AAV6 (e.g., modified in S663V and / or T492V), AAV8, AAV8.2, AAV9, and AAV rh 10. Pseudotype AAVs such as AAV2 / 8, AAV2 / 5, and AAV2 / 6 can also be used in the methods of this disclosure.

[0381] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by a hybrid virus, for example, a hybrid of one or more viruses described herein.

[0382] Packaging cells are used to form viral particles capable of infecting host or target cells. Examples of such cells include 293 cells, which can package adenoviruses, and ψ2 or PA317 cells, which can package retroviruses. Viral vectors used in gene therapy are typically created by producer cell lines that package nucleic acid vectors into viral particles. The vectors typically contain the minimum viral sequences necessary for packaging and subsequent integration into host or target cells (if applicable), with other viral sequences replaced by expression cassettes encoding the proteins to be expressed. For example, AAV vectors used in gene therapy typically contain only the reverse-ended repeat (ITR) sequences derived from the AAV genome necessary for packaging and gene expression in host or target cells. Missing viral function is trans-conjugated by the packaging cell line. The viral DNA is then packaged into the cell line, which contains helper plasmids encoding other AAV genes, namely rep and cap, but lacking ITR sequences. The cell line is also infected with adenovirus as a helper. Helper viruses promote the replication of AAV vectors from helper plasmids and the expression of AAV genes. Because helper plasmids lack ITR sequences, they are not packaged in large quantities. Adenovirus contamination can be reduced, for example, by heat treatment, which is more susceptible to adenoviruses than to AAV.

[0383] In some embodiments, the viral vector has cell type and / or tissue type recognition capabilities. For example, the viral vector may be pseudotype, having different / alternative viral envelope glycoproteins, being engineered with cell type-specific receptors (e.g., genetic modification of the viral envelope glycoprotein to take up target ligands such as peptide ligands, single-chain antibodies, and growth factors), and / or being engineered to have a bispecific molecular crosslink including one end that recognizes the viral glycoprotein and the other end that recognizes a portion of the target cell surface (e.g., ligand-receptor, monoclonal antibody, avidin-biotin, and chemical conjugation).

[0384] In some embodiments, viral vectors achieve cell-type specific expression. For example, tissue-specific promoters can be constructed to restrict the expression of transgenes (Cas9 and gRNA) only in target cells. Vector specificity can also be mediated by microRNA-dependent regulation of transgene expression. In some embodiments, viral vectors have high fusion efficiency with target cell membranes. For example, fusion proteins such as fusion-competent hemagglutinin (HA) can be incorporated to increase viral uptake into cells. In some embodiments, viral vectors have nuclear translocation capabilities. For example, viruses that do not infect non-dividing cells, requiring cell wall disruption (during cell division), can be modified to incorporate nuclear-translocation peptides into the viral matrix proteins, thereby enabling transduction of non-proliferating cells.

[0385] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by non-vector-based methods (e.g., using naked DNA or DNA complexes). For example, DNA can be delivered by, for instance, organically modified silica or silicates (Ormosil), electroporation, gene guns, sonoporation, magnetofection, lipid-mediated transfection, dendrimers, inorganic nanoparticles, calcium phosphate, or a combination thereof.

[0386] In some embodiments, Cas9 and / or gRNA-coding DNA are delivered by a combination of vector and non-vector-based methods. For example, virosomes may include liposomes combined with an inactivated virus (e.g., HIV or influenza virus), which may result in more efficient gene transfer, for example, in respiratory epithelial cells, compared to either the viral method or the liposome method alone.

[0387] In one embodiment, the delivery medium is a non-viral vector. In one embodiment, the non-viral vector is an inorganic nanoparticle (e.g., the payload is attached to the surface of the nanoparticle). Exemplary inorganic nanoparticles include, for example, magnetic nanoparticles (e.g., Fe lvlnO2) or silica. The outer surface of the nanoparticle can be conjugated with a positively charged polymer (e.g., polyethyleneimine, polylysine, polyserine), thereby enabling the attachment (e.g., conjugation or capture) of the payload. In one embodiment, the non-viral vector is an organic nanoparticle (e.g., capture of the payload into the interior of the nanoparticle). Exemplary organic nanoparticles include, for example, SNALP liposomes containing cationic lipids together with neutral helper lipids coated with polyethylene glycol (PEG) and protamine, and nucleic acid complexes coated with a lipid coating.

[0388] Examples of nucleic acids encoding CRISPR systems or components thereof, such as exemplary lipids and / or polymers for vector transfer, include those described in International Publication No. 2011 / 076807, International Publication No. 2014 / 136086, International Publication No. 2005 / 060697, International Publication No. 2014 / 140211, International Publication No. 2012 / 031046, International Publication No. 2013 / 103467, International Publication No. 2013 / 006825, International Publication No. 2012 / 006378, International Publication No. 2015 / 095340, and International Publication No. 2015 / 095346 (each of which is incorporated herein by reference in its entirety). In one embodiment, the medium has targeting modifications to increase the updating of nanoparticles and liposomes by target cells, such as cell-specific antigens, monoclonal antibodies, single-chain antibodies, aptamers, polymers, sugars, and cell-permeable peptides. In one embodiment, the medium uses fusionable and endosomal destabilizing peptides / polymers. In one embodiment, the medium undergoes conformational changes induced by acid (e.g., thereby accelerating cargo endosomal escape). In one embodiment, a polymer that can be cleaved by stimulation is used, for example, for release in an intracellular compartment. For example, a disulfide-based cationic polymer that is cleaved in a reducing cellular environment can be used.

[0389] In some embodiments, the delivery medium is a biological nonviral delivery medium. In some embodiments, the medium is attenuated bacteria (e.g., invasive but attenuated and naturally or artificially engineered to prevent disease development and transgene expression (e.g., Listeria monocytogenes, certain Salmonella strains, Bifidobacterium longum, and modified Escherichia coli), bacteria that target specific tissues with trophotropy and tissue-specific tropism, and bacteria that modify target tissue specificity with modified surface proteins). In some embodiments, the medium is a genetically modified bacteriophage (e.g., an engineered phage with high packaging capacity, low immunogenicity, containing a mammalian plasmid maintenance sequence, and having an integrated target ligand). In some embodiments, the medium is a mammalian virus-like particle. For example, modified virus particles can be created (e.g., by purifying "empty" particles and then assembling the virus with the desired cargo in exovivo). The medium can also be manipulated to incorporate target ligands and modify target tissue specificity. In one embodiment, the medium is a biological liposome. For example, the biological liposome is a phospholipid-based particle derived from human cells (e.g., erythrocyte ghosts, which are erythrocytes that are broken down into globular structures derived from the target (e.g., tissue targeting can be achieved by adding various tissue or cell-specific ligands)), or a target (i.e., patient) membrane-bound nanovesicle (30-100 nm) of secretory exosome-endocytosis origin (e.g., can be made from various cell types and thus can be taken up by cells without the need for target ligands).

[0390] In some embodiments, one or more nucleic acid molecules (e.g., DNA molecules) other than the components of the Cas system, e.g., the Cas9 molecular component and / or gRNA molecular component described herein, are delivered. In some embodiments, the nucleic acid molecules are delivered at the same time as the delivery of one or more components of the Cas system. In some embodiments, the nucleic acid molecules are delivered before or after the delivery of one or more components of the Cas9 system (e.g., about 30 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 9 hours, 12 hours, 1 day, 2 days, 3 days, 1 week, 2 weeks, or less than 4 weeks). In some embodiments, the nucleic acid molecules are delivered by means different from that used to deliver one or more components of the Cas9 system, e.g., the Cas9 molecular component and / or gRNA molecular component. The nucleic acid molecules can be delivered by any of the delivery methods described herein. For example, the nucleic acid molecules can be delivered by a viral vector, e.g., an embedded-deficient lentivirus, and the Cas9 molecular component and / or gRNA molecular component can be delivered by electroporation, which can reduce the toxicity caused by nucleic acids (e.g., DNA), for example. In one embodiment, the nucleic acid molecule encodes a therapeutic protein, such as a protein described herein. In another embodiment, the nucleic acid molecule encodes an RNA molecule, such as an RNA molecule described herein.

[0391] Delivery of RNA encoding the Cas9 molecule RNA encoding Cas9 molecules (e.g., active Cas9 molecules, inactive Cas9 molecules, or inactive Cas9 fusion proteins) and / or gRNA molecules can be delivered to cells, such as target cells as described herein, by methods known in the art or as described herein. For example, Cas9-coding RNA and / or gRNA-coding RNA can be delivered by microinjection, electroporation, lipid-mediated transfection, peptide-mediated delivery, or a combination thereof.

[0392] Delivery of the Cas9 molecule as a protein Cas9 molecules (e.g., active Cas9 molecules, inactive Cas9 molecules, or inactive Cas9 fusion proteins) can be delivered to cells by methods known in the art or as described herein. For example, Cas9 protein molecules can be delivered by microinjection, electroporation, lipid-mediated transfection, peptide-mediated delivery, cell squeezing or ablation (e.g., by nanoneedles) or a combination thereof. Delivery can be achieved by DNA encoding gRNA or by gRNA, for example, by pre-complexing gRNA and Cas9 protein into a ribonucleoprotein complex (RNP).

[0393] In one embodiment, the Cas9 molecule is delivered as a protein, for example, as described herein, and the gRNA molecule is delivered as one or more RNAs (for example, as dgRNA or sgRNA, as described herein). In an embodiment, the Cas9 protein is complexed with the gRNA molecule as a ribonucleoprotein complex ("RNP") before delivery to cells, for example, as described herein. In an embodiment, the RNP can be delivered to cells, for example, as described herein, by any method known in the art, for example, electroporation. As described herein, and not constrained by theory, it is preferable to use gRNA and Cas9 molecules that result in a high % editing rate (e.g., >85%, >90%, >95%, >98%, or >99%) of the target sequence in target cells, for example, as described herein, even if the concentration of RNP delivered to the cells is reduced. In this case, although not constrained by theory, delivering RNP containing gRNA molecules that produce a high % editing rate of the target sequence in target cells (including at low RNP concentrations) at reduced or low concentrations may be beneficial because it can reduce the frequency and number of off-target editing events. In one embodiment, when using low or reduced concentrations of RNP, RNP containing dgRNA molecules can be prepared using the following exemplary procedure: 1. Provide Cas9 molecules and tracr in solution at high concentrations (e.g., higher than the final RNP concentration delivered to the cells) and equilibrate these two components; 2. Provide crRNA molecules and equilibrate the components (thus forming a high-concentration RNP solution); 3. Dilute the RNP solution to the desired concentration; 4. The RNP at the desired concentration is delivered to the target cells, for example, by electroporation.

[0394] The above procedure can be modified to use sgRNA molecules by omitting step 2 above and equilibrating the components in step 1 by providing high concentrations of Cas9 molecules and sgRNA molecules in solution. In embodiments, Cas9 molecules and each gRNA component are provided in solution in a 1:2 ratio (Cas9:gRNA), for example, a 1:2 molar ratio of Cas9:gRNA molecules. When dgRNA molecules are used, the ratio, for example, the molar ratio is 1:2:2 (Cas9:tracr:crRNA). In embodiments, RNP is formed at a concentration of 20 μM or higher, for example, about 20 μM to about 50 μM. In embodiments, RNP is formed at a concentration of 10 μM or higher, for example, about 10 μM to about 30 μM. In embodiments, RNP is diluted to a final concentration of 10 μM or less (for example, about 0.01 μM to about 10 μM) in a delivery solution to the target cells (for example, those described herein), which contain the target cells. In an embodiment, RNP is diluted to a final concentration of 3 μM or less (e.g., a concentration of about 0.01 μM to about 3 μM) in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is diluted to a final concentration of 1 μM or less (e.g., a concentration of about 0.01 μM to about 1 μM) in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is diluted to a final concentration of 0.3 μM or less (e.g., a concentration of about 0.01 μM to about 0.3 μM) in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is provided at a final concentration of about 3 μM in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is provided at a final concentration of about 2 μM in a solution containing target cells (e.g., those described herein) for delivery to the target cells. In one embodiment, RNP is provided at a final concentration of about 1 μM in a delivery solution to the target cells (e.g., those described herein). In another embodiment, RNP is provided at a final concentration of about 0.3 μM in a delivery solution to the target cells (e.g., those described herein).In an embodiment, RNP is provided at a final concentration of about 0.1 μM in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is provided at a final concentration of about 0.05 μM in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is provided at a final concentration of about 0.03 μM in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is provided at a final concentration of about 0.01 μM in a delivery solution to the target cells (e.g., those described herein). In an embodiment, RNP is formulated in a medium suitable for electroporation. In an embodiment, RNP is delivered by electroporation to cells, for example, HSPC cells, as described herein, using, for example, electroporation conditions described herein.

[0395] In embodiments, components of a gene editing system (e.g., a CRISPR system) and / or nucleic acids encoding one or more components of a gene editing system (e.g., a CRISPR system) are introduced into cells by mechanically perturbing the cells, for example by passing the cells through pores or channels and constricting them. Such perturbation may be achieved in a solution containing components of a gene editing system (e.g., a CRISPR system) and / or nucleic acids encoding one or more components of a gene editing system (e.g., a CRISPR system), as described herein, for example. In embodiments, perturbation is achieved using a TRIAMF system, as described herein, for example, in the examples and in PCT patent application PCT / US17 / 54110 (which is incorporated herein by reference in whole).

[0396] Bimodal or differential delivery of components Delivering components of the Cas system, such as the Cas9 molecular component and the gRNA molecular component, separately, or more specifically, delivering the components in different ways, can enhance performance by improving, for example, tissue specificity and safety.

[0397] In some embodiments, Cas9 molecules and gRNA molecules are delivered in different manners, or as sometimes referred to herein as differential manners. Different manners or differential manners, as used herein, refer to delivery manners that impart different pharmacodynamic or pharmacokinetic properties to the target component molecules, e.g., Cas9 molecules, gRNA molecules, or template nucleic acids. For example, a delivery manner may result in different tissue distribution, different half-lives, or different temporal distributions in, for example, selected compartments, tissues, or organs.

[0398] Several delivery methods, such as delivery by nucleic acid vectors that persist in cells or their offspring, for example, through self-replication or insertion into cellular nucleic acids, result in more sustained expression and presence of the components.

[0399] X. Treatment method While not bound by theory, the present invention is in part based on unexpected findings regarding the relationship between WIZ gene expression / protein activity and hemoglobin F (HbF) production. As demonstrated in the examples and figures, knockdown or knockout of the WIZ gene or WIZ protein in cells (by various modalities / compositions described herein) significantly increased HbF induction in such cells, thereby treating HbF-related conditions and disorders (e.g., abnormal hemoglobin disorders, such as sickle cell anemia and β-thalassemia).

[0400] The Cas9 system described herein, for example, one or more gRNA molecules and one or more Cas9 molecules, is useful for treating diseases in mammals, such as humans. The terms “to treat,” “treated,” “being treated,” and “treatment” include preventing or delaying the onset of symptoms, complications, or biochemical signs of a disease, or reducing the symptoms of a disease, condition, or disorder, or preventing or inhibiting its further development, by administering the Cas9 system, for example, one or more gRNA molecules and one or more Cas9 molecules, to cells. Treatment may also include preventing or delaying the onset of symptoms, complications, or biochemical signs of a disease, or reducing the symptoms of a disease, condition, or disorder, or preventing or inhibiting its further development, by introducing the gRNA molecules (or two or more gRNA molecules) of the present invention, or by introducing the CRISPR system as described herein, or by administering one or more cells, for example, HSPCs (e.g., a population thereof), modified by any of the cell preparation methods described herein. Treatment may be preventive (thus preventing the disease, delaying its onset, or preventing the manifestation of its clinical or quasi-clinical symptoms), or it may be therapeutic suppression or mitigation of symptoms after the onset of the disease. Treatment can be measured by the therapeutic means described herein. Accordingly, the “treatment” methods of the present invention include curing, reducing the severity of, or improving one or more symptoms of a disease or condition, or extending the health or survival of the subject beyond what would be expected in the absence of such treatment, by administering cells modified by the introduction of a Cas9 system (e.g., one or more gRNA molecules and one or more Cas9 molecules) to the subject. For example, “treatment” includes a reduction of at least 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or more of the symptoms of the disease in question.

[0401] The Cas9 system, including gRNA molecules containing targeting domains as described herein (for example, Table 1), as well as the methods and cells (for example, as described herein), are useful for the treatment of hemoglobin disorders.

[0402] Delivery timing In one embodiment, one or more nucleic acid molecules (e.g., DNA molecules) other than the components of the Cas system, e.g., the Cas9 molecular component and / or gRNA molecular component described herein, are delivered. In one embodiment, the nucleic acid molecules are delivered at the same time as the delivery of one or more components of the Cas system. In one embodiment, the nucleic acid molecules are delivered before or after the delivery of one or more components of the Cas system (e.g., about 30 minutes, 1 hour, 2 hours, 3 hours, 6 hours, 9 hours, 12 hours, 1 day, 2 days, 3 days, 1 week, 2 weeks, or less than 4 weeks). In one embodiment, the nucleic acid molecules are delivered by means different from that used to deliver one or more components of the Cas system, e.g., the Cas9 molecular component and / or gRNA molecular component. The nucleic acid molecules can be delivered by any of the delivery methods described herein. For example, nucleic acid molecules can be delivered by a viral vector, such as an embedded-deficient lentivirus, and Cas9 molecular components and / or gRNA molecular components can be delivered by electroporation, which may reduce the toxicity caused by nucleic acids (e.g., DNA). In one embodiment, the nucleic acid molecule encodes a therapeutic protein, such as the protein described herein. In another embodiment, the nucleic acid molecule encodes an RNA molecule, such as the RNA molecule described herein.

[0403] Bimodal or differential delivery of components Delivering components of the Cas system, such as the Cas9 molecular component and the gRNA molecular component, separately, or more specifically, delivering the components in different ways, can enhance performance, for example, by improving tissue specificity and safety. In some embodiments, the Cas9 molecule and the gRNA molecule are delivered in different modes, or as sometimes referred to herein as differential modes. Different modes or differential modes, as used herein, refer to delivery modes that impart different pharmacodynamic or pharmacokinetic properties to the component molecules of interest, such as the Cas9 molecule, the gRNA molecule, the template nucleic acid, or the payload. For example, a delivery mode may result in different tissue distribution, different half-lives, or different temporal distributions in, for example, a selected compartment, tissue, or organ.

[0404] Certain delivery methods, such as delivery by nucleic acid vectors that persist in cells or their offspring, for example, by self-replication or insertion into cellular nucleic acids, result in increased persistence of component expression and presence. Examples include viral delivery, such as adeno-associated virus or lentiviral delivery.

[0405] For example, components, such as Cas9 molecules and gRNA molecules, can be delivered in different ways in terms of the half-life or persistence resulting from the delivered components in a living organism or in a specific compartment, tissue, or organ. In one embodiment, gRNA molecules can be delivered in this way. Cas9 molecular components can be delivered in a way that results in lower persistence or lower exposure to the living organism or a specific compartment, tissue, or organ.

[0406] More generally, in one embodiment, a first component is delivered using a first delivery method, and a second component is delivered using a second delivery method. The first delivery method imparts a first pharmacodynamic or pharmacokinetic property. The first pharmacodynamic property may be, for example, the distribution, persistence, or exposure of the component, or the nucleic acid encoding the component, in a living organism, compartment, tissue, or organ. The second delivery method imparts a second pharmacodynamic or pharmacokinetic property. The second pharmacodynamic property may be, for example, the distribution, persistence, or exposure of the component, or the nucleic acid encoding the component, in a living organism, compartment, tissue, or organ.

[0407] In one embodiment, the first pharmacodynamic or pharmacokinetic properties, such as distribution, duration, or exposure, are limited compared to the second pharmacodynamic or pharmacokinetic properties.

[0408] In one embodiment, the first delivery mode is selected to optimize pharmacodynamic or pharmacokinetic properties, such as distribution, duration, or exposure, for example, by minimizing them.

[0409] In one embodiment, the second delivery mode is selected to optimize, for example, maximize, pharmacokinetic properties such as distribution, duration, or exposure.

[0410] In one embodiment, the first delivery mode involves the use of a relatively persistent element, such as a nucleic acid, such as a plasmid, or a viral vector, such as AAV or a lentivirus. Because such vectors are relatively persistent, the products transcribed from them can also be relatively persistent.

[0411] In one embodiment, the second delivery mode includes a relatively transient element, such as RNA or a protein.

[0412] In one embodiment, the first component comprises gRNA, and the delivery method is relatively persistent; for example, the gRNA is transcribed from a plasmid or viral vector, such as AAV or lentivirus. In such gene transcription, the gene does not encode a protein product, and the gRNA cannot function independently, so it is thought to have little physiological effect. The second component, the Cas9 molecule, is delivered transiently, for example, as mRNA or as a protein, ensuring that a complete Cas9 molecule / gRNA molecule complex exists for only a short time and is active.

[0413] Furthermore, the components can be delivered in complementary molecular forms or different delivery vectors to enhance safety and tissue specificity.

[0414] Using differential delivery methods can enhance performance, safety, and efficacy. For example, it can reduce the possibility of accidental off-target modifications. Since peptides from Cas enzymes delivered by bacteria are presented on the cell surface by MHC molecules, immunogenic components, such as the Cas9 molecule, may have reduced immunogenicity if delivered in a less persistent manner. A two-part delivery system can mitigate these drawbacks.

[0415] A differential delivery mode allows components to be delivered to different but overlapping target regions. The formation of active complexes outside the overlap of target regions is minimized. Thus, in some embodiments, a first component, e.g., a gRNA molecule, is delivered by a first delivery mode that results in a first spatial distribution, e.g., a tissue distribution. A second component, e.g., a Cas9 molecule, is delivered by a second delivery mode that results in a second spatial distribution, e.g., a tissue distribution. In some embodiments, the first mode includes a first element selected from liposomes, nanoparticles, e.g., polymer nanoparticles, and nucleic acids, e.g., viral vectors. The second mode includes a second element selected from this group. In some embodiments, the first delivery mode includes a first targeting element, e.g., a cell-specific receptor or antibody, and the second delivery mode does not include that element. In some embodiments, the second delivery mode includes a second targeting element, e.g., a second cell-specific receptor or secondary antibody.

[0416] When Cas9 molecules are delivered via viral delivery vectors, liposomes, or polymer nanoparticles, delivery to multiple tissues and subsequent therapeutic activity may occur, even when targeting only a single tissue is desired. A two-part delivery system can overcome this challenge and enhance tissue specificity. When gRNA molecules and Cas9 molecules are packaged in separate delivery media with different but overlapping tissue tropisms, a fully functional complex is formed only in tissues targeted by both vectors.

[0417] Candidate Cas molecules, such as the Cas9 molecule, candidate gRNA molecules, candidate Cas9 / gRNA molecule complexes, and candidate CRISPR systems can be evaluated by methods known in the art or as described herein. For example, an exemplary method for evaluating the endonuclease activity of the Cas9 molecule is described, for example, in Jinek el al., SCIENCE 2012;337(6096):8 16-821.

[0418] Abnormal hemoglobinosis Hemoglobin disorders include several types of anemia with genetic causes that involve decreased production and / or increased destruction (hemolysis) of red blood cells (RBCs). These include genetic defects that result in the production of abnormal hemoglobin accompanied by a reduced ability to maintain oxygen levels. Some of these disorders involve the inability to produce sufficient amounts of normal β-globin, while others involve the complete inability to produce normal β-globin. These disorders related to the β-globin protein are generally referred to as β-hemoglobin disorders. For example, β-thalassemia is caused by a partial or complete deficiency in the expression of the β-globin gene, which leads to HbA deficiency or absence. Sickle cell anemia is caused by a point mutation in the β-globin structural gene, which leads to the production of abnormal (sickle-shaped) hemoglobin (HbS). HbS polymerizes particularly easily under deoxygenated conditions. HbS RBCs are more fragile and prone to hemolysis than normal RBCs, ultimately leading to anemia.

[0419] In some embodiments, Cas9 molecules and gRNA molecules described herein are used to target genes associated with abnormal hemoglobinopathy. Exemplary targets include, for example, genes involved in the regulation of the γ-globin gene. In some embodiments, the target is an undeleted HPFH region.

[0420] Fetal hemoglobin (also known as hemoglobin F, HbF, or α2γ2) is a tetramer of two adul...

Claims

1. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, wherein the composition comprises a guide RNA (gRNA) molecule, the gRNA molecule comprises tracr and crRNA, and the crRNA comprises a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers.

2. The pharmaceutical composition according to claim 1, wherein the targeting domain includes one of SEQ ID NOs: 1 to 3106.

3. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, The composition is (a) gRNA molecule and Cas9 molecule, (b) A nucleic acid containing a gRNA molecule and a nucleotide sequence encoding a Cas9 molecule, (c) A nucleic acid containing a nucleotide sequence encoding a gRNA molecule, and a Cas9 molecule, or (d) A nucleic acid containing a nucleotide sequence encoding a gRNA molecule and a nucleic acid containing a nucleotide sequence encoding a Cas9 molecule Includes, The gRNA molecule comprises tracr and crRNA, wherein the crRNA contains a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers. The aforementioned pharmaceutical composition.

4. The pharmaceutical composition according to claim 3, wherein the targeting domain includes one of SEQ ID NOs: 1 to 3106.

5. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, wherein the composition comprises a nucleic acid, the nucleic acid encoding a guide RNA (gRNA) molecule, the gRNA molecule comprising tracr and crRNA, and the crRNA comprising a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers.

6. The pharmaceutical composition according to claim 5, wherein the targeting domain includes one of Sequence IDs 1 to 3106.

7. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, wherein the composition comprises a vector, the vector comprises a nucleic acid encoding a guide RNA (gRNA) molecule, the gRNA molecule comprises tracr and crRNA, and the crRNA comprises a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers.

8. The pharmaceutical composition according to claim 7, wherein the targeting domain includes one of SEQ ID NOs: 1 to 3106.

9. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, The composition comprises cells, The cells are modified by a method of modifying a target sequence within the cells or in its vicinity. The method described above involves the cells and (a) gRNA molecule and Cas9 molecule, (b) Nucleic acids containing nucleotide sequences encoding gRNA molecules and Cas9 molecules, (c) A nucleic acid containing a nucleotide sequence encoding a gRNA molecule, and a Cas9 molecule, (d) Nucleic acids containing nucleotide sequences encoding gRNA molecules, and nucleic acids containing nucleotide sequences encoding Cas9 molecules This includes bringing the two into contact. The gRNA molecule comprises tracr and crRNA, wherein the crRNA contains a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers. The aforementioned pharmaceutical composition.

10. The pharmaceutical composition according to claim 9, wherein the targeting domain includes one of SEQ ID NOs: 1 to 3106.

11. A pharmaceutical composition for use in the treatment of abnormal hemoglobin disorders, wherein the composition comprises cells, the cells comprising a guide RNA (gRNA) molecule or a nucleic acid encoding a gRNA molecule, the gRNA molecule comprising tracr and crRNA, and the crRNA comprising a targeting domain complementary to the target sequence of a protein (WIZ) gene containing widely spaced zinc fingers.

12. The pharmaceutical composition according to claim 11, wherein the targeting domain includes one of sequence numbers 1 to 3106.

13. The pharmaceutical composition according to any one of claims 1, 3, 5, 7, 9, and 11, wherein the pharmaceutical composition reduces WIZ gene expression and / or WIZ protein activity.

14. The pharmaceutical composition according to any one of claims 2, 4, 6, 8, 10, and 12, wherein the targeted domain includes one of sequence number 1488, sequence number 1565, sequence number 2801, sequence number 2809, and sequence number 3071.

15. The aforementioned gRNA molecule (a) Sequence ID 3123; (b) Sequence ID 3159; or (c) Either (a) or (b) above, further comprising 1, 2, 3, 4, 5, 6, or 7 uracil (U) nucleotides at the 3' end. The sequence includes; any of the sequences (a) to (c) is positioned on the 3' side of the targeting domain. A pharmaceutical composition according to any one of claims 1 to 14.

16. The aforementioned gRNA molecule (a) Tracr containing Sequence ID No. 3152; or (b) Tracr containing Sequence ID No. 3109 or 3174 including, A pharmaceutical composition according to any one of claims 1 to 14.

17. The pharmaceutical composition according to any one of claims 3, 4, 9, and 10, wherein the Cas9 molecule is active or inactivated Streptococcus pyogenes Cas9.

18. The Cas9 molecule, (a) Sequence ID 3161; (b) Sequence ID 3162; (c) Sequence ID 3163; (d) Sequence ID 3164; (e) Sequence ID 3165; (f) Sequence ID 3166; (g) Sequence ID 3167; (h) Sequence ID 3168; (i) Sequence ID 3169; (j) Sequence ID 3170; (k) Sequence ID 3171; or (l) Sequence ID 3172 A pharmaceutical composition according to any one of claims 3, 4, 9, and 10, comprising the above.

19. The pharmaceutical composition according to claim 7 or 8, wherein the vector is selected from the group consisting of lentiviral vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, herpes simplex virus (HSV) vectors, plasmids, minicircles, nanoplasmides, and RNA vectors.

20. The pharmaceutical composition according to any one of claims 9 to 12, wherein the cells are HSPCs.