CRISPR / CAS-related methods and compositions for treating β-abnormal hemoglobinopathy
CRISPR/Cas-mediated genome editing increases gamma-globin gene expression to treat SCD and beta-thalassemia by altering regulatory elements, improving fetal hemoglobin levels and mitigating disease severity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- EDITAS MEDICINE INC
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-14
AI Technical Summary
Current therapies for sickle cell disease (SCD) and beta-thalassemia, such as gene therapy and hematopoietic stem cell transplantation, face challenges with long-term efficacy and safety, and the identification of compatible donors is difficult, necessitating improved methods for managing these hemoglobin disorders.
CRISPR/Cas-mediated genome editing is employed to increase the expression of gamma-globin genes by altering gamma-globin gene regulatory elements, using DNA repair mechanisms like NHEJ or HDR to delete, disrupt, or modify these elements, leading to spontaneous hereditary persistence of fetal hemoglobin (HPFH) mutations.
Enhances gamma-globin gene expression, potentially alleviating symptoms of SCD and beta-thalassemia by increasing fetal hemoglobin levels, reducing the severity of anemia and associated complications.
Smart Images

Figure 2026064992000020 
Figure 2026064992000021 
Figure 2026064992000022
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the benefits of U.S. Provisional Patent Application No. 62 / 308,190, filed on 14 March 2016, and U.S. Provisional Patent Application No. 62 / 456,615, filed on 8 February 2017, the contents of which are incorporated herein by reference in their entirety.
[0002] Sequence List This application includes a sequence listing submitted via EFS-Web in ASCII format, which is incorporated herein by reference in its entirety. This ASCII copy was created on 14 March 2017, named 8009WO00_SequenceListing.txt, and is 335KB in size.
[0003] The present invention relates to CRISPR / Cas-related methods and components for editing or regulating the expression of target nucleic acid sequences, as well as their applications in relation to β-abnormal hemoglobin disorders, including sickle cell disease and β-thalassemia. [Background technology]
[0004] Hemoglobin (Hb) carries oxygen from the lungs to erythrocytes or red blood cell (RBC) tissue. During prenatal development and immediately after birth, hemoglobin exists in the form of fetal hemoglobin (HbF), a tetrameric protein consisting of two alpha (α)-globin chains and two gamma (γ)-globin chains. HbF is largely replaced by adult hemoglobin (HbA), a tetrameric protein in which the γ-globin chains of HbF are replaced by beta (β)-globin chains, through a process known as globin switching. HbF is more efficient than HbA in oxygen transport. The average adult produces less than 1% HbF of total hemoglobin (Thein 2009). The α-hemoglobin gene is located on chromosome 16, and the β-hemoglobin gene (HBB), α-gamma (γ)-globin gene... A ) Globin chain (HBG1, also known as gamma globin A), and G gamma (γ G The globin chain (HBG2, also known as gamma globin G) is located on chromosome 11 within the globin gene cluster (i.e., the globin locus).
[0005] Mutations in the hemoglobin barb (HBB) can cause hemoglobin abnormalities (i.e., hemoglobin disorders), including sickle cell disease (SCD) and beta-thalassemia (β-thal). Approximately 93,000 people in the United States are diagnosed with hemoglobin disorders. Worldwide, 300,000 children are born with hemoglobin disorders each year (Angastiniotis 1998). Because these conditions are associated with HBB mutations, symptoms typically do not appear until after the globin switching from HbF to HbA.
[0006] Sickle cell disease (SCD) is the most common hereditary blood disorder in the United States, affecting approximately 80,000 people (Brousseau 2010). SCD is most common among people of African descent, with a prevalence of 1 in 500. In Africa, there are 15 million people with SCD (Aliyu 2008). SCD is also more common among people of Indian, Saudi Arabian, and Mediterranean descent. Among Hispanic Americans, the prevalence of sickle cell disease is 1 in 1,000 (Lewis 2014).
[0007] SCD is caused by a single homozygous mutation in the HBB gene, c.17A>T (HbS mutation). The sickle cell mutation is a point mutation (GAG→GTG) on the HBB, in which glutamic acid at amino acid position 6 of exon 1 is replaced with valine. The valine at position 6 of the β-hemoglobin chain is hydrophobic and causes a change in the conformation of the β-globin protein when it is not bound to oxygen. This conformational change causes polymerization of the HbS protein in the absence of oxygen, resulting in deformation of the RBC (i.e., sickle cell formation). SCD is inherited in an autosomal recessive manner, and therefore only patients with two HbS alleles will have this disease. Heterozygous subjects have the sickle cell phenotype and may suffer from anemia and / or pain attacks in cases of severe dehydration or oxygen deficiency.
[0008] Sickle red blood cells (RBCs) cause multiple symptoms, including anemia, sickle cell crisis, vascular occlusive crisis, myeloplastic crisis, and acute chest syndrome. Sickle RBCs are less elastic than wild-type RBCs and therefore cannot easily pass through the capillary bed, causing occlusion and ischemia (i.e., vascular occlusion). Vascular occlusive crisis occurs when sickle cells obstruct blood flow in the capillary bed of organs, leading to pain, ischemia, and necrosis. These episodes usually last 5–7 days. The spleen plays a role in removing dysfunctional red blood cells and is therefore usually enlarged in infancy and exposed to frequent vascular occlusive crises. By the end of childhood, the spleen of SCD patients often becomes infarcted, leading to autosplenectomy. Hemolysis is an invariant feature of SCD and causes anemia. Sickle cells survive in circulation for 10–20 days, while healthy RBCs survive for 90–120 days. SCD patients receive blood transfusions as needed to maintain adequate hemoglobin levels. Frequent transfusions expose patients to the risk of HIV, hepatitis B, and hepatitis C infection. Patients may also be at risk of acute chest crisis, as well as infarction of the limbs, peripheral organs, and central nervous system.
[0009] SCD patients have a shorter lifespan. The prognosis for SCD patients has steadily improved with careful lifelong management of disease crises and anemia. As of 2001, the average life expectancy for patients with sickle cell disease was in their mid-to-late 50s. Current SCD treatment involves hydration and pain management during disease crises, and blood transfusions as needed to improve anemia.
[0010] Thalassemia (e.g., β-Thal, δ-Thal, and β / δ-Thal) causes chronic anemia. β-Thal is estimated to affect about 1 in 100,000 people worldwide. Its prevalence is higher in certain populations, including those of European descent, where the prevalence is about 1 in 10,000. β-Thal major, a more severe form of the disease, is life-threatening unless treated with lifelong transfusions and chelation therapy. In the United States, there are approximately 3,000 cases of β-Thal major. β-Thal intermediates do not require transfusions but can cause growth retardation and significant systemic abnormalities, often requiring lifelong chelation therapy. HbA makes up the majority of hemoglobin in adult RBCs, but about 3% of adult hemoglobin is in the form of HbA2, an HbA variant in which two gamma-globin chains are replaced by two delta (Δ)-globin chains. δ-thal is associated with mutations in the δ-hemoglobin gene (HBD) that cause a deletion of HBD expression. Co-inheritance of HBD mutations can mask the diagnosis of β-thal (i.e., β / δ-thal) by lowering HbA2 levels to the normal range (Bouva 2006). β / δ-thal is usually caused by deletions of HBB and HBD sequences in both alleles. Homozygote (δ 0 / δ 0 β 0 / β 0 In these patients, HBG is expressed, and only HbF is produced.
[0011] Similar to SCD, β-Thal is caused by mutations in the HBB gene. The most common HBB mutations that result in β-Thal are: c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c.52A>T, c.59A>G, c.92+5G>C, c.124_127delTTCT, c.316-197C>T, c.-78A>G, c.52A>T, c.124_127delTTCT, c.316-197C>T, c.-138C>T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G, and c.316-2A>C. These and other mutations associated with β-Thal result in mutations or absence of the β-globin chain, disrupting the normal ratio of Hbα-hemoglobin to β-hemoglobin. Excess α-globin chains precipitate in the erythroid precursors of the bone marrow.
[0012] In β-Thal major, both alleles of HBB contain nonsense mutations, frameshift mutations, or splicing mutations that result in a complete lack of β-globin production (indicated as β 0 / β 0 ). β-Thal major results in a profound decrease in β-globin chains, significant precipitation of α-globin chains in erythroid cells, and more severe anemia.
[0013] β-Thal intermedia results from mutations in the 5' or 3' untranslated regions of HBB, mutations in the promoter region or polyadenylation signal of HBB, or splicing mutations within the HBB gene. The patient's genotype is indicated as β 0 / β + or β + / β + . β 0 represents the absence of β-globin chain expression; β +This indicates dysfunction, but β-globin chains are present. Phenotypic expression varies among patients. The β-thal intermediate reduces the precipitation of α-globin chains to erythrocyte progenitor cells because some β-globin is produced, resulting in less severe anemia than β-thal major. However, there are more serious consequences of erythrocyte lineage proliferation secondary to chronic anemia.
[0014] Patients with β-thal major between 6 months and 2 years of age suffer from stunted growth, fever, hepatosplenomegaly, and diarrhea. Appropriate treatment includes regular blood transfusions. Therapy for β-thal major also includes splenectomy and treatment with hydroxyurea. If patients receive regular transfusions, they develop normally until their early 20s. At that time, patients require chelation therapy (in addition to continuous transfusions) to prevent complications of iron overload. Iron overload can manifest as growth retardation or delayed sexual maturation. In adulthood, inadequate chelation therapy can lead to cardiomyopathy, cardiac arrhythmias, hepatic fibrosis and / or cirrhosis, diabetes, thyroid and parathyroid abnormalities, thrombosis, and osteoporosis. Frequent transfusions also expose patients to the risk of HIV, hepatitis B, and hepatitis C infection.
[0015] Patients with β-thal intermediates generally exist between the ages of 2 and 6. They typically do not require blood transfusions. However, bone abnormalities arise due to chronic hypertrophy of the red blood cell lineage to compensate for chronic anemia. Patients may experience fractures of long bones due to osteoporosis. Extramedullary erythropoiesis is common and leads to enlargement of the spleen, liver, and lymph nodes. Spinal cord compression and neurological problems can also occur. Patients also develop lower extremity ulcers and have an increased risk of thrombotic events, including stroke, pulmonary embolism, and deep vein thrombosis. Treatment for β-thal intermediates includes splenectomy, folate supplementation, hydroxyurea therapy, and radiotherapy for extramedullary tumors. Chelation therapy is used for patients who develop iron overload.
[0016] Patients with β-thal syndrome often have a shorter life expectancy. Patients with β-thal major syndrome who do not receive blood transfusions generally die in their 20s or 30s. Patients with β-thal major syndrome who receive regular blood transfusions and adequate chelation therapy can survive beyond their 50s. Heart failure secondary to iron toxicity is the leading cause of death in patients with β-thal major syndrome due to iron toxicity. [Overview of the Initiative] [Problems that the invention aims to solve]
[0017] Various new therapies for SCD and β-Thal are currently under development. The delivery of modified HBB genes via gene therapy is currently being studied in clinical trials. However, the long-term efficacy and safety of this approach remain unclear. While hematopoietic stem cell transplantation using HLA-matched allogeneic stem cell donors has been shown to cure SCD and β-Thal, this procedure carries risks, including those associated with ablation therapy to prepare the subject for transplantation, and the risk of graft-versus-host disease after transplantation. Furthermore, compatible allogeneic donors are often difficult to identify. Therefore, improved methods for managing these and other hemoglobin disorders are needed. [Means for solving the problem]
[0018] In certain embodiments, methods are provided herein for increasing the expression (i.e., transcriptional activity) of one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) in a target or cell using a genome editing system (e.g., a CRISPR / Cas-mediated genome editing system). In certain embodiments, these methods may utilize any repair mechanism to alter (e.g., delete, disrupt, or modify) all or part of one or more γ-globin gene regulatory elements. In certain embodiments, these methods may utilize a DNA repair mechanism, e.g., NHEJ or HDR, to delete or disrupt one or more γ-globin gene regulatory elements (e.g., silencers, enhancers, promoters, or insulators). In certain embodiments, these methods may utilize a DNA repair mechanism, e.g., HDR, to alter the sequence of one or more nucleotides in a γ-globin gene regulatory element (e.g., silencer, enhancer, promoter, or insulator), including mutating, inserting, deleting, or disrupting. In certain embodiments, these methods utilize one or more DNA repair mechanisms, such as a combination of NHEJ and HDR.In certain embodiments, these methods are, for example, HBG1 13bp del c.-114~-102;4bp del c.-225~-222;c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A; or HBG2 13bp del c.-114~-102;c.-109 G>T;c.-114 C>A;c.-114 C>T;c.-157 C>T;c.-158 This results in spontaneous HPFH mutations, including C>T;c.-167, C>T;c.-167, C>A;c.-175, T>C;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G; or c.-567 T>G, which result in mutations or mutations in the γ-globin regulatory elements associated with spontaneous HPFH mutations.
[0019] In certain embodiments, methods are provided herein for treating β-abnormal hemoglobinopathy in subjects requiring increased expression (i.e., transcriptional activity) of one or more gamma-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) using CRISPR / Cas-mediated genome editing. In certain embodiments, these methods utilize DNA repair mechanisms, e.g., NHEJ or HDR, to delete or disrupt one or more gamma-globin gene regulatory elements (e.g., silencers, enhancers, promoters, or insulators). In certain embodiments, these methods utilize DNA repair mechanisms, e.g., HDR, to alter one or more nucleotide sequences in gamma-globin gene regulatory elements (e.g., silencers, enhancers, promoters, or insulators), including mutation, insertion, deletion, or disruption. In certain embodiments, these methods utilize one or more DNA repair mechanisms, e.g., a combination of NHEJ and HDR. In certain embodiments, these methods are, for example, HBG1 13bp del c.-114~-102;4bp del c.-225~-222;c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A; or HBG2 13bp del c.-114~-102;c.-109 G>T;c.-114 C>A;c.-114 C>T;c.-157 C>T;c.-158 This results in spontaneous HPFH mutations, including C>T;c.-167, C>T;c.-167, C>A;c.-175, T>C;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G; or c.-567 T>G, which lead to mutations or alterations in the γ-globin regulatory elements associated with spontaneous HPFH mutations. In certain embodiments, β-abnormal hemoglobinopathy is SCD or β-Thal.
[0020] In certain embodiments, gRNAs are provided herein for use in CRISPR / Cas-mediated methods to increase the expression (i.e., transcriptional activity) of one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2). In certain embodiments, these gRNAs include a targeting domain comprising the nucleotide sequences specified in SEQ ID NOs. 251-901. In certain embodiments, these gRNAs further include one or more of a first complementarity domain, a second complementarity domain, a ligation domain, a 5' elongation domain, an adjacency domain, or a tail domain. In certain embodiments, the gRNAs are modular. In other embodiments, the gRNAs are monomolecules (or chimeric). [Brief explanation of the drawing]
[0021] [Figure 1A-1I] Some illustrative depictions of gRNAs. [Figure 1A] These sequences depict modular gRNA molecules, partially derived from (or partially modeled after) Streptococcus pyogenes (S. pyogenes), as double-stranded structures (SEQ ID NOs. 39 and 40, respectively, in order of appearance). [Figure 1B] A single gRNA molecule partially derived from S. pyogenes is depicted as a double-stranded structure (SEQ ID NO: 41). [Figure 1C] A single gRNA molecule partially derived from S. pyogenes is depicted as a double-stranded structure (SEQ ID NO: 42). [Figure 1D] This diagram depicts a single gRNA molecule partially derived from S. pyogenes as a double-stranded structure (SEQ ID NO: 43). [Figure 1E] This diagram depicts a single gRNA molecule partially derived from S. pyogenes as a double-stranded structure (SEQ ID NO: 44). [Figure 1F]Modular gRNA molecules partially derived from Streptococcus thermophilus (S. thermophilus) are depicted as double-stranded structures (Sequence IDs 45 and 46, in order of appearance). [Figure 1G] The alignment of modular gRNA molecules of S. pyogenes and S. thermophilus is depicted (sequence numbers 39, 45, 47, and 46, respectively, in order of appearance). [Figure 1H-1I] Another exemplary structure of a single gRNA molecule is depicted. [Figure 1H] An exemplary structure of a single gRNA molecule partially derived from S. pyogenes is shown as a double-stranded structure (SEQ ID NO: 42). [Figure 1I] An exemplary structure of a single gRNA molecule partially derived from S. aureus is shown as a double-stranded structure (SEQ ID NO: 38). [Figure 2A] The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2B]The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2C] The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2D]The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2E] The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2F]The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 2G] The alignment of the Cas9 sequence (Chylinski 2013) is depicted. The N-terminal RuvC-like domain is indicated by "Y" in the box. The other two RuvC-like domains are indicated by "B" in the box. The HNH-like domain is indicated by "G" in the box. Sm: S. mutans (SEQ ID NO: 1); Sp: S. pyogenes (SEQ ID NO: 2); St: S. thermophilus (SEQ ID NO: 4); Li: L. innocua (SEQ ID NO: 5). The "motif" (SEQ ID NO: 14) is a consensus sequence based on the four sequences. Residues conserved in all four sequences are indicated by single-letter amino acid abbreviations; "*" indicates any amino acid found at the corresponding position in any of the four sequences; and "-" indicates absence. [Figure 3A] Alignment of the N-terminal RuvC-like domain from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs. 52-95, 120-123). The last row of Figure 3B identifies four highly conserved residues. [Figure 3B] Alignment of the N-terminal RuvC-like domain from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs. 52-95, 120-123). The last row of Figure 3B identifies four highly conserved residues. [Figure 4A] After excluding outliers, the alignment of the N-terminal RuvC-like domain from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs. 52-123). The last row of Figure 4B identifies three highly conserved residues. [Figure 4B] After excluding outliers, the alignment of the N-terminal RuvC-like domain from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs. 52-123). The last row of Figure 4B identifies three highly conserved residues. [Figure 5A] Alignment of HNH-like domains from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs: 124-198). The last row of Figure 5C identifies the conserved residues. [Figure 5B] Alignment of HNH-like domains from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs: 124-198). The last row of Figure 5C identifies the conserved residues. [Figure 5C] Alignment of HNH-like domains from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs: 124-198). The last row of Figure 5C identifies the conserved residues. [Figure 6A] Excluding outliers, the alignment of HNH-like domains from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs: 124-141, 148, 149, 151-153, 162, 163, 166-174, 177-187, 194-198). The last row of Figure 6B identifies three highly conserved residues. [Figure 6B] Excluding outliers, the alignment of HNH-like domains from the Cas9 molecule disclosed in Chylinski 2013 is shown (SEQ ID NOs: 124-141, 148, 149, 151-153, 162, 163, 166-174, 177-187, 194-198). The last row of Figure 6B identifies three highly conserved residues. [Figure 7] This section explains gRNA domain terminology using an example gRNA sequence (SEQ ID NO: 42). [Figure 8A] This provides a schematic diagram of the domain structure of S. pyogenes Cas9. Figure 8A shows the structure of the Cas9 domain, including amino acid positions, with reference to the two lobes of Cas9 (recognition (REC) lobe and nuclease (NUC) lobe). Figure 8B shows the percentage homology of each domain across 83 Cas9 orthologues. [Figure 8B] This provides a schematic diagram of the domain structure of S. pyogenes Cas9. Figure 8A shows the structure of the Cas9 domain, including amino acid positions, with reference to the two lobes of Cas9 (recognition (REC) lobe and nuclease (NUC) lobe). Figure 8B shows the percentage homology of each domain across 83 Cas9 orthologues. [Figure 9] This provides schematic diagrams of the HBG1 and HBG2 genes(s) in the context of globin loci. The coding sequence (CDS), mRNA region, and gene are shown. (A) Targeted regions for gRNA design are shown (dashed lines and brackets indicating gene regions adjacent to the HBG1 and HBG2 genes). (B) Core promoter elements are shown. (C) Motifs in gene regulatory regions where transcription activators and transcriptional repressors can bind to regulate gene expression are shown. Note the overlap between motifs and the genomic regions targeted by gRNA design. Examples of deletions in the HBG1 and HBG2 gene regulatory regions that cause HPFH, along with their associated %HbF, are shown. [Figure 10A]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 10B]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 10C]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 10D]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 10E]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 10F]This shows data from gRNA screening for the incorporation of 13 bp del c.-114~-102 HPFH mutations in human K562 erythroleukemia cells. (A) Gene editing determined by T7E1 endonuclease assay analysis of HBG1 and HBG2 locus-specific PCR products amplified from genomic DNA extracted from K562 cells after electroporation with DNA encoding S. pyogenes-specific gRNA and plasmid DNA encoding S. pyogenes-Cas9. (B) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG1 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. (C) Gene editing determined by DNA sequence analysis of PCR products amplified from the HBG2 locus in genomic DNA extracted from K562 cells after electroporation with the indicated gRNA and Cas9 plasmid. For (B) and (C), the type of editing event (insertion, deletion) and the subtype of deletion (partially [12nt HPFH] or completely [13-26nt HPFH] deleted 13nt target, deleted other sequence [other deletion]) are indicated by different shaded / patterned bars. (D)~(F) Examples of deletions in the HBG1 gene regulatory region. [Figure 11A]Figure 11A shows the results of gene editing of human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNPs complexed with in vitro transcriptional S. pyogenes gRNAs (HBG gRNA Sp35 (including SEQ ID NO: 339) and Sp37 (including SEQ ID NO: 333)) that target specific 13nt sequences for deletion. Figure 11A shows the percentage of indels detected by T7E1 analysis of HBG1 and HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells treated with the indicated RNPs or donor-matched untreated control cells (n=3 CB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11B shows the percentage of indels detected by T7E1 analysis of HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells or adult CD34+ cells treated with the indicated RNP, or donor-matched untreated control cells (n=3 CB CD34+ cells, n=3 mPB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11C (top panel) shows the edits detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from human CB CD34+ cells electroporated with HBG Sp35 RNP or HBG Sp37 RNP + / - ssODN1 (SEQ ID NO: 906) or PhTx ssODN1 (SEQ ID NO: 909). Figure 11C (lower left panel) shows the level of gene editing determined by Sanger DNA sequencing analysis of gDNA from cells edited with HBG Sp37 RNP and ssODN1 and PhTx ssODN1. Figure 11C (lower right panel) shows specific types of deletions detected within the total deletions from the data presented in the lower left panel. [Figure 11B]Figure 11A shows the results of gene editing of human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNPs complexed with in vitro transcriptional S. pyogenes gRNAs (HBG gRNA Sp35 (including SEQ ID NO: 339) and Sp37 (including SEQ ID NO: 333)) that target specific 13nt sequences for deletion. Figure 11A shows the percentage of indels detected by T7E1 analysis of HBG1 and HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells treated with the indicated RNPs or donor-matched untreated control cells (n=3 CB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11B shows the percentage of indels detected by T7E1 analysis of HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells or adult CD34+ cells treated with the indicated RNP, or donor-matched untreated control cells (n=3 CB CD34+ cells, n=3 mPB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11C (top panel) shows the edits detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from human CB CD34+ cells electroporated with HBG Sp35 RNP or HBG Sp37 RNP + / - ssODN1 (SEQ ID NO: 906) or PhTx ssODN1 (SEQ ID NO: 909). Figure 11C (lower left panel) shows the level of gene editing determined by Sanger DNA sequencing analysis of gDNA from cells edited with HBG Sp37 RNP and ssODN1 and PhTx ssODN1. Figure 11C (lower right panel) shows specific types of deletions detected within the total deletions from the data presented in the lower left panel. [Figure 11C]Figure 11A shows the results of gene editing of human umbilical cord blood (CB) and human adult CD34+ cells after electroporation with RNPs complexed with in vitro transcriptional S. pyogenes gRNAs (HBG gRNA Sp35 (including SEQ ID NO: 339) and Sp37 (including SEQ ID NO: 333)) that target specific 13nt sequences for deletion. Figure 11A shows the percentage of indels detected by T7E1 analysis of HBG1 and HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells treated with the indicated RNPs or donor-matched untreated control cells (n=3 CB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11B shows the percentage of indels detected by T7E1 analysis of HBG2-specific PCR products amplified from gDNA extracted from CB CD34+ cells or adult CD34+ cells treated with the indicated RNP, or donor-matched untreated control cells (n=3 CB CD34+ cells, n=3 mPB CD34+ cells, 3 separate experiments). The data shown represent mean and error bars corresponding to the standard deviation in the three separate donor / experiment. Figure 11C (top panel) shows the edits detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from human CB CD34+ cells electroporated with HBG Sp35 RNP or HBG Sp37 RNP + / - ssODN1 (SEQ ID NO: 906) or PhTx ssODN1 (SEQ ID NO: 909). Figure 11C (lower left panel) shows the level of gene editing determined by Sanger DNA sequencing analysis of gDNA from cells edited with HBG Sp37 RNP and ssODN1 and PhTx ssODN1. Figure 11C (lower right panel) shows specific types of deletions detected within the total deletions from the data presented in the lower left panel. [Figure 12A]This shows gene editing of HBG1 and HBG2 in K562 erythroleukemia cells. Figure 12A shows NHEJ (indels) detected by T7E1 analysis of HBG1 and HBG2 PCR products amplified from gDNA extracted from K562 cells 3 days after nucleofection with RNP complexed with the indicated gRNA. Figure 12B shows Sanger DNA sequence analysis of PCR products amplified from the HBG1 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13nt HPFH sequences (Sp35 (including SEQ ID NO: 339), Sp36 (including SEQ ID NO: 338), Sp37 (including SEQ ID NO: 333)). Figure 12C shows Sanger DNA sequence analysis of PCR products amplified from the HBG2 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13bp HPFH sequences (Sp35, Sp36, Sp37). In Figures 12B and 12C, deletions were further divided into 13bp targeted deletions (HPFH deletions, 18–26nt deletions, >26nt deletions) and deletions that did not include 13bp deletions (<12nt deletions, other deletions, insertions). [Figure 12B]This shows gene editing of HBG1 and HBG2 in K562 erythroleukemia cells. Figure 12A shows NHEJ (indels) detected by T7E1 analysis of HBG1 and HBG2 PCR products amplified from gDNA extracted from K562 cells 3 days after nucleofection with RNP complexed with the indicated gRNA. Figure 12B shows Sanger DNA sequence analysis of PCR products amplified from the HBG1 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13nt HPFH sequences (Sp35 (including SEQ ID NO: 339), Sp36 (including SEQ ID NO: 338), Sp37 (including SEQ ID NO: 333)). Figure 12C shows Sanger DNA sequence analysis of PCR products amplified from the HBG2 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13bp HPFH sequences (Sp35, Sp36, Sp37). In Figures 12B and 12C, deletions were further divided into 13bp targeted deletions (HPFH deletions, 18–26nt deletions, >26nt deletions) and deletions that did not include 13bp deletions (<12nt deletions, other deletions, insertions). [Figure 12C]This shows gene editing of HBG1 and HBG2 in K562 erythroleukemia cells. Figure 12A shows NHEJ (indels) detected by T7E1 analysis of HBG1 and HBG2 PCR products amplified from gDNA extracted from K562 cells 3 days after nucleofection with RNP complexed with the indicated gRNA. Figure 12B shows Sanger DNA sequence analysis of PCR products amplified from the HBG1 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13nt HPFH sequences (Sp35 (including SEQ ID NO: 339), Sp36 (including SEQ ID NO: 338), Sp37 (including SEQ ID NO: 333)). Figure 12C shows Sanger DNA sequence analysis of PCR products amplified from the HBG2 locus of cells nucleofected with Cas9 protein complexed with gRNA targeting 13bp HPFH sequences (Sp35, Sp36, Sp37). In Figures 12B and 12C, deletions were further divided into 13bp targeted deletions (HPFH deletions, 18–26nt deletions, >26nt deletions) and deletions that did not include 13bp deletions (<12nt deletions, other deletions, insertions). [Figure 13A] Figure 13A shows gene editing of HBG in human adult recruited peripheral blood (mPB) CD34+ cells and induction of fetal hemoglobin in erythrocyte offspring of RNP-treated cells after electroporation of mPB CD34+ cells with HBG Sp37 RNP + / - ssODN encoding a 13bp deletion. Figure 13A shows the percentage of editing detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from RNP-treated mPB CD34+ cells or donor-matched untreated control cells. Figure 13B shows the plicatilization of HBG mRNA expression in erythroblasts differentiated from RNP-treated and untreated donor-matched control mPB CD34+ cells at day 7. mRNA levels are normalized against GAPDH and calibrated to the levels detected in the untreated control on the corresponding day of differentiation. [Figure 13B]Figure 13A shows gene editing of HBG in human adult recruited peripheral blood (mPB) CD34+ cells and induction of fetal hemoglobin in erythrocyte offspring of RNP-treated cells after electroporation of mPB CD34+ cells with HBG Sp37 RNP + / - ssODN encoding a 13bp deletion. Figure 13A shows the percentage of editing detected by T7E1 analysis of HBG2 PCR products amplified from gDNA extracted from RNP-treated mPB CD34+ cells or donor-matched untreated control cells. Figure 13B shows the plicatilization of HBG mRNA expression in erythroblasts differentiated from RNP-treated and untreated donor-matched control mPB CD34+ cells at day 7. mRNA levels are normalized against GAPDH and calibrated to the levels detected in the untreated control on the corresponding day of differentiation. [Figure 14A] Figure 14A shows the ex vivo differentiation potential of RNP-treated and untreated mPB CD34+ cells from the same donor. Figure 14A shows the capacity of hematopoietic bone marrow / erythrocyte colony-forming cells (CFCs), with the number and subtype of colonies indicated (GEMM: granulocyte-erythrocyte-monocyte-macrophage colony, E: erythrocyte colony, GM: granulocyte-macrophage colony, M: macrophage colony, G: granulocyte colony). Figure 14B shows the percentage of glycophorin A expressed over the time course of erythrocyte differentiation, determined by flow cytometry analysis for the indicated time points and samples. [Figure 14B] Figure 14A shows the ex vivo differentiation potential of RNP-treated and untreated mPB CD34+ cells from the same donor. Figure 14A shows the capacity of hematopoietic bone marrow / erythrocyte colony-forming cells (CFCs), with the number and subtype of colonies indicated (GEMM: granulocyte-erythrocyte-monocyte-macrophage colony, E: erythrocyte colony, GM: granulocyte-macrophage colony, M: macrophage colony, G: granulocyte colony). Figure 14B shows the percentage of glycophorin A expressed over the time course of erythrocyte differentiation, determined by flow cytometry analysis for the indicated time points and samples. [Modes for carrying out the invention]
[0022] Detailed explanation definition In the context of this specification, "domain" is used to refer to a segment of a protein or nucleic acid. Unless otherwise specified, a domain does not need to possess any particular functional properties.
[0023] The calculation of homology or sequence identity (terms used synonymously herein) between two sequences is performed as follows: These sequences are aligned for optimal comparison (for example, gaps may be inserted in one or both of the first and second amino acid or nucleic acid sequences for optimal alignment, and non-homologous sequences may be ignored for comparison). Optimal alignment is determined as the highest score using the GAP program of the GCG software package, which has a Blossum62 score matrix with gap penalty 12, gap extension penalty 4, and frameshift gap penalty 5. Next, amino acid residues or nucleotides at the corresponding amino acid or nucleotide positions are compared. If a position in the first sequence is occupied by the same amino acid residue or nucleotide as the corresponding position in the second sequence, the molecules are identical at that position. The percentage of identity between two sequences is a function of the number of identical positions shared by these sequences.
[0024] In the use of this specification, "polypeptide" refers to an amino acid polymer having fewer than 100 amino acid residues. In one embodiment, the polypeptide has fewer than 50, fewer than 20, or fewer than 10 amino acid residues.
[0025] "Alt-HDR," "alternative homology-directed repair," or "alternative HDR" as used herein refers to a DNA damage repair process that uses homologous nucleic acids (e.g., endogenous homologous sequences, e.g., sister chromatids, or exogenous nucleic acids, e.g., template nucleic acids). Alt-HDR differs from standard HDR in that the process utilizes a different pathway and can be inhibited by standard HDR mediators, RAD51, and BRCA2. Furthermore, alt-HDR uses single-stranded or nicked homologous nucleic acids for the repair of damage.
[0026] "Standard HDR" or "standard homologous recombination repair" as used herein refers to a DNA damage repair process using homologous nucleic acids (e.g., endogenous homologous sequences, e.g., sister chromatids, or exogenous nucleic acids, e.g., template nucleic acids). Standard HDR typically functions in the case of significant excision in double-strand breaks that form at least one single-stranded portion of DNA. In normal cells, HDR typically involves a series of steps, e.g., break recognition, break stabilization, excision, single-stranded DNA stabilization, DNA crossover intermediate formation, degradation of this crossover intermediate, and ligation. This process requires RAD51 and BRCA2, and the homologous nucleic acid is typically double-stranded.
[0027] Unless otherwise specified, the term "HDR" in this specification encompasses both standard HDR and alt-HDR.
[0028] In the context of this specification, “non-homologous end joining” or “NHEJ” refers to ligation-mediated repair and / or non-template-mediated repair, including standard NHEJ (cNHEJ), alternative NHEJ (altNHEJ), microhomologous end joining (MMEJ), single-strand annealing (SSA), and synthesis-dependent microhomologous end joining (SD-MMEJ).
[0029] In the context of this specification, "reference molecule" refers to the molecule on which a modified molecule or candidate molecule is compared. For example, a reference Cas9 molecule refers to the Cas9 molecule on which a modified or candidate Cas9 molecule is compared. Similarly, a reference gRNA refers to the gRNA molecule on which a modified or candidate gRNA molecule is compared. A modified molecule or candidate molecule can be compared to a reference molecule based on its sequence (e.g., the modified or candidate molecule may have X% sequence identity or homology with the reference molecule) or its activity (e.g., the modified or candidate molecule may have X% of the activity of the reference molecule). For example, if the reference molecule is a Cas9 molecule, the modified or candidate molecule can be characterized as having less than or equal to 10% of the nuclease activity of the reference Cas9 molecule. Examples of reference Cas9 molecules include naturally occurring, unmodified Cas9 molecules, such as those derived from S. pyogenes, S. aureus, S. thermophilus, or N. meningitidis. In certain embodiments, the reference Cas9 molecule is a naturally occurring Cas9 molecule that has the closest sequence identity or homology to the modified or candidate Cas9 molecule being compared. In certain embodiments, the reference Cas9 molecule is a parent molecule having a naturally occurring sequence or a known sequence that has been mutated to become the modified or candidate Cas9 molecule.
[0030] The term “genome editing system” refers to any system having RNA-induced DNA editing activity. The genome editing system of this disclosure comprises at least two components compatible with the naturally occurring CRISPR system: a guide RNA (gRNA) and an RNA-induced nuclease. These two components associate with a specific nucleic acid sequence to form a complex that can edit DNA in or near that nucleic acid sequence by causing one or more single-strand breaks (SSBs or nicks), double-strand breaks (DSBs), and / or point mutations.
[0031] In the context of this specification, "substitution" or "substituted" in relation to molecular modification simply indicates the presence of a replacement entity, without requiring any process restrictions.
[0032] In the context of this specification, "subject" may mean human, mouse, or non-human primate.
[0033] In the context of this specification, “treat,” “treating,” and “treatment” mean the treatment of a subject, for example, a human disease, including (a) inhibiting the disease, i.e., stopping or preventing its onset or progression; (b) alleviating the disease, i.e., regressing the disease state; (c) alleviating one or more symptoms of the disease; and (d) curing the disease. For example, “treating” SCD or β-Thal may mean, but are not limited to, preventing the onset or progression of SCD or β-Thal, alleviating one or more symptoms of SCD or β-Thal (e.g., anemia, sickle cell disease, occlusive crisis), or curing SCD or β-Thal.
[0034] In the context of this specification, “prevent,” “preventing,” and “prevention” mean the prevention of a disease in a subject, for example, a human disease, which includes (a) preventing or preventing a disease; (b) influencing the predisposition to a disease; and (c) preventing or delaying the onset of at least one symptom of a disease.
[0035] In the context of amino acid sequences, "X" refers to any amino acid (e.g., any 20 natural amino acids) unless otherwise specified.
[0036] In the context of this specification, a “regulatory region” refers to a DNA sequence containing one or more regulatory elements (e.g., silencers, enhancers, promoters, or insulators) that control or regulate the expression of a gene. For example, a γ-globin gene regulatory region contains one or more regulatory elements that control or regulate the expression of the γ-globin gene. In certain embodiments, the regulatory region is adjacent to the gene being controlled or regulated. For example, the γ-globin gene regulatory region may be adjacent to or bound to the γ-globin gene. In other embodiments, the regulatory region may be adjacent to or bound to another gene that may result in upregulation or downregulation of the gene whose expression is controlled or regulated. For example, the γ-globin gene regulatory region may be adjacent to a gene that expresses a repressor of γ-globin gene expression. In the case of HBG1, the regulatory region contains at least 1 to 2990 nucleotides of SEQ ID NO: 902. In the case of HBG2, the regulatory region contains at least 1 to 2914 nucleotides of SEQ ID NO: 903.
[0037] In the context of this specification, “HBG target site” refers to a location in the HBG1 or HBG2 regulatory region containing a target site (e.g., a target sequence that is deleted or mutated) (e.g., “HBG1 target site” and “HBG2 target site” respectively), which, when altered (e.g., disrupted or deleted by the introduction of a DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertion or deletion, or altered by a DNA repair mechanism-mediated (e.g., HDR-mediated) sequence alteration), results in increased expression (e.g., derepression) of the HBG1 or HBG2 gene product (i.e., γ-globin). In certain embodiments, the HBG target site is located within an HBG1 or HBG2 regulatory element (e.g., a silencer, enhancer, promoter, or insulator) in a regulatory region adjacent to HBG1 or HBG2. In some of these embodiments, alteration of the HBG target site results in decreased repressor binding, i.e., derepression, and consequently increased expression of HBG1 or HBG2. In other embodiments, the HBG target site is located within a regulatory element of a gene other than HBG1 or HBG2 that encodes a gene product involved in regulating HBG1 or HBG2 gene expression (e.g., a repressor of HBG1 or HBG2 gene expression). In specific embodiments, the HBG target site is a region of the HBG1 or HBG2 regulatory region with the highest density of binding motifs involved in regulating HBG1 or HBG2 expression. In specific embodiments, the method provided herein targets multiple HBG target sites simultaneously or sequentially.
[0038] In the context of this specification, "target sequence" refers to a nucleic acid sequence containing the HBG target site.
[0039] In the context of this specification, "Cas9 molecule" or "Cas9 polypeptide" refers to a molecule or polypeptide that interacts with a gRNA molecule and, in cooperation with the gRNA molecule, can localize to a site containing a target domain, and, in certain embodiments, to a PAM sequence. Cas9 molecules and Cas9 polypeptides include both naturally occurring Cas9 molecules and Cas9 polypeptides, as well as genetically engineered, altered, or modified Cas9 molecules or Cas9 polypeptides that differ from a reference sequence, e.g., the most similar naturally occurring Cas9 molecule, by, for example, at least one amino acid residue.
[0040] overview Methods for increasing the expression (i.e., transcriptional activity) of one or more gamma-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) using a genome editing system (e.g., CRISPR / Cas-mediated genome editing) are provided herein. These methods utilize a genome editing system (e.g., CRISPR / Cas-mediated genome editing) to alter (e.g., delete, disrupt, or modify) one or more gamma-globin gene regulatory regions to increase (e.g., derepress, promote) the expression of the gamma-globin gene. In some embodiments, these methods alter one or more regulatory elements (e.g., silencers, enhancers, promoters, or insulators) associated with the targeted gamma-globin gene. In other embodiments, these methods alter one or more regulatory elements of a gene other than the targeted gamma-globin gene (e.g., a gene encoding a gamma-globin gene repressor). In certain embodiments, a genome editing system (e.g., CRISPR / Cas-mediated genome editing) is used to alter regulatory elements of HBG1, HBG2, or both HBG1 and HBG2 (e.g., silencers, enhancers, promoters, or insulators).In certain embodiments, the genome editing system (e.g., CRISPR / Cas-mediated genome editing) may, for example, HBG1 13bp del c.-114~-102;4bp del c.-225~-222;c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A; or HBG2 13bp del c.-114~-102;c.-109 G>T;c.-114 C>A;c.-114 This results in spontaneous HPFH mutations, including C>T;c.-157, C>T;c.-158, C>T;c.-167, C>T;c.-167, C>A;c.-175, T>C;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G; or c.-567 T>G, which result in mutations or mutations in the γ-globin regulatory elements associated with spontaneous HPFH mutations.
[0041] In some embodiments, methods using genome editing systems described herein (e.g., CRISPR / Cas-mediated genome editing) can alter (e.g., delete, disrupt, or modify) all or part of one or more γ-globin gene regulatory elements by utilizing any repair mechanism. In certain embodiments, the method disrupts all or part of one or more γ-globin gene regulatory elements by utilizing DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertion or deletion. For example, the method can utilize DNA repair mechanisms (e.g., NHEJ or HDR) to delete all or part of the negative regulatory element (e.g., silencer) of the γ-globin gene, thereby resulting in inactivation of the negative regulatory element (e.g., loss of binding between silencer and repressor) and increased expression of the γ-globin gene. In other embodiments, the method disrupts all or part of one or more regulatory elements associated with the gene encoding the γ-globin gene repressor by utilizing DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertion or deletion. For example, this method can utilize a DNA repair mechanism (e.g., NHEJ or HDR) to delete all or part of the positive regulatory element (e.g., promoter) of the γ-globin repressor gene, thereby resulting in decreased repressor expression, reduced binding of the repressor to the γ-globin gene silencer, and increased γ-globin gene expression. In other embodiments, this method utilizes a DNA repair mechanism (e.g., HDR) to modify the sequence of one or more γ-globin gene regulatory elements (e.g., introducing mutations in the HBG1 and / or HBG2 regulatory elements corresponding to spontaneously occurring HPFH mutations, or deleting all or part of the HBG1 and / or HBG2 regulatory elements). In some embodiments, this method can use a combination of one or more DNA repair mechanisms (e.g., NHEJ and HDR). In certain embodiments, this method sustains the target HbF.This specification also provides compositions (e.g., gRNA, Cas9 polypeptides and molecules, template nucleic acids, vectors) and kits for use in these methods.
[0042] The transition from the expression of gamma-globin genes (i.e., HBG1, HBG2) to the expression of HBB (i.e., globin switching) is associated with the development of symptoms of β-abnormal hemoglobinopathy, including SCD and β-thal. Therefore, in certain embodiments, methods, compositions, and kits for treating or preventing β-abnormal hemoglobinopathy, including SCD and β-thal, by increasing the expression of one or more gamma-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) using CRISPR / Cas-mediated genome editing are provided herein. In some embodiments, the method alters one or more regulatory elements (e.g., silencers, enhancers, promoters, or insulators) associated with the targeted gamma-globin gene. In other embodiments, the method alters one or more regulatory elements in a gene other than the targeted gamma-globin gene (e.g., a gene encoding a gamma-globin gene repressor). In certain embodiments, CRISPR / Cas-mediated genome editing is used to alter regulatory elements (e.g., silencers, enhancers, promoters, or insulators) of HBG1, HBG2, or both HBG1 and HBG2. In some embodiments, this method utilizes DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertions or deletions to disrupt all or part of one or more γ-globin gene regulatory elements. For example, this method can utilize DNA repair mechanisms (e.g., NHEJ or HDR) to delete all or part of the negative regulatory element (e.g., silencer) of the γ-globin gene, thereby resulting in inactivation of the negative regulatory element (e.g., loss of binding between silencer and repressor) and increased expression of the γ-globin gene. In other embodiments, this method utilizes DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertions or deletions to disrupt all or part of one or more regulatory elements associated with the gene encoding the γ-globin gene repressor.For example, this method can utilize a DNA repair mechanism (e.g., NHEJ or HDR) to delete all or part of the positive regulatory element (e.g., promoter) of the γ-globin repressor gene, thereby resulting in decreased repressor expression, reduced binding of the repressor to the γ-globin gene silencer, and increased γ-globin gene expression. In other embodiments, this method utilizes a DNA repair mechanism (e.g., HDR) to modify the sequence of one or more γ-globin gene regulatory elements (e.g., introducing mutations into the HBG1 and / or HBG2 regulatory elements corresponding to a spontaneously occurring HPFH mutation, or deleting all or part of the HBG1 and / or HBG2 regulatory elements). In some embodiments, this method can use a combination of one or more DNA repair mechanisms (e.g., NHEJ and HDR). In certain embodiments, this method sustains the target HbF.
[0043] In certain embodiments, increasing the expression of one or more γ-globin genes (e.g., HBG1, HBG2) using the methods provided herein results in the preferential production of HbF over HbA and / or an increase in the HbF level as a percentage of total hemoglobin. Therefore, methods for increasing the total HbF level, increasing the HbF level as a percentage of total hemoglobin, or increasing the ratio of HbF to HbA in a subject by increasing the expression of one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) using CRISPR / Cas-mediated genome editing are further provided herein. Similarly, in certain embodiments, increasing the expression of one or more γ-globin genes results in the preferential production of HbF over HbS and / or a decrease in the percentage of HbS as a percentage of total hemoglobin. Therefore, CRISPR / Cas-mediated genome editing is used to increase the expression of one or more γ-globin genes (e.g., HBG1, HBG2, or HBG1 and HBG2) to reduce total HbS levels in a subject, reduce HbS levels as a percentage of total hemoglobin levels, or increase the ratio of HbF to HbS.
[0044] gRNAs for use in the methods disclosed herein are provided herein in certain embodiments. In certain embodiments, these gRNAs include a targeting domain that is complementary or partially complementary to a target domain located at or adjacent to an HBG target site. In certain embodiments, the targeting domain consists of or is essentially derived from a nucleotide sequence, including the nucleotide sequence shown in one of Sequence IDs 251-901.
[0045] Genomic studies have identified several genes that regulate globin switching, including BCL11A, Kruppel-like factor 1 (KLF1), MYB, and genes within the β-globin locus. Certain mutations in these genes can result in inhibited or incomplete globin switching, also known as hereditary hyperfetal hemoglobinopathy (HPFH). HPFH mutations can be deletions or non-deletions (e.g., point mutations). HPFH subjects express HbF throughout their lives, i.e., without symptoms of anemia, and experience either no globin switching or only partial globin switching. Heterozygous subjects exhibit 20–40% pancellular HbF, and co-inheritance results in mitigation of β-abnormal hemoglobinopathy (Thein 2009; Akinbami 2016). Subjects who are compound heterozygotes for abnormal hemoglobin disorders and HPFH, such as SCD and HPFH, β-Thal and HPFH, sickle cell phenotype and HPFH, or delta-β-Thal and HPFH, experience milder disease and symptoms than subjects without HPFH mutations. Homozygous patients with HbS who also co-inherit HPFH mutations, such as mutations that induce HbF expression through disinhibition of HBG1 or HBG2, do not develop SCD or β-Thal symptoms (Steinberg et al., Disorders of Hemoglobin, Cambridge Univ. Press, 2009, p.570). HPFH is clinically benign (Chassanidis 2009).
[0046] While HPFH is rare in global populations, it is more common in populations with high prevalence of hemoglobin disorders, including those of Southern European, South American, and African descent. In these populations, the prevalence of HPFH can reach 1-2 per 1,000 people (Costa 2002: Ahern 1973). Theoretically, HPFH mutations persist in these populations to improve the conditions associated with hemoglobin disorders.
[0047] The most common spontaneous HPFH mutation is a deletion within the β-globin locus. Common examples of HPFH-deficient mutations include French HPFH (23kb deletion), Caucasian HPFH (19kb deletion), HPFH-1 (84kb deletion), HPFH-2 (84kb deletion), and HPFH-3 (50kb deletion). In subjects with these mutations, β-globin synthesis is reduced, and γ-globin synthesis is secondarily increased.
[0048] Other HPFH mutations are located in the regulatory region of the γ-globin gene. One such mutation is a 13-nucleotide deletion located upstream of both the HBG1 and HBG2 genes (13 base pairs (bp) del c.-114~-102; CAATAGCCTTGAC del, a sequence based on the reverse complement of HBG1 / HBG2). This deletion disrupts a silencer element that normally prevents the expression of HBG1 / HBG2, and heterozygous adult subjects with this deletion exhibit approximately 30% HbF. Another HPFH mutation is a 4-nucleotide deletion (4 base pairs (bp) del c.-225~-222 (AGCA del)). Other HPFH mutations found in both HBG1 and HBG2 regulatory elements include, for example, non-del HPFH mutations such as c.-114 C>T;c.-158 C>T;c.-167 C>T; and c.-175 T>C.
[0049] Examples of non-del HPFH mutations associated with the HBG1 regulatory element include c.-117 G>A; c.-170 G>A; c.-175 T>G; c.-195 C>G; c.-196 C>T; c.-198 T>C; c.-201 C>T; c.-251 T>C; and c.-499 T>A.
[0050] Examples of non-del HPFH mutations associated with HBG2 regulatory elements include c.-109 G>T;c.-114 C>A;c.-157 C>T;c.-167 C>A;c.-202 C>G;c.-211 C>T;c.-228 T>C;c.-255 C>G; and c.-567 T>G.
[0051] Further polymorphisms in the HBG1 and HBG2 promoter regions were identified in a cohort of Brazilian SCD patients with corrected HbF levels >5% (Barbosa 2010). These polymorphisms include c.-309 A>G and c.-369 C>G in the HBG2 promoter.
[0052] Examples of HBG1 and HBG2 promoter elements that can be modified to reproduce HPFH mutations include the erythrocyte Kruppel-like factor (EKLF-2) and fetal Kruppel-like factor (FKLF) transcription factor binding motif (CTCCACCCA), the CP1 / Coup TFII binding motif (CCAATAGC), the GATA1 binding motif (CTATCT, ATATCT), or the step selection element (SSE) binding motif. Examples of HBG1 and HBG2 enhancer elements that can be modified to reproduce HPFH mutations include the SOX binding motif, e.g., SOX14, SOX2, or SOX1 (CCAATAGCCTTGA).
[0053] In certain embodiments of the methods provided herein, CRISPR / Cas-mediated alteration is used to alter one regulatory element or motif in the γ-globin gene regulatory region, for example, a silencer sequence in the HBG1 or HBG2 regulatory region, or a promoter or enhancer sequence associated with a gene encoding an HBG1 or HBG2 repressor. In other embodiments, CRISPR / Cas-mediated alteration is used to alter two or more (e.g., three, four, or five or more) regulatory elements or motifs in the γ-globin gene regulatory region, for example, an HBG1 or HBG2 silencer sequence and an HBG1 or HBG2 enhancer sequence; an HBG1 or HBG2 silencer sequence and a promoter or enhancer sequence associated with a gene encoding an HBG1 or HBG2 repressor; or an HBG1 or HBG2 silencer sequence and a promoter or enhancer sequence associated with a gene encoding an HBG1 or HBG2 repressor. The introduction of multiple mutations into the regulatory region of a single gene, or the introduction of one mutation into the regulatory regions of two or more genes, is referred to herein as “multiplexing.” Thus, multipleplexing constitutes (a) modification of two or more locations in the regulatory region of one gene within one or more identical cells, or (b) modification of one location in the regulatory regions of two or more genes.
[0054] In certain embodiments of the methods provided herein, CRISPR / Cas-mediated alterations of one or more γ-globin gene regulatory elements produce a phenotype identical or similar to that associated with spontaneously occurring HPFH mutations. In certain embodiments, the CRISPR / Cas-mediated alteration results in a γ-globin gene regulatory element containing alterations corresponding to spontaneously occurring HPFH mutations. In other embodiments, alterations of one or more γ-globin gene regulatory elements result in alterations not observed in spontaneously occurring HPFH mutations (i.e., mutations not naturally occurring).
[0055] In certain embodiments of the methods provided herein, for example, HBG1 13bp del c.-114~-102;4bp del c.-225~-222;c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A; or HBG2 13bp del c.-114~-102;c.-109 G>T;c.-114 C>A;c.-114 C>T;c.-157 C>T;c.-158 CRISPR / Cas-mediated alterations of one or more gamma-globin gene regulatory elements, including C>T;c.-167, C>T;c.-167, C>A;c.-175, T>C;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G; or c.-567 T>G, cause mutations or mutations in gamma-globin regulatory elements associated with spontaneously occurring HPFH mutations.
[0056] In certain embodiments, the methods provided herein involve altering one or more transcription factor binding motifs (e.g., gene regulatory motifs) in the γ-globin gene regulatory element. These transcription factor binding motifs include, for example, binding motifs occupied by transcription factors (TFs), TF complexes, and transcriptional repressors within the promoter regions of HBG1 and / or HBG2. In certain embodiments of the methods provided herein, the introduction of CRISPR / Cas-mediated alterations in one or more γ-globin gene regulatory elements alters the binding of transcription factors, e.g., repressors, in one, two, three, or more motifs. In certain embodiments, the introduction of CRISPR / Cas-mediated alterations in one or more γ-globin gene regulatory elements promotes the initiation of transcription by RNA polymerase II near or in the vicinity of the γ-globin gene promoter region, for example, by decreasing the binding of repressors in the silencer region, or by increasing the binding of transcription factors to the enhancer region.
[0057] In certain embodiments, the methods provided herein utilize DNA repair mechanism-mediated deletion (e.g., NHEJ-mediated or HDR-mediated) to delete all or part of nucleotides -114 to -102 in one or both alleles of HBG1, HBG2, or both alleles of HBG1 and HBG2, thereby producing an HPFH phenotype identical or similar to the phenotype associated with a spontaneously occurring 13bp del c.-114 to -102 mutation. In other embodiments, DNA repair mechanism-mediated deletion (e.g., NHEJ-mediated or HDR-mediated) is used to delete all or part of nucleotides -225 to -222 in one or both alleles of HBG1, thereby producing an HPFH phenotype identical or similar to the phenotype associated with a spontaneously occurring HBG1 4bp del -225 to -222 mutation. In other embodiments, DNA repair mechanism-mediated deletions (e.g., NHEJ-mediated or HDR-mediated) are used to delete all or part of nucleotides -225 to -222 in one or both alleles of HBG2.
[0058] In certain embodiments, the methods provided herein utilize DNA repair mechanism-mediated deletion (e.g., NHEJ-mediated or HDR-mediated) to delete all or part of nucleotides -114 to -102 in one or both alleles of HBG1 and one or both alleles of HBG2.
[0059] In certain embodiments, the methods provided herein utilize DNA repair mechanism-mediated deletion (e.g., NHEJ-mediated or HDR-mediated) to delete all or part of nucleotides -225 to -222 in one or both alleles of HBG1, and all or part of nucleotides -114 to -102 in one or both alleles of HBG2. In other embodiments, DNA repair mechanisms (e.g., NHEJ-mediated or HDR-mediated deletion) are utilized to delete all or part of nucleotides -225 to -222 in one or both alleles of HBG1, and all or part of nucleotides -114 to -102 in one or both alleles of HBG1.
[0060] In embodiments in which DNA repair mechanism-mediated deletions (e.g., NHEJ-mediated or HDR-mediated) are used to delete one or more nucleotides from HBG1, HBG2, or the HBG1 and HBG2 regulatory elements, the deletion may be identical to a deletion observed in a spontaneously occurring HPFH mutation, i.e., the deletion may consist of nucleotides -114 to -102 of HBG1 or HBG2, or nucleotides -225 to -222 of HBG1. In other embodiments, DNA repair mechanism-mediated deletions (e.g., NHEJ-mediated or HDR-mediated) remove only a portion of these nucleotides, for example, deleting 12 or fewer nucleotides in the range of -114 to -102 of HBG1 or HBG2, or deleting 3 or fewer nucleotides in the range of -225 to -222 of HBG1. In certain embodiments, in addition to all or some of the nucleotides within the spontaneously occurring HPFH mutation deletion boundary, one or more nucleotides can be knocked out on either side of this spontaneously occurring deletion boundary (i.e., outside -114 to -102 or -225 to -222).
[0061] In certain embodiments, the methods provided herein utilize DNA repair mechanism-mediated (e.g., NHEJ-mediated or HDR-mediated) insertion to insert one or more nucleotides into the HBG1 regulatory region, the HBG2 regulatory region, or a region spanning nucleotides -114 to -102 of both the HBG1 and HBG2 regulatory regions, or a region spanning nucleotides -225 to -222 of the HBG1 regulatory region, in order to disrupt a repressor binding site.
[0062] In certain embodiments, the method provided herein utilizes a DNA repair mechanism (e.g., HDR) to induce a single nucleotide change (i.e., a non-deletion mutation) corresponding to a spontaneous mutation associated with HPFH. For example, in certain embodiments, the method utilizes a DNA repair mechanism (e.g., HDR) to induce a single nucleotide change in the HBG1 regulatory region corresponding to a spontaneous mutation associated with HPFH, including, for example, c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A. In other embodiments, DNA repair mechanisms (e.g., HDR) are used to induce single nucleotide changes in the HBG2 regulatory region corresponding to spontaneous mutations associated with HPFH, including, for example, c.-109 G>T;c.-114 C>A;c.-114 C>T;c.-157 C>T;c.-158 C>T;c.-167 C>T;c.-167 C>A;c.-175 T>C;c.-202 C>G;c.-211 C>T;c.-228 T>C;c.-255 C>G;c.-309 A>G;c.-369 C>G;c.-567 T>G.
[0063] In certain embodiments, DNA repair mechanisms (e.g., HDR) are used to induce single-nucleotide changes in the HBG1 regulatory region that correspond to spontaneously occurring HPFH mutations found in the HBG2 regulatory region but not in the HBG1 regulatory region. Examples of such changes include c.-109 G>T; c.-114 C>A; c.-157 C>T; c.-167 C>A; c.-202 C>G; c.-211 C>T; c.-228 T>C; c.-255 C>G; c.-309 A>G; c.-369 C>G; or c.-567 T>G.
[0064] Similarly, in certain embodiments, DNA repair mechanisms (e.g., HDR) are used to induce single-nucleotide changes in the HBG2 regulatory region that correspond to spontaneously occurring HPFH mutations found in the HBG1 regulatory region but not in the HBG2 regulatory region. Examples of such changes include c.-117 G>A; c.-170 G>A; c.-175 T>G; c.-195 C>G; c.-196 C>T; c.-198 T>C; c.-201 C>T; c.-251 T>C; or c.-499 T>A.
[0065] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-114 C>T into the HBG1 and / or HBG2 regulatory regions by a DNA repair mechanism (e.g., HDR).
[0066] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-158 C>T (i.e., rs7482144 or XmnI-HBG2 mutation) into the HBG1 and / or HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0067] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-167 C>T into the HBG1 and / or HBG2 regulatory regions by a DNA repair mechanism (e.g., HDR).
[0068] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-175 T>C (i.e., a T→C substitution at position c.-175 of the conserved octamer [ATGCAAAT] sequence) into the HBG1 regulatory region by a DNA repair mechanism (e.g., HDR). This mutation, associated with 40% of HbF, has been shown to disable the ability of the ubiquitous octamer-binding nucleoprotein to bind to the HBG promoter fragment while simultaneously increasing the ability of the two erythrocyte-specific proteins to bind to the same fragment by 3-5 times (Mantovani 1988).
[0069] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-175 T>C into the HBG2 regulatory region via a DNA repair mechanism (e.g., HDR). This mutation is associated with 20–30% HbF expression.
[0070] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-117 G>A into the HBG1 regulatory region via a DNA repair mechanism (e.g., HDR). This mutation, referred to as the "Greek type," is the most common non-deletion HPFH mutation and maps two nucleotides upstream from the distal CCAAT box (Waber 1986). HBG1 c.-117 G>A significantly reduces the binding of erythrocyte-specific factors, rather than ubiquitous proteins, to the CCAAT box region fragment and is associated with 10-20% HbF (Mantovani 1988). This mutation is thought to interfere with the binding of nuclear factor E (NF-E), which likely plays a role in repressing γ-globin transcription in adult erythroid cells (Superti-Furga 1988). In other embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-117 G>A into the HBG2 regulatory region, thereby producing a non-native HPFH mutation.
[0071] In certain embodiments, the method provided herein includes introducing a non-deletion HPFH mutation c.-170 G>A into the HBG1 regulatory region by a DNA repair mechanism (e.g., HDR).
[0072] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-175 T>G into the HBG1 regulatory region by a DNA repair mechanism (e.g., HDR).
[0073] In certain embodiments, the method provided herein includes introducing the non-deletion HPFH mutation c.-195 C>G into the HBG1 regulatory region.
[0074] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-196 C>T into the HBG1 regulatory region via a DNA repair mechanism (e.g., HDR). This mutation is associated with 10–20% HbF.
[0075] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-198 T>C into the HBG1 regulatory region via a DNA repair mechanism (e.g., HDR). This mutation is associated with 18–21% HbF.
[0076] In certain embodiments, the method provided herein includes introducing the non-deletion HPFH mutation c.-201C>T into the HBG1 regulatory region.
[0077] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-251 T>C into the HBG1 regulatory region by a DNA repair mechanism (e.g., HDR).
[0078] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-499 T>A into the HBG1 regulatory region by a DNA repair mechanism (e.g., HDR).
[0079] In certain embodiments, the method provided herein involves introducing a non-deletion HPFH mutation c.-109 G>T ("Hellenic mutation") into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR). This mutation is located at the 3' end of the HBG2 CCAAT box within the promoter region (Chassanidis 2009).
[0080] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-114 C>A into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0081] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-157 C>T into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0082] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-167 C>A into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0083] In certain embodiments, the methods provided herein involve introducing a non-deletion HPFH mutation c.-202 C>G into the HBG2 regulatory region via a DNA repair mechanism (e.g., HDR). This mutation is associated with 15–25% HbF expression.
[0084] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-211 C>T into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0085] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-228 T>C into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0086] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-255 C>G into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0087] In certain embodiments, the method provided herein includes introducing a non-deletion HPFH mutation c.-309 A>G into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0088] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-369 C>G into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0089] In certain embodiments, the methods provided herein include introducing a non-deletion HPFH mutation c.-567 T>G into the HBG2 regulatory region by a DNA repair mechanism (e.g., HDR).
[0090] In certain embodiments, the methods provided herein involve deletion, disruption, or mutation of the BCL11a core-binding motif (i.e., GGCCGG) located at position c.-56 and / or another position in the γ-globin gene regulatory region for HBG1 and / or HBG2.
[0091] In certain embodiments, the methods provided herein involve altering one or more nucleotides within a GATA (e.g., GATA1) motif. In some of these embodiments, a DNA repair mechanism (e.g., HDR) is used to introduce a T>C mutation into the HBG1 GATA-binding motif within the sequence AAATATCTGT, thereby resulting in the altered sequence AAACATCTGT. This spontaneously occurring T>C HPFH mutation is associated with 40% HbF.
[0092] In certain embodiments, the methods provided herein utilize one or more DNA repair mechanisms (e.g., both NHEJ and HDR) approaches. For example, in certain embodiments, this method involves the introduction of 13 bp del c.-114~-102 into one or both alleles of HBG1 and / or HBG2, combined with an NHEJ-mediated deletion, e.g., an HDR-mediated single nucleotide alteration, and / or the introduction of 4 bp del c.-225~-222 into one or both alleles of HBG1, e.g., c.-109 G>T; c.-114 C>A; c.-114 C>T; c.-117 G>A; c.-157 C>T; c.-158 C>T; c.-167 C>T; c.-167 C>A; c.-170 G>A; c.-175 T>C; c.-175 T>G; c.-195 C>G; c.-196 C>T; c.-198 T>C; c.-201 This utilizes the introduction of one or more HBG1 and / or HBG2 alleles in the following configurations: C>T;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-251, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G;c.-499, T>A; or c.-567.
[0093] In certain embodiments, this method involves the introduction of 13 bp del c.-114~-102 into one or both alleles of HBG1 and / or HBG2, combined with HDR-mediated deletions, e.g., HDR-mediated single nucleotide alterations, and / or the introduction of 4 bp del c.-225~-222 into one or both alleles of HBG1, e.g., c.-109 G>T; c.-114 C>A; c.-114 C>T; c.-117 G>A; c.-157 C>T; c.-158 C>T; c.-167 C>T; c.-167 C>A; c.-170 G>A; c.-175 T>C; c.-175 T>G; c.-195 C>G; c.-196 C>T; c.-198 T>C; c.-201 This utilizes the introduction of one or more HBG1 and / or HBG2 alleles in the following configurations: C>T;c.-202, C>G;c.-211, C>T;c.-228, T>C;c.-251, T>C;c.-255, C>G;c.-309, A>G;c.-369, C>G;c.-499, T>A; or c.-567.
[0094] While I don't want to be bound by theory, introducing 4bp del c.-225~-222 into the HBG1 gene regulatory region would reduce 70% of γ A - Globin (γ-globin product of the HBG1 gene) and 30% γ G - The normal ratio with globin (γ-globin product of the HBG2 gene) is reversed, and γ-globin is approximately 30% A - Globin and about 70% gamma G -It is produced as globin. I don't want to be bound by theory, but γ G - Globin and gamma A - The reversal of the ratio with globin is in the subject γ G - Increases globin production. Although we do not wish to be bound by theory, simultaneous introduction of 4 bp del c.-225~-222 into the HBG1 gene regulatory region and 13 bp del c.-114~-102 into the HBG2 gene regulatory region increases the transcriptional activity of HBG2 in the target, γ G- This leads to increased globin production and an increase in HbF. While we do not wish to be bound by theory, simultaneous (a) introduction of a 4bp del c.-225~-222 into the HBG1 gene regulatory region, for example via NHEJ-mediated deletion or HDR-mediated deletion, and (b) introduction of non-deletion HPFH mutations, for example c.-109G>T;c.-114 C>T;c.-114 C>A;c.-157 C>T;c.-158 C>T;c.-167 C>T;c.-167 C>A;c.-175 T>C;c.-202 C>G;c.-211 C>T;c.-228 T>C;c.-255 C>G;c.-309 A>G;c.-369 C>G;c-567 T>G into the HBG2 gene regulatory region may increase the transcriptional activity of HBG2 in the target, γ G - This leads to increased globin production and an increase in HbF.
[0095] While I don't want to be bound by theory, introducing 4bp del c.-225~-222 into the HBG1 gene regulatory region, γ A - γ for globin production (γ-globin product of the HBG1 gene) G -Globin production (γ-globin product of the HBG2 gene) may decrease, A - Globin is γ G - More globin is produced. Although we do not wish to be bound by theory, simultaneous introduction of 4 bp del c.-225~-222 into the HBG2 gene regulatory region and 13 bp del c.-114~-102 into the HBG1 gene regulatory region may increase the transcriptional activity of HBG1, thus affecting the target γ A-Globin production increases, and HbF increases. While we do not wish to be bound by theory, simultaneous (a) introduction of 4bp del c.-225~-222 into the HBG2 gene regulatory region, for example, via NHEJ-mediated deletion or HDR-mediated deletion, and (b) introduction of non-deletion HPFH mutations, for example, c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>G;c.-175 T>C;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A into the HBG1 gene regulatory region increases the transcriptional activity of HBG1 in the subject, γ A - This can lead to increased globin production and an increase in HbF.
[0096] While I don't want to be bound by theory, the simultaneous (a) introduction of a 13bp del c.-114~-102 into the HBG1 gene regulatory region, for example, via NHEJ-mediated deletion or HDR-mediated deletion, and (b) non-deletion HPFH mutations, for example, via HDR, such as c.-109 G>T; c.-114 C>T; c.-114 C>A; c.-157 C>T; c.-158 C>T; c.-167 C>A; c.-167 C>T; c.-175 T>C; c.-202 C>G; c.-211 C>T; c.-228 T>C; c.-255 C>G; c.-309 A>G; c.-369 C>G; or c.-567 Introducing T>G into the HBG2 gene regulatory region increases the transcriptional activity of HBG2 in the target, γ G - This leads to increased globin production and an increase in HbF.
[0097] While we do not wish to be bound by theory, simultaneous (a) introduction of a 13bp del c.-114~-102 into the HBG2 gene regulatory region, for example via NHEJ-mediated deletion or HDR-mediated deletion, and (b) introduction of non-deletion HPFH mutations, such as c.-114 C>T;c.-117 G>A;c.-158 C>T;c.-167 C>T;c.-170 G>A;c.-175 T>C;c.-175 T>G;c.-195 C>G;c.-196 C>T;c.-198 T>C;c.-201 C>T;c.-251 T>C; or c.-499 T>A into the HBG1 gene regulatory region may increase the transcriptional activity of HBG1 in the target, γ A - This leads to increased globin production and an increase in HbF.
[0098] Simultaneous (a) siRNA-mediated knockdown of BCL11A and (b) siRNA-mediated knockdown of SOX6 result in increased expression of HBG1 and HBG2 (Xu 2010). In certain embodiments, the methods provided herein include inhibiting the action of BCL11A, SOX6, or BCL11A and SOX6 on HBG1 and HBG2 expression by using DNA repair mechanism (e.g., HDR, NHEJ, or NHEJ and HDR) modification of the HBG1 and HBG2 promoter regions and the erythrocyte-specific enhancer of BCL11A, either alone or in combination. In certain embodiments, the methods provided herein include reducing BCL11A expression by inhibiting the function of the erythrocyte-specific enhancer of its introns by NHEJ and HDR, and simultaneously inducing HPFH mutations for a synergistic effect on HbF production.
[0099] The embodiments described herein can be used for all classes of vertebrates, including but not limited to primates, mice, rats, rabbits, pigs, dogs, and cats.
[0100] Timing and target selection For example, in subjects who are considered to be at risk of developing β-abnormal hemoglobin disorders (e.g., SCD, β-thal) based on genetic testing, family history, or other factors, but who have not yet shown any signs or symptoms of the disease, treatment using the methods disclosed herein can be initiated before the onset of the disease. In some embodiments, treatment may be initiated before the switching of naturally occurring globins, i.e., before the transition from primarily HbF to primarily HbA. In other embodiments, treatment may be initiated after the switching of naturally occurring globins has occurred.
[0101] In certain embodiments, treatment is initiated after the onset of the disease, for example, one, two, three, four, five, six, seven, eight, nine, ten, twelve, sixteen, twenty-four, thirty-six, or forty-eight months or later, following the onset of SCD or β-thal or one or more symptoms associated therewith. In some of these embodiments, treatment is initiated in the early stages of disease progression, for example, when the subject exhibits only mild symptoms or only a portion of symptoms. Exemplary symptoms include, but are not limited to, anemia, diarrhea, fever, stunted growth, sickle cell crisis, occlusive crisis, myeloplastic crisis, and acute chest syndrome anemia, occlusive disease, hepatomegaly, thrombosis, pulmonary embolism, stroke, foot ulcers, cardiomyopathy, cardiac arrhythmias, splenomegaly, bone growth retardation and / or delayed puberty, as well as evidence of extramedullary erythropoiesis. In other embodiments, treatment is initiated well after the onset of the disease or at a more advanced stage of disease progression, for example, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 12 months, 16 months, 24 months, 36 months, or 48 months or later after the onset of SCD or β-thal. While we do not wish to be bound by theory, this treatment is considered effective when the subject is sufficiently advanced in the disease.
[0102] In certain embodiments, the methods provided herein prevent or delay the onset of one or more symptoms associated with a disease being treated. In certain embodiments, the methods provided herein result in the prevention or delay of disease progression compared to an untreated subject. In certain embodiments, the methods provided herein result in a complete cure of the disease.
[0103] In certain embodiments, the methods provided herein are performed one at a time. In other embodiments, the methods provided herein utilize multiple doses.
[0104] In certain embodiments, the subject treated using the methods provided herein is transfusion-dependent.
[0105] In certain embodiments, the method provided herein involves altering the expression of one or more gamma-globin genes (e.g., HBG1, HBG2) in vivo using CRISPR / Cas-mediated genome editing in cells. In other embodiments, the method provided herein involves altering the expression of one or more gamma-globin genes in ex vivo cells using CRISPR / Cas-mediated genome editing, and then transplanting these cells into a subject. In some of these embodiments, the cells are originally derived from the subject. In certain embodiments, the cells undergoing alteration are adult erythroid cells. In other embodiments, the cells are hematopoietic stem cells (HSCs).
[0106] In certain embodiments, the method provided herein includes the delivery of one or more gRNA molecules and one or more Cas9 polypeptides or nucleic acid sequences encoding Cas9 polypeptides to cells. In certain embodiments, the method further includes the delivery of one or more nucleic acids, for example, HDR donor templates.
[0107] In certain embodiments, one or more of these components (i.e., one or more gRNA molecules, one or more Cas9 polypeptides or nucleic acid sequences encoding Cas9 polypeptides, and one or more nucleic acids, e.g., HDR donor templates) are delivered using one or more AAV vectors, lentiviral vectors, nanoparticles, or a combination thereof.
[0108] In certain embodiments, the methods provided herein are applied to subjects having one or more mutations in the HBB gene, including one or more mutations associated with β-abnormal hemoglobinopathy, such as SCD or β-thal. Examples of such mutations include, but are not limited to, c.17A>T, c.-136C>G, c.92+1G>A, c.92+6T>C, c.93-21G>A, c.118C>T, c.316-106C>G, c.25_26delAA, c.27_28insG, c.92+5G>C, c.118C>T, c.135delC, c.315+1G>A, c.-78A>G, c .52A>T, c.59A>G, c.92+5G>C, c.124_127delTTCT, c.316-197C>T, c.-78A>G, c.52A>T, c.124_127delT TCT, c.316-197C>T, c.-138C>T, c.-79A>G, c.92+5G>C, c.75T>A, c.316-2A>G, and c.316-2A>C.
[0109] NHEJ-mediated introduction of indels into the γ-globin gene regulatory element. In certain embodiments, the methods provided herein utilize NHEJ-mediated insertions or deletions to disrupt all or part of the regulatory elements of the gamma-globin gene in order to increase the expression of the gamma-globin gene (e.g., HBG1, HBG2, or HBG1 and HBG2).
[0110] In certain embodiments, the methods provided herein utilizing NHEJ include deleting or disrupting all or part of the HBG1 or HBG2 silencer element by NHEJ, thereby resulting in silencer inactivation and subsequent increased HBG1 and / or HBG2 expression. In certain embodiments, the NHEJ-mediated deletion results in the removal of all or part of c.-114~-102 or -225~-222 in one or both alleles of HBG1, and / or the removal of all or part of c.-114~-102 in one or both alleles of HBG2. In some of these embodiments, one or more nucleotides on the 5' or 3' side of these regions are also deleted.
[0111] In certain embodiments, the methods provided herein utilizing NHEJ include introducing one or more cleavages (e.g., single-strand or double-strand breaks) into the γ-globin gene regulatory region, in some embodiments thereof, the one or more cleavages are sufficiently close to an HBG target site where it can be reasonably expected that the cleavage-inducing indels will extend to all or part of the HBG target site.
[0112] In certain embodiments, the targeting domain of the first gRNA molecule is configured to provide a cleavage event, e.g., a double-strand break or a single-strand break, sufficiently close to the HBG target site to enable NHEJ-mediated insertion or deletion at the HBG target site. In certain embodiments, the gRNA targeting domain is configured so that the cleavage event, e.g., a double-strand break or a single-strand break, is located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the HBG target site. The cleavage, e.g., a double-strand break or a single-strand break, may be located upstream or downstream of the HBG target site.
[0113] In certain embodiments, the second gRNA molecule, including the second targeting domain, is configured to provide a cleavage event, e.g., a double-strand break or a single-strand break, sufficiently close to the HBG target site, thereby enabling NHEJ-mediated insertion or deletion at the HBG target site, either alone or in combination with a cleavage positioned by the first gRNA molecule. In certain embodiments, the targeting domains of the first and second gRNA molecules are configured so that the cleavage event, e.g., a double-strand break or a single-strand break, is located independently of each gRNA molecule within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the target site. In certain embodiments, the cleavage, e.g., a double-strand break or a single-strand break, is located on both sides of the nucleotide at the HBG target site. In other embodiments, the cleavage, such as a double-strand break or a single-strand break, is located on one side of the nucleotide at the HBG target site, for example, upstream or downstream.
[0114] In certain embodiments, as discussed below, the single-strand break is accompanied by a further single-strand break positioned by a second gRNA molecule. For example, the gRNA target domain can be configured such that the break event, e.g., two single-strand breaks, are located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the HBG target site. In certain embodiments, the first and second gRNA molecules are configured such that, upon inducing Cas9 nickase, the single-strand break is accompanied by a further single-strand break positioned by a second gRNA located sufficiently close to each other, resulting in a change of the HBG target site. In certain embodiments, the first and second gRNA molecules are configured such that, for example, if Cas9 is a nickase, the single-strand break placed by the second gRNA is within 10, 20, 30, 40, or 50 nucleotides of the break placed by the first gRNA molecule. In certain embodiments, the two gRNA molecules are configured such that the breaks are at the same position on different strands or within a few nucleotides of each other, for example, substantially mimicking a double-strand break.
[0115] In certain embodiments, the double-strand break may be accompanied by further double-strand breaks positioned by the second gRNA molecule, as discussed below. For example, the targeting domain of the first gRNA molecule is configured such that the double-strand break is located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site; and the targeting domain of the second gRNA molecule is configured such that the double-strand break is located downstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site.
[0116] In certain embodiments, the double-strand break may be accompanied by two further single-strand breaks positioned by the second and third gRNA molecules. For example, the targeting domain of the first gRNA molecule is configured such that the double-strand break is located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site; and The targeting domains of the second and third gRNA molecules are configured such that two single-strand breaks are located downstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site. In certain embodiments, the targeting domains of the first, second, and third gRNA molecules are configured such that the cleavage events, e.g., double-strand breaks or single-strand breaks, are located independently of each gRNA molecule.
[0117] In certain embodiments, the first and second single-strand breaks may be accompanied by two further single-strand breaks positioned by the third and fourth gRNA molecules. For example, the targeting domains of the first and second gRNA molecules are configured such that the two single-strand breaks are located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site. Furthermore, the targeting domains of the third and fourth gRNA molecules are configured such that two single-strand breaks are located downstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site.
[0118] In certain embodiments, the method provided herein involves introducing an NHEJ-mediated deletion in a genomic sequence containing an HBG target site. In certain embodiments, the method involves introducing two double-strand breaks, one of which is located 5' to the HBG target site and the other 3' to the HBG target site (i.e., adjacent). Two gRNAs, e.g., a single (or chimeric) gRNA molecule or a modular gRNA molecule, are configured to position the two double-strand breaks opposite the HBG target site. In certain embodiments, the first double-strand break is located upstream of the mutation, and the second double-strand break is located downstream of the mutation. In certain embodiments, the two double-strand breaks are positioned to remove all or part of HBG1 c.-114~-102, HBG1 4bp del -225~-222. In one embodiment, the breaks (i.e., the two double-strand breaks) are positioned to avoid repeat elements, e.g., undesirable target chromosome elements such as Alu repeats, or endogenous splice sites.
[0119] In other embodiments, the method involves introducing two sets of breaks, one double-strand break and a pair of single-strand breaks. These two sets are adjacent to the HBG target site, i.e., one set is on the 5' side of the HBG target site and the other is on the 3' side of the HBG target site. Two gRNAs, e.g., a single (or chimeric) gRNA molecule or a modular gRNA molecule, are configured to position the two sets of breaks (either double-strand breaks or a pair of single-strand breaks) on opposite sides of the HBG target site. In certain embodiments, the breaks (i.e., two sets of breaks (either double-strand breaks or a pair of single-strand breaks)) are positioned to avoid repeat elements, e.g., unwanted target chromosome elements such as Alu repeats, or endogenous splice sites.
[0120] In other embodiments, the method involves introducing two pairs of single-strand breaks, one of which is located 5' to the HBG target site and the other 3' to the HBG target site (i.e., adjacent). Two gRNAs, e.g., a single (or chimeric) gRNA molecule or a modular gRNA molecule, are configured to place the two sets of breaks on opposite sides of the HBG target site. In certain embodiments, the breaks (i.e., two pairs of single-strand breaks) are positioned to avoid repeat elements, e.g., undesirable target chromosome elements such as Alu repeats, or endogenous splice sites.
[0121] HDR-mediated introduction of sequence changes in γ-globin gene regulatory elements In certain embodiments, the methods provided herein utilize HDRs to modify one or more nucleotides in the regulatory element of a γ-globin gene (e.g., HBG1, HBG2, or HBG1 and HBG2) in order to increase the expression of the γ-globin gene. In some of these embodiments, HDRs are used to incorporate one or more nucleotide modifications corresponding to spontaneous mutations associated with HPFH. For example, in certain embodiments, HDRs are used to incorporate one or more of the following single nucleotide changes into the HBG1 regulatory region: c.-114 C>T; c.-117 G>A; c.-158 C>T; c.-167 C>T; c.-170 G>A; c.-175 T>C; c.-175 T>G; c.-195 C>G; c.-196 C>T; c.-198 T>C; c.-201 C>T; c.-251 T>C; or c-499 T>A. In other embodiments, HDR is used to incorporate one or more of the following single nucleotide changes into the HBG2 regulatory region: c.-109 G>T; c.-114 C>A; c.-114 C>T; c.-157 C>T; c.-158 C>T; c.-167 C>T; c.-167 C>A; c.-175 T>C; c.-202 C>G; c.-211 C>T; c.-228 T>C; c.-255 C>G; c.-309 A>G; c.-369 C>G; c.-567 T>G.
[0122] In certain embodiments, the methods provided herein utilize HDR-mediated alterations (e.g., insertions or deletions) to disrupt all or part of the regulatory elements of the γ-globin gene (e.g., HBG1, HBG2, or HBG1 and HBG2) in order to increase the expression of the γ-globin gene.
[0123] In certain embodiments, the methods provided herein utilizing HDR include deletion or disruption of all or part of the HBG1 or HBG2 silencer element by HDR, resulting in silencer inactivation and subsequent increased HBG1 and / or HBG2 expression. In certain embodiments, the HDR-mediated deletion results in the removal of all or part of c.-l14~-102 or -225~-222 in one or both alleles of HBG1, and / or the removal of all or part of c.-l14~-102 in one or both alleles of HBG2. In some of these embodiments, one or more nucleotides on the 5' or 3' side of these regions are also deleted.
[0124] In certain embodiments, the methods provided herein utilizing HDR include introducing one or more cleavages (e.g., single-strand or double-strand breaks) into the γ-globin gene regulatory region, in some embodiments thereof, the one or more cleavages are sufficiently close to HBG target sites where it is reasonably foreseeable that cleavage-induced changes will extend to all or part of the HBG target sites.
[0125] In certain embodiments, HDR-mediated transformation may include the use of a template nucleic acid.
[0126] In certain embodiments, the HDR-mediated gene alteration is incorporated into one gamma-globin gene allele (e.g., one allele of HBG1 and / or HBG2). In other embodiments, the gene alteration is incorporated into both alleles (e.g., both alleles of HBG1 and / or HBG2). In either scenario, the treated subject exhibits increased gamma-globin gene expression (e.g., expression of HBG1, HBG2, or both HBG1 and HBG2).
[0127] In certain embodiments, the methods provided herein that utilize HDR include introducing one or more breaks (e.g., single-strand or double-strand breaks) sufficiently close to the target location (e.g., either 5' or 3') that enable HDR-related changes at the HBG target location.
[0128] In certain embodiments, the targeting domain of the first gRNA molecule is configured to provide a cleavage event, e.g., a double-strand break or a single-strand break, sufficiently close to the HBG target site to enable HDR-related changes at this target site. In certain embodiments, the gRNA targeting domain is configured so that the cleavage event, e.g., a double-strand break or a single-strand break, is located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the HBG target site. The cleavage, e.g., a double-strand break or a single-strand break, can be located upstream or downstream of the HBG target site.
[0129] In certain embodiments, the second, third, and / or fourth gRNA molecules are configured to provide a cleavage event, e.g., a double-strand break or a single-strand break, sufficiently close to the HBG target site (e.g., either 5' or 3') to enable HDR-related changes at this target site. In certain embodiments, the gRNA targeting domain is configured so that the cleavage event, e.g., a double-strand break or a single-strand break, is located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the HBG target site. The cleavage, e.g., a double-strand break or a single-strand break, can be located upstream or downstream of the target site.
[0130] In certain embodiments, a single-strand break is accompanied by a further single-strand break positioned by a second, third, and / or fourth gRNA molecule. For example, the gRNA targeting domain can be configured such that the break event is located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of the HBG target site. In certain embodiments, the first and second gRNA molecules are configured such that when Cas9 nickase is induced, the single-strand break is accompanied by a further single-strand break positioned by the second gRNA, which is sufficiently close to the break in the first strand, resulting in a change in the HBG target site. In certain embodiments, the first and second gRNA molecules are configured such that, for example, if Cas9 is a nickase, the single-strand breaks placed by the second gRNA are within 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides of the breaks placed by the first gRNA molecule. In certain embodiments, the two gRNA molecules are configured such that the breaks are at the same position on different strands or within a few nucleotides of each other, for example, substantially mimicking double-strand breaks.
[0131] In certain embodiments, the double-strand break may be accompanied by further double-strand breaks positioned by second, third, and / or fourth gRNA molecules. For example, the targeting domain of the first gRNA molecule can be configured such that the double-strand break is located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site; and the targeting domain of the second gRNA molecule can be configured such that the double-strand break is located downstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site.
[0132] In certain embodiments, the double-strand break may be accompanied by two further single-strand breaks positioned by the second and third gRNA molecules. For example, the targeting domain of the first gRNA molecule can be configured such that the double-strand break is located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site; and The targeting domains of the second and third gRNA molecules can be configured so that two single-strand breaks are located downstream of this target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site. In certain embodiments, the targeting domains of the first, second, and third gRNA molecules are configured so that the cleavage events, for example, a double-strand break or a single-strand break, are located independently of each gRNA molecule.
[0133] In certain embodiments, the first and second single-strand breaks may be accompanied by two further single-strand breaks positioned by the third and fourth gRNA molecules. For example, the targeting domains of the first and second gRNA molecules can be configured so that the two single-strand breaks are located upstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site. Furthermore, the targeting domains of the third and fourth gRNA molecules may be configured such that two single-strand breaks are located downstream of the HBG target site, for example, within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, or 500 nucleotides of this target site.
[0134] Guide RNA (gRNA) molecule As used herein, the term gRNA molecule refers to a nucleic acid that facilitates the specific targeting or homing of a gRNA molecule / Cas9 molecule complex to a target nucleic acid. A gRNA molecule may be unimolecule (having a single RNA molecule) (e.g., a chimera) or modular (multiple, typically containing two separate RNA molecules). The gRNA molecules provided herein include a targeting domain consisting of, or essentially consisting of, a nucleic acid sequence that is fully or partially complementary to the target domain. In certain embodiments, the gRNA molecule further includes one or more additional domains, including, for example, a first complementarity domain, a binding domain, a second complementarity domain, an adjacency domain, a tail domain, and a 5' elongation domain. Each of these domains is described in more detail below. In certain embodiments, one or more domains in the gRNA molecule include a nucleotide sequence that is identical to, or sequence-homographic to, a naturally occurring sequence derived from, for example, S. pyogenes, S. aureus, or S. thermophilus.
[0135] Several exemplary gRNA structures are shown in Figures 1A–1I. The three-dimensional morphology of the active form of the gRNA, or highly complementary regions with respect to intra-chain or inter-chain interactions, may be shown in the double helix in Figures 1A–1I and other depictions provided herein. Figure 7 shows gRNA domain nomenclature using the gRNA sequence SEQ ID NO: 42, which contains one hairpin loop within the tracrRNA-derived region. In certain embodiments, the gRNA may contain two or more (e.g., two, three, or more) hairpin loops in this region (e.g., Figures 1H–1I).
[0136] In certain embodiments, the single-molecule gRNA or chimeric gRNA is preferably 5' to 3': Complementary targeting domain to the target domain in the γ-globin gene A regulatory region, for example, a targeting domain from any of sequence numbers 251-901; First complementary domain; Linked domains; Second complementary domain (complementary to the first complementary domain); Adjacent domains; and Includes optional tail domains.
[0137] In a particular embodiment, modular gRNA is: The first chain, preferably from 5' to 3': A targeting domain complementary to the target domain in the γ-globin gene regulatory region, for example, a targeting domain from any of sequence numbers 251-901; and The first chain, including the first complementarity domain; and The second chain, preferably from 5' to 3': Optionally selected 5' elongated domain; Second complementary domain; Adjacent domains; and Includes a second strand containing an optional tail domain.
[0138] Targeted domains The targeting domain (sometimes referred to instead as the guide sequence or complementary region) comprises, or essentially comprises, a nucleic acid sequence that is complementary or partially complementary to the target nucleic acid sequence in the γ-globin gene regulatory region. The nucleic acid sequence in the γ-globin gene regulatory region that is complementary or partially complementary to all or part of the targeting domain is referred to herein as the target domain. In certain embodiments, the target domain includes an HBG target site. In other embodiments, the HBG target site is located outside (i.e., upstream or downstream of) the target domain. In certain embodiments, the target domain is located entirely within the γ-globin gene regulatory region, for example, within a regulatory element associated with the γ-globin gene or a regulatory element associated with a gene encoding a repressor of γ-globin gene expression. In other embodiments, all or part of the target domain is located outside the γ-globin gene regulatory region, for example, within the HBG1 or HBG2 coding region, exon, or intron.
[0139] Methods for selecting targeting domains are known in the art (see, e.g., Fu 2014; Sternberg 2014). Examples of suitable targeting domains for use in the methods, compositions, and kits described herein include those listed in Sequence IDs 251-901.
[0140] The strand of target nucleic acid containing the target domain is referred to herein as the “complementary strand” because it is complementary to the targeting domain sequence. Since the targeting domain is part of a gRNA molecule, it contains the base uracil (U) rather than thymine (T); conversely, any DNA molecule encoding a gRNA molecule will contain thymine rather than uracil. In a targeting domain / target domain pair, the uracil base in the targeting domain will pair with the adenine base in the target domain. In certain embodiments, the degree of complementarity between the targeting domain and the target domain is sufficient to enable the Cas9 molecule to target the target nucleic acid.
[0141] In certain embodiments, the targeting domain includes a core domain and an optional secondary domain. In some of these embodiments, the core domain is located on the 3' side of the secondary domain, and in some of these embodiments, the core domain is located on or near the 3' side of the targeting domain. In some of these embodiments, the core domain consists of or is composed of approximately 8 to approximately 13 nucleotides at the 3' end of the targeting domain. In certain embodiments, only the core domain is complementary to, or partially complementary to, the corresponding portion of the target domain, and in some of these embodiments, the core domain is fully complementary to the corresponding portion of the target domain. In other embodiments, the secondary domain is also complementary to, or partially complementary to, a portion of the target domain. In certain embodiments, the core domain is complementary to, or partially complementary to, the core domain target in the target domain, while the secondary domain is complementary to, or partially complementary to, the secondary domain target in the target domain. In certain embodiments, the core domain and the secondary domain have a degree of complementarity with respect to their respective corresponding portions of the target domain. In other embodiments, the degree of complementarity between the core domain and its targets and the degree of complementarity between the secondary domain and its targets may differ. In some of these embodiments, the core domain has a higher degree of complementarity with respect to its targets than the secondary domain, while in other embodiments, the secondary domain has a higher degree of complementarity than the core domain.
[0142] In certain embodiments, the targeting domain and / or the core domain within the targeting domain is 3–100, 5–100, 10–100, or 20–100 nucleotides long, and in some of these embodiments, the targeting domain or core domain is 3–15, 3–20, 5–20, 10–20, 15–20, 5–50, or 20–50 nucleotides long. In certain embodiments, the targeting domain and / or the core domain within the targeting domain is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides long. In certain embodiments, the targeting domain and / or the core domain within the targeting domain is 6+ / -2, 7+ / -2, 8+ / -2, 9+ / -2, 10+ / -2, 10+ / -4, 10+ / -5, 11+ / -2, 12+ / -2, 13+ / -2, 14+ / -2, 15+ / -2, or 16+ / -2, 20+ / -5, 30+ / -5, 40+ / -5, 50+ / -5, 60+ / -5, 70+ / -5, 80+ / -5, 90+ / -5, or 100+ / -5 nucleotides long.
[0143] In certain embodiments where the targeting domain includes a core domain, the core domain is 3 to 20 nucleotides long, and in some of these embodiments, the core domain is 5 to 15 or 8 to 13 nucleotides long. In certain embodiments where the targeting domain includes a secondary domain, the secondary domain is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides long. In certain embodiments where the targeting domain includes a core domain of 8 to 13 nucleotides long, the targeting domain is 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, or 16 nucleotides long, and the secondary domain is 13 to 18, 12 to 17, 11 to 16, 10 to 15, 9 to 14, 8 to 13, 7 to 12, 6 to 11, 5 to 10, 4 to 9, or 3 to 8 nucleotides long.
[0144] In certain embodiments, the targeting domain is fully complementary to the target domain. Similarly, if the targeting domain includes a core domain and / or secondary domains, in certain embodiments, one or both of the core domain and / or secondary domains are fully complementary to the corresponding portion of the target domain. In other embodiments, the targeting domain is partially complementary to the target domain, and in some of these embodiments, where the targeting domain includes a core domain and / or secondary domains, one or both of the core domain and / or secondary domains are partially complementary to the corresponding portion of the target domain. In some of these embodiments, the nucleic acid sequence of the targeting domain, or the core domain or targeting domain within the targeting domain, is at least 80, 85, 90, or 95% complementary to the target domain or its corresponding portion. In certain embodiments, the targeting domain and / or the core or secondary domain within the targeting domain contains one or more nucleotides that are not complementary to the target domain or any part thereof, and in some of these embodiments, the targeting domain and / or the core or secondary domain within the targeting domain contains one, two, three, four, five, six, seven, or eight nucleotides that are not complementary to the target domain. In certain embodiments, the core domain contains one, two, three, four, or five nucleotides that are not complementary to the corresponding portion of the target domain. In certain embodiments where the targeting domain contains one or more nucleotides that are not complementary to the target domain, the one or more of the non-complementary nucleotides are located within five nucleotides of the 5' or 3' end of the targeting domain. In some of these embodiments, the targeting domain contains one, two, three, four, or five nucleotides within its 5' end, 3' end, or within five nucleotides of both its 5' and 3' ends that are not complementary to the target domain. In certain embodiments, the targeting domain includes two or more nucleotides that are not complementary to the targeting domain, the two or more non-complementary nucleotides are adjacent to each other, and in some of these embodiments, the two or more consecutive non-complementary nucleotides are located within five nucleotides of the 5' or 3' end of the targeting domain.In other embodiments, each of the two or more consecutive non-complementary nucleotides is located at least six nucleotides from the 5' and 3' ends of the targeting domain.
[0145] In certain embodiments, the targeting domain, core domain, and / or secondary domain are unmodified. In other embodiments, the targeting domain, core domain, and / or secondary domain, or one or more nucleotides within them, have modifications, but are not limited to, those described below. In certain embodiments, one or more nucleotides of the targeting domain, core domain, and / or secondary domain may include 2' modifications (e.g., modifications at the 2' position on ribose), e.g., 2-acetylation, e.g., 2'-methylation. In certain embodiments, the backbone of the targeting domain may be modified with a phosphorothioate. In certain embodiments, modifications to one or more nucleotides of the targeting domain, core domain, and / or secondary domain make the targeting domain and / or the gRNA containing the targeting domain less susceptible to degradation or more biocompatible, e.g., less immunogenic. In certain embodiments, the targeting domain and / or core or secondary domain includes one, two, three, four, five, six, seven, or eight or more modifications, and in some of these embodiments, the targeting domain and / or core or secondary domain includes one, two, three, or four modifications within the five nucleotides of each 5' end and / or one, two, three, or four modifications within the five nucleotides of each 3' end. In certain embodiments, the targeting domain and / or core or secondary domain includes modifications in two or more consecutive nucleotides.
[0146] In certain embodiments where the targeting domain includes a core and secondary domains, the core and secondary domains contain the same number of modifications. In some of these embodiments, neither domain contains any modifications. In other embodiments, the core domain contains more modifications than the secondary domains, or vice versa.
[0147] In certain embodiments, modifications to one or more nucleotides in the targeting domain, including within the core or secondary domain, are selected so as not to impede targeting efficiency, which can be evaluated by testing candidate modifications using the system described below. gRNAs having candidate targeting domains with selected lengths, sequences, degrees of complementarity, or degrees of modification can be evaluated using the system described below. Candidate targeting domains can be placed alone or together with one or more other candidate modifications in a gRNA / Cas9 molecular system known to be functional for a selected target and evaluated.
[0148] In certain embodiments, all of the modified nucleotides are complementary to and can hybridize with the corresponding nucleotides present in the target domain. In other embodiments, one, two, three, four, five, six, seven, or eight or more nucleotides are complementary to and can hybridize with the corresponding nucleotides present in the target domain.
[0149] Figures 1A to 1I provide examples of the arrangement of targeting domains within gRNA.
[0150] First and second complementary domains The first and second complementarity domains (sometimes referred to as the crRNA-derived hairpin sequence and the tracrRNA-derived hairpin sequence, respectively) are fully or partially complementary to each other. In certain embodiments, the degree of complementarity is sufficient for the two domains to form a double-stranded region under at least some physiological conditions. In certain embodiments, the degree of complementarity between the first and second complementarity domains, along with other properties of the gRNA, is sufficient to enable the targeting of the Cas9 molecule to the target nucleic acid. Examples of the first and second complementarity domains are shown in Figures 1A–1G.
[0151] In certain embodiments (see, for example, Figures 1A-1B), the first and / or second complementary domains contain one or more nucleotides that are not complementary to the corresponding complementary domain. In certain embodiments, the first and / or second complementary domains contain 1, 2, 3, 4, 5, or 6 nucleotides that are not complementary to the corresponding complementary domain. For example, the second complementary domain may contain 1, 2, 3, 4, 5, or 6 nucleotides that do not pair with the corresponding nucleotide in the first complementary domain. In some embodiments, the nucleotides on the first or second complementary domain that are not complementary to the corresponding complementary domain loop out from the double helix formed between the first and second complementary domains. In some of these embodiments, the unpaired loop-out is located on the second complementary domain, and in some of these embodiments, the unpaired region begins at the 1st, 2nd, 3rd, 4th, 5th, or 6th nucleotide from the 5' end of the second complementary domain.
[0152] In certain embodiments, the first complementary domain is 5-30, 5-25, 7-25, 5-24, 5-23, 7-22, 5-22, 5-21, 5-20, 7-18, 7-15, 9-16, or 10-14 nucleotides long, and in some of these embodiments, the first complementary domain is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In certain embodiments, the second complementary domain is 5-27, 7-27, 7-25, 5-24, 5-23, 5-22, 5-21, 7-20, 5-20, 7-18, 7-17, 9-16, or 10-14 nucleotides long, and in some of these embodiments, the second complementary domain is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides long. In certain embodiments, the first and second complementary domains are each independently 6+ / -2, 7+ / -2, 8+ / -2, 9+ / -2, 10+ / -2, 11+ / -2, 12+ / -2, 13+ / -2, 14+ / -2, 15+ / -2, 16+ / -2, 17+ / -2, 18+ / -2, 19+ / -2, or 20+ / -2, 21+ / -2, 22+ / -2, 23+ / -2, or 24+ / -2 nucleotide lengths. In certain embodiments, the second complementary domain is, for example, 2, 3, 4, 5, or 6 nucleotides longer than the first complementary domain.
[0153] In certain embodiments, the first and / or second complementary domains each independently comprise three subdomains, which are a 5' subdomain, a central subdomain, and a 3' subdomain, in the 5' to 3' direction. In certain embodiments, the 5' and 3' subdomains of the first complementary domain are fully or partially complementary to the 3' and 5' subdomains of the second complementary domain, respectively.
[0154] In certain embodiments, the 5' subdomain of the first complementary domain is 4–9 nucleotides long, and in some of these embodiments, the 5' subdomain is 4, 5, 6, 7, 8, or 9 nucleotides long. In certain embodiments, the 5' subdomain of the second complementary domain is 3–25, 4–22, 4–18, or 4–10 nucleotides long, and in some of these embodiments, the 5' domain is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In certain embodiments, the central subdomain of the first complementary domain is 1, 2, or 3 nucleotides long. In certain embodiments, the central subdomain of the second complementary domain is 1, 2, 3, 4, or 5 nucleotides long. In certain embodiments, the 3' subdomain of the first complementary domain is 3–25, 4–22, 4–18, or 4–10 nucleotides long, and in some of these embodiments, the 3' subdomain is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides long. In certain embodiments, the 3' subdomain of the second complementary domain is 4–9, for example, 4, 5, 6, 7, 8, or 9 nucleotides long.
[0155] The first and / or second complementary domains may share homology with or derive from a naturally occurring or reference first and / or second complementary domain. In some of these embodiments, the first and / or second complementary domains have at least 50%, 60%, 70%, 80%, 85%, 90%, or 95% homology with a naturally occurring or reference first and / or second complementary domain, or differ from it by 1, 2, 3, 4, 5, or 6 or fewer nucleotides. In some of these embodiments, the first and / or second complementary domains may have at least 50%, 60%, 70%, 80%, 85%, 90%, or 95% homology with S. pyogenes or S. aureus.
[0156] In certain embodiments, the first and / or second complementary domains are unmodified. In other embodiments, the first and / or second complementary domains, or one or more nucleotides within them, have modifications, but are not limited to, those described below. In certain embodiments, one or more nucleotides of the first and / or second complementary domains may include 2' modifications (e.g., modifications at the 2' position on ribose), such as 2-acetylation or 2-methylation. In certain embodiments, the backbone of the targeting domain can be modified with a phosphorothioate. In certain embodiments, modifications to one or more nucleotides of the first and / or second complementary domains make the gRNA containing the first and / or second complementary domains less susceptible to degradation or more biocompatible, for example, less immunogenic. In certain embodiments, the first and / or second complementary domains include one, two, three, four, five, six, seven, or eight or more modifications, and in some of these embodiments, the first and / or second complementary domains each independently include one, two, three, or four modifications within the five nucleotides of their respective 5' end, 3' end, or both of their 5' and 3' ends. In other embodiments, the first and / or second complementary domains each independently include no modifications within the five nucleotides of their respective 5' end, 3' end, or both of their 5' and 3' ends. In certain embodiments, one or both of the first and second complementary domains include modifications in two or more consecutive nucleotides.
[0157] In certain embodiments, modifications to one or more nucleotides in the first and / or second complementarity domains are selected so as not to impede targeting efficiency, which can be evaluated by testing the candidate modifications in the systems described below. gRNA molecules having candidate first or second complementarity domains having selected lengths, sequences, degrees of complementarity, or degrees of modification can be evaluated using the systems described below. Candidate complementarity domains can be placed alone or together with one or more other candidate modifications in a gRNA / Cas9 molecular system known to be functional for a selected target and evaluated.
[0158] In certain embodiments, the double-stranded region formed by the first and second complementary domains is, for example, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 base pairs long (excluding any loop-out or unpaired nucleotides).
[0159] In certain embodiments, the first and second complementary domains, when doubled, contain 11 pair-forming nucleotides (see, for example, the gRNA of SEQ ID NO: 48). In certain embodiments, the first and second complementary domains, when doubled, contain 15 pair-forming nucleotides (see, for example, the gRNA of SEQ ID NO: 50). In certain embodiments, the first and second complementary domains, when doubled, contain 16 pair-forming nucleotides (see, for example, the gRNA of SEQ ID NO: 51). In certain embodiments, the first and second complementary domains, when doubled, contain 21 pair-forming nucleotides (see, for example, the gRNA of SEQ ID NO: 29).
[0160] In certain embodiments, one or more nucleotides are exchanged between the first and second complementary domains to remove the poly-U sequence. For example, nucleotides 23 and 48 or nucleotides 26 and 45 of the gRNA of SEQ ID NO: 48 can be exchanged to produce the gRNA of SEQ ID NO: 49 or 31, respectively. Similarly, nucleotides 23 and 39 of the gRNA of SEQ ID NO: 29 can be exchanged with nucleotides 50 and 68 to produce the gRNA of SEQ ID NO: 30.
[0161] Linked domains In monomolecule or chimeric gRNAs, the ligation domain is located between the first and second complementary domains and functions to bind them together. Figures 1B-1E provide examples of ligation domains. In certain embodiments, part of the ligation domain is derived from a crRNA region, and another part is derived from a tracrRNA region.
[0162] In certain embodiments, the linking domain covalently connects the first and second complementary domains. In certain embodiments, the linking domain consists of or includes covalent bonds. In other embodiments, the linking domain non-covalently connects the first and second complementary domains. In certain embodiments, the linking domain has a nucleotide length of 10 or less, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In other embodiments, the linking domain has a nucleotide length greater than 10, e.g., 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 or more nucleotides. In certain embodiments, the linked domains are 2-50, 2-40, 2-30, 2-20, 2-10, 2-5, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, 10-20, 10-15, 20-100, 20-90, 20-80, 20-70, 20-60, 20-50, 20-40, 20-30, or 20-25 nucleotides long. In certain embodiments, the linked domain is 10+ / -5, 20+ / -5, 30+ / -5, 30+ / -10, 40+ / -5, 40+ / -10, 50+ / -5, 50+ / -10, 60+ / -5, 60+ / -10, 70+ / -5, 70+ / -10, 80+ / -5, 80+ / -10, 90+ / -5, 90+ / -10, 100+ / -5, or 100+ / -10 nucleotides long.
[0163] In certain embodiments, the binding domain shares homology with or derives from a naturally occurring sequence, for example, the tracrRNA sequence located at the 5' end of the second complementarity domain. In certain embodiments, the binding domain has at least 50%, 60%, 70%, 80%, 90%, or 95% homology with the binding domains described herein, for example, the binding domains of Figures 1B–1E, or differs from them by only 1, 2, 3, 4, 5, or 6 nucleotides.
[0164] In certain embodiments, the linking domain contains no modifications whatsoever. In other embodiments, the linking domain or one or more nucleotides within it have modifications, but are not limited to, those described below. In certain embodiments, one or more nucleotides of the linking domain may include 2' modifications (e.g., modifications at the 2' position on ribose), e.g., 2-acetylation, e.g., 2'-methylation. In certain embodiments, the backbone of the linking domain can be modified with a phosphorothioate. In certain embodiments, modifications to one or more nucleotides of the linking domain make the linking domain and / or the gRNA containing the linking domain less susceptible to degradation or more biocompatible, e.g., less immunogenic. In certain embodiments, the linking domain contains one, two, three, four, five, six, seven, or eight or more modifications, and in some of these embodiments, the linking domain contains one, two, three, or four modifications within five nucleotides from its 5' end and / or 3' end. In certain embodiments, the linking domain contains modifications to two or more consecutive nucleotides.
[0165] In certain embodiments, modifications to one or more nucleotides in the ligation domain are selected so as not to impede targeting efficiency, which can be evaluated by testing the candidate modifications in the systems described below. gRNAs having candidate ligation domains having selected lengths, sequences, degrees of complementarity, or degrees of modification can be evaluated in the systems described below. Candidate ligation domains can be placed alone or together with one or more other candidate modifications in a gRNA / Cas9 molecular system known to be functional for a selected target and evaluated.
[0166] In certain embodiments, the linking domain typically includes a double-stranded region adjacent to, or within 1, 2, or 3 nucleotides from, the 3' end of the first complementary domain and / or the 5' end of the second complementary domain. In some of these embodiments, the double-stranded region of the linking domain is 10+ / -5, 15+ / -5, 20+ / -5, or 30+ / -5 base pairs long. In certain embodiments, the double-stranded region of the linking domain is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 base pairs long. In certain embodiments, the sequences forming the double-stranded region are perfectly complementary. In other embodiments, one or both of the sequences forming the double-stranded region include one or more nucleotides (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 nucleotides) that are not complementary to the other double-stranded sequence.
[0167] 5' elongated domain In one embodiment, the modular gRNA disclosed herein includes a 5' elongation domain, i.e., one or more additional nucleotides on the 5' side of a second complementarity domain (see, for example, Figure 1A). In certain embodiments, the 5' elongation domain is 2 to 10 or more, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, or 2 to 4 nucleotides long, and in some of these embodiments, the 5' elongation domain is 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more nucleotides long.
[0168] In certain embodiments, the 5' elongation domain nucleotides do not include modifications such as those of the type provided in Section VIII. However, in certain embodiments, the 5' elongation domain includes one or more modifications, such as modifications that make it less susceptible to degradation or more biocompatible, such as lower immunogenicity. As an example, the backbone of the 5' elongation domain may be modified with a phosphorothioate or with other modifications described below. In certain embodiments, the nucleotides of the 5' elongation domain may include 2' modifications (e.g., modifications at the 2' position on ribose), such as 2-acetylation, 2' methylation, or other modifications described below.
[0169] In certain embodiments, the 5' elongation domain may contain approximately 1, 2, 3, 4, 5, 6, 7, or 8 modifications. In certain embodiments, the 5' elongation domain may contain approximately 1, 2, 3, or 4 modifications within the first 5 nucleotides of its 5' end, for example, in a modular gRNA molecule. In certain embodiments, the 5' elongation domain may contain approximately 1, 2, 3, or 4 modifications within the first 5 nucleotides of its 3' end, for example, in a modular gRNA molecule.
[0170] In certain embodiments, the 5' elongation domain includes modifications to, for example, two consecutive nucleotides located within five nucleotides of the 5' end of the 5' elongation domain, within five nucleotides of the 3' end of the 5' elongation domain, or more than five nucleotides away from one or both ends of the 5' elongation domain. In certain embodiments, two consecutive nucleotides located within five nucleotides of the 5' end of the 5' elongation domain, within five nucleotides of the 3' end of the 5' elongation domain, or more than five nucleotides away from one or both ends of the 5' elongation domain are not modified. In certain embodiments, creotides located within five nucleotides of the 5' end of the 5' elongation domain, within five nucleotides of the 3' end of the 5' elongation domain, or more than five nucleotides away from one or both ends of the 5' elongation domain are not modified.
[0171] Modifications in the 5' elongation domain may be selected so as not to interfere with the gRNA molecular efficiency, which is assessed by testing candidate modifications in the systems described in Section IV. gRNAs having candidate 5' elongation domains with selected lengths, sequences, degrees of complementarity, or degrees of modification may be evaluated in the systems described below. Candidate 5' elongation domains may be evaluated either alone or in combination with one or more other candidate modifications in gRNA / Cas9 molecular systems known to be functional for selected targets.
[0172] In certain embodiments, the 5' extension domain has at least 60%, 70%, 80%, 85%, 90% or 95% homology with a 5' extension domain of natural origin, such as S. pyogenes, S. aureus, or S. thermophilus, and a reference 5' extension domain described herein, such as the 5' extension domain described in FIGS. 1A - 1G, or differs from it by 1, 2, 3, 4, 5, or 6 nucleotides or fewer.
[0173] Adjacent domain FIGS. 1A - 1G provide examples of adjacent domains. In certain embodiments, the adjacent domain is 5 - 20 nucleotides in length, such as 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides in length. In some of these embodiments, the adjacent domain is 6 + / - 2, 7 + / - 2, 8 + / - 2, 9 + / - 2, 10 + / - 2, 11 + / - 2, 12 + / - 2, 13 + / - 2, 14 + / - 2, 14 + / - 2, 16 + / - 2, 17 + / - 2, 18 + / - 2, 19 + / - 2, or 20 + / - 2 nucleotides in length. In certain embodiments, the adjacent domain is 5 - 20, 7 - 18, 9 - 16, or 10 - 14 nucleotides in length.
[0174] In certain embodiments, the adjacent domain may share homology with or be derived from an adjacent domain of natural origin. In some of these embodiments, the adjacent domain has at least 50%, 60%, 70%, 80%, 85%, 90%, or 95% homology with an adjacent domain of natural origin, such as those disclosed herein, including those described in FIGS. 1A - 1G, such as S. pyogenes, S. aureus, or S. thermophilus, or differs from it by 1, 2, 3, 4, 5, or 6 nucleotides or fewer.
[0175] In certain embodiments, the adjacent domain contains no modifications at all. In other embodiments, the adjacent domain or one or more nucleotides therein contain modifications, including but not limited to, modifications such as those described herein. In certain embodiments, one or more nucleotides of the adjacent domain may include 2'-modifications (e.g., modifications at the 2'-position on the ribose), such as 2-acetylation, such as 2'-methylation. In certain embodiments, the backbone of the adjacent domain can be modified with phosphorothioate. In certain embodiments, the modification to one or more nucleotides of the adjacent domain renders the adjacent domain and / or the gRNA comprising the adjacent domain less susceptible to degradation or more biocompatible, e.g., less immunogenic. In certain embodiments, the adjacent domain contains 1, 2, 3, 4, 5, 6, 7, or 8 or more modifications, and in some of these embodiments, the adjacent domain contains 1, 2, 3, or 4 modifications within 5 nucleotides from its 5'-end and / or 3'-end. In certain embodiments, the adjacent domain contains modifications in 2 or more consecutive nucleotides.
[0176] In certain embodiments, the modification to one or more nucleotides in the adjacent domain is selected so as not to interfere with the targeting efficiency, which can be evaluated by testing candidate modifications in the systems described below. A gRNA having a candidate adjacent domain with a selected length, sequence, degree of complementarity, or degree of modification can be evaluated in the systems described below. The candidate adjacent domain can be placed and evaluated alone or together with one or more other candidate changes in a gRNA molecule / Cas9 molecule system that is known to be functional for the selected target.
[0177] Tailed domain A variety of tailed domains are suitable for use in the gRNA molecules disclosed herein. Figures 1A and 1C-1G provide examples of such tailed domains.
[0178] In certain embodiments, the tail domain is absent. In other embodiments, the tail domain is 1 to 100 nucleotides long, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long. In certain embodiments, the tail domain is 1 to 5, 1 to 10, 1 to 15, 1 to 20, 1 to 50, 10 to 100, 20 to 100, 10 to 90, 20 to 90, 10 to 80, 20 to 80, 10 to 70, 20 to 70, 10 to 60, 20 to 60, 10 to 50, 20 to 50, 10 to 40, 20 to 40, 10 to 30, 20 to 30, 20 to 25, 10 to 20, or 10 to 15 nucleotides long. In certain embodiments, the tail domain has a nucleotide length of 5+ / -5, 10+ / -5, 20+ / -5, 25+ / -10, 30+ / -10, 30+ / -5, 40+ / -10, 40+ / -5, 50+ / -10, 50+ / -5, 60+ / -10, 60+ / -5, 70+ / -10, 70+ / -5, 80+ / -10, 80+ / -5, 90+ / -10, 90+ / -5, 100+ / -10, or 100+ / -5.
[0179] In certain embodiments, the tail domain may share homology with or derive homology from a naturally occurring tail domain or the 5' end of a naturally occurring tail domain. In some of these embodiments, the adjacent domain has at least 50%, 60%, 70%, 80%, 85%, 90%, or 95% homology with a naturally occurring tail domain disclosed herein, such as those described in Figures 1A and 1C-1G, including the tail domains of S. pyogenes, S. aureus, or S. thermophilus, or differs from them by 1, 2, 3, 4, 5, or 6 or fewer nucleotides.
[0180] In certain embodiments, the tail domains include sequences that are complementary to each other and form a double-stranded region under at least some physiological conditions. In some of these embodiments, the tail domains include tail double-stranded domains that can form a tail double-stranded region. In certain embodiments, the tail double-stranded region is 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 base pairs long. In certain embodiments, the tail domains include a single-stranded domain on the 3' side of a tail double-stranded domain that does not form a double helix. In some of these embodiments, the single-stranded domain is 3 to 10 nucleotides long, for example, 3, 4, 5, 6, 7, 8, 9, 10, or 4 to 6 nucleotides long.
[0181] In certain embodiments, the tail domain contains no modifications whatsoever. In other embodiments, the tail domain or one or more nucleotides within it contain modifications, including, but not limited to, those described herein. In certain embodiments, one or more nucleotides of the tail domain may contain 2' modifications (e.g., modifications at the 2' position on ribose), e.g., 2-acetylation, e.g., 2'-methylation. In certain embodiments, the backbone of the tail domain can be modified with a phosphorothioate. In certain embodiments, modifications to one or more nucleotides of the tail domain make the tail domain and / or the gRNA containing the tail domain less susceptible to degradation or more biocompatible, e.g., less immunogenic. In certain embodiments, the tail domain contains one, two, three, four, five, six, seven, or eight or more modifications, and in some of these embodiments, the tail domain contains one, two, three, or four modifications within five nucleotides from its 5' end and / or 3' end. In certain embodiments, the tail domain contains modifications to two or more consecutive nucleotides.
[0182] In certain embodiments, modifications to one or more nucleotides in the tail domain are selected so as not to impede targeting efficiency, which can be evaluated by testing candidate modifications as described below. gRNAs having candidate tail domains with selected lengths, sequences, degrees of complementarity, or degrees of modification can be evaluated using the systems described below. Candidate tail domains can be placed alone or together with one or more other candidate modifications in a gRNA / Cas9 molecular system known to be functional for a selected target and evaluated.
[0183] In certain embodiments, the tail domain includes 3' terminal nucleotides relevant to in vitro or in vivo transcription methods. When the T7 promoter is used for in vitro transcription of gRNA, these nucleotides may be any nucleotides present prior to the 3' end of the DNA template. When the U6 promoter is used for in vivo transcription, these nucleotides may be a UUUUU sequence. When the H1 promoter is used for transcription, these nucleotides may be a UUUU sequence. When an alternating pol-III promoter is used, these nucleotides may be, for example, a variety of uracil bases depending on the termination signal of the pol-III promoter, or they may consist of alternating bases.
[0184] In certain embodiments, the adjacent and tail domains together include, are composed of, or are substantially composed of, the sequences described in Sequence ID No. 32, 33, 34, 35, 36, or 37.
[0185] Exemplary single-molecule / chimeric gRNAs In certain embodiments, the monomolecule or chimeric gRNA disclosed herein has the structure 5'[targeting domain]-[first complementarity domain]-[linking domain]-[second complementarity domain]-[adjacent domain]-[tail domain]-3', During the ceremony, The targeting domain includes a core domain and optionally a secondary domain, and is 10 to 50 nucleotides long; The first complementarity domain is 5 to 25 nucleotides long and, in certain embodiments, has at least 50, 60, 70, 80, 85, 90, or 95% homology with the first complementarity domain of reference disclosed herein; The linking domain is 1 to 5 nucleotides long; The second complementarity domain is 5 to 27 nucleotides long and, in certain embodiments, has at least 50, 60, 70, 80, 85, 90, or 95% homology with the second complementarity domain of reference disclosed herein; The adjacent domains are 5 to 20 nucleotides long and, in certain embodiments, have at least 50, 60, 70, 80, 85, 90, or 95% homology with the reference adjacent complementarity domain disclosed herein; and The tail domain is absent, or the nucleotide sequence is 1 to 50 nucleotides long and, in certain embodiments, has at least 50, 60, 70, 80, 85, 90, or 95% homology to the reference tail domain disclosed herein.
[0186] In certain embodiments, the single gRNA disclosed herein preferably comprises: a targeting domain comprising, for example, 10 to 50 nucleotides in the 5' to 3 directions; a first complementarity domain comprising, for example, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides; a linking domain; a second complementarity domain; an adjacent domain; and a tail domain. During the ceremony, (a) The adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) The second complementary domain has at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide; or (c) The first complementary domain has at least 16, 19, 21, 26, 31, 32, 36, 41, 46, 50, 51, or 54 nucleotides on the 3' side of the last nucleotide of the second complementary domain which is complementary to its corresponding nucleotide.
[0187] In certain embodiments, the sequences from (a), (b), and / or (c) have at least 50%, 60%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% homology with the corresponding sequences of naturally occurring gRNAs or gRNAs described herein.
[0188] In certain embodiments, the adjacent and tail domains together comprise at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides.
[0189] In certain embodiments, the 3' end of the last nucleotide of the second complementary domain contains at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides.
[0190] In certain embodiments, there are at least 16, 19, 21, 26, 31, 32, 36, 41, 46, 50, 51, or 54 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to the corresponding nucleotide of the first complementary domain.
[0191] In certain embodiments, the targeting domain comprises, consists of, or consists essentially of 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 contiguous nucleotides) that are complementary or partially complementary to the target domain or a portion thereof. For example, the targeting domain is 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides in length. In some of these embodiments, the targeting domain is complementary to the target domain over the entire length of the targeting domain, the entire length of the target domain, or both.
[0192] In certain embodiments, the single molecule or chimeric gRNA molecules (comprising a targeting domain, a first complementary domain, a linker domain, a second complementary domain, an adjacent domain, and optionally a tail domain) disclosed herein comprise the nucleotide sequence set forth in SEQ ID NO: 42, where the targeting domain is listed as 20 Ns (residues 1 - 20), which may range in length from 16 - 26 nucleotides, wherein the last six residues (residues 97 - 102) represent the termination signal of the U6 promoter, which may be absent or present in fewer numbers. In certain embodiments, the single molecule or chimeric gRNA molecule is a S. pyogenes gRNA molecule.
[0193] In certain embodiments, the monomolecule or chimeric gRNA molecule disclosed herein (including a targeting domain, a first complementarity domain, a linking domain, a second complementarity domain, an adjacency domain, and optionally, a tail domain) comprises the nucleotide sequence described in Sequence ID No. 38, where the targeting domain is listed as 20 N(residues 1-20), which may be in the range of 16-26 nucleotides in length, with the last six residues (residues 97-102) representing the termination signal of the U6 promoter, which may be absent or fewer in number. In certain embodiments, the monomolecule or chimeric gRNA molecule is a S. aureus gRNA molecule.
[0194] The sequences and structures of exemplary chimeric gRNAs are also shown in Figures 1H-1I.
[0195] Exemplary modular gRNA In certain embodiments, the modular gRNA disclosed herein preferably comprises a first strand comprising: a first strand comprising: a first strand comprising: a first strand comprising: optionally: a 5' extension domain; a second strand comprising: a second complementation domain; an adjacent domain; and a tail domain; in a third direction from 5'; During the ceremony, (a) The adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) The second complementary domain has at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide; or (c) The first complementary domain has at least 16, 19, 21, 26, 31, 32, 36, 41, 46, 50, 51, or 54 nucleotides on the 3' side of the last nucleotide of the second complementary domain which is complementary to its corresponding nucleotide.
[0196] In certain embodiments, the sequences from (a), (b), or (c) have at least 60, 75, 80, 85, 90, 95, or 99% homology with the corresponding sequences of naturally occurring gRNAs or with the gRNAs described herein.
[0197] In certain embodiments, the adjacent and tail domains together comprise at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides.
[0198] In certain embodiments, the 3' of the second complementary domain contains at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides.
[0199] In certain embodiments, the first complementary domain has at least 16, 19, 21, 26, 31, 32, 36, 41, 46, 50, 51, or 54 nucleotides at 3' relative to the last nucleotide of the second complementary domain which is complementary to its corresponding nucleotide. In certain embodiments, the targeting domain comprises, has, or consists of 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 consecutive nucleotides) which are complementary to the target domain, for example, the targeting domain is 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides long.
[0200] In certain embodiments, the targeting domain comprises, consists of, or substantially comprises 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, or 26 consecutive nucleotides) that are complementary to the target domain or a portion thereof. In some of these embodiments, the targeting domain is complementary to the target domain over its entire length, over its entire length, or both.
[0201] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 16 nucleotides (e.g., 16 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 16 nucleotides long. In certain embodiments of these embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0202] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 17 nucleotides (e.g., 17 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 17 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0203] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 18 nucleotides (e.g., 18 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 18 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0204] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 19 nucleotides (e.g., 19 consecutive nucleotides) complementary to the target domain, for example, the targeting domain is 19 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0205] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 20 nucleotides (e.g., 20 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 20 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0206] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 21 nucleotides (e.g., 21 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 21 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0207] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 22 nucleotides (e.g., 22 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 22 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0208] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 23 nucleotides (e.g., 23 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 23 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain complementary to its corresponding nucleotide of the first complementary domain.
[0209] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 24 nucleotides (e.g., 24 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 24 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0210] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 25 nucleotides (e.g., 25 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 25 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0211] In certain embodiments, the targeting domain contains, consists of, or substantially consists of 26 nucleotides (e.g., 26 consecutive nucleotides) complementary to the targeting domain, for example, the targeting domain is 26 nucleotides long. In certain embodiments, (a) the adjacent and tail domains together contain at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides; (b) there are at least 15, 18, 20, 25, 30, 31, 35, 40, 45, 49, 50, or 53 nucleotides on the 3' side of the last nucleotide of the second complementary domain that is complementary to its corresponding nucleotide of the first complementary domain.
[0212] gRNA delivery In certain embodiments of the methods provided herein, the methods involve the delivery of one or more (e.g., two, three, or four) gRNA molecules as described herein. In some of these embodiments, the gRNA molecules are delivered by intravenous injection, intramuscular injection, subcutaneous injection, or inhalation.
[0213] How to design gRNA Methods for selecting, designing, and validating targeting domains for use with the gRNAs described herein are provided. Exemplary targeting domains for incorporation into gRNAs are also provided herein.
[0214] Methods for target sequence selection and validation, as well as off-target analysis, have already been described (see, e.g., Mali 2013; Hsu 2013; Fu 2014; Heigwer 2014; Bae 2014; Xiao 2014). For example, software tools can be used to optimize the selection of potential targeting domains corresponding to the user's target sequence, minimizing overall off-target activity across the genome. Off-target activity may be other than cleavage. For each possible targeting domain selection using S. pyogenes Cas9, this tool can identify all off-target sequences (preceding either NAG or NGG PAM) across the genome, including up to a certain number (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10) of mismatched base pairs. The cleavage efficiency at each off-target sequence can be predicted, for example, using an experimentally induced weighting scheme. Next, each possible targeting domain is ranked according to its total predicted off-target cleavage; the highest-ranking targeting domain is likely to have the most on-target cleavage and the fewest off-target cleavage. Other functions, such as automated reagent design for CRISPR construction, primer design for on-target surveyor assays, and primer design for high-throughput detection and quantification of off-target cleavage by next-generation sequencing, can also be included in this tool. Candidate targeting domains and gRNAs containing these targeting domains can be functionally evaluated using methods known in the art and / or methods described herein.
[0215] As a non-limiting example, targeting domains for use with gRNAs used with S. pyogenes Cas9 and S. aureus Cas9 were identified using DNA sequence search algorithms. 17-mer and 20-mer targeting domains were designed for S. pyogenes targets, while 18-mer, 19-mer, 20-mer, 21-mer, 22-mer, 23-mer, and 24-mer targeting domains were designed for S. aureus targets. gRNA design was performed using custom gRNA design software based on the publicly available tool cas-offinder (Bae 2014). This software scores the guides after calculating their genome-wide off-target propensity. Typically, matches ranging from perfect matches to seven mismatches are considered for guides of lengths 17–24. Once off-target sites are computationally determined, an aggregate score is calculated for each guide and summarized in a tabular output using a web interface. In addition to identifying possible target sites adjacent to PAM sequences, the software also identifies all PAM-adjacent sequences that differ from the selected target sites by one, two, three, or four or more nucleotides. Genomic DNA sequences of the HBG1 and HBG2 regulatory regions were obtained from the UCSC Genome Browser, and these sequences were screened for repeat elements using the publicly available RepeatMasker program. RepeatMasker searches the input DNA sequence for repeat clements and low-complexity regions. The output is a detailed annotation of repeats present in a given query sequence.
[0216] After identification, target domains were ranked hierarchically based on their distance to their target sites, their orthogonality, and the presence of 5'G (based on the identification of close matches in the human genome, including relevant PAMs, e.g., the NGG PAM of S. pyogenes, or the NNGRRT (SEQ ID NO: 204) or NNGRRV (SEQ ID NO: 205) PAM of S. aureus). Orthogonality refers to the number of sequences in the human genome that contain the minimum number of mismatches to the target sequence. "High level of orthogonality" or "good orthogonality" may refer to a 20-mer targeting domain that, for example, has no identical sequences in the human genome other than the intended target, nor any sequences that contain one or two mismatches in the target sequence. Targeting domains with good orthogonality are selected to minimize off-target DNA cleavage.
[0217] Targeting domains were identified in both single-gRNA nuclease cleavage and double-gRNA pairing "nickase" strategies. The criteria for selecting targeting domains, and the determination of which targeting domains can be used in the double-gRNA pairing "nickase" strategy, are based on two considerations: (1) The targeting domain pair should be oriented on the DNA such that the PAM faces outward and produces a 5' overhang upon cleavage by D10A Cas9 nickase; and (2) The hypothesis is that cleavage by double nickase pairs leads to deletion of the entire intervening sequence at a sufficiently high frequency. However, cleavage using a double nickase pair can also result in indel mutations at only one site of the gRNA. Candidate pair members can be tested for how efficiently they remove an entire sequence, compared to inducing an indel mutation at the target site of a single targeting domain.
[0218] Targeting domain for use in deletion of HBG1 c.-114~-102 In conjunction with the methods disclosed herein, we identified targeting domains for use in gRNA due to deletions of c.-114~-102 in HBG1 and ranked them into four hierarchical levels for S. pyogenes and S. aureus.
[0219] In the case of S. pyogenes, the tier 1 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, specifically within 400 bp from any end of this target site, (2) a high level of orthogonality, and (3) the presence of a 5'G. The tier 2 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, in particular within 400 bp from any end of this target site, and (2) a high level of orthogonality. The tier 3 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, in particular within 400 bp from any end of this target site, and (2) the presence of a 5'G. The targeting domains of tier 4 were selected based on the distance from either end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, particularly within 400 bp of either end of this target site.
[0220] In the case of S. aureus, the tier 1 targeting domains were selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, particularly within 400 bp of any end of this target site, (2) a high level of orthogonality, (3) the presence of 5'G, and (4) the presence of the PAM having the sequence NNGRRT (sequence number: 204). The tier 2 targeting domains were selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, particularly within 400 bp of any end of this target site, (2) a high level of orthogonality, and (3) the presence of the PAM having the sequence NNGRRT (sequence number: 204). The targeting domains of tier 3 were selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, in particular within 400 bp of any end of this target site, and (2) the PAM having the sequence NNGRRT (sequence number 204). The targeting domains of tier 4 were selected based on (1) the distance from any end of the target site (i.e., HBG1 c.-114~-102) to the upstream or downstream, in particular within 400 bp of any end of this target site, and (2) the PAM having the sequence NNGRRV (sequence number 205).
[0221] Please note that the hierarchy is not exhaustive (each targeting domain is listed only once for strategic purposes). In some cases, a targeting domain could not be identified based on the criteria of a particular hierarchy. The identified targeting domains are summarized in Table 6 below.
[0222] [Table 1]
[0223] Targeting domain for use in deletion of HBG2 c.-114~-102 In conjunction with the methods disclosed herein, we identified targeting domains for use in gRNA due to deletions of c.-114~-102 in HBG2 and ranked them into four hierarchical levels for S. pyogenes and S. aureus.
[0224] In the case of S. pyogenes, the tier 1 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, specifically within 400 bp from any end of this target site, (2) a high level of orthogonality, and (3) the presence of a 5'G. The tier 2 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, in particular within 400 bp from any end of this target site, and (2) a high level of orthogonality. The tier 3 targeting domain was selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, in particular within 400 bp from any end of this target site, and (2) the presence of a 5'G. The targeting domains of tier 4 were selected based on the distance from either end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, particularly within 400 bp of either end of this target site.
[0225] In the case of S. aureus, the tier 1 targeting domains were selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, particularly within 400 bp of any end of this target site, (2) a high level of orthogonality, (3) the presence of 5'G, and (4) PAM having the sequence NNGRRT (sequence number: 204). The tier 2 targeting domains were selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, particularly within 400 bp of any end of this target site, (2) a high level of orthogonality, and (3) PAM having the sequence NNGRRT (sequence number: 204). The targeting domains of tier 3 were selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, in particular within 400 bp of any end of this target site, and (2) the PAM having the sequence NNGRRT (sequence number 204). The targeting domains of tier 4 were selected based on (1) the distance from any end of the target site (i.e., HBG2 c.-114~-102) to the upstream or downstream, in particular within 400 bp of any end of this target site, and (2) the PAM having the sequence NNGRRV (sequence number 205).
[0226] Please note that the hierarchy is not exhaustive (each targeting domain is listed only once for strategic purposes). In some cases, a targeting domain could not be identified based on the criteria of a particular hierarchy. The identified targeting domains are summarized in Table 7 below.
[0227] [Table 2]
[0228] In certain embodiments, two or more (e.g., three or four) gRNA molecules are used with one Cas9 molecule. In other embodiments, when two or more (e.g., three or four) gRNA molecules are used with two or more Cas9 molecules, at least one Cas9 molecule is derived from a different species than the other Cas9 molecules. For example, when two gRNA molecules are used with two Cas9 molecules, one Cas9 molecule may be derived from one species, and the other Cas9 molecules may be derived from different species. Both Cas9 species are used to produce single-strand or double-strand breaks as desired.
[0229] Any of the targeting domains in the table described herein can be used with a Cas9 molecule that induces single-strand breaks (i.e., S. pyogenes or S. aureus Cas9 nickase) or a Cas9 molecule that induces double-strand breaks (i.e., S. pyogenes or S. aureus Cas9 nuclease).
[0230] If two gRNAs are designed to be used with two Cas9 molecules, the two Cas9 molecules can be different species. Both Cas9 species can be used to produce single-strand or double-strand breaks as desired.
[0231] It is intended herein that any upstream gRNA described herein can be paired with any downstream gRNA described herein. When an upstream gRNA designed for use with one type of Cas9 is paired with a downstream gRNA designed for use with a different type of Cas9, both types of Cas9 are used to produce single-strand or double-strand breaks as desired.
[0232] RNA-induced nuclease RNA-inducing nucleases as used in this disclosure include, but are not limited to, naturally occurring class 2 CRISPR nucleases, such as Cas9 and Cpf1, and other nucleases derived from or obtained from these. Functionally, RNA-inducing nucleases are defined as nucleases that (a) interact with gRNA (e.g., complex with it); and (b) associate with gRNA, (i) a sequence complementary to the gRNA's targeting domain, and optionally (ii) a target region of DNA containing a PAM, and optionally cleave or modify it. RNA-inducing nucleases can be broadly defined by their PAM specificity and cleavage activity, even if variations exist among individual RNA-inducing nucleases that share the same PAM specificity or cleavage activity. Those skilled in the art will understand that some aspects of this disclosure relate to systems, methods, and compositions that can be carried out using any suitable RNA-inducing nuclease having a specific PAM specificity and / or cleavage activity. Therefore, unless otherwise specified, the term RNA-induced nuclease should be understood as a general term and not limited to specific types of RNA-induced nucleases (e.g., Cas9 vs. Cpfl), species (e.g., S. pyogenes vs. S. aureus), or variations (e.g., full length vs. cleavage or splitting; spontaneous PAM specificity vs. genetically engineered PAM specificity, etc.).
[0233] The name of a PAM sequence derives from its sequence relationship with a “protospacer” sequence (or “spacer”) that is complementary to the gRNA targeting domain. Together with the protospacer sequence, the PAM sequence defines the target region or sequence of a particular RNA-inducible nuclease / gRNA combination.
[0234] Various RNA-induced nucleases may require different sequence relationships between the PAM and the protospacer. Generally, Cas9 recognizes the PAM sequence at 3' of the visualized protospacer compared to the upper or complementary strand: [ka] On the other hand, Cpf1 generally recognizes the PAM sequence, which is the 5' of the protospacer: [ka]
[0235] In addition to recognizing the specific orientation of PAM and protospacer sequences, RNA-induced nucleases can also recognize specific PAM sequences. For example, S. aureus Cas9 recognizes the NNGRRT or NNGRRV PAM sequence located 3' immediately to the region where the N residue is recognized by the gRNA targeting domain. S. pyogenes Cas9 recognizes the NGG PAM sequence, and F. nobicida Cpfl recognizes the TTN PAM sequence. PAM sequences have been identified in various RNA-induced nucleases, and strategies for identifying novel PAM sequences are described in Shmakov 2015. Furthermore, it should be noted that genetically modified RNA-induced nucleases may have PAM specificity different from that of the reference molecule (for example, in the case of genetically modified RNA-induced nucleases, the reference molecule may be a naturally occurring variant from which the RNA-induced nuclease originates, or a naturally occurring variant that has the highest amino acid sequence homology to the genetically modified RNA-induced nuclease).
[0236] RNA-induced nucleases can be characterized by their DNA cleavage activity in addition to their PAM specificity: naturally occurring RNA-induced nucleases typically form double-sided subunits (DSBs) at their target nucleic acids, but genetically engineered mutants have been created that produce only single-sided subunits (SSBs) (as referenced herein by reference, Ran 2013, above) or do not cleave at all.
[0237] Cas9 molecule Cas9 molecules from a variety of biological species may be used in the methods and compositions described herein. While S. pyogenes and S. aureus Cas9 molecules are the subject of much of the disclosure herein, Cas9 molecules from, derived from, or based on the Cas9 proteins of other biological species listed herein may also be used. These include, for example, Acidovorax avenae, Actinobacillus pleuropneumoniae, Actinobacillus succinogenes, Actinobacillus suis, species of the genus Actinomyces, cycliphilus denitrificans, Aminomonas paucivorans, Bacillus cereus, Bacillus smithii, Bacillus thuringiensis, species of the genus Bacteroides, and Blastopirellula marina. marina), Bradyrhizobium species, Brevibacillus laterosporus, Campylobacter coli, Campylobacter jejuni, Campylobacter lari, Candidatus Puniceispirillum, Clostridium cellulolyticum, Clostridium perfringens, Corynebacterium accolens, Corynebacterium diphtheriadiphtheria), Corynebacterium matruchotii, Dinoroseobacter shibae, Eubacterium dolichum, gammaproteobacteria, Gluconacetobacter diazotrophicus, Haemophilus parainfluenzae, Haemophilus sputorum, Helicobacter canadensis, Helicobacter cinaedi, Helicobacter mustelae, Ilyobacter polytropus Polytropus), Kingella kingae, Lactobacillus crispatus, Listeria ivanovii, Listeria monocytogenes, Listeriaceae bacteria, Methylocystis species, Methylosinus trichosporium, Mobiluncus mulieris, Neisseria bacilliformis, Neisseria cinerea, Neisseria flavescens, Neisseria lactamica, Neisseria species, Neisseria wazwashii (Neisseria wadsworthii), Nitrosomonas species, Parvibaculum lavamentivoranslavamentivorans), Pasteurella multocida, Phascolarctobacterium succinatutens, Ralstonia syzygii, Rhodopseudomonas palustris, Rhodovulum species, Simonsiella muelleri, Sphingomonas species, Sporolactobacillus vineae, Staphylococcus rugdunensis Contains Cas9 molecules from species of the genera *Streptococcus lugdunensis*, *Streptococcus*, *Subdoligranulum*, *Tistrella mobilis*, *Treponema*, or *Verminephrobacter eiseniae*.
[0238] Cas9 domain The crystal structures were determined for two different naturally occurring bacterial Cas9 molecules (Jinek 2014) and S. pyogenes Cas9 (Nishimasu 2014; Anders 2014) using guide RNA (e.g., synthetic fusion of crRNA and tracrRNA).
[0239] The naturally occurring Cas9 molecule comprises two lobes: a recognition (REC) lobe and a nuclease (NUC) lobe; these each further contain the domains described herein. Figures 8A and 8B provide schematic diagrams of the mechanism of the key Cas9 domains in the main structure. The domain nomenclature and numbering of amino acid residues contained within each domain used throughout this disclosure are as previously described (Nishimasu 2014). The amino acid residue numbering refers to Cas9 derived from S. pyogenes.
[0240] The REC lobe contains an arginine-rich cross-linking helix (BH), the REC1 domain, and the REC2 domain. The REC lobe does not have structural similarities to other known proteins, suggesting that it is a Cas9-specific functional domain. The BH domain is a long α-helical and arginine-rich region containing amino acids 60-93 of S. pyogenes Cas9 (SEQ ID NO: 2). The REC1 domain is important for the recognition of repeat:anti-repeat double strands, such as gRNA or atracrRNA, and is therefore important for Cas9 activity by recognizing target sequences. The REC1 domain contains two REC1 motifs at amino acids 94-179 and 308-717 of S. pyogenes Cas9 (SEQ ID NO: 2). These two REC1 domains are separated by the REC2 domain in the linear primary structure, but are constructed in the tertiary structure to form the REC1 domain. The REC2 domain or a portion thereof may also play a role in the recognition of repeat:anti-repeat double strands. The REC2 domain contains amino acids 180-307 of the S. pyogenes Cas9 sequence.
[0241] The NUC lobe contains a RuvC domain, an HNH domain, and a PAM interaction (PI) domain. The RuvC domain has structural similarities to components of the retroviral integrase superfamily and cleaves single strands, such as the non-complementary strand of a target nucleic acid molecule. The RuvC domain is constructed from three splintered RuvC motifs located at amino acids 1-59, 718-769, and 909-1098 of S. pyogenes Cas9 (SEQ ID NO: 2), respectively (RuvCI, RuvCII, and RuvCIII, commonly referred to in the art as the RuvCI domain or the N-terminal RuvC domain, RuvCII domain, and RuvCIII domain). Similar to the REC1 domain, the three RuvC motifs are linearly isolated by other domains in the primary structure. However, in the tertiary structure, the three RuvC motifs are constructed to form the RuvC domain. The HNH domain has structural similarities to HNH endonucleases and, for example, cleaves single strands such as the complementary strand of a target nucleic acid molecule. The HNH domain is located between the RuvCII-III motifs and contains amino acids 775-908 of S. pyogenes Cas9 (SEQ ID NO: 2). The PI domain interacts with the PAM of the target nucleic acid molecule and contains amino acids 1099-1368 of S. pyogenes Cas9 (SEQ ID NO: 2).
[0242] RuvC domain and HNH domain In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an HNH-like domain and a RuvC-like domain, and in these particular embodiments, the cleavage activity depends on the RuvC-like domain and the HNH-like domain. The Cas9 molecule or Cas9 polypeptide may contain one or more domains of the RuvC-like domain and the HNH-like domain. In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises, for example, the RuvC-like domain described below, and / or the RuvC-like domain such as the HNH-like domain described below.
[0243] RuvC domain In certain embodiments, the RuvC-like domain cleaves a single strand, such as the non-complementary strand of a target nucleic acid molecule. A Cas9 molecule or Cas9 polypeptide may contain two or more RuvC-like domains (e.g., one, two, or three or more RuvC-like domains). In certain embodiments, the RuvC-like domain is at least 5, 6, 7, or 8 amino acid long, but no more than 20, 19, 18, 17, 16, or 15 amino acid long. In certain embodiments, a Cas9 molecule or Cas9 polypeptide contains an N-terminal RuvC-like domain of about 10 to 20 amino acids, such as about 15 amino acids long.
[0244] N-terminal RuvC-like domain Some natural Cas9 molecules contain two or more RuvC-like domains, and cleavage depends on the N-terminal RuvC-like domain. Therefore, Cas9 molecules or Cas9 polypeptides may contain an N-terminal RuvC-like domain. Representative N-terminal RuvC-like domains are described below.
[0245] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is a compound of formula I, It contains an N-terminal RuvC-like domain, which includes the amino acid sequence D-X1-G-X2-X3-X4-X5-G-X6-X7-X8-X9 (SEQ ID NO: 20), During the ceremony, X1 is selected from I, V, M, L, and T (for example, selected from I, V, and L); X2 is selected from T, I, V, S, N, Y, E, and L (for example, selected from T, V, and I); X3 is selected from N, S, G, A, D, T, R, M, and F (for example, A or N); X4 is selected from S, Y, N, and F (for example, S); X5 is selected from V, I, L, C, T, and F (for example, selected from V, I, and L); X6 is selected from W, F, V, Y, S, and L (for example, W); X7 is selected from A, S, C, V, and G (for example, selected from A and S); X8 is selected from V, I, L, A, M, and H (for example, selected from V, I, M, and L); X9 is selected from any amino acid or is absent (for example, selected from T, V, I, L, Δ, F, S, A, Y, M, and R, or for example, selected from T, V, I, L, and Δ).
[0246] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of Sequence ID No. 20 by approximately one, but two, three, four, or five or fewer residues.
[0247] In certain embodiments, the N-terminal RuvC-like domain is capable of cleaving. In other embodiments, the N-terminal RuvC-like domain is not capable of cleaving.
[0248] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is derived from formula II: It contains an N-terminal RuvC-like domain, which includes the amino acid sequence D-X1-G-X2-X3-S-X5-G-X6-X7-X8-X9 (SEQ ID NO: 21), During the ceremony, X1 is selected from I, V, M, L, and T (for example, selected from I, V, and L); X2 is selected from T, I, V, S, N, Y, E, and L (for example, selected from T, V, and I); X3 is selected from N, S, G, A, D, T, R, M, and F (for example, A or N); X5 is selected from V, I, L, C, T, and F (for example, selected from V, I, and L); X6 is selected from W, F, V, Y, S, and L (for example, W); X7 is selected from A, S, C, V, and G (for example, selected from A and S); X8 is selected from V, I, L, A, M, and H (for example, selected from V, I, M, and L); X9 is selected from any amino acid or is absent (for example, selected from T, V, I, L, Δ, F, S, A, Y, M, and R, or for example, selected from T, V, I, L, and Δ).
[0249] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of Sequence ID No. 21 by approximately one, but two, three, four, or five or fewer residues.
[0250] In certain embodiments, the N-terminal RuvC-like domain is represented by formula III: DIG-X2-X3-SVGWA-X8-X9 (Sequence No. 22) It contains the amino acid sequence, During the ceremony, X2 is selected from T, I, V, S, N, Y, E, and L (for example, selected from T, V, and I); X3 is selected from N, S, G, A, D, T, R, M, and F (for example, A or N); X8 is selected from V, I, L, A, M, and H (for example, selected from V, I, M, and L); X9 is selected from any amino acid or is absent (for example, selected from T, V, I, L, Δ, F, S, A, Y, M, and R, or for example, selected from T, V, I, L, and Δ).
[0251] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of Sequence ID No. 22 by approximately one, but two, three, four, or five or fewer residues.
[0252] In certain embodiments, the N-terminal RuvC-like domain is represented by formula IV: DIGTNSVGWAVX (Sequence ID 23) It comprises the amino acid sequence, During the ceremony, X is a nonpolar alkyl amino acid or hydroxyl amino acid, for example, X is selected from V, I, L, and T (for example, the Cas9 molecule may contain an N-terminal RuvC-like domain as shown in Figures 2A-2G (indicated as Y)).
[0253] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of Sequence ID No. 23 by approximately one, but two, three, four, or five or fewer residues.
[0254] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of the N-terminal RuvC-like domain disclosed herein, for example, in Figures 3A-3B, by about one but two, three, four, or five or fewer residues. In one embodiment, one, two, three, or all of the highly conserved residues identified in Figures 3A-3B are present.
[0255] In certain embodiments, the N-terminal RuvC-like domain differs from the sequence of the N-terminal RuvC-like domain disclosed herein by approximately one, but by two, three, four, or five or fewer residues, as shown in Figures 4A-4B. In one embodiment, one, two, three, or all of the highly conserved residues identified in Figures 4A-4B are present.
[0256] Additional RuvC domain In addition to the N-terminal RuvC-like domain, a Cas9 molecule or Cas9 polypeptide may contain one or more additional RuvC-like domains. In certain embodiments, a Cas9 molecule or Cas9 polypeptide may contain two additional RuvC-like domains. Preferably, the additional RuvC-like domains have an amino acid length of at least 5 amino acids, for example, less than 15 amino acids, for example, 5 to 10 amino acids, for example, 8 amino acids.
[0257] The additional RuvC-like domain is expression V: I-X1-X2-E-X3-ARE (Sequence ID 15) It may contain the following amino acid sequence: During the ceremony, X1 is either V or H; X2 is I, L, or V (for example, I or V); X3 is either M or T.
[0258] In certain embodiments, the additional RuvC-like domain is given by formula VI: IV-X2-EMARE (Sequence ID 16) It comprises the amino acid sequence, During the ceremony, X2 is I, L, or V (for example, I or V) (for example, a Cas9 molecule or Cas9 polypeptide may contain an additional RuvC-like domain, as shown in Figures 2A-2G (shown as B)).
[0259] The additional RuvC-like domain is given by formula VII: HHA-X1-DA-X2-X3 (Sequence ID 17) It may also contain the following amino acid sequence: During the ceremony, X1 is either H or L; X2 is either R or V; X3 is either E or V.
[0260] In certain embodiments, additional RuvC-like domains are HHAHDAYL (Sequence ID 18) It contains the amino acid sequence.
[0261] In certain embodiments, the additional RuvC-like domain differs from the sequences of SEQ ID NOs. 15-18 by approximately one, but two, three, four, or five or fewer residues.
[0262] In some embodiments, the sequence located on the side of the N-terminal RuvC-like domain has the amino acid sequence of formula VIII:K-X1'-Y-X2'-X3'-X4'-ZTD-X9'-Y (SEQ ID NO: 19), During the ceremony, X1' is selected from K and P. X2' is selected from V, L, I, and F (e.g., V, I, and L); X3' is selected from G, A, and S (for example, G); X4' is selected from L, I, V, and F (for example, L); X9' is selected from D, E, N, and Q; Z is, for example, an N-terminal RuvC-like domain having 5 to 12 amino acids, as described above.
[0263] HNH Domain In certain embodiments, the HNH-like domain cleaves a single-stranded complementary domain, such as the complementary strand of a double-stranded nucleic acid molecule. In certain embodiments, the HNH-like domain is at least 15, 20, or 25 amino acids long, but 40, 35, or 30 or less in length, for example, 20-35 amino acids long, or for example, 25-30 amino acids long. Typical HNH-like domains are described below.
[0264] In one embodiment, the Cas9 molecule or Cas9 polypeptide is represented by the formula IX:X1-X2-X3-H-X4-X5-P-X6-X7-X8-X9-X 10 -X 11 -X 12 -X 13 -X 14 -X 15 -NX 16 -X 17 -X 18 -X 19 -X 20 -X 21 -X 22 -X 23 It contains an HNH-like domain having the amino acid sequence -N (SEQ ID NO: 25), During the ceremony, X1 is selected from D, E, Q, and N (for example, D and E); X2 is selected from L, I, R, Q, V, M, and K; X3 is selected from D and E; X4 is selected from I, V, T, A, and L (for example, A, I, and V); X5 is selected from V, Y, I, L, F, and W (for example, V, I, and L); X6 is selected from Q, H, R, K, Y, I, L, F, and W; X7 is selected from S, A, D, T, and K (for example, S and A); X8 is selected from F, L, V, K, Y, M, I, R, A, E, D, and Q (for example, F); X9 is selected from L, R, T, I, V, S, C, Y, K, F, and G; X 10 is selected from K, Q, Y, T, F, L, W, M, A, E, G, and S; X 11 is selected from D, S, N, R, L, and T (for example, D); X 12 It is selected from D, N, and S; X 13 is selected from S, A, T, G, and R (for example, S); X 14 is selected from I, L, F, S, R, Y, Q, W, D, K, and H (for example, I, L, and F); X 15 is selected from D, S, I, N, E, A, H, F, L, Q, M, G, Y, and V; X 16 is selected from K, L, R, M, T, and F (for example, L, R, and K); X 17 is selected from V, L, I, A, and T; X 18 is selected from L, I, V, and A (for example, L and I); X 19 is selected from T, V, C, E, S, and A (for example, T and V); X 20 is selected from R, F, T, W, E, L, N, C, K, V, S, Q, I, Y, H, and A; X 21 The letter is selected from S, P, R, K, N, A, H, Q, G, and L; X 22 is selected from D, G, T, N, S, K, A, I, E, L, Q, R, and Y; X 23The following are selected from K, V, A, E, Y, I, C, L, S, T, G, K, M, D, and F.
[0265] In certain embodiments, the HNH-like domain differs from the sequence of Sequence ID No. 25 by at least one but two, three, four, or five or fewer residues.
[0266] In certain embodiments, the HNH-like domain has cleavage capability. In certain embodiments, the HNH-like domain does not have cleavage capability.
[0267] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is a compound of formula X, X1-X2-X3-H-X4-X5-P-X6-S-X8-X9-X 10 -DDSX 14 -X 15 -NKVLX 19 -X 20 -X 21 -X 22 -X 23 -N(Sequence ID 26) It contains an HNH-like domain that includes the amino acid sequence, During the ceremony, X1 is selected from D and E; X2 is selected from L, I, R, Q, V, M, and K; X3 is selected from D and E; X4 is selected from I, V, T, A, and L (for example, A, I, and V); X5 is selected from V, Y, I, L, F, and W (for example, V, I, and L); X6 is selected from Q, H, R, K, Y, I, L, F, and W; X8 is selected from F, L, V, K, Y, M, I, R, A, E, D, and Q (for example, F); X9 is selected from L, R, T, I, V, S, C, Y, K, F, and G; X 10 is selected from K, Q, Y, T, F, L, W, M, A, E, G, and S; X 14is selected from I, L, F, S, R, Y, Q, W, D, K, and H (for example, I, L, and F); X 15 is selected from D, S, I, N, E, A, H, F, L, Q, M, G, Y, and V; X 19 is selected from T, V, C, E, S, and A (for example, T and V); X 20 is selected from R, F, T, W, E, L, N, C, K, V, S, Q, I, Y, H, and A; X 21 The letter is selected from S, P, R, K, N, A, H, Q, G, and L; X 22 is selected from D, G, T, N, S, K, A, I, E, L, Q, R, and Y; X 23 The following are selected from K, V, A, E, Y, I, C, L, S, T, G, K, M, D, and F.
[0268] In certain embodiments, the HNH-like domain differs from the sequence of Sequence ID No. 26 by one, two, three, four, or five residues.
[0269] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is a compound of formula XI, X1-V-X3-HIVP-X6-S-X8-X9-X 10 -DDSX 14 -X 15 -NKVLTX 20 -X 21 -X 22 -X 23 Contains an HNH-like domain with the amino acid sequence -N (sequence number 27). During the ceremony, X1 is selected from D and E; X3 is selected from D and E; X6 is selected from Q, H, R, K, Y, I, L, and W; X8 is selected from F, L, V, K, Y, M, I, R, A, E, D, and Q (for example, F); X9 is selected from L, R, T, I, V, S, C, Y, K, F, and G; X 10 is selected from K, Q, Y, T, F, L, W, M, A, E, G, and S; X 14 is selected from I, L, F, S, R, Y, Q, W, D, K, and H (e.g., I, L, and F); X 15 is selected from D, S, I, N, E, A, H, F, L, Q, M, G, Y, and V; X 20 is selected from R, F, T, W, E, L, N, C, K, V, S, Q, I, Y, H, and A; X 21 is selected from S, P, R, K, N, A, H, Q, G, and L; X 22 is selected from D, G, T, N, S, K, A, I, E, L, Q, R, and Y; X 23 is selected from K, V, A, E, Y, I, C, L, S, T, G, K, M, D, and F.
[0270] In certain embodiments, the HNH-like domain differs from the sequence of SEQ ID NO: 27 by 1, 2, 3, 4, or 5 residues.
[0271] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an HNH-like domain having the amino acid sequence of formula XII, D-X2-D-H-I-X5-P-Q-X7-F-X9-X 10 -D-X 12 -S-I-D-N-X 16 -V-L-X 19 -X 20 -S-X 22 -X 23 -N (SEQ ID NO: 28) wherein, wherein, X2 is selected from I and V; X5 is selected from I and V; X7 is selected from A and S; X9 is selected from I and L; X 10 is selected from K and T; X12 This is selected from D and N; X 16 is selected from R, K, and L; X 19 is selected from T and V; X 20 It is selected from S and R; X 22 It is selected from K, D, and A; X 23 This is selected from E, K, G, and N (for example, the eaCas9 molecule or eaCas9 polypeptide may contain an HNH-like domain as described herein).
[0272] In one embodiment, the HNH-like domain differs from the sequence of Sequence ID No. 28 by approximately one, but two, three, four, or five or fewer residues.
[0273] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is a compound of formula XIII, LYYLQNG-X1'-DMY-X2'-X3'-X4'-X5'-LDI-X6'-X7'-LS-X8'-YZNR-X9'-KX 10 '-DX 11 -VP(sequence number 24) It contains the amino acid sequence. During the ceremony, X1' is selected from K and R; X2' is selected from V and T; X3' is selected from G and D; X4' is selected from E, Q, and D; X5' is selected from E and D; X6' is selected from D, N, and H; X7' is selected from Y, R, and N; X8' is selected from Q, D, and N; X9' is selected from G and E; X 10 ' is selected from S and G; X 11 ' is selected from D and N; and Z is, for example, an HNH-like domain as described above.
[0274] In certain embodiments, the Cas9 molecule or Cas9 polypeptide includes an amino acid sequence that differs from the sequence of SEQ ID NO: 24 by approximately one but two, three, four, or five or fewer residues.
[0275] In certain embodiments, the HNH-like domain differs from the sequences of HNH-like domains disclosed herein, such as in Figures 5A-5C, by approximately one but no more than two, three, four, or five residues. In certain embodiments, one or both of the highly conserved residues identified in Figures 5A-5C are present.
[0276] In certain embodiments, the HNH-like domain differs from the sequences of HNH-like domains disclosed herein, such as in Figures 6A-6B, by approximately one, but two, three, four, or five or fewer residues. In one embodiment, one, two, or all three of the highly conserved residue sacs identified in Figures 6A-6B are present.
[0277] Cas9 function In certain embodiments, a Cas9 molecule or Cas9 polypeptide has the ability to cleave a target nucleic acid molecule. Typically, a wild-type Cas9 molecule cleaves both strands of a target nucleic acid molecule. Cas9 molecules and Cas9 polypeptides can be genetically engineered to alter their nuclease cleavage (or other properties) to provide a Cas9 molecule or Cas9 polypeptide that is, for example, a nickase or lacks the ability to cleave a target nucleic acid. A Cas9 molecule or Cas9 polypeptide that has the ability to cleave a target nucleic acid molecule is referred to herein as an eaCas9 molecule (enzymatically active Cas9) or an eaCas9 polypeptide.
[0278] In certain embodiments, the eaCas9 molecule or eaCas9 polypeptide includes one or more of the following functions: (1) Nickase activity, that is, the ability to cleave single strands, such as the non-complementary or complementary strand of a nucleic acid molecule; (2) In certain embodiments, the presence of two nickase activities, i.e., double-strand nuclease activity, i.e., the ability to cleave both strands of a double-stranded nucleic acid to produce a double-strand break; (3) Endonuclease activity; (4) Exonuclease activity; and (5) Helicase activity, that is, the ability to unwind the helical structure of double-stranded nucleic acids.
[0279] In certain embodiments, the eaCas9 molecule or eaCas9 polypeptide cleaves both DNA strands, resulting in a double-strand break. In certain embodiments, the eaCas9 molecule or eaCas9 polypeptide cleaves only one strand, such as the strand to which the gRNA hybridizes, or a strand complementary to the strand to which the gRNA hybridizes. In one embodiment, the eaCas9 molecule or eaCas9 polypeptide includes cleavage activity associated with the HNH domain. In one embodiment, the eaCas9 molecule or eaCas9 polypeptide includes cleavage activity associated with the RuvC domain. In one embodiment, the eaCas9 molecule or eaCas9 polypeptide includes cleavage activity associated with both the HNH domain and the RuvC domain. In one embodiment, the eaCas9 molecule or eaCas9 polypeptide comprises an active or cleavable HNH domain and an inactive or incapable RuvC domain.
[0280] Targeting and PAM The Cas9 molecule or Cas9 polypeptide can interact with gRNA molecules and, in cooperation with the gRNA molecule, localize to a target domain and, in certain embodiments, to a site containing the aPAM sequence.
[0281] In certain embodiments, the ability of an eaCas9 molecule or polypeptide to interact with and cleave a target nucleic acid is PAM sequence-dependent. The PAM sequence is a sequence in the target nucleic acid. In one embodiment, cleavage of the target nucleic acid occurs upstream of the PAM sequence. eaCas9 molecules from different bacterial species may recognize different sequence motifs (e.g., PAM sequences). In one embodiment, the eaCas9 molecule from S. pyogenes recognizes the sequence motifs NGG, NAG, NGA and induces cleavage of the target nucleic acid sequence 1-10 bp upstream of that sequence, for example, 3-5 bp (see, e.g., Mali 2013). In one embodiment, the eaCas9 molecule of S. thermophilus recognizes the sequence motif NGGN (SEQ ID NO: 199) and / or NNAGAAW (W=A or T) (SEQ ID NO: 200) and induces cleavage of the target nucleic acid sequence 1 to 10 bp upstream of these sequences, for example, 3 to 5 bp (see, e.g., Horvath 2010; Deveau 2008). In another embodiment, the eaCas9 molecule of S. mutans recognizes the sequence motif NGG and / or NAAR (R=A or G) (SEQ ID NO: 201) and induces cleavage of the target nucleic acid sequence 1 to 10 bp upstream of this sequence, for example, 3 to 5 bp (see, e.g., Deveau 2008). In one embodiment, the Cas9 molecule of S. aureus recognizes the sequence motif NNGRR (R=A or G) (SEQ ID NO: 202) and induces the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp upstream of that sequence. In another embodiment, the Cas9 molecule of S. aureus recognizes the sequence motif NNGRRN (R=A or G) (SEQ ID NO: 203) and induces the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp upstream of that sequence. In yet another embodiment, the Cas9 molecule of S. aureus recognizes the sequence motif NNGRRT (R=A or G) (SEQ ID NO: 204) and induces the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp upstream of that sequence.In one embodiment, the eaCas9 molecule of S. aureus recognizes the sequence motif NNGRRV (R=A or G, V=A, G or C) (SEQ ID NO: 205) and induces cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp upstream from that sequence. The ability of the Cas9 molecule to recognize PAM sequences can be determined, for example, using the transformation assay described above (Jinek 2012). In each of the aforementioned embodiments (i.e., SEQ ID NOs: 199 to 205), N may be any nucleotide residue, for example, A, G, C, or T.
[0282] As discussed herein, the Cas9 molecule can be genetically engineered to alter its PAM specificity.
[0283] Representative naturally occurring Cas9 molecules are described above (see, for example, Chylinski 2013). Such Cas9 molecules include cluster 1 bacteriaceae, cluster 2 bacteriaceae, cluster 3 bacteriaceae, cluster 4 bacteriaceae, cluster 5 bacteriaceae, cluster 6 bacteriaceae, cluster 7 bacteriaceae, cluster 8 bacteriaceae, cluster 9 bacteriaceae, cluster 10 bacteriaceae, cluster 11 bacteriaceae, cluster 12 bacteriaceae, cluster 13 bacteriaceae, cluster 14 bacteriaceae, cluster 15 bacteriaceae, cluster 16 bacteriaceae, cluster 17 bacteriaceae, cluster 18 bacteriaceae, cluster 19 bacteriaceae, cluster 20 bacteriaceae, cluster 21 bacteriaceae, cluster 22 bacteriaceae, cluster 23 bacteriaceae, cluster 24 bacteriaceae, cluster 25 bacteriaceae, cluster 26 bacteriaceae, cluster 27 bacteriaceae, cluster 28 bacteriaceae, cluster 29 bacteriaceae, cluster 30 bacteriaceae, cluster 31 bacteriaceae, cluster 32 bacteriaceae, cluster 33 bacteriaceae, cluster 34 bacteriaceae, cluster 35 bacteriaceae, cluster 36 bacteriaceae, cluster 37 bacteriaceae, cluster 38 bacteriaceae, cluster 39 bacteriaceae, cluster The following are examples of Cas9 molecules from the Bacteriaceae families 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, or 78.
[0284] Representative naturally occurring Cas9 molecules include those from the Cluster 1 Bacteriaceae. Examples include S. aureus, S. pyogenes (e.g., SF370, MGAS10270, MGAS10750, MGAS2096, MGAS315, MGAS5005, MGAS6180, MGAS9429, NZ131, SSI-1 strain), S. thermophilus (e.g., LMD-9 strain), and S. pseudoporcinus (e.g., SPIN strain). S. mutans (e.g., UA159, NN2025), S. macacae (e.g., NCTC11558), S. gallolyticus (e.g., UCN34, ATCCBAA-2069), S. equines (e.g., ATCC9812, MGCS124), S. dysdalactiae (e.g., GGS124), S. bovis (e.g., ATCC700338), S. anginosus (e.g., F0211), S. agalactiae (e.g., NEM316, A909), Listeria monocytogenes Examples include Cas9 molecules from monocytogenes (e.g., strain F6854), Listeria innocua (L. innocua), e.g., strain Clip11262), Enterococcus italicus (e.g., strain DSM 15952), or Enterococcus faecium (e.g., strain 1,231,408).
[0285] In certain embodiments, the Cas9 molecule or Cas9 polypeptide is Having 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology with sequence numbers 1, 2, 4-6, or 12; When compared, they differ by 2, 5, 10, 15, 20, 30, or 40% or less of amino acid residues; At least 1, 2, 5, 10 or 20 amino acids, but 100, 80, 70, 60, 50, 40 or 30 or fewer different amino acids; or This includes any Cas9 molecular sequence described herein, or an amino acid sequence identical to a natural Cas9 molecular sequence, such as those enumerated herein or derived from a species (e.g., described in SEQ ID NOs: 1-4 or Chylinski 2013). In one embodiment, the Cas9 molecule or Cas9 polypeptide includes one or more of the following functions: nickase activity; double-strand cleavage activity (e.g., endonuclease and / or exonuclease activity); helicase activity; or the ability to localize to a target nucleic acid together with a gRNA molecule.
[0286] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises one of the amino acid sequences of the consensus sequences shown in Figures 2A-2G, where "*" indicates any amino acid present at the corresponding position in the amino acid sequence of the Cas9 molecule of S. pyogenes, S. thermophilus, S. mutans, or L. innocua, and "-" indicates absence. In one embodiment, the Cas9 molecule or Cas9 polypeptide differs from the consensus sequences disclosed in Figures 2A-2G by at least one, but no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residues. In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises the amino acid sequence of Sequence ID No. 2. In other embodiments, the Cas9 molecule or Cas9 polypeptide differs from the sequence of SEQ ID NO: 2 by at least one, but no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residues.
[0287] A comparison of the sequences of several Cas9 molecules suggests that certain regions are conserved. These are identified as follows: Region 1 (residues 1-180, or residues 120-180 if it is Region 1) Region 2 (residues 360-480); Region 3 (residues 660-720); Region 4 (residues 817-900); and Region 5 (residues 900-960).
[0288] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises regions 1 to 5 and, together with sufficient additional Cas9 molecular sequences, provides a bioactive molecule such as, for example, a Cas9 molecule having at least one activity as described herein. In certain embodiments, each of regions 1 to 5 independently has 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology with the corresponding residues of the Cas9 molecule or Cas9 polypeptide described herein, such as the sequences from Figures 2A to 2G (SEQ ID NOs: 1, 2, 4, 5, 14).
[0289] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 1: The amino acid sequence of Cas9 in S. pyogenes (SEQ ID NO: 2), amino acids 1-180, exhibits 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology (numbering is based on the motif sequence in Figure 2; 52% of the residues in the four Cas9 sequences in Figures 2A-2G are conserved); The amino acids 1-180 of the Cas9 amino acid sequence of S. pyogenes, S. thermophilus, S. mutans, or Listeria innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differ from at least 1, 2, 5, 10, or 20 amino acids, but no more than 90, 80, 70, 60, 50, 40, or 30 amino acids; or It is identical to amino acids 1-180 of the Cas9 amino acid sequence of S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0290] In one embodiment, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 1', and this region is The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) exhibit 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology between amino acids 120-180 and 50% of the Cas9 sequences (55% of the residues in the four Cas9 sequences in Figure 2 are conserved); The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differ from amino acids 120-180 by at least 1, 2, or 5 amino acids, but with no more than 35, 30, 25, 20, or 10 amino acids; or, It is identical to amino acid sequences 120-180 of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0291] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 2, which is: The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) exhibit 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology between amino acids 360–480 and 50%, respectively (52% of the residues in the four Cas9 sequences in Figure 2 are conserved); The amino acid sequence of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differs from amino acids 360-480 by at least 1, 2, or 5 amino acids, but not more than 35, 30, 25, 20, or 10 amino acids; or, It is identical to amino acid sequences 360-480 of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0292] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 3, which is: The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) exhibit 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology between amino acids 660-720 and those of the Cas9 sequence (56% of the residues in the four Cas9 sequences in Figure 2 are conserved); The amino acid sequence of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differs from amino acids 660-720 by at least 1, 2, or 5 amino acids, but not more than 35, 30, 25, 20, or 10 amino acids; or, It is identical to amino acids 660-720 of the Cas9 amino acid sequence in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0293] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 4, which is: The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) exhibit 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology between amino acids 817-900 and those of the four Cas9 sequences (55% of the residues in the four Cas9 sequences in Figures 2A-2G are conserved); The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differ from amino acids 817-900 by at least 1, 2, or 5 amino acids, but with no more than 35, 30, 25, 20, or 10 amino acids; or, It is identical to amino acids 817-900 of the Cas9 amino acid sequence in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0294] In certain embodiments, the Cas9 molecule or Cas9 polypeptide comprises an amino acid sequence referred to as region 5, which is: The amino acid sequences of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) exhibit 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% homology between amino acids 900-960 and 96%, respectively (60% of the residues in the four Cas9 sequences in Figures 2A-2G are conserved); The amino acid sequence of Cas9 in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively) differs from amino acids 900-960 by at least 1, 2, or 5 amino acids, but no more than 35, 30, 25, 20, or 10 amino acids; or, It is identical to amino acids 900-960 of the Cas9 amino acid sequence in S. pyogenes, S. thermophilus, S. mutans, or L. innocua (SEQ ID NOs. 2, 4, 1, and 5, respectively).
[0295] Genetically modified or altered Cas9 molecules and Cas9 polypeptides The Cas9 molecules and Cas9 polypeptides described herein may possess any of several properties, including nuclease activity (e.g., endonuclease and / or exonuclease activity); helicase activity; the ability to functionally bind to gRNA molecules; and the ability to target (or localize to) sites on nucleic acids (e.g., PAM recognition and specificity). In certain embodiments, the Cas9 molecule or Cas9 polypeptide may include all or some of these properties. In typical embodiments, the Cas9 molecule or Cas9 polypeptide has the ability to interact with gRNA molecules and localize to sites on nucleic acids in cooperation with gRNA molecules. Other functions, such as PAM specificity, cleavage activity, or helicase activity, may vary more broadly in the Cas9 molecule and Cas9 polypeptide.
[0296] The Cas9 molecule includes genetically engineered Cas9 molecules and genetically engineered Cas9 polypeptides (in this context, genetic engineering simply means that the Cas9 molecule or Cas9 polypeptide differs from the reference sequence, and does not imply any process or origin limitations). Genetically engineered Cas9 molecules or Cas9 polypeptides may include, for example, modified nuclease activity (compared to natural or other standard Cas9 molecules) or modified enzymatic properties such as modified helicase activity. As discussed herein, genetically engineered Cas9 molecules or Cas9 polypeptides may have nickase activity (in contrast to double-stranded nuclease activity). In certain embodiments, genetically engineered Cas9 molecules or Cas9 polypeptides may have size-altering modifications, such as deletions of amino acid sequences that reduce their size, without significant effect on one or more Cas9 activities. In certain embodiments, genetically engineered Cas9 molecules or Cas9 polypeptides may include modifications that affect PAM recognition. For example, a genetically modified Cas9 molecule may be modified to recognize PAM sequences other than those recognized by the endogenous wild-type PI domain. In certain embodiments, the Cas9 molecule or Cas9 polypeptide may have a different sequence from the natural Cas9 molecule, but have no significant modifications to one or more Cas9 functions.
[0297] Cas9 molecules or Cas9 polypeptides with desired properties can be generated in several ways, such as by modifying a parent polypeptide, such as a natural Cas9 molecule or Cas9 polypeptide, to provide a modified Cas9 molecule or Cas9 polypeptide with the desired properties. For example, one or more mutations or differences can be introduced from a parent Cas9 molecule, such as a natural or genetically engineered Cas9 molecule. Such mutations and differences include substitutions (e.g., conservative substitutions or substitutions of non-essential amino acids); insertions; or deletions. In one embodiment, the Cas9 molecule or Cas9 polypeptide may contain one or more mutations or differences, such as at least 1, 2, 3, 4, 5, 10, 15, 20, 30, 40, or 50 mutations compared to a standard, such as a parent Cas9 molecule, but fewer than 200, 100, or 80 mutations.
[0298] In certain embodiments, one or more mutations have no substantial effect on Cas9 activity, such as the Cas9 activity described herein. In other embodiments, one or more mutations have a substantial effect on Cas9 activity, such as the Cas9 activity described herein.
[0299] Non-cutting and modified cutting Cas9 In one embodiment, a Cas9 molecule or Cas9 polypeptide may have different cleavage properties from a natural Cas9 molecule, for example, it may differ from a natural Cas9 molecule with the closest homology. For example, a Cas9 molecule or Cas9 polypeptide may differ from a natural Cas9 molecule, such as the S. pyogenes Cas9 molecule, for example, in the following ways: for example, its ability to modulate a decrease or increase in double-strand nucleic acid cleavage (endonuclease and / or exonuclease activity) compared to a natural Cas9 molecule (e.g., the S. pyogenes Cas9 molecule); for example, its ability to modulate a decrease or increase in single-strand nucleic acid cleavage (nickase activity), such as the non-complementary or complementary strand of a nucleic acid molecule, compared to a natural Cas9 molecule (e.g., the S. pyogenes Cas9 molecule); or, for example, its ability to cleave nucleic acid molecules, such as double-stranded or single-stranded nucleic acid molecules, may be eliminated.
[0300] In certain embodiments, the eaCas9 molecule or eaCas9 polypeptide includes one or more of the following functions: cleavage activity related to the N-terminal RuvC-like domain; cleavage activity related to the HNH-like domain; and cleavage activity related to the HNH domain and cleavage activity related to the N-terminal RuvC-like domain.
[0301] In certain embodiments, the eaCas9 molecule or eaCas9 polypeptide comprises an active or cleaving HNH-like domain (e.g., the HNH-like domains described herein, e.g., SEQ ID NOs. 24-28) and an inactive or inactive N-terminal RuvC-like domain. A typical inactive or inactive N-terminal RuvC-like domain may have an aspartic acid mutation in the N-terminal RuvC-like domain, for example, the aspartic acid at position 9 of the consensus sequence disclosed in Figures 2A-2G, or the aspartic acid at position 10 of SEQ ID NO: 2, may be substituted with alanine, for example. In one embodiment, the eaCas9 molecule or eaCas9 polypeptide, unlike the wild type, either does not cleave the target nucleic acid in the N-terminal RuvC-like domain, or cleaves it with significantly lower efficiency, such as less than 20%, 10%, 5%, 1%, or 0.1% of the cleavage activity of the reference Cas9 molecule, as measured by assays described herein. The reference Cas9 molecule may be a naturally occurring, unmodified Cas9 molecule, such as a Cas9 molecule from S. pyogenes, S. aurelius, or S. thermophilus. In one embodiment, the reference Cas9 molecule is a naturally occurring Cas9 molecule with the closest sequence identity or homology.
[0302] In one embodiment, the eaCas9 molecule or eaCas9 polypeptide comprises an inactive or non-cleaving HNH domain and an active or cleaving N-terminal RuvC-like domain (e.g., RuvC-like domains described herein, such as SEQ ID NOs. 15-23). A typical inactive or non-cleaving HNH-like domain may have one or more mutations: for example, histidine in an HNH-like domain such as histidine shown at position 856 of the consensus sequence disclosed in Figures 2A-2G may be substituted with alanine; for example, one or more asparagines in an HNH-like domain such as asparagine shown at position 870 and / or position 879 of the consensus sequence disclosed in Figures 2A-2G may be substituted with alanine. In one embodiment, eaCas9, unlike the wild type, either does not cleave the target nucleic acid in its HNH-like domain, or cleaves it with significantly lower efficiency, such as less than 20, 10, 5, 1, or 0.1% of the cleavage activity of a reference Cas9 molecule, as measured by the assays described herein. The reference Cas9 molecule may be a naturally occurring, unmodified Cas9 molecule, such as a Cas9 molecule from S. pyogenes, S. aurelius, or S. thermophilus. In one embodiment, the reference Cas9 molecule is a naturally occurring Cas9 molecule with the closest sequence identity or homology.
[0303] In certain embodiments, typical Cas9 functions include nine or more PAM specificity, cleavage activity, and helicase activity. Mutations may be located in one or more RuvC domains, such as the N-terminal RuvC domain; an HNH domain; or in regions outside the RuvC and HNH domains. In one embodiment, the mutation is located in the RuvC domain. In one embodiment, the mutation(s) is located in the HNH domain. In one embodiment, the mutation is located in both the RuvC and HNH domains.
[0304] Representative mutations that may occur in the RuvC domain or HNH domain of the S. pyogenes Cas9 sequence include D10A, E762A, H840A, N854A, N863A, and / or D986A. An exemplary mutation that may occur in the RuvC domain of the S. aureus Cas9 sequence is N580A (see, for example, SEQ ID NO: 11).
[0305] For example, whether a particular sequence, such as a substitution, may affect one or more activities, such as targeting activity or cleavage activity, can be evaluated or predicted, for example, by assessing whether the mutation is conserved. In one embodiment, “non-essential” amino acid residues, as used in the context of a Cas9 molecule, are residues that can be modified from the wild-type sequence of a Cas9 molecule, such as a naturally occurring Cas9 molecule like the eaCas9 molecule, without inactivating or, more preferably, substantially altering Cas9 activity (e.g., cleavage activity), while changes to “essential” amino acid residues result in a substantial loss of activity (e.g., cleavage activity).
[0306] In one embodiment, the Cas9 molecule has different cleavage properties from the natural Cas9 molecule, for example, it differs from the natural Cas9 molecule with the closest homology. For example, a Cas9 molecule may differ from a natural Cas9 molecule, such as the Cas9 molecule from S. aureus or S. pyogenes, in the following ways: for example, its ability to regulate the decrease or increase of cleavage of a double-stranded break (endonuclease and / or exonuclease activity) compared to a natural Cas9 molecule (e.g., the Cas9 molecule from S. aureus or S. pyogenes); for example, its ability to regulate the decrease or increase of single-strand breaks of nucleic acids, such as the non-complementary or complementary strand of a nucleic acid molecule (nickase activity), compared to a natural Cas9 molecule (e.g., the Cas9 molecule from S. aureus or S. pyogenes); or, for example, its ability to cleave nucleic acid molecules, such as double-stranded or single-stranded nucleic acid molecules, may be removed. In certain embodiments, the nickase is a S. aureus Cas9-derived nickase containing the sequence of sequence number 10 (D10A) or sequence number 11 (N580A) (Friedland 2015).
[0307] In one embodiment, the modified Cas9 molecule is an eaCas9 molecule that includes one or more of the following functions: cleavage activity related to the RuvC domain; cleavage activity related to the HNH domain; and cleavage activity related to both the HNH domain and the RuvC domain.
[0308] In certain embodiments, the modified Cas9 molecule or Cas9 polypeptide is as follows: The sequences corresponding to the fixed sequences of the consensus sequences disclosed in Figures 2A-2G differ from the fixed residues in the consensus sequences disclosed in Figures 2A-2G by 1, 2, 3, 4, 5, 10, 15, or 20% or less; The sequences corresponding to the residues identified by "*" in the consensus sequences disclosed in Figures 2A-2G include sequences that differ by 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, or 40% or less from the corresponding "*" residues in natural Cas9 molecules such as the S. pyogenes, S. thermophilus, S. mutans, or L. innocua Cas9 molecules.
[0309] In one embodiment, the modified Cas9 molecule or Cas9 polypeptide is an eaCas9 molecule or eaCas9 polypeptide comprising the amino acid sequence of S. pyogenes Cas9 (SEQ ID NO: 2) disclosed in Figures 2A-2G, wherein one or more residues (e.g., 2, 3, 5, 10, 15, 20, 30, 50, 70, 80, 90, 100, or 200 amino acid residues) have one or more amino acids (e.g., substitutions) that are different from the sequence of S. pyogenes, indicated by "*" in the consensus sequence (SEQ ID NO: 14) disclosed in Figures 2A-2G.
[0310] In one embodiment, the modified Cas9 molecule or Cas9 polypeptide is an eaCas9 molecule or eaCas9 polypeptide comprising the amino acid sequence of S. thermophilus Cas9 (SEQ ID NO: 4) disclosed in Figures 2A-2G, wherein one or more residues (e.g., 2, 3, 5, 10, 15, 20, 30, 50, 70, 80, 90, 100, or 200 amino acid residues) have one or more amino acids (e.g., substitutions) that are different from the sequence of S. thermophilus, as indicated by "*" in the consensus sequence (SEQ ID NO: 14) disclosed in Figures 2A-2G.
[0311] In one embodiment, the modified Cas9 molecule or Cas9 polypeptide is an eaCas9 molecule or eaCas9 polypeptide comprising the amino acid sequence of S. mutans Cas9 (SEQ ID NO: 5) disclosed in Figures 2A-2G, wherein one or more residues (e.g., 2, 3, 5, 10, 15, 20, 30, 50, 70, 80, 90, 100, or 200 amino acid residues) have one or more amino acids (e.g., substitutions) that are different from the sequence of S. mutans, as indicated by "*" in the consensus sequence (SEQ ID NO: 14) disclosed in Figures 2A-2G.
[0312] In one embodiment, the modified Cas9 molecule or Cas9 polypeptide is an eaCas9 molecule or eaCas9 polypeptide comprising the amino acid sequence of L. innocua Cas9 disclosed in Figures 2A-2G, wherein one or more residues (e.g., 2, 3, 5, 10, 15, 20, 30, 50, 70, 80, 90, 100, or 200 amino acid residues) have one or more amino acids (e.g., substitutions) that are different from the L. innocua sequence, as indicated by "*" in the consensus sequence disclosed in Figures 2A-2G.
[0313] In certain embodiments, the modified Cas9 molecule or Cas9 polypeptide, for example, the eaCas9 molecule or eaCas9 polypeptide, may be, for example, a fusion of two or more different Cas9 molecules, for example, two or more naturally occurring Cas9 molecules from different species. For example, a natural Cas9 molecule fragment from one species may be fused with a Cas9 molecule fragment from a second species. As an example, a fragment of the S. pyogenes Cas9 molecule containing an N-terminal RuvC-like domain may be fused with a fragment of the Cas9 molecule from a species other than S. pyogenes (for example, S. thermophilus) containing an HNH-like domain.
[0314] Cas9 with altered PAM recognition or no PAM recognition Naturally occurring Cas9 molecules can recognize specific PAM sequences, such as the PAM recognition sequences described above, in species such as S. pyogenes, S. thermophilus, S. mutans, and S. aureus.
[0315] In certain embodiments, the Cas9 molecule or Cas9 polypeptide has the same PAM specificity as the natural Cas9 molecule. In other embodiments, the Cas9 molecule or Cas9 polypeptide has PAM specificity unrelated to the natural Cas9 molecule with the closest sequence homology, or PAM specificity unrelated to the natural Cas9 molecule. For example, the natural Cas9 molecule may be modified, for example, by modifying PAM recognition, for example, by modifying the PAM sequence recognized by the Cas9 molecule or Cas9 polypeptide, to reduce off-target sites and / or improve specificity; or to eliminate the PAM recognition requirement. In certain embodiments, the Cas9 molecule or Cas9 polypeptide may be modified, for example, by increasing the length of the PAM recognition sequence and / or improving Cas9 specificity to a high level of identity (e.g., 98%, 99%, or 100% match between the gRNA and PAM sequences), for example, by reducing off-target sites and / or increasing specificity. In certain embodiments, the length of the PAM recognition sequence is at least 4, 5, 6, 7, 8, 9, 10, or 15 amino acids. In one embodiment, Cas9 specificity requires at least 90%, 95%, 97%, 98%, or 99% homology between the gRNA and the PAM sequence. Cas9 molecules or Cas9 polypeptides that recognize different PAM sequences and / or have reduced off-target activity can be constructed using directed evolution. Exemplary methods and systems that can be used for directed evolution of Cas9 molecules are described (see, for example, Esvelt 2011). Candidate Cas9 molecules may be evaluated by methods, for example, described below.
[0316] Size Optimization Cas9 Genetically modified Cas9 molecules and genetically modified Cas9 polypeptides described herein include, for example, Cas9 molecules or Cas9 polypeptides that contain deletions that reduce molecular size while still retaining desired Cas9 properties such as essentially native conformation, Cas9 nuclease activity, and / or target nucleic acid molecule recognition. Provided herein are Cas9 molecules or Cas9 polypeptides comprising one or more deletions and optionally one or more linkers, the linkers being positioned between amino acid residues located on the side of the deletion. Methods for identifying appropriate deletions in a reference Cas9 molecule, methods for generating Cas9 molecules having deletions and linkers, and methods for using such Cas9 molecules will be apparent to those skilled in the art by examining this literature.
[0317] For example, Cas9 molecules with deletions, such as those found in S. aureus or S. pyogenes Cas9 molecules, are smaller than the corresponding natural Cas9 molecules, for example, having a reduced number of amino acids. The smaller size of Cas9 molecules increases the flexibility of delivery methods, thereby enhancing the usefulness of genome editing. Cas9 molecules may contain one or more deletions that do not substantially affect or reduce the resulting activity of the Cas9 molecules described herein. Functions maintained in Cas9 molecules containing the deletions described herein include nickase activity, i.e., the ability to cleave single strands, such as the non-complementary or complementary strands of nucleic acid molecules; Nickase activity, i.e., the ability to cleave single strands, such as the non-complementary or complementary strand of a nucleic acid molecule; double-strand nuclease activity, i.e., the ability to cleave both strands of a double-stranded nucleic acid to produce a double-strand break, which in one embodiment is the presence of two nickase activities; endonuclease activity; exonuclease activity; helicase activity, i.e., the ability to unwind the helical structure of a double-stranded nucleic acid; and recognition activity of nucleic acid molecules, such as target nucleic acids or gRNA.
[0318] The activity of the Cas9 molecules described herein can be evaluated using activity assays described herein or in the art.
[0319] Identify regions suitable for deletion. Regions suitable for Cas9 molecule deletion can be identified by a variety of methods. Natural orthologous Cas9 molecules from various bacterial species, e.g., any one of those listed in Table 1, can be modeled to the crystal structure of S. pyogenes Cas9 (Nishimasu 2014), and the level of conservation of the entire selected Cas9 orthologue can be examined by comparing it to the three-dimensional structure of the protein. For example, low-conservation or non-conservation regions spatially distal to areas involved in Cas9 activity, such as interfaces with target nucleic acid molecules and / or gRNA, correspond to candidate deletion regions or domains that do not substantially affect or reduce Cas9 activity.
[0320] Nucleic acid encoding the Cas9 molecule For example, nucleic acids encoding Cas9 molecules or Cas9 polypeptides, such as the eaCas9 molecule or eaCas9 polypeptide, are provided herein. Representative nucleic acids or Cas9 polypeptides encoding Cas9 molecules are described above (see, for example, Cong 2013; Wang 2013; Mali 2013; Jinek 2012).
[0321] In one embodiment, the nucleic acid encoding the Cas9 molecule or Cas9 polypeptide may be a synthetic nucleic acid sequence. For example, the synthetic nucleic acid molecule may be chemically modified, for example, as described herein. In one embodiment, the Cas9 mRNA has one or more properties (for example, all of the following): it is capped, polyadenylated, and substituted with 5-methylcytidine and / or pseudouridine.
[0322] In addition or alternatively, the synthetic nucleic acid sequence may be codon-optimized, for example, by substituting at least one less common codon with a common codon. For example, the synthetic nucleic acid may induce the synthesis of optimized messenger mRNA, optimized for expression in the mammalian expression system described herein.
[0323] Furthermore, or alternatively, the nucleic acid encoding the aCas9 molecule or Cas9 polypeptide may include a nuclear localization sequence (NLS), which is known in the art.
[0324] An exemplary codon-optimized nucleic acid sequence encoding the Cas9 molecule of S. pyogenes is described in Sequence ID No. 3. The corresponding amino acid sequence of the S. pyogenes Cas9 molecule is described in Sequence ID No. 2.
[0325] Exemplary codon-optimized nucleic acid sequences encoding the S. aureus Cas9 molecule are described in SEQ ID NOs. 7-9. The amino acid sequence of the S. aureus Cas9 molecule is described in SEQ ID NO. 6.
[0326] It is understood that if any of the above Cas9 sequences fuse with a peptide or polypeptide at the C-terminus, the stop codon is excluded.
[0327] Other Cas molecules and Cas polypeptides Various types of Cas molecules or Cas polypeptides may be used to carry out the inventions disclosed herein. In some embodiments, Cas molecules of type II Cas systems are used. In other embodiments, Cas molecules of other Cas systems are used. For example, type I or type III Cas molecules may be used. Representative Cas molecules (and Cas systems) are described above (see, for example, Haft 2005 and Makarova 2011). Representative Cas molecules (and Cas systems) are also shown in Table 2.
[0328] [Table 3]
[0329] [Table 4]
[0330] [Table 5]
[0331] Cpfl molecule The crystal structure of Acidaminococcus sp. Cpf1, which forms a complex with crRNA and double-stranded (ds)DNA targets (including the TTTN PAM sequence), was elucidated by Yamano 2016, which is incorporated herein by reference. Like Cas9, Cpf1 has two lobes: a REC (recognition) lobe and a NUC (nuclease) lobe. The REC lobe contains REC1 and REC2 domains and is not similar to any known protein structure. The NUC lobe, on the other hand, contains three RuvC domains (RuvC-1, RuvC-II, and RuvC-III) and a BH domain. However, in contrast to Cas9, the Cpfl REC lobe lacks an HNH domain and also contains other domains that similarly lack similarity to known protein structures: a structurally unique PI domain, three Wedge (WED) domains (WED-1, WED-II, and WED-III), and a nuclease (Nuc) domain.
[0332] While Cas9 and Cpfl share structural and functional similarities, it should be understood that certain Cpfl activities are mediated by structural domains that are not similar to any Cas9 domain. For example, cleavage of the complementary strand of target DNA appears to be mediated by a Nuc domain that is sequence- and spatially distinct from the HNH domain of Cas9. In addition, the non-target region (handle) of Cpfl gRNA adopts a pseudoknot structure rather than the stem-loop structure formed by the repeat:anti-repeat double helix in Cas9 gRNA.
[0333] Modification of RNA-induced nucleases While the RNA-inducing nucleases described above possess activities and properties that may be useful in a variety of applications, those skilled in the art will understand that, in some cases, RNA-inducing nucleases can also be modified to alter their cleavage activity, PAM specificity, or other structural or functional characteristics.
[0334] First, regarding modifications that alter cleavage activity, mutations that reduce or eliminate the activity of domains within the NUC lobe are described above. Exemplary mutations that can occur in the RuvC domain, Cas9 HNH domain, or Cpfl Nuc domain are described in Ran 2013 and Yamano 2016, and Cotta-Ramusino 2016. Generally, mutations that reduce or eliminate the activity of one of two nuclease domains produce RNA-inducible nucleases with nickase activity, but it should be noted that the type of nickase activity differs depending on which domain is inactivated. As an example, inactivation of the RuvC domain of Cas9 produces a nickase that cleaves the complementary or upper strand, as shown below (where C indicates the cleavage site). [ka]
[0335] On the other hand, inactivation of the Cas9 HNH domain generates a nickase that cleaves the bottom chain or non-complementary chain: [ka]
[0336] Modifications of PAM specificity for naturally occurring Cas9 reference molecules are described in Kleinstiver 2015a for both S. pyogenes and S. aureus (Kleinstiver 2015b). Kleinstiver et al. also describe modifications that improve the targeting fidelity of Cas9 (Kleinstiver 2016). Each of these references is incorporated herein by reference.
[0337] RNA-induced nucleases are divided into two or more parts, as described in Zetsche 2015 and Fine 2015, as referenced herein.
[0338] RNA-induced nucleases can be size-optimized or cleaved in certain embodiments by, for example, gRNA association, target and PAM recognition, and by one or more deletions that reduce the size of the nuclease while maintaining cleavage activity. In certain embodiments, RNA-induced nucleases can optionally be covalently or noncovalently bound to another polypeptide, nucleotide, or other structure using a linker. Exemplary binding nucleases and linkers are described in Guilinger 2014, which are incorporated by reference for all purposes herein.
[0339] RNA-induced nucleases also optionally include tags, for example, nuclear localization signals that facilitate the translocation of the RNA-induced nuclease protein into the nucleus. In certain embodiments, RNA-induced nucleases may incorporate C-terminal and / or N-terminal nuclear localization signals. Nuclear localization sequences are publicly known in the art and are described in Maeder 2015 and elsewhere.
[0340] The aforementioned list of modifications is intended to be illustrative in nature, and those skilled in the art will understand from this disclosure that other modifications may be possible or desirable in specific applications. Therefore, for brevity, the exemplary systems, methods, and compositions of this disclosure are presented with reference to specific RNA-inducing nucleases, but it should be understood that the RNA-inducing nucleases used can be modified without altering their operating principle. Such modifications are within the scope of this disclosure.
[0341] nucleic acids encoding RNA-induced nucleases RNA-induced nucleases, such as Cas9, Cpfl, or nucleic acids encoding functional fragments thereof, are provided herein. Exemplary nucleic acids encoding RNA-induced nucleases have already been described (see, for example, Cong 2013; Wang 2013; Mali 2013; Jinek 2012).
[0342] In some cases, the nucleic acid encoding the RNA-induced nuclease may be a synthetic nucleic acid sequence. For example, synthetic nucleic acid molecules can be chemically modified. In certain embodiments, the mRNA encoding the RNA-induced nuclease may have one or more (e.g., all) of the following properties: the mRNA may be capped; polyadenylated; and substituted with 5-methylcytidine and / or pseudouridine.
[0343] Synthetic nucleic acid sequences can also be codon-optimized, for example, by substituting at least one non-common or less common codon with a common codon. For example, synthetic nucleic acids can be optimized for expression in mammalian expression systems, such as those described herein, by inducing the synthesis of optimized messenger mRNA. An example of a codon-optimized Cas9 coding sequence is shown in Cotta-Ramusino 2016.
[0344] In addition, or alternatively, nucleic acids encoding RNA-induced nucleases may include nuclear localization sequences (NLSs). Nuclear localization sequences are known in the art.
[0345] Functional analysis of candidate molecules Candidate Cas9 molecules, candidate gRNA molecules, and candidate Cas9 / gRNA molecule complexes can be evaluated by methods known in the art or as described herein. For example, exemplary methods for evaluating the endonuclease activity of Cas9 molecules have already been described (Jinek 2012).
[0346] Binding and cleavage assay: Testing the endonuclease activity of the Cas9 molecule. The ability of the Cas9 molecule / gRNA molecule complex to bind to and cleave target nucleic acids can be evaluated by plasmid cleavage assays. In this assay, synthetic or in vitro transcribed gRNA molecules are pre-annealed prior to the reaction by heating to 95°C and slowly cooling to room temperature. Natural or restriction digested linearized plasmid DNA (300 ng (approximately 8 nM)) is incubated with purified Cas9 protein molecules (50–500 nM) and gRNA (50–500 nM, 1:1) in Cas9 plasmid cleavage buffer (20 mM HEPES pH 7.5, 150 mM KCl, 0.5 mM DTT, 0.1 mM EDTA) for 60 minutes at 37°C in or without 10 mM MgCl2. The reaction is stopped with 5× DNA loading buffer (30% glycerol, 1.2% SDS, 250 mM EDTA), separated by 0.8 or 1% agarose gel electrophoresis, and visualized by ethidium bromide staining. The resulting degradation products indicate whether the Cas9 molecule cleaves both DNA strands or only one strand of the double helix. For example, a linear DNA product indicates cleavage of both DNA strands. A nicked, open-circular product indicates cleavage of only one strand of the double helix.
[0347] Alternatively, the ability of the Cas9 molecule / gRNA molecule complex to bind to and cleave target nucleic acids can be evaluated using an oligonucleotide DNA cleavage assay. In this assay, DNA oligonucleotides (10 pmol) are radiolabeled by incubation at 37°C for 30 minutes in 50 μL of reaction mixture with 5 units of T4 polynucleotide kinase and approximately 3-6 pmol (approximately 20-40 mCi) of [γ-32P]-ATP in 1×T4 polynucleotide kinase reaction buffer. After heat inactivation (65°C for 20 minutes), the reaction mixture is purified by passing it through a column to remove any unincorporated labels. Annealing of the labeled oligonucleotide with an equimolar amount of unlabeled complementary oligonucleotide at 95°C for 3 minutes generates a double-stranded substrate (100 nM), followed by slow cooling to room temperature. In the cleavage assay, gRNA molecules are annealed by heating at 95°C for 30 seconds, followed by slow cooling to room temperature. Cas9 (final concentration 500 nM) is pre-incubated in a total volume of 9 μl with annealed gRNA molecules (500 nM) in cleavage assay buffer (20 mM HEPES pH 7.5, 100 mM KCl, 5 mM MgCl2, 1 mM DTT, 5% glycerol). The reaction is initiated by adding 1 μL of target DNA (10 nM) and incubated at 37°C for 1 hour. The reaction is quenched by adding 20 μL of loading dye (5 mM EDTA, 0.025% SDS, 5% glycerol in formamide) and heated at 95°C for 5 minutes. The degradation products are separated on a 12% denatured polyacrylamide gel containing 7 M urea and visualized by phosphorography. The resulting degradation products indicate whether the complementary strand, the non-complementary strand, or both were cleaved.
[0348] One or both of these assays can be used to evaluate the compatibility of candidate gRNA molecules or candidate Cas9 molecules.
[0349] Binding assay: Tests the binding of the Cas9 molecule to target DNA. A typical method for evaluating the binding of the Cas9 molecule to target DNA is described above (Jinek 2012).
[0350] For example, in electrophoretic mobility shift analysis, target DNA double strands are formed by mixing each strand (10 nmol) in deionized water, heating at 95°C for 3 minutes, and slowly cooling to room temperature. All DNA is purified on an 8% natural gel containing 1×TBE. DNA bands are visualized and excised by UV shadowing and eluted by immersing gel fragments in DEPC-treated H2O. The eluted DNA is precipitated with ethanol and dissolved in DEPC-treated H2O. DNA samples are 5' end-labeled with T4 polynucleotide kinase using [γ-32P]-ATP over 30 minutes at 37°C. The polynucleotide kinase is heat-denatured at 65°C for 20 minutes, and any unintegrated radiolabels are removed using a column. The binding assay is performed in a buffer containing 20 mM HEPES pH 7.5, 100 mM KCl, 5 mM MgCl2, 1 mM DTT, and 10% glycerol, in a total volume of 10 μl. The Cas9 protein molecule is programmed with equimolar amounts of pre-annealed gRNA molecules and titrated from 100 pM to 1 μM. Radiolabeled DNA is added to a final concentration of 20 pM. The sample is incubated at 37°C for 1 hour and separated at 4°C on an 8% natural polyacrylamide gel containing 1 × TBE and 5 mM MgCl2. The gel is dried, and the DNA is visualized by phosphorography.
[0351] Differential scanning fluorescence (DSF) The thermal stability of the Cas9-gRNA ribonucleoprotein (RNP) complex can be measured by DSF. This technique measures the thermal stability of the protein, which can be increased under favorable conditions such as the addition of binding RNA molecules, such as gRNA.
[0352] This assay can be performed using two different protocols: one for testing the best stoichiometric ratio of gRNA:Cas9 protein, and another for determining the best solution conditions for RNP formation.
[0353] To determine the optimal conditions for RNP complex formation, a 2 μM aqueous solution of Cas9 and 10 × SYPRO Orange® (Life Technologies catalog number S-6650) are dispensed into a 384-well plate. Next, equimolar amounts of gRNA diluted in solutions of various pH levels, and salts are added. After incubation at room temperature for 10 minutes and brief centrifugation to remove any bubbles, a gradient from 20°C to 90°C is performed using a Bio-Rad CFX384® Real-Time System C1000 Touch® thermal cycler with a temperature increase of 1°C every 10 seconds, along with Bio-Rad CFX Manager software.
[0354] Assay 2 The second assay consists of mixing gRNA molecules at various concentrations with 2 μM Cas9 in the optimal buffer from assay 1 and incubating the mixture in a 384-well plate at room temperature for 10 minutes. Equivolutes of the optimal buffer and 10 × SYPRO Orange® (Life Technologies catalog no. S-6650) are added, and the plate is sealed with Microseal® B adhesive (MSB-1001). Following a short centrifugation to remove any air bubbles, a gradient from 20°C to 90°C is performed using a Bio-Rad CFX384® Real-Time System C1000 Touch® thermal cycler with a temperature increase of 1°C every 10 seconds, along with Bio-Rad CFX Manager software.
[0355] NHEJ approach for gene targeting In certain embodiments of the methods provided herein, NHEJ-mediated deletion is used to delete all or part of the negative regulatory elements (e.g., silencers) of gamma-globin genes (e.g., HBG1, HBG2). As described herein, nuclease-induced NHEJ can be used to specifically knock out all or part of the regulatory elements. In other embodiments, NHEJ-mediated insertion is used to insert a sequence into the negative regulatory elements of gamma-globin genes, resulting in inactivation of the regulatory elements.
[0356] While we do not wish to be bound by theory, in certain embodiments, genomic alterations associated with the methods described herein are thought to depend on the erroneous nature of nuclease-induced NHEJ and the NHEJ repair pathway. NHEJ repair double-strand breaks in DNA by joining two ends into one; however, generally, the original sequence is restored only when two matching ends, identical to the ends formed by the double-strand break, are perfectly joined. The DNA ends of a double-strand break are often a challenge to enzymatic processing, with nucleotides being added or removed on one or both strands before the ends can be rejoined. As a result, insertion and / or deletion (indel) mutations exist in the DNA sequence of the NHEJ repair site. Two-thirds of these mutations typically alter the leading frame and thus produce non-functional proteins. In addition, mutations that maintain the leading frame but insert or delete a significant amount of sequence can disrupt protein function. This is locus-dependent, as mutations in critical functional domains are more likely to be unacceptable than mutations in non-critical regions of the protein.
[0357] Indel mutations generated by NHEJ are inherently unpredictable; however, at a given cleavage site, certain indel sequences are preferred and overrepresented in that population, likely due to minute homology in small regions. Deletion lengths can vary widely; this length is most commonly in the range of 1–50 bp, but can reach over 100–200 bp. Insertions tend to be shorter and often involve short duplications of sequences adjacent to and surrounding the cleavage site. However, large insertions are possible, and in such cases, the inserted sequence is often traced to other regions of the genome or to plasmid DNA present in cells.
[0358] Because NHEJ is a mutagenic process, it can also be used to delete small sequence motifs (e.g., motifs less than 50 nucleotides long) unless the generation of a specific final sequence is required. When double-strand breaks are targeted near the target sequence, the deletion mutations caused by NHEJ repair often extend to undesirable nucleotides, thus removing them. In the case of deletions of larger DNA segments, an NHEJ can occur between the ends by introducing two double-strand breaks, one on each side of the sequence, removing the entire intervening sequence. In this way, large DNA segments of several hundred kilobases can be deleted. Using both of these approaches, specific DNA sequences can be deleted; however, the erroneous nature of NHEJ still allows for indel mutations at the repair site.
[0359] NHEJ-mediated indels can be generated by using both double-strand cleaved eaCas9 molecules and single-stranded or nickase-mediated eaCas9 molecules in the methods and compositions described herein. NHEJ-mediated indels targeted to a desired regulatory region can be used to disrupt or delete the targeted regulatory element.
[0360] Arrangement of double-strand or single-strand breaks relative to the target location. In certain embodiments in which a gRNA and Cas9 nuclease induce a double-strand break for the purpose of inducing an NHEJ-mediated indel, the gRNA, for example, a single (or chimeric) gRNA molecule or a modular gRNA molecule, is configured to position a single double-strand break near a nucleotide at a target site. In one embodiment, the cleavage site is located 0 to 30 bp away from the target site (e.g., less than 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 bp away from the target site).
[0361] In certain embodiments, two gRNAs that complex with Cas9 nickase induce two single-strand breaks for the purpose of inducing an NHEJ-mediated indel, the two gRNAs, for example, independent monomolecular (or chimeric) gRNAs or modular gRNAs, are configured to position the two single-strand breaks to provide NHEJ repair at target site nucleotides. In certain embodiments, the gRNAs are configured to position the breaks at the same site on different strands or within a few nucleotides of each other, essentially mimicking a double-strand break. In certain embodiments, the closer nick is 0–30 bp away from the target site (e.g., less than 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 bp away from the target site), and the two nicks are within 25–55 bp of each other (e.g., 25–50, 25–45, 25–40, 25–35, 25–30, 50–55, 45–55, 40–55, 35–55, 30–55, 30–50, 35–50, 40–50, 45–50, 35–45, or 40–45 bp) and less than 100 bp away from each other (e.g., less than 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp). In certain embodiments, the gRNA is configured to place single-strand breaks on both sides of the nucleotide at the target site.
[0362] Both double-strand break eaCas9 molecules and single-strand or nickase, along with the eaCas9 molecule, can be used in the methods and compositions described herein to induce breaks on both sides of a target site. Double-strand breaks or paired single-strand breaks can be induced on both sides of a target site to remove a nucleic acid sequence with two breaks (e.g., a region between the two breaks is deleted). In certain embodiments, two gRNAs, e.g., independent monomolecular (or chimeric) gRNAs or modular gRNAs, are configured to place double-strand breaks on both sides of a target site. In other embodiments, three gRNAs, e.g., independent monomolecular (or chimeric) gRNAs or modular gRNAs, are configured to place a double-strand break (i.e., one gRNA complexing with the Cas9 nuclease) and two single-strand breaks or paired single-strand breaks (i.e., two gRNAs complexing with the Cas9 nickase) on both sides of a target site. In yet another embodiment, four gRNAs, for example, independent monomolecular (or chimeric) gRNAs or modular gRNAs, are configured such that two pairs of single-strand breaks (i.e., two pairs of gRNAs complex with Cas9 nicks) occur on either side of the target site. The closer of the two single-strand breaks, or pairs of single-strand breaks, is ideally located within 0 to 500 bp of the target site (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target site). When nickase is used, the two nicks in a pair are within 25-55 bp of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 50-55, 45-55, 40-55, 35-55, 30-55, 30-50, 35-50, 40-50, 45-50, 35-45, or 40-45 bp) and are less than 100 bp apart from each other (e.g., less than 90, 80, 70, 60, 50, 40, 30, 20, or 10 bp).
[0363] HDR repair, HDR-mediated knock-in, knock-out, or deletion, and template nucleic acids In certain embodiments of the methods provided herein, HDR-mediated sequence alteration is used to alter (e.g., delete, disrupt, or modify) one or more nucleotide sequences in the regulatory region of the gamma-globin gene (e.g., HBG1, HBG2) using an externally provided template nucleic acid (also referred to herein as a donor construct). While we do not wish to be bound by theory, it is assumed that HDR-mediated alteration of the HBG target site within the gamma-globin gene regulatory region occurs by HDR using an externally provided donor template or template nucleic acid. For example, the donor construct or template nucleic acid provides alteration of the HBG target site. In one embodiment, the endogenous genome donor sequence is located on the same chromosome as the target sequence. Furthermore, in another embodiment, the endogenous genome donor sequence is located on a different chromosome than the target sequence. Alteration of the HBG target site by the endogenous genome donor sequence relies on cleavage by the Cas9 molecule. Cleavage by Cas9 may include a double-strand break or two single-strand breaks.
[0364] In certain embodiments of the methods provided herein, HDR-mediated alterations are used to knock out or delete all or part of the negative regulatory elements (e.g., silencers) of gamma-globin genes (e.g., HBG1, HBG2). As described herein, HDR can be used to knock out or delete all or part of the regulatory elements in a target-specific manner.
[0365] In other embodiments, HDR-mediated sequence alterations are used to alter the sequence of one or more nucleotides in the regulatory region of the gamma-globin gene (e.g., HBG1, HBG2) without using an externally provided template nucleic acid. While we do not wish to be bound by theory, it is thought that HDRs with endogenous genomic donor sequences cause alterations to the HBG target site. For example, an endogenous genomic donor sequence results in an alteration to the HBG target site. In one embodiment, the endogenous genomic donor sequence is intended to be located on the same chromosome as the target sequence. In another embodiment, the endogenous genomic donor sequence is further intended to be located on a different chromosome than the target sequence. The alteration to the HBG target site by the endogenous genomic donor sequence depends on cleavage by the Cas9 molecule. Cleavage by Cas9 may include a double-strand break or two single-strand breaks.
[0366] In certain embodiments of the methods provided herein, HDR-mediated alterations are used to alter a single nucleotide in the regulatory region of the γ-globin gene. These embodiments may utilize either one double-strand break or two single-strand breaks. Certain embodiments include single-nucleotide alterations using (1) one double-strand break, (2) two single-strand breaks, (3) two double-strand breaks occurring on each side of the target site, (4) one double-strand break and two single-strand breaks occurring on each side of the target site, (5) four single-strand breaks occurring on each side of the target site, or (6) a single single-strand break.
[0367] In certain embodiments where single-stranded template nucleic acids are used, the target location can be altered by alternative HDR.
[0368] In certain embodiments of the methods provided herein, HDR-mediated alterations are used to introduce one or more nucleotide alterations (e.g., deletions) in the γ-globin gene regulatory region. In certain embodiments, the γ-globin gene regulatory region may be an HBG target site. In certain embodiments, the alteration (e.g., deletion) may be introduced at a target site within the HBG target site. In certain embodiments, the alteration (e.g., deletion) may be selected from one or more of HBG1 13bp del c.-114~-102, HBG1 4bp del c.-225~-222, and HBG1 13bp del c.-114~-102. In certain embodiments, the target site may be selected from one or more of the following: HBG1 c.-114~-102 (e.g., nucleotides 2824~2836 of SEQ ID NO: 902(HBG1)), HBG1 c.-225~-222 (e.g., nucleotides 2716~2719 of SEQ ID NO: 902(HBG1)), and HBG2 c.-114~-102 (e.g., nucleotides 2748~2760 of SEQ ID NO: 903(HBG2)).
[0369] Modifications performed using HBG target site donor templates rely on cleavage by the Cas9 molecule. Cas9 cleavage can include nicks, double-strand breaks, or two single-strand breaks, for example, one on each strand of the target nucleic acid. Following the introduction of cleavage into the target nucleic acid, excision occurs at the cleavage ends, resulting in single strands overlapping the DNA region.
[0370] Standard HDR introduces a double-stranded donor template containing a homologous sequence to the target nucleic acid, which is either directly incorporated into the target nucleic acid or used as a template to modify the target nucleic acid sequence. Following excision in a break, repair can proceed via various pathways, such as the double Holliday junction model (or double-strand break repair, DSBR pathway) or the synthesis-dependent strand annealing (SDSA) pathway. In the double Holliday junction model, strand infiltration occurs due to two single-stranded overhangs of the target nucleic acid relative to the homologous sequence in the donor template, forming an intermediate containing two Holliday junctions. These junctions migrate as new DNA is synthesized from the ends of the infiltrated strand, filling the gap created by the excision. The ends of the newly synthesized DNA ligate with the excised ends, the junctions are degraded, and modification of the target nucleic acid occurs, such as the incorporation of the HPFH mutation sequence of the donor template at the corresponding target site. Cross-linking with the donor template nucleic acid can occur during junction degradation. In the SDSA pathway, a single single-stranded overhang infiltrates the donor template, and new DNA is synthesized from the end of the infiltrating strand to fill the gap created by the excision. The newly synthesized DNA then anneals with the remaining single-stranded overhang, and after new DNA is synthesized to fill the gap, the strands are joined to form a DNA double helix.
[0371] In alternative HDRs, a single-stranded donor template, such as a template nucleic acid, is introduced. Nicks, single-strand breaks, or double-strand breaks in the target nucleic acid to modify the desired HBG target site are mediated by a Cas9 molecule, such as those described herein, where excision occurs in the breaks, exposing a single-stranded overhang. The incorporation of the template nucleic acid sequence for modifying the HBG target site typically occurs via the SDSA pathway, as described above.
[0372] Further details about template nucleic acids are described in Section IV of International Application PCT / US2014 / 057905, entitled “Template nucleic acids.”
[0373] In certain embodiments, double-strand breaks are performed by a Cas9 molecule having cleavage activity associated with an HNH-like domain and cleavage activity associated with a RuvC-like domain, e.g., an N-terminal RuvC-like domain, e.g., wild-type Cas9. Such embodiments require a single gRNA.
[0374] In certain embodiments, a single-strand break, or nick, is carried out by a Cas9 molecule having nickase activity, such as a Cas9 nickase described herein. The nick target nucleic acid can serve as a substrate for alt-HDR.
[0375] In other embodiments, the two single-strand breaks, or nicks, are carried out by a Cas9 molecule having nickase activity, such as cleavage activity associated with the HNH-like domain or cleavage activity associated with the N-terminal RuvC-like domain. Such embodiments typically require two gRNAs, one for each single-strand break configuration. In one embodiment, the nickase-active Cas9 molecule cleaves the strand into which the gRNA hybridizes, but not the strand complementary to the strand into which the gRNA hybridizes. In another embodiment, the nickase-active Cas9 molecule does not cleave the strand into which the gRNA hybridizes, but instead cleaves the strand complementary to the strand into which the gRNA hybridizes.
[0376] In certain embodiments, the nickase is a Cas9 molecule having HNH activity, such as a Cas9 molecule with inactivated RuvC activity, for example, a Cas9 molecule with a mutation at D10, for example, a D10A mutation (e.g., SEQ ID NO: 10). D10A inactivates RuvC; therefore, the Cas9 nickase has HNH activity (only) and cleaves on the strand to which the gRNA hybridizes (e.g., a complementary strand that does not have NGG PAM on it). In other embodiments, a Cas9 molecule having an H840 mutation, for example, H840A, can be used as the nickase. H840A inactivates HNH; therefore, the Cas9 nickase has RuvC activity (only) and cleaves on the non-complementary strand (e.g., a strand that has NGG PAM and whose sequence is identical to that of the gRNA). In other embodiments, a Cas9 molecule having an N863 mutation, for example, an N863A mutation, can be used as the nickase. Since N863A inactivates HNH, Cas9 nickase possesses only RuvC activity and cleaves on the non-complementary strand (the strand that has NGG PAM and whose sequence is identical to that of gRNA).
[0377] In a particular embodiment where two single-stranded nicks are positioned using nickase and two gRNAs, one nick is on the + strand and the other nick is on the target nucleic acid strand. The PAM may face outward. These gRNAs can be selected so that they are separated by approximately 0-50, 0-100, or 0-200 nucleotides. In one embodiment, there is no overlap between the targeting domains and complementary target sequences of the two gRNAs. In another embodiment, the gRNAs are separated by approximately 50, 100, or 200 nucleotides without overlap. In one embodiment, the use of two gRNAs can increase specificity, for example, by reducing off-target binding (Ran 2013).
[0378] In certain embodiments, a single nick can be used to induce an HDR, such as an alt-HDR. This specification explores the possibility of increasing the HR-to-NHEJ ratio at a given cleavage site using a single nick. In certain embodiments, the targeting domain of the gRNA forms a single-strand break in a target nucleic acid strand that is complementary to it. In other embodiments, the targeting domain of the gRNA forms a single-strand break in a target nucleic acid strand other than the complementary strand.
[0379] Arrangement of double-strand or single-strand breaks relative to the target location A double-strand or single-strand break on one side of the strand should be sufficiently close to the HBG target site where the alteration, such as the incorporation of an HPFH mutation, occurs in the desired region. In certain embodiments, the distance is 50, 100, 200, 300, 350, or 400 nucleotides or less from the HBG target site. While we do not wish to be bound by theory, in certain embodiments, the break should be sufficiently close to the HBG target site so that the HBG target site is within the region that undergoes exonuclease-mediated removal during end-cleavage. If the distance between the HBG target site and the break is too large, the sequence to be altered does not have to be included in the end-cleavage and therefore does not have to be altered as a donor sequence, which is either an externally provided donor sequence or an endogenous genomic donor sequence; in some embodiments, the donor sequence is used solely to alter the sequence within the end-cleavage region.
[0380] In certain embodiments, the methods described herein introduce one or more cuts near the enhancer region(s), such as the silencer region(s), of the gamma-globin gene regulatory region(s), such as the HGB1 and / or HGB2 gene(s). In some of these embodiments, two or more cuts are introduced adjacent to at least a portion of the enhancer region(s), such as the silencer region(s), of the gamma-globin gene(s). These two or more cuts remove (e.g., delete) a genomic sequence that includes at least a portion of the enhancer region(s), such as the silencer region(s), of the gamma-globin gene(s). All methods described herein result in alterations to the regulatory region(s), e.g., enhancer region(s), e.g., silencer region(s), e.g., of the HGBl and / or HGB2 gene(s).
[0381] In certain embodiments, the gRNA targeting domain is configured such that a cleavage event, e.g., a double-strand or single-strand break, is located within 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 nucleotides of the region to be altered, e.g., a mutation. The cleavage, e.g., a double-strand or single-strand break, can be located upstream or downstream of the mutation in the region to be modified. In some embodiments, the cleavage is located within the region to be modified, a region defined by at least two mutant nucleotides. In some embodiments, the cleavage is located directly adjacent to the region to be modified, e.g., immediately upstream or downstream of the mutation.
[0382] In certain embodiments, a single-strand break is accompanied by an additional single-strand break positioned by a second gRNA molecule, as discussed below. For example, the targeting domain is configured such that a cleavage event, such as two single-strand breaks, is located within nucleotides 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, or 200 of the HBG target site. In one embodiment, the first and second gRNA molecules are configured such that, upon induction of Cas9 nickase, the single-strand break is accompanied by an additional single-strand break positioned sufficiently close to each other by the second gRNA, resulting in modification of the desired region. In one embodiment, the first and second gRNA molecules are configured such that, for example, if Cas9 is a nickase, the single-strand break placed by the second gRNA is within 10, 20, 30, 40, or 50 nucleotides of the break placed by the first gRNA molecule. In one embodiment, the two gRNA molecules are designed to place breaks at the same position on different strands or within a small number of nucleotides of each other, essentially mimicking, for example, a double-strand break.
[0383] In a specific embodiment in which a gRNA (monomolecule (or chimeric) or modular gRNA) and Cas9 nuclease induce double-strand breaks to induce HDR-mediated sequence modification, the break site is located 0 to 200 base pairs away from the HBG target site (e.g., 0 to 175, 0 to 150, 0 to 125, 0 to 100, 0 to 75, 0 to 50, 0 to 25, 25 to 200, 25 to 175, 25 to 150, 25 to 125, 25 to 100, 25 to 75, 25 to 50, 50 to 200, 50 to 175, 50 to 150, 50 to 125, 50 to 100, 50 to 75, 75 to 200, 75 to 175, 75 to 150, 75 to 125, 75 to 100 base pairs). In certain embodiments, the cleavage site is located 0 to 100 base pairs away from the HBG target site (e.g., 0 to 75, 0 to 50, 0 to 25, 25 to 100, 25 to 75, 25 to 50, 50 to 100, 50 to 75, or 75 to 100 base pairs).
[0384] In certain embodiments, HDR can be promoted by using nickase to induce breaks including overhangs. The single-strand nature of the overhang, for example, can increase the likelihood that the cell will repair the break by HDR, in contrast to NHEJ. In particular, in some embodiments, HDR is promoted by selecting a first gRNA that targets a first target sequence with a first nickase, and a second gRNA that targets a second target sequence on the DNA strand opposite to the first target sequence and is offset from the first nick with a second nickase.
[0385] In certain embodiments, the targeting domain of the gRNA molecule is configured to position the cleavage event far enough away from a pre-selected nucleotide so that the nucleotide is not altered. In certain embodiments, the targeting domain of the gRNA molecule is designed so that the intron cleavage event is located far enough away from the intron / exon boundary or a naturally occurring splice signal to avoid alteration of the exon sequence or unwanted splicing events. The gRNA molecule may be a first, second, third, and / or fourth gRNA molecule, as described below.
[0386] Arrangement of the first cut and the second cut In certain embodiments, the double-strand break may be accompanied by an additional double-strand break positioned by a second gRNA molecule, as discussed below.
[0387] In certain embodiments, the double-strand break may be accompanied by two additional double-strand breaks positioned by a second gRNA molecule and a third gRNA molecule.
[0388] In certain embodiments, the first and second single-strand breaks may be accompanied by two additional double-strand breaks positioned by a third gRNA molecule and a fourth gRNA molecule.
[0389] When using two or more gRNAs to place two or more cleavage events, such as double-strand or single-strand breaks, within a target nucleic acid, it is being investigated whether the same or different Cas9 proteins can carry out these two or more cleavage events. For example, when using two gRNAs to place two double-strand breaks, a single Cas9 nuclease can be used to generate both double-strand breaks. When using two or more gRNAs to place two or more single-strand breaks (nicks), a single Cas9 nickase can be used to generate two or more nicks. When using two or more gRNAs to place at least one double-strand break and at least one single-strand break, two Cas9 proteins, such as one Cas9 nuclease and one Cas9 nickase, can be used. When using two or more Cas9 proteins, it is being investigated whether sequential delivery of the two or more Cas9 proteins can control the specificity of double-strand breaks to single-strand breaks at desired locations in the target nucleic acid.
[0390] In some embodiments, the targeting domains of the first gRNA molecule and the second gRNA molecule are complementary to the counterstrand of the target nucleic acid molecule. In some embodiments, the gRNA molecule and the second gRNA molecule are designed so that the PAM faces outward.
[0391] In certain embodiments, two gRNAs are selected to instruct Cas9-mediated cleavage at two locations at a predetermined distance from each other. In certain embodiments, the two cleavage sites are on the opposite strand of the target nucleic acid. In some embodiments, the two cleavage sites form a blunt-end cleavage, while in other embodiments, they are offset such that the DNA ends include one or two overhangs (e.g., one or more 5' overhangs and / or one or more 3' overhangs). In some embodiments, each cleavage event is a nick. In certain embodiments, the nicks are close enough to each other to form a cleavage recognized by a double-strand break mechanism (e.g., opposite to recognition by an SSBr mechanism). In certain embodiments, the nicks are far enough apart so that they generate overhangs which are substrates for HDRs, i.e., the arrangement of the cleavage mimics a DNA substrate that has undergone several excisions. For example, in some embodiments, the nicks are separated so that they generate overhangs which are substrates for processive excision. In some embodiments, the two cleavage sites are separated by no more than 25–65 nucleotides from each other. The two cleavages may, for example, consist of approximately 25, 30, 35, 40, 45, 50, 55, 60, or 65 nucleotides from each other. The two cleavages may, for example, consist of at least approximately 25, 30, 35, 40, 45, 50, 55, 60, or 65 nucleotides from each other. The two cleavages may, for example, consist of at most approximately 30, 35, 40, 45, 50, 55, 60, or 65 nucleotides from each other. In certain embodiments, the two cleavages may consist of approximately 25-30, 30-35, 35-40, 40-45, 45-50, 50-55, 55-60, or 60-65 nucleotides from each other.
[0392] In some embodiments, cuts that mimic a resected cut include a 3' end overhang (e.g., generated by a DSB and a nick, where the nick leaves a 3' overhang), a 5' overhang (e.g., generated by a DSB and a nick, where the nick leaves a 5' overhang), 3' and 5' overhangs (e.g., generated by three cuts), two 3' overhangs (e.g., generated by two staggered nicks), or two 5' overhangs (e.g., generated by two staggered nicks).
[0393] In certain embodiments in which two gRNAs (independently, single-molecule (or chimeric) or modular gRNAs) that complex with Cas9 nickase induce two single-strand breaks for the purpose of inducing HDR-mediated modification, the closer nicks are 0 to 200 base pairs (e.g., 0 to 175, 0 to 150, 0 to 125, 0 to 100, 0 to 75, 0 to 50, 0 to 25, 25 to 200, 25 to 175, 25 to 150, 25 to 125, 25 to 100, 25 to 75, 25 to 50, 50 to 200, 50 to 175, 50 to 150, 50 to 125, 50 to 100, 50 to 75, 75 to 200, 75 to 175, 75 to The two nicks are ideally within 25-65 base pairs of each other (e.g., 25-50, 25-45, 25-40, 25-35, 25-30, 30-55, 30-50, 30-45, 30-40, 30-35, 35-55, 35-50, 35-45, 35-40, 40-55, 40-50, 40-45 base pairs, 45-50 base pairs, 50-55 base pairs, 55-60 base pairs, or 60-65 base pairs), and are less than 100 base pairs apart from each other (e.g., 90, 80, 70, 60, 50, 40, 30, 20, 10, or 5 base pairs apart from each other). In certain embodiments, the cleavage site is located 0 to 100 base pairs away from the HBG target site (e.g., 0 to 75, 0 to 50, 0 to 25, 25 to 100, 25 to 75, 25 to 50, 50 to 100, 50 to 75, or 75 to 100 base pairs).
[0394] In certain embodiments, two gRNAs, for example, independently, monomolecular (or chimeric) or modular gRNAs, are designed to place double-strand breaks on either side of the target site. In alternative embodiments, three gRNAs, for example, independently, monomolecular (or chimeric) or modular gRNAs, are designed to place a double-strand break (i.e., a complex of Cas9 nuclease and one gRNA) and two single-strand breaks or paired single-strand breaks (i.e., a complex of Cas9 nickase and two gRNAs) on either side of the target site. In other embodiments, four gRNAs, for example, independent monomolecular (or chimeric) gRNAs or modular gRNAs, are configured to produce two pairs of single-strand breaks (i.e., two pairs of gRNAs that complex with Cas9 nickase) on either side of the target site. Ideally, double-strand breaks (or multiple breaks) or proximity of a pair of single-strand nicks should be within 0 to 500 bp of the HBG target site (e.g., 450, 400, 350, 300, 250, 200, 150, 100, 50, or 25 bp or less from the target site). When using nickase, the two nicks in a pair are, in several embodiments, within 25 to 65 base pairs of each other (e.g., 25 to 55, 25 to 50, 25 to 45, 25 to 40, 25 to 35, 25 to 30, 50 to 55, 45 to 55, 40 to 55, 35 to 55, 30 to 55, 30 to 50, 35 to 50, 40 to 50, 45 to 50, 35 to 45, 40 to 45 base pairs, 45 to 50 base pairs, 50 to 55 base pairs, 55 to 60 base pairs, or 60 to 65 base pairs) and separated from each other by no more than 100 base pairs (e.g., no more than 90, 80, 70, 60, 50, 40, 30, or 20 or 10 base pairs).
[0395] When targeting a Cas9 molecule for cleavage using two gRNA molecules, various combinations of Cas9 molecules are possible. In some embodiments, the first gRNA is used to target the first Cas9 molecule at a first target site, and the second gRNA is used to target the second Cas9 molecule at a second target site. In some embodiments, the first Cas9 molecule generates a nick on the first strand of the target nucleic acid, and the second Cas9 molecule generates a nick on the opposing strand, thereby resulting in a double-strand break (e.g., a blunt-end break or a break including an overhang).
[0396] Various combinations of nickase can be selected to target one single-strand break on one strand and a second single-strand break on the opposing strand. When selecting a combination, it can be considered that there are nickase with one active RuvC-like domain and nickase with one active HNH domain. In certain embodiments, the RuvC-like domain cleaves the non-complementary strand of the target nucleic acid molecule. In certain embodiments, the HNH-like domain cleaves the single-strand complementary domain, for example, the complementary strand of a double-stranded nucleic acid molecule. Generally, both Cas9 molecules have the same active domain (e.g., both have an active RuvC domain, or both have an active HNH domain) and will select two gRNAs to bind to the opposing strand of the target. More specifically, in some embodiments, the first gRNA is complementary to the first strand of the target nucleic acid and binds to a nickase having an active RuvC-like domain, causing the nickase to cleave the strand complementary to the first gRNA, i.e., the second strand of the target nucleic acid; the second gRNA is complementary to the second strand of the target nucleic acid and binds to a nickase having an active RuvC-like domain, causing the nickase to cleave the strand complementary to the second gRNA, i.e., the first strand of the target nucleic acid. Conversely, in some embodiments, the first gRNA is complementary to the first strand of the target nucleic acid and binds to a nickase having an active HNH domain, causing the nickase to cleave the strand complementary to the first gRNA, i.e., the first strand of the target nucleic acid; the second gRNA is complementary to the second strand of the target nucleic acid and binds to a nickase having an active HNH domain, causing the nickase to cleave the strand complementary to the second gRNA, i.e., the second strand of the target nucleic acid. In another embodiment, if one Cas9 molecule has an active RuvC-like domain and the other Cas9 molecule has an active HNH domain, the gRNAs of both Cas9 molecules can be complementary to the same strand of the target nucleic acid. Therefore, the Cas9 molecule with the active RuvC-like domain will cleave the non-complementary strand, and the Cas9 molecule with the HNH domain will cleave the complementary strand, resulting in a double-strand break.
[0397] Homology arm of donor template The homology arm should extend at least as far as the region where end excision may occur, so that, for example, an excised single-stranded overhang can find a complementary region within the donor template. Its total length may be limited by parameters such as plasmid size or viral packaging limits. In one embodiment, the homology arm does not extend into repeat elements, such as Alu repeat sequences or LINE repeat sequences.
[0398] Examples of homology arm lengths include at least 50, 100, 250, 500, 750, 1000, 2000, 3000, 4000, or 5000 nucleotides. In some embodiments, the homology arm lengths are 50-100, 100-250, 250-500, 500-750, 750-1000, 1000-2000, 2000-3000, 3000-4000, or 4000-5000 nucleotides.
[0399] In the use herein, a template nucleic acid refers to a nucleic acid sequence that can be used with Cas9 and gRNA molecules to alter (e.g., delete, disrupt, or modify) the structure of an HBG target site. In certain embodiments, the HBG target site may be two nucleotides on a target nucleic acid to which one or more nucleotides have been added, e.g., a site between adjacent nucleotides. Alternatively, the HBG target site may comprise one or more nucleotides to be altered by the template nucleic acid. In certain embodiments, the alteration (e.g., deletion) may be introduced at a target site within the HBG target site. In certain embodiments, the alteration (e.g., deletion) may be selected from one or more of HBG1 13bp del c.-114~-102, HBG1 4bp del c.-225~-222, and HBG1 13bp del c.-114~-102. In certain embodiments, the target site may be selected from one or more of the following: HBG1 c.-114~-102 (e.g., nucleotides 2824~2836 of SEQ ID NO: 902(HBG1)), HBG1 c.-225~-222 (e.g., nucleotides 2716~2719 of SEQ ID NO: 902(HBG1)), and HBG2 c.-114~-102 (e.g., nucleotides 2748~2760 of SEQ ID NO: 903(HBG2)).
[0400] In certain embodiments, the target nucleic acid is modified so that part or all of the sequence of the template nucleic acid is typically located near or adjacent to a cleavage site(s). In certain embodiments, the template nucleic acid is single-stranded. In other embodiments, the template nucleic acid is double-stranded. In certain embodiments, the template nucleic acid is DNA, e.g., double-stranded DNA. In other embodiments, the template nucleic acid is single-stranded DNA. In one embodiment, the template nucleic acid is encoded on the same vector backbone as Cas9 and gRNA, e.g., AAV genome, plasmid DNA. In certain embodiments, the template nucleic acid is excised from the vector backbone in vivo and is adjacent to, for example, a gRNA recognition sequence. In certain embodiments, the template nucleic acid includes an endogenous genomic sequence.
[0401] In certain embodiments, the template nucleic acid alters the structure of the target site by participating in an HDR event. In certain embodiments, the template nucleic acid alters the sequence of the target site. In certain embodiments, the template nucleic acid results in the incorporation of modified or non-natural bases into the target nucleic acid.
[0402] In certain embodiments, the template nucleic acid results in the deletion of one or more nucleotides in the target nucleic acid. In certain embodiments, the template nucleic acid results in the deletion of one or more nucleotides at the HBG target site. In certain embodiments, the alteration (e.g., deletion) can be introduced at the target site within the HBG target site. In certain embodiments, the alteration (e.g., deletion) may be selected from one or more of HBG1 13bp del c.-114~-102, HBG1 4bp del c.-225~-222, and HBG1 13bp del c.-114~-102. In certain embodiments, the target site may be selected from one or more of the following: HBG1 c.-114~-102 (e.g., nucleotides 2824~2836 of SEQ ID NO: 902(HBG1)), HBG1 c.-225~-222 (e.g., nucleotides 2716~2719 of SEQ ID NO: 902(HBG1)), and HBG2 c.-114~-102 (e.g., nucleotides 2748~2760 of SEQ ID NO: 903(HBG2)).
[0403] Typically, a template sequence undergoes cleavage-mediated or catalytic recombination by a target sequence. In certain embodiments, the template nucleic acid includes a sequence corresponding to one site on the target sequence that is cleaved by an eaCas9-mediated cleavage event. In certain embodiments, the template nucleic acid includes a sequence corresponding to both a first site on the target sequence that is cleaved by a first Cas9-mediated event and a second site on the target sequence that is cleaved by a second Cas9-mediated event.
[0404] The structure of the regulatory region can be altered by using a template nucleic acid homologous to the HBG target site in the γ-globin gene regulatory region. For example, one or more nucleotides at the HBG target site can be deleted by using a template nucleic acid homologous to the 5' and 3' regions of the HBG target site in the γ-globin gene regulatory region.
[0405] Template nucleic acids typically consist of the following components: [5' homology arm]-[substitution sequence]-[3' homology arm] Includes.
[0406] The homology arm provides recombination into the chromosome, and therefore replacement of undesirable elements, such as mutations or signatures, by the substitution sequence. The homology arm is a region homologous to a DNA region within or near (e.g., adjacent or contacting) the target nucleic acid to be cleaved. In certain embodiments, the homology arm is adjacent to the most distal cleavage site.
[0407] In certain embodiments, a template nucleic acid can be used to remove (e.g., delete) a genomic sequence containing at least a portion of the enhancer region(s) of the γ-globin gene regulatory region(s), e.g., the HGBl and / or HGB2 gene(s), e.g., the silencer region(s). In certain embodiments, a template nucleic acid can be used to delete one or more nucleotides at an HBG target site, i.e., to introduce a change (e.g., deletion) at an HBG target site. In certain embodiments, the change (e.g., deletion) can be introduced at a target site within the HBG target site. In certain embodiments, the change (e.g., deletion) may be selected from one or more of HBG1 13bp del c.-114~-102, HBG1 4bp del c.-225~-222, and HBG1 13bp del c.-114~-102. In certain embodiments, the target site may be selected from one or more of the following: HBG1 c.-114~-102 (e.g., nucleotides 2824~2836 of SEQ ID NO: 902(HBG1)), HBG1 c.-225~-222 (e.g., nucleotides 2716~2719 of SEQ ID NO: 902(HBG1)), and HBG2 c.-114~-102 (e.g., nucleotides 2748~2760 of SEQ ID NO: 903(HBG2)).
[0408] Substitution sequences in donor templates are described elsewhere, including Cotta-Ramusino 2016, which are incorporated herein by reference. The substitution sequences can be of any suitable length. In certain embodiments, the substitution sequences may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more sequence modifications to the in-cellular, naturally occurring sequence to be edited.
[0409] In certain embodiments, if the desired repair outcome is a deletion of the target nucleic acid, the substitution sequence may be 0 nucleotides or 0 bp. In certain embodiments, the template nucleic acid excludes sequences homologous to the target nucleic acid sequence to be deleted. If the substitution sequence is 0 nucleotides or 0 bp, the target nucleic acid sequence located between the 5' homology arm and the 3' homology arm annealing to the template nucleic acid is deleted.
[0410] In certain embodiments, the 3' end of the 5' homology arm is located adjacent to the 5' end of the substitution sequence. In certain embodiments, the 5' homology arm can be extended at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, or 5000 nucleotides toward the 5' end of the substitution sequence. In certain embodiments, if the substitution sequence is 0 nucleotides or 0 bp, the 3' end of the 5' homology arm is located adjacent to the 5' end of the 3' homology arm. In certain embodiments where the substitution sequence is 0 nucleotides or 0 bp, the 5' homology arm can be extended by at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, or 5000 nucleotides toward the 5' end of the 3' homology arm.
[0411] In certain embodiments, the 5' end of the 3' homology arm is located adjacent to the 3' end of the substitution sequence. In one embodiment, the 3' homology arm can be extended at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, or 5000 nucleotides toward the 3' end of the substitution sequence. In certain embodiments where the substitution sequence is 0 nucleotides or 0 bp, the 5' end of the 3' homology arm is located adjacent to the 3' end of the 5' homology arm. In one embodiment, the 3' homology arm can be extended by at least 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, or 5000 nucleotides toward the 3' end of the 5' homology arm.
[0412] In certain embodiments, homology arms, e.g., 5' and 3' homology arms, may each contain a sequence of approximately 1000 bp adjacent to the most distal gRNA (e.g., a 1000 bp sequence on either side of the HBG target site) in order to alter one or more nucleotides at the HBG target site.
[0413] In this specification, it is considered that one or both homology arms may be shortened to avoid the inclusion of certain sequence repeating elements, such as Alu repeating elements or LINE elements. For example, the 5' homology arm may be shortened to avoid sequence repeating elements. In other embodiments, the 3' homology arm may be shortened to avoid sequence repeating elements. In some embodiments, both the 5' and 3' homology arms may be shortened to avoid the inclusion of certain sequence repeating elements.
[0414] In this specification, it is intended that template nucleic acids for altering the sequence of HBG target sites can be designed to be used as single-stranded oligonucleotides, such as single-stranded oligodeoxynucleotides (ssODNs). When using ssODNs, the 5' and 3' homology arms may be in the range of up to approximately 200 nucleotide lengths, e.g., at least 25, 50, 75, 100, 125, 150, 175, or 200 bp. Longer homology arms are also being considered for ssODNs as improvements in oligonucleotide synthesis are continuously being made. In some embodiments, longer homology arms are achieved by methods other than chemical synthesis, e.g., by denaturing a long double-stranded nucleic acid and then purifying one of the strands by affinity to a strand-specific sequence immobilized on a solid support, for example.
[0415] While we do not wish to be bound by theory, in certain embodiments, alt-HDR proceeds more efficiently when the template nucleic acid has homology extended to the 5' side of the nick (i.e., in the 5' direction of the strand containing the nick) or the target site (i.e., in the 5' direction of the target site). Thus, in some embodiments, the template nucleic acid has a long homology arm and a short homology arm, where the longer homology arm can anneal to the 5' side of the nick or target site. In some embodiments, the arm that can anneal to the 5' side of the nick or target site is at least 25, 50, 75, 100, 125, 150, 175, or 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000, 4000, or 5000 nucleotides from the 5' or 3' end of the nick or target site or substitution sequence. In some embodiments, the arm that can anneal to the 5' side of the nick or target site is at least 10%, 20%, 30%, 40%, or 50% longer than the arm that can anneal to the 3' side of the nick or target site. In some embodiments, the arm that can anneal to the 5' side of the nick or target site is at least 2, 3, 4, or 5 times longer than the arm that can anneal to the 3' side of the nick or target site. Depending on whether the ssDNA template can anneal to an intact strand or a cleaved strand, the homology arm that anneals to the 5' side of the nick or target site may be located at the 5' end or the 3' end of the ssDNA template, respectively.
[0416] Similarly, in some embodiments, the template nucleic acid has a 5' homology arm, a substitution sequence, and a 3' homology arm such that the template nucleic acid has homology extending to the 5' side of the nick. For example, the 5' and 3' homology arms may be substantially the same length, but the substitution sequence may extend further on the 5' side of the nick than on the 3' side. In some embodiments, the substitution sequence extends at ...
Claims
1. A gRNA molecule containing a targeting domain that includes a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
2. The gRNA molecule according to claim 1, wherein the HBG1 or HBG2 regulatory region is adjacent to the HBG1 gene or the HBG2 gene, respectively.
3. The gRNA molecule according to claim 2, wherein the HBG1 regulatory region is located in the region spanning nucleotides 1 to 2990 of SEQ ID NO:
902.
4. The gRNA molecule according to claim 2, wherein the HBG2 regulatory region is located in the region spanning nucleotides 1 to 2914 of sequence number 903.
5. The gRNA molecule according to any one of claims 1 to 4, wherein the targeting domain is configured to provide a cleavage event selected from double-strand breaks and single-strand breaks within 500, 400, 300, 200, 100, 50, 25, or 10 nucleotides of the HBG target site.
6. The gRNA molecule according to claim 1, wherein the target domain is entirely located within the HBG1 or HBG2 regulatory region.
7. The gRNA molecule according to any one of claims 1 to 6, wherein the targeting domain is configured to target a transcriptional regulatory element in the HBG1 or HBG2 regulatory region.
8. The gRNA molecule according to claim 7, wherein the transcriptional regulatory element is a promoter.
9. The gRNA molecule according to claim 8, wherein the promoter controls the transcription of one or more HBG1 and HBG2.
10. The gRNA molecule according to claim 7, wherein the transcriptional regulatory element is a silencer.
11. The gRNA molecule according to any one of claims 1 to 10, wherein the targeting domain includes a nucleotide sequence that is identical to, or differs from, a nucleotide sequence of 1, 2, 3, 4, or 5 or fewer nucleotides from, any one of the nucleotide sequences shown in any one of sequence numbers 251 to 901.
12. The gRNA molecule according to claim 11, wherein the targeting domain includes the same nucleotide sequence as the nucleotide sequence shown in any one of SEQ ID NOs. 251 to 901.
13. The gRNA molecule according to any one of claims 1 to 12, wherein the gRNA molecule is a modular gRNA molecule.
14. The gRNA molecule according to any one of claims 1 to 12, wherein the gRNA molecule is a single gRNA molecule.
15. The gRNA molecule according to any one of claims 1 to 12, wherein the gRNA molecule is a chimeric gRNA molecule.
16. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 16 nucleotides or longer.
17. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 17 nucleotides or longer.
18. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 18 nucleotides or longer.
19. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 19 nucleotides or longer.
20. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 20 nucleotides or longer.
21. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 21 nucleotides or longer.
22. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 22 nucleotides or longer.
23. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 23 nucleotides or longer.
24. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 24 nucleotides or longer.
25. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 25 nucleotides or longer.
26. The gRNA molecule according to any one of claims 1 to 12, wherein the targeting domain is 26 nucleotides or longer.
27. A gRNA molecule according to any one of claims 1 to 12, further comprising one or more of a first complementarity domain, a linking domain, a second complementarity domain, an adjacent domain, a 5' extension domain, and a tail domain.
28. The gRNA molecule according to claim 27, comprising: a targeting domain; a first complementarity domain; a linking domain; a second complementarity domain; and an adjacent domain from 5' to 3'.
29. The gRNA molecule according to claim 28, further comprising a tail domain.
30. A gRNA molecule according to any one of claims 1 to 29, comprising: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 20 nucleotides together; and a targeting domain consisting of 17 or 18 nucleotides.
31. A gRNA molecule according to any one of claims 1 to 29, comprising: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 25 nucleotides together; and a targeting domain consisting of 17 or 18 nucleotides.
32. A gRNA molecule according to any one of claims 1 to 29, comprising: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 30 nucleotides together; and a targeting domain consisting of 17 nucleotides.
33. A gRNA molecule according to any one of claims 1 to 37, comprising: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 40 nucleotides together; and a targeting domain consisting of 17 nucleotides.
34. (a) A nucleic acid composition comprising a nucleotide sequence encoding a gRNA molecule, which includes a targeting domain comprising a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
35. The nucleic acid composition according to claim 34, wherein the gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
36. The nucleic acid composition according to claim 34 or 35, wherein the targeting domain is configured to provide a cleavage event selected from double-strand breaks and single-strand breaks within 500, 400, 300, 200, 100, 50, 25, or 10 nucleotides of the HBG target site.
37. The nucleic acid composition according to any one of claims 34 to 36, wherein the targeting domain includes a nucleotide sequence that is identical to, or differs from, a nucleotide sequence in any one of SEQ ID NOs. 251 to 901 by only 1, 2, 3, 4, or 5 or fewer nucleotides.
38. The nucleic acid composition according to claim 37, wherein the targeting domain comprises a nucleotide sequence that is identical to the nucleotide sequence shown in any one of SEQ ID NOs. 251 to 901.
39. The nucleic acid composition according to any one of claims 34 to 38, wherein the gRNA molecule is a modular gRNA molecule.
40. The nucleic acid composition according to any one of claims 34 to 38, wherein the gRNA molecule is a single gRNA molecule.
41. The nucleic acid composition according to any one of claims 34 to 38, wherein the gRNA molecule is a chimeric gRNA molecule.
42. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 16 nucleotides or longer.
43. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 17 nucleotides or longer.
44. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 18 nucleotides or longer.
45. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 19 nucleotides or longer.
46. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 20 nucleotides or longer.
47. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 21 nucleotides or longer.
48. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 22 nucleotides or longer.
49. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 23 nucleotides or longer.
50. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 24 nucleotides or longer.
51. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 25 nucleotides or longer.
52. The nucleic acid composition according to any one of claims 34 to 41, wherein the targeting domain is 26 nucleotides or longer.
53. The nucleic acid composition according to any one of claims 34 to 52, wherein the gRNA molecule further comprises one or more of the following: a targeting domain; a first complementarity domain; a linking domain; a second complementarity domain; and an adjacent domain.
54. The nucleic acid composition according to claim 53, wherein the gRNA molecule comprises, from 5' to 3': a targeting domain; a first complementarity domain; a ligation domain; a second complementarity domain; and an adjacent domain.
55. The nucleic acid composition according to claim 54, further comprising a tail domain in the gRNA.
56. The nucleic acid composition according to any one of claims 34 to 55, wherein the gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 20 nucleotides together; and a targeting domain of 17 or 18 nucleotides.
57. The nucleic acid composition according to any one of claims 34 to 55, wherein the gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 25 nucleotides together; and a targeting domain having a length of 17 or 18 nucleotides.
58. The nucleic acid composition according to any one of claims 34 to 55, wherein the gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 30 nucleotides together; and a targeting domain 17 nucleotides long.
59. The nucleic acid composition according to any one of claims 34 to 55, wherein the gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 40 nucleotides together; and a targeting domain 17 nucleotides long.
60. (b) A nucleic acid composition according to any one of claims 34 to 59, further comprising a nucleotide sequence encoding an RNA-induced nuclease.
61. The nucleic acid composition according to claim 60, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
62. The nucleic acid composition according to claim 61, wherein the Cas9 molecule is an enzymatically active Cas9 (eaCas9) molecule.
63. The nucleic acid composition according to claim 62, wherein the eaCas9 molecule contains a nickase molecule.
64. The nucleic acid composition according to claim 62 or 63, wherein the eaCas9 molecule causes a double-strand break in the target nucleic acid.
65. The nucleic acid composition according to claim 62 or 63, wherein the eaCas9 molecule causes a single-strand break in the target nucleic acid.
66. The nucleic acid composition according to claim 65, wherein the single-strand break occurs in a target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
67. The nucleic acid composition according to claim 66, wherein the single-strand break occurs in a target nucleic acid strand other than the target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
68. The nucleic acid composition according to any one of claims 61 to 67, wherein the eaCas9 molecule has HNH-like domain cleavage activity but does not have N-terminal RuvC-like domain cleavage activity or significant N-terminal RuvC-like domain cleavage activity.
69. The nucleic acid composition according to claim 68, wherein the eaCas9 molecule is an HNH-like domain niccasase.
70. The nucleic acid composition according to claim 68 or 69, wherein the eaCas9 molecule contains a mutation at D10.
71. The nucleic acid composition according to any one of claims 65 to 70, wherein the eaCas9 molecule has N-terminal RuvC-like domain cleavage activity but does not have HNH-like domain cleavage activity or significant HNH-like domain cleavage activity.
72. The nucleic acid composition according to claim 70, wherein the eaCas9 molecule is an N-terminal RuvC-like domain niccasse.
73. The nucleic acid composition according to claim 71 or claim 72, wherein the eaCas9 molecule contains a mutation at H840 or N863.
74. (c) A nucleic acid composition according to any one of claims 34 to 73, further comprising a nucleotide sequence encoding a second gRNA molecule, the nucleotide sequence comprising a targeting domain comprising a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
75. The nucleic acid composition according to claim 74, wherein the second gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
76. The nucleic acid composition according to claim 74 or 75, wherein the targeting domain of the second gRNA molecule is configured to provide a cleavage event selected from double-strand breaks and single-strand breaks within 500, 400, 300, 200, 100, 50, 25, or 10 nucleotides of the HBG target site.
77. The nucleic acid composition according to any one of claims 74 to 76, wherein the targeting domain of the second gRNA molecule includes a nucleotide sequence that is identical to, or differs from, a nucleotide sequence of 1, 2, 3, 4, or 5 or fewer nucleotides in any one of SEQ ID NOs: 251 to 901.
78. The nucleic acid composition according to claim 77, wherein the targeting domain of the second gRNA molecule contains a nucleotide sequence identical to the nucleotide sequence shown in any one of SEQ ID NOs. 251 to 901.
79. The nucleic acid composition according to any one of claims 74 to 78, wherein the second gRNA molecule is a single gRNA molecule.
80. The nucleic acid composition according to any one of claims 74 to 78, wherein the second gRNA molecule is a modular gRNA molecule.
81. The nucleic acid composition according to any one of claims 74 to 78, wherein the second gRNA molecule is a chimeric gRNA molecule.
82. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 16 nucleotides or longer.
83. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 17 nucleotides or longer.
84. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 18 nucleotides or longer.
85. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 19 nucleotides or longer.
86. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 20 nucleotides or longer.
87. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 21 nucleotides or longer.
88. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 22 nucleotides or longer.
89. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 23 nucleotides or longer.
90. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 24 nucleotides or longer.
91. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 25 nucleotides or longer.
92. The nucleic acid composition according to any one of claims 74 to 81, wherein the targeting domain of the second gRNA molecule is 26 nucleotides or longer.
93. The nucleic acid composition according to any one of claims 74 to 92, wherein the second gRNA molecule further comprises one or more of a targeting domain; a first complementarity domain; a linking domain; a second complementarity domain; and an adjacent domain.
94. The nucleic acid composition according to claim 93, wherein the second gRNA molecule comprises, from 5' to 3': a targeting domain; a first complementarity domain; a linking domain; a second complementarity domain; and an adjacent domain.
95. The nucleic acid composition according to claim 94, wherein the second gRNA further comprises a tail domain.
96. The nucleic acid composition according to any one of claims 74 to 95, wherein the second gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 20 nucleotides together; and a targeting domain of 17 or 18 nucleotides.
97. The nucleic acid composition according to any one of claims 74 to 95, wherein the second gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 25 nucleotides together; and a targeting domain having a length of 17 or 18 nucleotides.
98. The nucleic acid composition according to any one of claims 74 to 95, wherein the second gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 30 nucleotides together; and a targeting domain 17 nucleotides long.
99. The nucleic acid composition according to any one of claims 74 to 95, wherein the second gRNA molecule comprises: a ligation domain containing 25 or fewer nucleotides; adjacent domains and a tail domain containing at least 40 nucleotides together; and a targeting domain 17 nucleotides long.
100. (d) A nucleic acid composition according to any one of claims 74 to 99, further comprising a nucleotide sequence encoding a third gRNA molecule, the nucleotide sequence comprising a targeting domain comprising a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
101. (f) A nucleotide sequence encoding a fourth gRNA molecule comprising a targeting domain that is complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region, the nucleic acid composition according to claim 100.
102. (g) A nucleic acid composition according to any one of claims 34 to 101, further comprising a template nucleic acid.
103. The nucleic acid composition according to claim 102, wherein the template nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
104. The nucleic acid composition according to claim 103, wherein the template nucleic acid comprises a 5' homology arm, a substitution sequence, and a 3' homology arm.
105. The nucleic acid composition according to claim 104, wherein the 5' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence has a length of 0 nucleotides; and the 3' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.
106. The nucleic acid composition according to claim 105, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of the target site within the HBG target position, and the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 3' side of the target site within the HBG target position.
107. The nucleic acid composition according to claim 106, wherein the target site is selected from the group consisting of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)); HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)); and HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
108. The nucleic acid composition according to claim 107, wherein the target site is HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
109. The nucleic acid composition according to claim 108, wherein the 5' homology arm is essentially composed of or consists of SEQ ID NO: 904 (i.e., ssODN1 5' homology arm), which includes SEQ ID NO:
904.
110. The nucleic acid composition according to claim 108 or 109, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
111. The nucleic acid composition according to any one of claims 108 to 110, wherein the 3' homology arm is essentially composed of or consists of SEQ ID NO: 905 (i.e., ssODN1 3' homology arm), wherein the 3' homology arm is SEQ ID NO:
905.
112. The nucleic acid composition according to claim 107, wherein the target site is HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
113. The nucleic acid composition according to claim 112, wherein the 5' homology arm essentially consists of or comprises SEQ ID NO: 904 (i.e., ssODN1 5' homology arm).
114. The nucleic acid composition according to claim 112 or 113, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
115. The nucleic acid composition according to any one of claims 112 to 114, wherein the 3' homology arm is essentially composed of or consists of SEQ ID NO: 905 (i.e., ssODN1 3' homology arm), wherein the 3' homology arm is SEQ ID NO:
905.
116. The nucleic acid composition according to any one of claims 108 to 115, wherein the template nucleic acid comprises SEQ ID NO: 906 (i.e., ssODN1), essentially consisting of SEQ ID NO: 906, or comprising SEQ ID NO:
906.
117. The nucleic acid composition according to claim 107, wherein the target site is HBG1 c.-225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG1 c.-225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
118. The nucleic acid composition according to claim 117, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
119. The nucleic acid composition according to any one of claims 103 to 118, wherein the ssODN comprises a 5' phosphorothioate modification.
120. The nucleic acid composition according to any one of claims 103 to 118, wherein the ssODN comprises a 3'-phosphorothioate modification.
121. The nucleic acid composition according to any one of claims 103 to 118, wherein the ssODN comprises a 5' phosphorothioate modification and a 3' phosphorothioate modification.
122. The nucleic acid composition according to any one of claims 34 to 121, which does not include (c) a nucleotide sequence encoding a second gRNA molecule, (d) a nucleotide sequence encoding a third gRNA molecule, or (e) a nucleotide sequence encoding a fourth gRNA molecule.
123. The nucleic acid composition according to any one of claims 34 to 122, wherein (a) and (b) are present on a single nucleic acid molecule.
124. The nucleic acid composition according to any one of claims 101 to 122, wherein (a), (b), and (g) are present on a single nucleic acid molecule.
125. The nucleic acid composition according to claim 123 or 124, wherein the nucleic acid molecule is an AAV vector or an LV vector.
126. The nucleic acid composition according to any one of claims 34 to 122, wherein (a) is present on a first nucleic acid molecule and (b) is present on a second nucleic acid molecule.
127. The nucleic acid composition according to claim 126, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
128. The nucleic acid composition according to any one of claims 74 to 122, wherein (a) and (c) are present on a single nucleic acid molecule.
129. The nucleic acid composition according to any one of claims 101 to 122, wherein (a) and (g) are present on a single nucleic acid molecule.
130. The nucleic acid composition according to claim 128 or 129, wherein the nucleic acid molecule is an AAV vector or an LV vector.
131. The nucleic acid composition according to any one of claims 74 to 122, wherein (a) is present on a first nucleic acid molecule; and (c) is present on a second nucleic acid molecule.
132. The nucleic acid composition according to any one of claims 101 to 122, wherein (a) is present on a first nucleic acid molecule; and (g) is present on a second nucleic acid molecule.
133. The nucleic acid composition according to claim 130 or 131, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
134. The nucleic acid composition according to any one of claims 60 to 122, wherein (a), (b), and (c) are present on a single nucleic acid molecule.
135. The nucleic acid composition according to any one of claims 101 to 122, wherein (a), (b), and (g) are present on a single nucleic acid molecule.
136. The nucleic acid composition according to claim 134 or 135, wherein the nucleic acid molecule is an AAV vector or an LV vector.
137. The nucleic acid composition according to any one of claims 60 to 122, wherein one of (a), (b), and (c) is encoded on a first nucleic acid molecule; and the second and third of (a), (b), and (c) are encoded on a second nucleic acid molecule.
138. The nucleic acid composition according to any one of claims 101 to 122, wherein one of (a), (b), and (g) is encoded on a first nucleic acid molecule; and the second and third of (a), (b), and (g) are encoded on a second nucleic acid molecule.
139. The nucleic acid composition according to claim 137 or 138, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
140. The nucleic acid composition according to claim 137 or 139, wherein (a) is present on a first nucleic acid molecule; and (b) and (c) are present on a second nucleic acid molecule.
141. The nucleic acid composition according to claim 138 or 139, wherein (a) is present on a first nucleic acid molecule; and (b) and (g) are present on a second nucleic acid molecule.
142. The nucleic acid composition according to claim 140 or 141, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
143. The nucleic acid composition according to claim 137 or 139, wherein (b) is present on a first nucleic acid molecule; and (a) and (c) are present on a second nucleic acid molecule.
144. The nucleic acid composition according to claim 138 or 139, wherein (b) is present on a first nucleic acid molecule; and (a) and (g) are present on a second nucleic acid molecule.
145. The nucleic acid composition according to claim 143 or 144, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
146. The nucleic acid composition according to claim 137 or 139, wherein (c) is present on the first nucleic acid molecule; and (b) and (a) are present on the second nucleic acid molecule.
147. The nucleic acid composition according to claim 138 or 139, wherein (g) is present on the first nucleic acid molecule; and (b) and (a) are present on the second nucleic acid molecule.
148. The nucleic acid composition according to claim 146 or 147, wherein the first and second nucleic acid molecules are AAV vectors or LV vectors.
149. The nucleic acid composition according to any one of claims 126, 131, 132, 137, 139, 140, 141, 143, 144, 146, or 147, wherein the first nucleic acid molecule is other than an AAV vector and the second nucleic acid molecule is an AAV vector.
150. A nucleic acid composition according to any one of claims 34 to 149, comprising a promoter operably coupled to (a).
151. A nucleic acid composition according to any one of claims 74 to 149, comprising a second promoter operably coupled to (c).
152. The nucleic acid composition according to any one of claims 151, wherein the promoter and the second promoter are different from each other.
153. The nucleic acid composition according to any one of claims 151, wherein the promoter and the second promoter are the same.
154. The nucleic acid composition according to any one of claims 60 to 149, comprising a promoter operably coupled to (b).
155. A composition comprising (a) a gRNA molecule according to any one of claims 1 to 33.
156. (b) The composition according to claim 155, further comprising an RNA-induced nuclease.
157. The composition according to claim 156, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
158. The composition according to claim 157, wherein the Cas9 molecule is an enzymatically active Cas9 (eaCas9) molecule.
159. The composition according to claim 158, wherein the eaCas9 molecule comprises a nickase molecule.
160. The composition according to claim 158 or 159, wherein the eaCas9 molecule causes a double-strand break in the target nucleic acid.
161. The composition according to claim 158 or 159, wherein the eaCas9 molecule causes a single-strand break in the target nucleic acid.
162. The composition according to claim 161, wherein the single-strand break occurs in a target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
163. The composition according to claim 161, wherein the single-strand break occurs in a target nucleic acid strand other than the target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
164. The composition according to any one of claims 158 to 163, wherein the eaCas9 molecule has HNH-like domain cleavage activity but does not have N-terminal RuvC-like domain cleavage activity or significant N-terminal RuvC-like domain cleavage activity.
165. The composition according to claim 164, wherein the eaCas9 molecule is an HNH-like domain niccasase.
166. The composition according to claim 164 or 165, wherein the eaCas9 molecule comprises a mutation at D10.
167. The composition according to any one of claims 158 to 163, wherein the eaCas9 molecule has N-terminal RuvC-like domain cleavage activity but does not have HNH-like domain cleavage activity or significant HNH-like domain cleavage activity.
168. The composition according to claim 167, wherein the eaCas9 molecule is an N-terminal RuvC-like domain niccasse.
169. The composition according to claim 167 or 168, wherein the eaCas9 molecule comprises a mutation at H840 or N863.
170. (c) The composition according to any one of claims 156 to 169, further comprising a second gRNA molecule having a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
171. The composition according to claim 170, wherein the second gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
172. (d) The composition according to any one of claims 156 to 171, further comprising a third gRNA molecule.
173. (e) The composition according to claim 172, further comprising a fourth gRNA molecule.
174. (g) The composition according to any one of claims 155 to 173, further comprising a template nucleic acid.
175. The composition according to claim 174, wherein the template nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
176. The composition according to claim 175, wherein the template nucleic acid comprises a 5' homology arm, a substitution sequence, and a 3' homology arm.
177. The composition according to claim 176, wherein the 5' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence has a length of 0 nucleotides; and the 3' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.
178. The composition according to claim 177, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of the target portion within the HBG target position, and the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 3' side of the target portion within the HBG target position.
179. The composition according to claim 178, wherein the target site is selected from the group consisting of HBG1 c. -114 to -102 (e.g., nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)); HBG1 c. -225 to -222 (e.g., nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)); and HBG2 c. -114 to -102 (e.g., nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
180. The composition according to claim 179, wherein the target site is HBG1 c.-114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG1 c.-114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
181. The composition according to claim 180, wherein the 5' homology arm is essentially composed of or consists of Sequence ID No. 904, which includes Sequence ID No. 904 (i.e., ssODN1 5' homology arm).
182. The composition according to claim 180 or 181, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
183. The composition according to any one of claims 179 to 182, wherein the 3' homology arm is essentially composed of or consists of SEQ ID NO: 905 (i.e., ssODN1 3' homology arm), which includes SEQ ID NO:
905.
184. The composition according to claim 179, wherein the target site is HBG2 c.-114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG2 c.-114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
185. The composition according to claim 184, wherein the 5' homology arm is essentially composed of or consists of Sequence ID No. 904, which includes Sequence ID No. 904 (i.e., ssODN1 5' homology arm).
186. The composition according to claim 184 or 185, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of sequence number 903 (HBG2)).
187. The composition according to any one of claims 184 to 186, wherein the 3' homology arm is essentially composed of or consists of SEQ ID NO: 905 (i.e., ssODN1 3' homology arm), which includes SEQ ID NO:
905.
188. The composition according to any one of claims 175 to 187, wherein the template nucleic acid comprises SEQ ID NO: 906 (i.e., ssODN1), essentially consisting of SEQ ID NO: 906, or comprising SEQ ID NO:
906.
189. The composition according to claim 179, wherein the target site is HBG1 c.-225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG1 c.-225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
190. The composition according to claim 189, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
191. The composition according to any one of claims 175 to 190, wherein the ssODN comprises a 5'-phosphorothioate modification.
192. The composition according to any one of claims 175 to 190, wherein the ssODN comprises a 3'-phosphorothioate modification.
193. The composition according to any one of claims 175 to 190, wherein the ssODN comprises a 5'-phosphorothioate modification and a 3'-phosphorothioate modification.
194. A method of altering cells, (a) the gRNA molecule according to any one of claims 1 to 33; and (b) A method comprising the step of bringing the cells into contact with an RNA-induced nuclease.
195. The method according to claim 194, further comprising the step of contacting the cells with (c) a second gRNA molecule comprising a targeting domain comprising a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
196. The method according to claim 194 or 195, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
197. The method according to claim 196, wherein the second gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
198. The method according to claim 195 or 196, wherein the Cas9 molecule is an enzymatically active Cas9 (eaCas9) molecule.
199. The method according to claim 198, wherein the eaCas9 molecule comprises a nickase molecule.
200. The method according to claim 198 or 199, wherein the eaCas9 molecule causes a double-strand break in the target nucleic acid.
201. The method according to claim 198 or 199, wherein the eaCas9 molecule causes a single-strand break in the target nucleic acid.
202. The method according to claim 201, wherein the single-strand break occurs in a target nucleic acid strand whose targeting domain is complementary to the gRNA molecule.
203. The method according to claim 201, wherein the single-strand break occurs in a target nucleic acid strand other than the target nucleic acid strand in which the targeting domain of the gRNA molecule is complementary.
204. The method according to any one of claims 198 to 203, wherein the eaCas9 molecule has HNH-like domain cleavage activity but does not have N-terminal RuvC-like domain cleavage activity or significant N-terminal RuvC-like domain cleavage activity.
205. The method according to claim 204, wherein the eaCas9 molecule is an HNH-like domain niccasse.
206. The method according to claim 204 or 205, wherein the eaCas9 molecule comprises a mutation at D10.
207. The method according to any one of claims 198 to 203, wherein the eaCas9 molecule has N-terminal RuvC-like domain cleavage activity but does not have HNH-like domain cleavage activity or significant HNH-like domain cleavage activity.
208. The method according to claim 207, wherein the eaCas9 molecule is an N-terminal RuvC-like domain niccasse.
209. The method according to claim 207 or 208, wherein the eaCas9 molecule comprises a mutation at H840 or N863.
210. The method according to any one of claims 196 to 209, further comprising the step of bringing the cells into contact with (d) a third gRNA molecule.
211. The method according to claim 210, further comprising the step of bringing the cells into contact with (e) a fourth gRNA molecule.
212. The method according to any one of claims 194 to 211, wherein the cells are derived from a subject suffering from β-abnormal hemoglobinopathy.
213. The method according to claim 212, wherein the β-abnormal hemoglobinopathy is selected from the group consisting of SCD and β-Thal.
214. The method according to any one of claims 194 to 213, wherein the cells are erythrocytes.
215. The method according to claim 214, wherein the cell is an erythroblast.
216. The method according to any one of claims 194 to 215, wherein the contact step is performed in vivo.
217. The method according to any one of claims 194 to 216, comprising the step of obtaining knowledge of the sequence of HBG target sites in the cells.
218. The method according to any one of claims 194 to 217, comprising the step of introducing an indel at an HBG target site.
219. The method according to claim 218, wherein the indel is selected from the group consisting of HBG1 13bp del -114 to -102; HBG1 4bp del -225 to -222; and HBG2 13bp del -114 to -102.
220. The method according to claim 218 or 219, wherein the indel is introduced using an NHEJ.
221. The method according to any one of claims 194 to 220, comprising the step of introducing a single nucleotide alteration to an HBG target site.
222. The single nucleotide change is as follows: HBG1 c. -114 C>T; c. -117 G>A; c. -158 C>T; c. -167 C>T; c. -170 G>A; c. -175 T>G; c. -175 T>C; c. -195 C>G; c. -196 C>T; c. -198 T>C; c. -201 C>T; c. -251 T>C; or c. -499 T>A; and HBG2 c. -109 G>T; c. -114 C>A; c. -114 C>T; c. -157 C>T; c. -158 C>T; c. The method according to claim 221, selected from the group consisting of -167 C>T;c. -167 C>A;c. -175 T>C;c. -202 C>G;c. -211 C>T;c. -228 T>C;c. -255 C>G;c. -309 A>G;c. -369 C>G; or c. -567 T>G.
223. The method according to claim 221 or 222, wherein the single nucleotide change is introduced using HDR.
224. The method according to any one of claims 194 to 223, comprising the step of introducing a change to a target site at an HBG target location.
225. The method according to claim 224, wherein the change is selected from the group consisting of HBG1 13bp del -114 to -102; HBG1 4bp del -225 to -222; and HBG2 13bp del -114 to -102.
226. The method according to claim 224 or 225, wherein the change is introduced using HDR.
227. The method according to claim 226, further comprising the step of contacting the cells with (g) a template nucleic acid.
228. The method according to claim 227, wherein the template nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
229. The method according to claim 228, wherein the template nucleic acid comprises a 5' homology arm, a substitution sequence, and a 3' homology arm.
230. The method according to claim 229, wherein the 5' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence has a length of 0 nucleotides; and the 3' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.
231. The method according to claim 230, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of the target portion, and the 3' homology arm has homology of approximately 50 to 100 bp, for example 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 3' side of the target portion.
232. The method according to claim 231, wherein the change is HBG1 13bp del -114 to -102, and the target site is HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1).
233. The method according to claim 232, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of sequence number 902 (HBG1)).
234. The method according to any one of claims 230 to 233, wherein the 3' homology arm has homology of about 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
235. The method according to claim 234, wherein the change is HBG2 13bp del -114 to -102, and the target site is HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2).
236. The method according to claim 235, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of sequence number 903 (HBG2)).
237. The method according to claim 234 or 235, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
238. The method according to any one of claims 232 to 237, wherein the 5' homology arm is essentially composed of or consists of sequence number 904, which includes sequence number 904 (ssODN1 5' homology arm).
239. The method according to any one of claims 232 to 238, wherein the 3' homology arm is essentially composed of or consists of Sequence ID No. 905, which includes Sequence ID No. 905 (ssODN1 3' homology arm).
240. The method according to any one of claims 232 to 239, wherein the template nucleic acid includes, essentially consists of, or comprises SEQ ID NO: 906 (ssODN1).
241. The method according to claim 231, wherein the change is HBG1 4bp del -225 to -222, and the target site is HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1).
242. The method according to claim 241, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
243. The method according to claim 241 or 242, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
244. The method according to any one of claims 228 to 243, wherein the ssODN includes a 5' phosphorothioate modification.
245. The method according to any one of claims 228 to 243, wherein the ssODN includes a 3'-phosphorothioate modification.
246. The method according to any one of claims 228 to 243, wherein the ssODN comprises a 5'-phosphorothioate modification and a 3'-phosphorothioate modification.
247. The method according to any one of claims 195 to 246, wherein the contacting step includes the step of contacting the cells with a nucleic acid composition encoding at least one of (a), (b), and (c).
248. The method according to any one of claims 227 to 246, wherein the contact step includes the step of contacting the cells with a nucleic acid composition encoding (a), (b), (g), and optionally (c).
249. The method according to claim 248 or 249, wherein the contact step includes a step of contacting the cells with the nucleic acid composition according to any one of claims 34 to 154.
250. The method according to any one of claims 195 to 249, wherein the contact step includes delivering the cells to a nucleic acid composition encoding (b) and (a).
251. The method according to claim 250, wherein the nucleic acid composition further encodes (c).
252. The method according to claim 250 or 251, wherein the nucleic acid composition further encodes (g).
253. The method according to any one of claims 195 to 251, wherein the contact step includes a step of delivering to the cells (a) and (b).
254. The method according to any one of claims 195 to 251, wherein the contact step includes a step of delivering nucleic acid compositions encoding the cells (a) and (b).
255. The method according to claim 253 or 254, wherein the contact step further comprises the step of delivering to the cell (c).
256. The method according to any one of claims 227 to 255, wherein the contact step further comprises the step of delivering to the cell (g).
257. A method for treating β-abnormal hemoglobinopathy in those who require it: The subject or the cells of the subject: (a) the gRNA molecule according to any one of claims 1 to 33; and (b) A method comprising the step of contacting with an RNA-induced nuclease.
258. The method according to claim 257, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
259. The method according to claim 257 or 258, wherein the β-abnormal hemoglobinopathy is selected from the group consisting of SCD and β-Thal.
260. The method according to any one of claims 257 to 259, further comprising the step of contacting the subject or the cells of the subject with (c) a second gRNA molecule comprising a targeting domain comprising a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
261. The method according to claim 260, wherein the second gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
262. The method according to any one of claims 258 to 261, wherein the Cas9 molecule is an enzymatically active Cas9 (eaCas9) molecule.
263. The method according to claim 262, wherein the eaCas9 molecule comprises a nickase molecule.
264. The method according to claim 262 or 263, wherein the eaCas9 molecule causes a double-strand break in the target nucleic acid.
265. The method according to claim 262 or 263, wherein the eaCas9 molecule causes a single-strand break in the target nucleic acid.
266. The method according to claim 265, wherein the single-strand break occurs in a target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
267. The method according to claim 265, wherein the single-strand break occurs in a target nucleic acid strand other than the target nucleic acid strand to which the targeting domain of the gRNA molecule is complementary.
268. The method according to any one of claims 263 to 267, wherein the eaCas9 molecule has HNH-like domain cleavage activity but does not have N-terminal RuvC-like domain cleavage activity or significant N-terminal RuvC-like domain cleavage activity.
269. The method according to claim 268, wherein the eaCas9 molecule is an HNH-like domain niccasse.
270. The method according to claim 268 or 269, wherein the eaCas9 molecule comprises a mutation at D10.
271. The method according to any one of claims 262 to 270, wherein the eaCas9 molecule has N-terminal RuvC-like domain cleavage activity but does not have HNH-like domain cleavage activity or significant HNH-like domain cleavage activity.
272. The method according to claim 271, wherein the eaCas9 molecule is an N-terminal RuvC-like domain niccasse.
273. The method according to any one of claims 271 or 272, wherein the eaCas9 molecule comprises a mutation at H840 or N863.
274. The method according to any one of claims 257 to 273, further comprising the step of bringing the subject or the cells of the subject into contact with (d) a third gRNA molecule.
275. The method according to claim 274, further comprising the step of bringing the subject or the cells of the subject into contact with a fourth gRNA molecule.
276. The method according to any one of claims 257 to 275, comprising the step of introducing a single nucleotide alteration at an HBG target site.
277. The single nucleotide change is as follows: HBG1 c. -114 C>T; c. -117 G>A; c. -158 C>T; c. -167 C>T; c. -170 G>A; c. -175 T>G; c. -175 T>C; c. -195 C>G; c. -196 C>T; c. -198 T>C; c. -201 C>T; c. -251 T>C; or c. -499 T>A; and HBG2 c. -109 G>T; c. -114 C>A; c. -114 C>T; c. -157 C>T; c. -158 C>T; c. The method according to claim 276, selected from the group consisting of -167 C>T;c. -167 C>A;c. -175 T>C;c. -202 C>G;c. -211 C>T;c. -228 T>C;c. -255 C>G;c. -309 A>G;c. -369 C>G; or c. -567 T>G.
278. The method according to claim 276 or 277, wherein the single nucleotide change is introduced using HDR.
279. The method according to any one of claims 257 to 278, comprising the step of introducing an indel into a target site within an HBG target location.
280. The method according to claim 279, wherein the indel is selected from the group consisting of HBG1 13bp del -114 to -102; HBG1 4bp del -225 to -222; and HBG2 13bp del -114 to -102.
281. The method according to claim 279 or 280, wherein the change is introduced using HDR.
282. The method according to claim 281, further comprising the step of contacting the subject or the cells of the subject with (g) a template nucleic acid.
283. The method according to claim 282, wherein the template nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
284. The method according to claim 283, wherein the template nucleic acid comprises a 5' homology arm, a substitution sequence, and a 3' homology arm, the substitution sequence being 0 nucleotides.
285. The method according to claim 284, wherein the 5' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence has a length of 0 nucleotides; and the 3' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.
286. The method according to claim 285, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of the target portion within the HBG target position, and the 3' homology arm has homology of approximately 50 to 100 bp, for example 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 3' side of the target portion within the HBG target position.
287. The method according to claim 286, wherein the indel is HBG1 13bp del -114 to -102, and the target site is HBG1 -114 to -102 of sequence number 902.
288. The method according to claim 287, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of sequence number 902 (HBG1)).
289. The method according to claim 287 or 288, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG1 c. 114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
290. The method according to claim 286, wherein the indel is HBG2 13bp del -114 to -102, and the target site is HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2).
291. The method according to claim 290, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of sequence number 903 (HBG2)).
292. The method according to claim 290 or 291, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
293. The method according to any one of claims 290 to 292, wherein the 5' homology arm is essentially composed of or consists of sequence number 904, which includes sequence number 904 (ssODN1 5' homology arm).
294. The method according to any one of claims 290 to 293, wherein the 3' homology arm is essentially composed of or consists of Sequence ID No. 905, which includes Sequence ID No. 905 (ssODN1 3' homology arm).
295. The method according to any one of claims 283 to 294, wherein the template nucleic acid includes, essentially consists of, or comprises SEQ ID NO: 906 (ssODN1).
296. The method according to claim 256, wherein the indel is HBG1 4bp del -225 to -222, and the target site is HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1).
297. The method according to claim 296, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 5' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
298. The method according to claim 296 or 297, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
299. The method according to any one of claims 283 to 298, wherein the ssODN includes a 5' phosphorothioate modification.
300. The method according to any one of claims 283 to 298, wherein the ssODN includes a 3'-phosphorothioate modification.
301. The method according to any one of claims 258 to 298, wherein the ssODN comprises a 5'-phosphorothioate modification and a 3'-phosphorothioate modification.
302. The method according to any one of claims 257 to 301, wherein the contact step is performed in vivo.
303. The method according to any one of claims 257 to 302, wherein the contact step includes intravenous injection.
304. The method according to any one of claims 260 to 303, wherein the contacting step includes bringing the object or the cells of the object into contact with a nucleic acid composition encoding at least one of (a), (b), and (c).
305. The method according to any one of claims 282 to 303, wherein the contacting step includes bringing the object or the cells of the object into contact with a nucleic acid composition encoding at least one of (a), (b), (c), and (g).
306. The method according to any one of claims 257 to 305, wherein the contact step includes a step of contacting the target or the cells of the target with the nucleic acid composition according to any one of claims 34 to 154.
307. The method according to any one of claims 257 to 305, wherein the contact step includes the step of delivering a nucleic acid composition encoding (b) and (a) to the object or the cells of the object.
308. The method according to claim 307, wherein the nucleic acid composition further encodes (c).
309. The method according to claim 307 or 308, wherein the nucleic acid composition further codes for (g).
310. The method according to any one of claims 257 to 305, wherein the contact step includes the step of delivering (a) and (b) to the object or the cells of the object.
311. The method according to any one of claims 257 to 305, wherein the contact step includes the step of delivering a nucleic acid composition encoding (a) and (b) to the object or the cells of the object.
312. The method according to claim 310 or 311, wherein the contact step further includes the step of delivering (c) to the object or the cells of the object.
313. The method according to any one of claims 282 to 312, wherein the contact step further comprises the step of delivering (g) to the object or the cells of the object.
314. (a) a gRNA molecule according to any one of claims 1 to 33, a nucleic acid composition according to any one of claims 34 to 154, or a composition according to any one of claims 155 to 193; and A reaction mixture containing target cells suffering from β-abnormal hemoglobinopathy.
315. (a) a gRNA molecule according to any one of claims 1 to 33, or a nucleic acid composition encoding the gRNA molecule, and: (b) RNA-induced nucleases; (c) A second gRNA molecule comprising a targeting domain containing a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region; and A kit comprising one or more nucleic acid compositions encoding one or more of (d), (b), and (c).
316. The kit according to claim 315, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
317. The kit according to claim 315 or 316, wherein the second gRNA molecule is the gRNA molecule described in any one of claims 1 to 33.
318. A kit according to any one of claims 315 to 317, comprising a nucleic acid composition encoding one or more of (a), (b), and (c).
319. The kit according to any one of claims 315 to 318, further comprising a third gRNA molecule having a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
320. The kit according to claim 319, further comprising a fourth gRNA molecule comprising a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
321. (g) A kit according to any one of claims 315 to 320, further comprising a template nucleic acid.
322. A gRNA molecule according to any one of claims 1 to 33, for use in the treatment of β-abnormal hemoglobinopathy in a person requiring it.
323. The gRNA molecule according to claim 291, wherein the gRNA molecule is used in combination with (b) an RNA-induced nuclease.
324. The gRNA molecule according to claim 323, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
325. The gRNA molecule according to any one of claims 322 to 324, wherein the gRNA molecule is used in combination with (c) a second gRNA molecule comprising a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
326. The gRNA molecule according to any one of claims 322 to 325, wherein the gRNA molecule is used in combination with (g) template nucleic acid.
327. Use of a gRNA molecule according to any one of claims 1 to 33 in the manufacture of a drug for treating β-abnormal hemoglobinopathy in a person requiring it.
328. The use according to claim 327, wherein the drug further comprises (b) an RNA-inducing nuclease.
329. The use according to claim 328, wherein the RNA-induced nuclease is a Cas9 molecule.
330. The use according to any one of claims 327 to 329, wherein the agent further comprises (c) a second gRNA molecule comprising a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within the HBG1 or HBG2 regulatory region.
331. The use according to any one of claims 327 to 330, wherein the drug further comprises (g) template nucleic acid.
332. (a) the gRNA molecule according to any one of claims 1 to 33; and (b) A genome editing system including an RNA-induced nuclease.
333. The genome editing system according to claim 332, wherein the RNA-induced nuclease is a Cas9 molecule or a Cas9 fusion protein.
334. The genome editing system according to claim 333, wherein the Cas9 molecule is an enzymatically active Cas9 (eaCas9) molecule.
335. The genome editing system according to claim 334, wherein the eaCas9 molecule comprises a nickase molecule.
336. The genome editing system according to claim 334 or 335, wherein the eaCas9 molecule causes a double-strand break in the target nucleic acid.
337. The genome editing system according to claim 334 or 335, wherein the eaCas9 molecule causes a single-strand break in the target nucleic acid.
338. The genome editing system according to claim 337, wherein the single-strand break occurs in a target nucleic acid strand whose targeting domain is complementary to the gRNA molecule.
339. The genome editing system according to claim 337, wherein the single-strand break occurs in a target nucleic acid strand other than the target nucleic acid strand in which the targeting domain of the gRNA molecule is complementary.
340. The genome editing system according to any one of claims 334 to 339, wherein the eaCas9 molecule has HNH-like domain cleavage activity but does not have N-terminal RuvC-like domain cleavage activity or significant N-terminal RuvC-like domain cleavage activity.
341. The genome editing system according to claim 340, wherein the eaCas9 molecule is an HNH-like domain niccasse.
342. The genome editing system according to claim 340 or 341, wherein the eaCas9 molecule comprises a mutation at D10.
343. The genome editing system according to any one of claims 334 to 342, wherein the eaCas9 molecule has N-terminal RuvC-like domain cleavage activity but does not have HNH-like domain cleavage activity or significant HNH-like domain cleavage activity.
344. The genome editing system according to claim 343, wherein the eaCas9 molecule is an N-terminal RuvC-like domain niccasse.
345. The genome editing system according to claim 343 or 344, wherein the eaCas9 molecule comprises a mutation at H840 or N863.
346. (c) A genome editing system according to any one of claims 332 to 345, further comprising a second gRNA molecule having a targeting domain having a nucleotide sequence complementary or partially complementary to a target domain located entirely or partially within an HBG1 or HBG2 regulatory region.
347. The genome editing system according to claim 34, wherein the second gRNA molecule is a gRNA molecule according to any one of claims 1 to 33.
348. (d) A genome editing system according to any one of claims 332 to 347, further comprising a third gRNA molecule.
349. (e) The genome editing system according to claim 348, further comprising a fourth gRNA molecule.
350. (g) A genome editing system according to any one of claims 332 to 349, further comprising a template nucleic acid.
351. The genome editing system according to claim 350, wherein the template nucleic acid is a single-stranded oligodeoxynucleotide (ssODN).
352. The genome editing system according to claim 351, wherein the template nucleic acid comprises a 5' homology arm, a substitution sequence, and a 3' homology arm.
353. The genome editing system according to claim 352, wherein the 5' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides; the substitution sequence has a length of 0 nucleotides; and the 3' homology arm has a length of about 200 nucleotides, for example, at least 25, 50, 75, 100, 125, 150, 175, or 200 nucleotides.
354. The genome editing system according to claim 353, wherein the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of the target site within the HBG target position, and the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 3' side of the target site within the HBG target position.
355. The genome editing system according to claim 354, wherein the target site is selected from the group consisting of HBG1 c. -114 to -102 (e.g., nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)); HBG1 c. -225 to -222 (e.g., nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)); and HBG2 c. -114 to -102 (e.g., nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
356. The genome editing system according to claim 355, wherein the target site is HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
357. The genome editing system according to claim 356, wherein the 5' homology arm essentially consists of or comprises SEQ ID NO: 904 (i.e., ssODN1 5' homology arm).
358. The genome editing system according to claim 180 or 181, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG1 c. -114 to -102 (for example, nucleotides 2824 to 2836 of SEQ ID NO: 902 (HBG1)).
359. The genome editing system according to any one of claims 355 to 358, wherein the 3' homology arm is essentially composed of or consists of SEQ ID NO: 905 (i.e., ssODN1 3' homology arm), wherein the 3' homology arm is SEQ ID NO:
905.
360. The genome editing system according to claim 355, wherein the target site is HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)), and the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp on the 5' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
361. The genome editing system according to claim 360, wherein the 5' homology arm essentially consists of or comprises SEQ ID NO: 904 (i.e., ssODN1 5' homology arm).
362. The genome editing system according to claim 360 or 361, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp, to the 3' side of HBG2 c. -114 to -102 (for example, nucleotides 2748 to 2760 of SEQ ID NO: 903 (HBG2)).
363. The genome editing system according to any one of claims 356 to 362, wherein the 3' homology arm essentially consists of or comprises SEQ ID NO: 905 (i.e., ssODN1 3' homology arm).
364. The genome editing system according to any one of claims 356 to 363, wherein the template nucleic acid comprises SEQ ID NO: 906 (i.e., ssODN1), essentially consisting of SEQ ID NO: 906, or comprising SEQ ID NO:
906.
365. The genome editing system according to claim 355, wherein the target site is HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)), and the 5' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 5' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
366. The genome editing system according to claim 365, wherein the 3' homology arm has homology of approximately 50 to 100 bp, for example, 55 to 95 bp, 60 to 90 bp, 70 to 90 bp, or 80 to 90 bp to the 3' side of HBG1 c. -225 to -222 (for example, nucleotides 2716 to 2719 of SEQ ID NO: 902 (HBG1)).
367. The genome editing system according to any one of claims 351 to 366, wherein the ssODN includes a 5' phosphorothioate modification.
368. The genome editing system according to any one of claims 351 to 366, wherein the ssODN includes a 3' phosphorothioate modification.
369. The genome editing system according to any one of claims 351 to 366, wherein the ssODN comprises a 5' phosphorothioate modification and a 3' phosphorothioate modification.