RNA-guided human genome engineering
Patent Information
- Application Number
- JP2025158193
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-03-13
- Filing Date
- 2025-09-24
- Publication Date
- 2026-02-10
AI Technical Summary
Existing methods for genome editing in eukaryotic cells, such as ZFNs and TALENs, are inefficient and can cause toxicity, while bacterial CRISPR systems require complex RNA processing machinery.
A two-component system comprising RNA complementary to genomic DNA and an enzyme, like Cas9, is expressed in eukaryotic cells to specifically target and modify genomic DNA, avoiding bacterial RNA processing and reducing toxicity.
This approach enables efficient, site-specific genome editing with reduced toxicity, allowing for precise modification of multiple genomic sites and integration of foreign DNA, demonstrated by high HR efficiency and low NHEJ rates.
Smart Images

Figure 00000021_0000 
Figure 00000021_0001 
Figure 00000022_0000
Abstract
Description
[Technical Field]
[0001] Related application data This application claims priority to U.S. Provisional Patent Application No. 61 / 779,169, filed March 13, 2013, and U.S. Provisional Patent Application No. 61 / 738,355, filed December 17, 2012, each of which is incorporated by reference in its entirety for all purposes.
[0002] Statement of Government Interest This invention was made with government support under P50HG005550 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]
[0003] Bacterial and archaeal CRISPR systems rely on crRNA complexed with Cas proteins to degrade complementary sequences present in invading viral and plasmid DNA (References 1-3). Recently, in vitro reconstitution of the S. pyogenes type II CRISPR system demonstrated that crRNA fused to the normally trans-encoded tracrRNA is sufficient to direct the Cas9 protein to sequence-specifically cleave target DNA sequences that match the crRNA (Reference 4). Summary of the Invention [Means for solving the problem]
[0004] This disclosure refers to documents by number, which are listed at the end of this disclosure, and the documents corresponding to the numbers are incorporated herein by reference as supplemental documents to the corresponding numbers, as if they were cited in full.
[0005] According to one embodiment of the present disclosure, a two-component system containing RNA complementary to genomic DNA and an enzyme that interacts with the RNA is transfected into a eukaryotic cell. The RNA and enzyme are expressed by the cell. The RNA of the RNA / enzyme complex then binds to the complementary genomic DNA. The enzyme then performs a function, such as cleaving the genomic DNA. The RNA comprises about 10 to about 250 nucleotides. The RNA comprises about 20 to about 100 nucleotides. According to one embodiment, the enzyme can perform any desired function in a site-specific manner, and the enzyme is modified for this purpose. According to one embodiment, the eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. According to one embodiment, the enzyme cleaves the genomic sequence targeted by the RNA sequence (see References (4-6)), thereby creating a genomically modified eukaryotic cell.
[0006] According to one embodiment, the present disclosure provides a method for genetically modifying human cells by including in the genome of the cell a nucleic acid that encodes RNA that is complementary to genomic DNA, and including in the genome of the cell a nucleic acid that encodes an enzyme that performs a desired function on genomic DNA. According to one embodiment, the RNA and the enzyme are expressed. According to one embodiment, the RNA is hybridized with complementary genomic DNA. According to one embodiment, when the RNA is hybridized with complementary genomic DNA, the enzyme is activated, and performs a desired function, such as cutting, in a site-specific manner. According to one embodiment, the RNA and the enzyme are components of bacterial type II CRISPR system.
[0007] According to one embodiment, a method for modifying a eukaryotic cell is provided, comprising transfecting a eukaryotic cell with a nucleic acid encoding an RNA complementary to the genomic DNA of the eukaryotic cell and a nucleic acid encoding an enzyme that interacts with the RNA and site-specifically cleaves the genomic DNA, wherein the cell expresses the RNA and the enzyme, the RNA binds to the complementary genomic DNA, and the enzyme site-specifically cleaves the genomic DNA. According to one embodiment, the enzyme is Cas9, modified Cas9, or a Cas9 homolog. According to one embodiment, the eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. According to one embodiment, the RNA comprises from about 10 nucleotides to about 250 nucleotides. According to one embodiment, the RNA comprises from about 20 nucleotides to about 100 nucleotides.
[0008] According to one embodiment, a method for modifying a human cell is provided, comprising transfecting a human cell with a nucleic acid encoding an RNA complementary to the genomic DNA of a eukaryotic cell and a nucleic acid encoding an enzyme that interacts with the RNA and site-specifically cleaves the genomic DNA, wherein the human cell expresses the RNA and the enzyme, the RNA binds to the complementary genomic DNA, and the enzyme site-specifically cleaves the genomic DNA. According to one embodiment, the enzyme is Cas9, modified Cas9, or a Cas9 homolog. According to one embodiment, the RNA comprises from about 10 nucleotides to about 250 nucleotides. According to another embodiment, the RNA comprises from about 20 nucleotides to about 100 nucleotides.
[0009] According to one embodiment, a method for modifying a eukaryotic cell at multiple genomic DNA sites is provided, comprising transfecting a eukaryotic cell with multiple nucleic acids encoding RNAs complementary to different sites on the genomic DNA of the eukaryotic cell, and transfecting the eukaryotic cell with a nucleic acid encoding an enzyme that interacts with the RNAs and cleaves the genomic DNA in a site-specific manner, wherein the cell expresses the RNAs and the enzymes, the RNAs bind to the complementary genomic DNA, and the enzymes cleave the genomic DNA in a site-specific manner. According to one embodiment, the enzyme is Cas9. According to one embodiment, the eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. According to one embodiment, the RNA comprises from about 10 nucleotides to about 250 nucleotides. According to one embodiment, the RNA comprises from about 20 nucleotides to about 100 nucleotides. [Brief explanation of the drawings]
[0010] [Figure 1A] FIG. 1 shows genome editing in human cells using a modified Type II CRISPR system. [Figure 1B] FIG. 1 shows genome editing in human cells using a modified Type II CRISPR system. [Figure 1C-1] FIG. 1C shows genome editing in human cells using a modified type II CRISPR system. [Figure 1C-2] FIG. 1C shows genome editing in human cells using a modified type II CRISPR system. [Figure 2A] FIG. 1 shows RNA-induced genome editing of the native AAVS1 locus in multiple cell types. [Figure 2B-1] FIG. 2B shows RNA-guided genome editing of the native AAVS1 locus in multiple cell types. [Figure 2B-2] FIG. 2B shows RNA-guided genome editing of the native AAVS1 locus in multiple cell types. [Figure 2B-3]FIG. 2B shows RNA-guided genome editing of the native AAVS1 locus in multiple cell types. [Figure 2C] FIG. 1 shows RNA-induced genome editing of the native AAVS1 locus in multiple cell types. [Figure 2D] FIG. 1 shows RNA-induced genome editing of the native AAVS1 locus in multiple cell types. [Figure 2E] FIG. 1 shows RNA-induced genome editing of the native AAVS1 locus in multiple cell types. [Figure 2F] FIG. 1 shows RNA-induced genome editing of the native AAVS1 locus in multiple cell types. [Figure 3A-1] Figure 3A shows the expression format and complete sequence of the cas9 gene insert. [Figure 3A-2] Figure 3A shows the expression format and complete sequence of the cas9 gene insert. [Figure 3B] FIG. 1 shows a U6 promoter-based expression scheme for guide RNAs and the predicted secondary structures of the RNA transcripts. [Figure 3C] FIG. 1 shows the seven gRNAs used. [Figure 4] FIG. 1 shows testing of all possible combinations of repair DNA donor, Cas9 protein, and gRNA for their ability to successfully perform HR in 93T. [Figure 5A] FIG. 1 shows analysis of gRNA- and Gas9-mediated genome editing. [Figure 5B] FIG. 1 shows analysis of gRNA- and Gas9-mediated genome editing. [Figure 6A] FIG. 1 shows three 293T stable expressing cell lines, each carrying a distinct GFP reporter construct. [Figure 6B] FIG. 1 shows three 293T stable expressing cell lines, each carrying a distinct GFP reporter construct. [Figure 7]FIG. 1B shows two new gRNAs targeting the adjacent GFP sequence of the reporter described in FIG. 1B in 293T. [Figure 8A] FIG. 1 shows two 293T stable expressing cell lines, each carrying a distinct GFP reporter construct. [Figure 8B] FIG. 1 shows two 293T stable expressing cell lines, each carrying a distinct GFP reporter construct. [Figure 9A] FIG. 1 shows RNA-induced NHEJ in human iPS cells. [Figure 9B] FIG. 1 shows RNA-induced NHEJ in human iPS cells. [Figure 9C] FIG. 1 shows RNA-induced NHEJ in human iPS cells. [Figure 10A] FIG. 1 shows RNA-induced NHEJ in K562 cells. [Figure 10B] FIG. 1 shows RNA-induced NHEJ in K562 cells. [Figure 11A] FIG. 1 shows RNA-induced NHEJ in 293T cells. [Figure 11B] FIG. 1 shows RNA-induced NHEJ in 293T cells. [Figure 12A] FIG. 1 shows HR at the endogenous AAVS1 locus using either dsDNA or short oligonucleotide donors. [Figure 12B] FIG. 1 shows HR at the endogenous AAVS1 locus using either dsDNA or short oligonucleotide donors. [Figure 12C] FIG. 1 shows HR at the endogenous AAVS1 locus using either dsDNA or short oligonucleotide donors. [Figure 13A-1] FIG. 13A shows a method for multiplex synthesis of guide RNAs targeting genes in the human genome, recovery, and cloning into U6 expression vectors. [Figure 13A-2]FIG. 13A shows a method for multiplex synthesis of guide RNAs targeting genes in the human genome, recovery, and cloning into U6 expression vectors. [Figure 13B] FIG. 1 shows a method for multiplex synthesis of guide RNAs targeting genes in the human genome, recovery, and cloning into U6 expression vectors. [Figure 14A] FIG. 1 shows CRISPR-mediated RNA-induced transcriptional activation. [Figure 14B] FIG. 1 shows CRISPR-mediated RNA-induced transcriptional activation. [Figure 14C] FIG. 1 shows CRISPR-mediated RNA-induced transcriptional activation. [Figure 14D] FIG. 1 shows CRISPR-mediated RNA-induced transcriptional activation. [Figure 15A] FIG. 1 shows the flexibility of gRNA sequences. [Figure 15B] FIG. 1 shows the flexibility of gRNA sequences. DETAILED DESCRIPTION OF THE INVENTION
[0011] In one embodiment, a human codon-optimized Cas9 protein with a C-terminal SV40 nuclear localization signal is synthesized and cloned into a mammalian expression system (Figure 1A and Figure 3A). Thus, Figure 1 relates to genome editing in human cells using a modified type II CRISPR system. As shown in Figure 1A, RNA-guided gene targeting in human cells involves coexpressing a Cas9 protein with a C-terminal SV40 nuclear localization signal with one or more guide RNAs (gRNAs) expressed from a human U6 polymerase III promoter. Once the gRNA recognizes the target sequence, Cas9 unwinds the DNA duplex and cleaves both strands only if the correct protospacer adjacent motif (PAM) is present at the 3' end. In principle, GN 20Any genomic sequence in the form of GG can be targeted. As shown in Figure 1B, the GFP coding sequence integrated into the genome was disrupted by the insertion of a stop codon and a 68-bp genomic fragment from the AAVS1 locus. Restoration of the GFP sequence by homologous recombination (HR) with an appropriate donor sequence resulted in GFP expression, which could be quantified by FACS. + T1gRNA and T2gRNA target sequences within the AAVS1 fragment. The binding sites for each TAL effector nuclease heterodimer (TALEN) are underlined. As shown in Figure 1C, the bar graph shows the HR efficiency induced by T1-, T2-, or TALEN-mediated nuclease activity at the target site, as measured by FACS. Representative FACS plots and microscopic images of target cells are shown below (scale bar is 100 microns). Data are mean ± SEM (N=3).
[0012] In one embodiment, a crRNA-tracrRNA fusion transcript (hereafter referred to as guide RNA (gRNA)) is expressed from a human U6 polymerase III promoter to direct Cas9 to cleave the target sequence. In one embodiment, the gRNA is directly transcribed by the cell. This embodiment preferably avoids reconstitution of the RNA processing machinery used by bacterial CRISPR systems (Figures 1A and 3B) (see References (4, 7-9)). In one embodiment, a method is provided for modifying genomic DNA using U6 transcription starting with G and the PAM (protospacer adjacent motif) sequence NGG following the 20 bp crRNA target. In this embodiment, the target genomic site is GN 20 GG form (see Figure 3C).
[0013] In one embodiment, to test the functionality of the genome modification methods described herein, a GFP reporter assay (Figure 1B) in 293T cells was developed, similar to a previously reported assay (see Reference (10)). In one embodiment, a stable cell line was established with the GFP coding sequence integrated into its genome disrupted by the insertion of a 68-bp genomic fragment from the AAVS1 locus and a stop codon that renders the expressed protein fragment non-fluorescent. Homologous recombination (HR) using an appropriate repair donor restored the normal GFP sequence, thereby allowing the resulting GFP to be expressed. + Cells can be quantified by flow activated cell sorting (FACS).
[0014] In one embodiment, a method for homologous recombination (HR) is provided. Two gRNAs, T1 and T2, targeting the intervening AAVS1 fragments are constructed (Fig. 1b). Their activity was compared with that of a previously reported TAL effector nuclease heterodimer (TALEN) targeting the same region (see reference (11)). Successful HR events were observed with all three targeting reagents, with gene correction rates approaching 3% and 8% using T1gRNA and T2gRNA, respectively (Fig. 1C). This RNA-mediated editing process was remarkably fast, resulting in the first detectable GFP signal. +Cells emerged approximately 20 hours after transfection, compared with approximately 40 hours with the AAVS1 TALEN. HR was observed only when the repair donor, Cas9 protein, and gRNA were simultaneously introduced, confirming that all components were required for genome editing (Figure 4). No obvious toxicity was associated with Cas9 / crRNA expression, although studies with ZFNs and TALENs have shown that nicking only one strand further reduces toxicity. We therefore tested the Cas9D10A mutant, which is known to function as a nickase in vitro, and found similar HR but a lower rate of non-homologous end joining (NHEJ) (Figure 5) (see references 4, 5). Consistent with the findings that related Cas9 proteins cleave both strands 6 bp upstream of the PAM (Reference 4), the NHEJ data confirmed that most deletions or insertions occurred at the 3′ end of the target sequence (Figure 5B). Furthermore, mutations in the target genomic site were confirmed to prevent gRNA-mediated HR at this site, demonstrating that CRISPR-mediated genome editing is sequence-specific (Figure 6). Two gRNA target sites in the GFP gene and three additional gRNA target fragments derived from homologous regions of the DNA methyltransferase 3a (DNMT3a) and DNMT3b genes were shown to be capable of inducing significant HR in a sequence-specific manner in the engineered reporter cell line (Figures 7 and 8). Overall, these results support the robust HR induction at multiple target sites by RNA-guided genome targeting in human cells.
[0015] In one embodiment, the native locus was modified. Using gRNA, the AAVS1 locus (Figure 2A), located in the PPP1R12C gene on chromosome 19, which is widely expressed in most tissues, was targeted in 293T cells, K562 cells, and PGP1 human iPS cells (see Reference (12)). The results were analyzed by next-generation sequencing of the targeted locus. Thus, Figure 2 illustrates RNA-guided genome editing of the native AAVS1 locus in multiple cell types. As shown in Figure 2A, T1gRNA (red) and T2gRNA (green) target sequences in the intron of the PPP1R12C gene in the AAVS1 locus on chromosome 19. As shown in Figure 2B, the total count and location of NHEJ-induced deletions in 293T cells, K562 cells, and PGP1 iPS cells after expression of Cas9 and T1gRNA or T2gRNA, as quantified by next-generation sequencing, are shown. The red and green dashed lines indicate the boundaries of the T1gRNA and T2gRNA target sites. The NHEJ frequencies of T1gRNA and T2gRNA were 10% and 25% in 293T cells, 13% and 38% in K562 cells, and 2% and 4% in PGP1 iPS cells, respectively. Figure 2C shows the structure of the DNA donor for HR at the AAVS1 locus, as well as the location of sequencing primers (arrows) for detecting successful targeting. As shown in Figure 2D, PCR assays 3 days after transfection demonstrated that only cells expressing the donor, Cas9, and T2gRNA exhibited successful HR. As shown in Figure 2E, successful HR was confirmed by Sanger sequencing of the PCR amplification products, demonstrating the presence of the expected DNA bases at both the genome-donor and donor-insert boundaries. As shown in Figure 2F, clones of successfully targeted 293T cells were selected with puromycin for 2 weeks. Microscopic images of two representative GFP+ clones are shown (scale bar is 100 microns).
[0016] Consistent with the GFP reporter assay results, numerous NHEJ events were observed at endogenous loci in all three cell types. The two gRNAs, T1 and T2, achieved NHEJ rates of 10% and 25% in 293T cells, 13% and 38% in K562 cells, and 2% and 4% in PGP1-iPSCs, respectively (Figure 2B). No obvious toxicity was observed in any of these cell types due to the expression of Cas9 and crRNA, which are required for NHEJ induction (Figure 9). As expected, NHEJ deletions for T1 and T2 were centered around the target site, further validating the sequence specificity of this targeting process (Figures 9, 10, and 11). Simultaneous introduction of both T1 and T2 gRNAs resulted in highly efficient deletion of the 19-bp fragment located between them (Figure 10), demonstrating the feasibility of multiplexed editing of genomic loci using this approach.
[0017] In some embodiments, HR was used to integrate a dsDNA donor construct (see reference (13)) or an oligo donor into the native AAVS1 locus (Figures 2C, 12). HR-mediated integration was confirmed using both PCR (Figures 2D, 12) and Sanger sequencing (Figure 2E). 293T or iPS clones were easily derived from the engineered cell population using puromycin selection for two weeks (Figures 2F, 12). These results demonstrate that Cas9 can efficiently integrate foreign DNA at endogenous loci in human cells. Accordingly, certain embodiments of the present disclosure include methods for integrating foreign DNA into the genome of cells using homologous recombination and Cas9.
[0018] In one embodiment, an RNA-guided genome editing system is provided that can be easily applied to modify other genomic sites by simply modifying the sequence of the gRNA expression vector to match the compatible sequence in the target site. In this embodiment, 190,000 gRNA-specific targetable sequences were created that target approximately 40.5% of the exons of genes in the human genome. These target sequences were assembled into a 200-bp format compatible with multiplex synthesis on a DNA array (see Reference (14)) (Figure 13). In this embodiment, a ready-to-use genome-wide reference of potential target sites in the human genome and a method for multiplex gRNA synthesis are provided.
[0019] In some embodiments, methods are provided for multiplexing genome modification in cells by modifying the cell's genome at multiple locations using one or more or multiple RNA / enzyme systems described herein. In some embodiments, the target site perfectly matches the PAM sequence NGG and the 8-12 base "seed sequence" at the 3' end of the gRNA. In some embodiments, a perfect match of the remaining 8-12 bases is not required. In some embodiments, Cas9 will function even with a single mismatch at the 5' end. In some embodiments, the chromatin structure and epigenetic state of the target site can affect the efficiency of Cas9 function. In some embodiments, Cas9 homologs with higher specificity are included as useful enzymes. Those skilled in the art will be able to identify and modify suitable Cas9 homologs. In some embodiments, CRISPR-targetable sequences include sequences with different PAM requirements (see Reference (9)) or sequences with directed evolution. In some embodiments, inactivating one of the Cas9 nuclease domains can increase the rate of HR to NHEJ and reduce toxicity (Figure 3A, Figure 5) (References 4, 5), while inactivating both domains can enable Cas9 to function as a retargetable DNA-binding protein. Embodiments of the present disclosure are broadly useful in synthetic biology (see References (21, 22)), directed and multiplexed perturbation of gene networks (see References (13, 23)), and ex vivo targeted gene therapy (see References (24-26)) and in vivo targeted gene therapy (see Reference (27)).
[0020] In one embodiment, a "re-engineerable organism" is provided as a model system for biological discovery and in vivo screening. In one embodiment, a "re-engineerable mouse" with an inducible Cas9 transgene is provided, allowing for the localized delivery (e.g., using adeno-associated virus) of a library of gRNAs targeting multiple genes or regulatory elements to screen for tumor-causing mutations in target tissue types. The use of Cas9 homologs or nuclease-null mutants with effector domains (e.g., activators) allows for the activation or repression of multiple genes in vivo. This embodiment allows for the screening of factors that enable phenotypes such as tissue regeneration and transdifferentiation. In one embodiment, (a) the use of DNA arrays allows for the multiplex synthesis of defined gRNA libraries (see Figure 13); and (b) small gRNAs (see Figure 3b) are packaged and delivered using a number of non-viral or viral delivery methods.
[0021] In one embodiment, the lower toxicity observed with "nickases" for genome engineering applications is achieved by inactivating one of the Cas9 nuclease domains, which either nick the DNA strand base-paired with RNA or nick its complementary strand. Inactivating both domains allows Cas9 to function as a retargetable DNA-binding protein. In one embodiment, Cas9, a retargetable DNA-binding protein, binds to:
[0022] (a) transcriptional activation or repression domains for regulating expression of target genes, including, but not limited to, chromatin remodeling, histone modification, silencing, blocking, and direct interaction with the transcriptional machinery; (b) a nuclease domain such as FokI to enable “highly specific” genome editing through the dimerization of adjacent gRNA-Cas9 complexes; (c) fluorescent proteins for visualizing genomic loci and chromosomal dynamics; or (d) protein- or nucleic acid-conjugated organic fluorophores, quantum dots, molecular beacons, and other fluorescent molecules such as echo probes or molecular beacon substitutes; (e) Multivalent ligand-binding protein domains that enable programmable manipulation of genome-wide three-dimensional structure.
[0023] In some embodiments, the transcriptional activation and repression components can utilize natural or synthetic orthogonal CRISPR systems, such that gRNAs bind only to Cas activators or repressors, allowing multiple sets of gRNAs to regulate multiple targets.
[0024] According to one embodiment, the use of gRNAs offers the ability to multiplex beyond mRNAs, in part because RNAs are smaller in size than mRNAs (100 nucleotides versus 2000 nucleotides, respectively). This is particularly useful when nucleic acid delivery is size-limited, such as in viral packaging. This allows for multiple instances of cleavage, nicking, activation, or repression, or combinations thereof. The ability to easily target multiple regulatory targets allows for coarse- or fine-tuning or regulatory networks that are not limited to preexisting control circuits downstream of a particular regulator (e.g., four mRNAs are used to reprogram fibroblasts to iPSCs). Example applications of multiplexing include:
[0025] 1. Establishing (major or minor) histocompatibility alleles, haplotypes, and genotypes for human (or animal) tissue / organ transplantation. This embodiment provides, for example, an HLA homozygous cell line or humanized animal breed, or a set of gRNAs that can add such HLA alleles to an otherwise desirable cell line or breed.
[0026] 2. Mutation of multiple cis-regulatory elements (CREs = signals for transcription, splicing, translation, RNA and protein folding, degradation, etc.) in single cells (or collections of cells) can be used to efficiently study the complex sets of regulatory interactions that can occur during normal development, or in pathological, synthetic, or pharmaceutical scenarios. According to certain embodiments, CREs are (or can be) somewhat orthogonal (i.e., have little crosstalk), so many can be tested in one setting, such as an expensive time series of animal embryos. One application is with RNA fluorescence in situ sequencing (FISSeq).
[0027] 3. Multiplexed combinations of CRE mutations and / or epigenetic activation or repression of CREs can be used to modify or reprogram iPSCs or ESCs or other stem or non-stem cells into any cell type or combination of cell types for use in organs-on-chips or other cell and organ cultures for testing pharmaceuticals (small molecules, proteins, RNA, cells, animal cells, plant cells, or microbial cells, aerosols, and other delivery methods), transplantation strategies, personalization strategies, etc.
[0028] 4. Generation of multiply mutant human cells for use in diagnostic testing (and / or DNA sequencing) for medical genetics. To the extent that the chromosomal location and composition of human genomic alleles (or epigenetic markers) can affect the accuracy of clinical genetic diagnosis, it is important that the alleles are present at the correct location in the reference genome, rather than in an ectopic (i.e., transgenic) location or in a separate synthetic DNA fragment. One embodiment is a series of independent cell lines, or structural variants, one for each diagnostic human SNP. Alternatively, one embodiment includes multiple sets of alleles in the same cell. In some cases, multiple alterations in a gene (or multiple genes) are desirable, assuming independent testing. In other cases, the combination of alleles in a particular haplotype allows for testing sequencing (genotyping) methods that accurately establish the haplotype phase (i.e., whether one or both copies of the gene are affected in an individual or individual somatic cell type).
[0029] 5. Modified Cas+gRNA systems can be used to target repetitive or endogenous viral elements in microbial, plant, animal, or human cells to reduce deleterious transpositions or to assist in sequencing or other analytical genomic / transcriptomic / proteomic / diagnostic tools (where near-identical copies can be problematic).
[0030] The following references, identified by number in the above section, are incorporated by reference in their entirety: 1. Wiedenheft B, Sternberg SH, Doudna JA, Nature 482, 331 (Feb 16, 2012). D. Bhaya, M. Davison, R. Barrangou, Annual Review of Genetics 45, 273 (2011).3. MP Terns, RM Terns, Current Opinion in Microbiology 14, 321 (June, 2011). 4. M. Jinek et al, Science 337, 816 (Aug 17, 2012). 5. G. Gasiunas, R. Barrangou, P. Horvath, V. Siksnys, Proceedings of the National Academy of Sciences of the United States of America 109, E2579 (September 25, 2012).6. R. Sapranauskas et al, Nucleic Acids Research 39, 9275 (Nov, 2011). 7. TR Brummelkamp, R. Bernards, R. Agami, Science 296, 550 (Apr 19, 2002). 8. M. Miyagishi, K. Taira, Nature biotechnology 20, 497 (May, 2002). 9. E. Deltcheva et al, Nature 471, 602 (Mar 31, 2011). 10. Zou J, Mali P, Huang X, Dowey SN, Cheng L, Blood 118, 4599 (Oct 27, 2011). 11. NE Sanjana et al, Nature Protocols 7, 171 (Jan, 2012). 12. JH Lee et al, PLoS Genet 5, EL 000718 (Nov, 2009). 13. D. Hockemeyer et al, Nature biotechnology 27, 851 (September, 2009). 14. S. Kosuri et al, Nature biotechnology 28, 1295 (Dec, 2010). 15. Pattanayak V, Ramirez CL, Joung JK, Liu DR, Nature methods 8, 765 (September, 2011). [ PubMed ] 16. NM King, O. Cohen-Haguenauer, Molecular therapy: the journal of the American Society of Gene Therapy 16, 432 (Mar, 2008). https: / / doi.org / 10.1016 / S0006-9381(08)00011-0 , Google Scholar Crossref , CAS 17. YG Kim, J. Cha, S. Chandrasegaran Science of the United States of America 93, 1156 (February 6, 1996). 18. EJ Rebar, CO Turkey, Science 263, 671 (Feb 4, 1994). Science 326, 1509, 19. J. Boch et al. 20. Moscow MJ, Bogdanove AJ, Science 326, 1501 (Dec 11, 2009). [ PubMed ] 21. Khalil AS, Collins JJ, Nature reviews. Genetics 11, 367 (May, 2010). 22. PE Purnick, R. Weiss, Nature reviews. Molecular cell biology 10, 410 (Jun, 2009). 23. J. Zou et al, Cell stem cell 5, 97 (Jul 2, 2009). 24. N. Holt et al, Nature biotechnology 28, 839 (Aug, 2010). 25. FD Urnov et al, Nature 435, 646 (Jun 2, 2005). 26. A. Lombardo et al, Nature biotechnology 25, 1298 (Nov, 2007). 27. H. Li et al, Nature 475, 217 (Jul 14, 2011). [Example]
[0031] The following examples are provided as representative examples of the present disclosure, and should not be construed as limiting the scope of the disclosure, as other equivalent embodiments will be apparent in light of the disclosure, the drawings, and the appended claims.
[0032] Example 1 Type II CRISPR-Cas system According to certain aspects, embodiments of the present disclosure use short RNAs to identify foreign nucleic acids for nuclease activity in eukaryotic cells. According to certain aspects of the present disclosure, eukaryotic cells are modified to contain nucleic acids in their genomes encoding one or more short RNAs and one or more nucleases that are activated by the binding of the short RNAs to target DNA sequences. According to certain aspects, exemplary short RNA / enzyme systems can be identified in bacteria or archaea, such as (CRISPR) / CRISPR-associated (Cas) systems that use short RNAs to guide the degradation of foreign nucleic acids. CRISPR ("clustered regularly interspaced short palindromic repeats") defense involves the acquisition and integration of a novel targeting "spacer" from invading viral or plasmid DNA into the CRISPR locus, the expression and processing of a short guide CRISPR RNA (crRNA) consisting of spacer repeat units, and the cleavage of nucleic acids (most commonly DNA) complementary to the spacer.
[0033] Three classes of CRISPR systems are commonly known, referred to as Type I, Type II, or Type III. According to certain embodiments, a particularly useful enzyme for cleaving dsDNA according to the present disclosure is the single-effector enzyme Cas9, which is common to Type II (see Reference (1)). In bacteria, Type II effector systems consist of a long pre-crRNA transcribed from a spacer-containing CRISPR locus, a multifunctional Cas9 protein, and a tracrRNA, which is important for gRNA processing. When tracrRNA hybridizes to the repeat regions separating the spacers of the pre-crRNA, dsRNA cleavage by endogenous RNase III is initiated, followed by a second cleavage event within each spacer by Cas9, generating mature crRNAs that remain bound to the tracrRNA and Cas9. According to one embodiment, the eukaryotic cells of the present disclosure are modified to avoid the use of RNase III and crRNA processing in general. See reference (2).
[0034] According to certain embodiments, enzymes of the present disclosure, such as Cas9, unwind the DNA duplex and search for and cleave sequences that match the crRNA. Target recognition occurs upon detection of complementarity between the "protospacer" sequence in the target DNA and the remaining spacer sequence in the crRNA. Importantly, Cas9 cleaves DNA only if the correct protospacer adjacent motif (PAM) is also present at the 3' end. According to certain embodiments, different protospacer adjacent motifs can be used. For example, the S. pyogenes system requires the NGG sequence (where N can be any nucleotide). The S. thermophilus type II system requires NGGNG (see reference (3)) and NNAGAAW (see reference (4)), respectively, while various S. mutans systems tolerate NGG or NAAR (see reference (5)). Bioinformatics analyses have generated large databases of CRISPR loci in various bacteria that may aid in identifying additional useful PAMs and expanding the set of CRISPR-targetable sequences (see references (6, 7)). In S. thermophilus, Cas9 generates a blunt-ended double-strand break 3 bp before the 3′ end of the protospacer (see reference (8)), a process mediated by two catalytic domains in the Cas9 protein: an HNH domain that cleaves the complementary strand of DNA and a RuvC-like domain that cleaves the non-complementary strand (see Figure 1A and Figure 3). Although the S. pyogenes system has not been characterized with the same precision, a DSB also occurs near the 3′ end of the protospacer. When one of the two nuclease domains is inactivated, Cas9 functions as a nickase in vitro (see reference (2)) and in human cells (see Figure 5).
[0035] In some embodiments, the specificity of gRNA-guided Cas9 cleavage is used as a genome engineering mechanism in eukaryotic cells. In some embodiments, gRNA hybridization does not need to be 100% for the enzyme to recognize the gRNA / DNA hybrid and affect cleavage. Some off-target activity may occur. For example, the S. pyogenes system tolerates mismatches in the first six bases of the 20-bp mature spacer sequence in vitro. In some embodiments, higher stringency may be beneficial in vivo if potential off-target sites matching NGG (in the last 14 bp) exist in the human reference genome for the gRNA. The effects of mismatches and enzyme activity in general are described in references (9), (2), (10), and (4).
[0036] According to certain aspects, specificity can be improved. If interference is affected by the melting temperature of the gRNA-DNA hybrid, AT-rich target sequences may have fewer off-target sites. Careful target site selection to avoid spurious sites with at least 14 bp of matching sequence elsewhere in the genome can improve specificity. The frequency of off-target sites can be reduced by using Cas9 mutants that require longer PAM sequences. Directed evolution can improve Cas9 specificity to a level sufficient to completely prevent off-target activity, but ideally, a perfect match of the 20 bp gRNA with a minimal PAM is required. Therefore, modification of the Cas9 protein is a representative embodiment of the present disclosure. Therefore, novel methods are envisioned that allow multiple rounds of evolution in a short time frame (see Reference (11)). CRISPR systems useful in the present disclosure are described in References (12, 13).
[0037] Example 2 Plasmid construction The Cas9 gene sequence was human codon-optimized and constructed by hierarchical fusion PCR assembly of nine 500-bp gBlocks ordered from IDT. Figure 3A, which depicts the modified type II CRISPR system for human cells, shows the expression format and full sequence of the Cas9 gene insert. The RuvC-like and HNH motifs, as well as the C-terminal SV40NLS, are highlighted in blue, brown, and orange, respectively. Cas9_D10A was constructed similarly. The resulting full-length product was cloned into the pcDNA3.3-TOPO vector (Invitrogen). Target gRNA expression constructs were ordered directly from IDT as individual 455-bp gBlocks and cloned into the pCR-BluntII-TOPO vector (Invitrogen) or PCR-amplified. Figure 3B shows the U6 promoter-based expression scheme for guide RNAs and the predicted secondary structure of the RNA transcripts. Because the use of the U6 promoter restricts the first position of the RNA transcript to "G," this approach was used to express GN. 20 The seven gRNAs used are shown in Figure 3C.
[0038] Vectors for HR reporter assays containing interrupted GFP were constructed by fusion PCR assembly of the GFP sequence containing a stop codon and a 68-bp AAVS1 fragment (or its variants; see Figure 6) or 58-bp fragments derived from the DNMT3a and DNMT3b genomic loci (see Figure 8) in the Addgene EGIP lentivector (plasmid #26777). These lentivectors were then used to establish stable GFP reporter cell lines. The TALENs used in this study were constructed using the protocol described in reference (14). All DNA reagents developed in this study are available from Addgene.
[0039] Example 3 cell culture PGP1 iPS cells were maintained in mTeSR1 (Stemcell Technologies) on plates coated with Matrigel (BD Biosciences). They were subcultured every 5–7 days using TrypLE Express (Invitrogen). K562 cells were cultured and maintained in RPMI (Invitrogen) containing 15% FBS. HEK293T cells were cultured in Dulbecco's Modified Eagle's Medium (DMEM, Invitrogen) high glucose supplemented with 10% fetal bovine serum (FBS, Invitrogen), penicillin / streptomycin (pen / strep, Invitrogen), and non-essential amino acids (NEAA, Invitrogen). All cells were maintained in a humidified incubator at 37°C and 5% CO2.
[0040] Example 4 Gene targeting of PGP1 iPS, K562, and 293T PGP1 iPS cells were cultured in a Rho kinase (ROCK) inhibitor (Calbiochem) for 2 hours before nucleofection. Cells were harvested using TrypLE Express (Invitrogen) and diluted to 2 × 10 6 Cells were resuspended in P3 reagent (Lonza) with 1 μg of Cas9 plasmid, 1 μg of gRNA, and / or 1 μg of DNA donor plasmid and nucleofected according to the manufacturer's instructions (Lonza). Cells were then plated on mTeSRl-coated plates in mTeSRl medium supplemented with ROCK inhibitor for the first 24 hours. For K562, 2 × 10 cells were used. 6 0.1 × 10 cells were resuspended in SF Reagent (Lonza) with 1 μg of Cas9 plasmid, 1 μg of gRNA, and / or 1 μg of DNA donor plasmid and nucleofected according to the manufacturer's instructions (Lonza). For 293T, 0.1 × 10 cells were nucleofected using Lipofectamine 2000 according to the manufacturer's protocol. 6Each cell was transfected with 1 µg of Cas9 plasmid, 1 µg of gRNA, and / or 1 µg of DNA donor plasmid. The DNA donors used for targeting endogenous AAVS1 were either dsDNA donors (Figure 2C) or 90-base-long oligonucleotides. The former contains a SA-2A-puromycin-CaGGS-eGFP cassette and flanking short homology arms to enrich for successfully targeted cells.
[0041] Targeting efficiency was evaluated as follows: Cells were harvested 3 days after nucleofection and diluted to approximately 1 × 10 cells using prepGEM (ZyGEM). 6 Genomic DNA was extracted from the cells. PCR was performed to amplify the targeted region using the genomic DNA from the cells, and the amplified product was deep sequenced using a MiSeq Personal Sequencer (Illumina) with a coverage of >200,000 reads. The sequencing data was analyzed to estimate NHEJ efficiency. The reference AAVS1 sequence analyzed is as follows:
[0042] CACTTCAGGACAGCATGTTTGCTGCCTCCAGGGATCCTGTGTCCCCGAGCTGGGACCACCTTATATCCCAGGGCCGGTTAATGTGGCTCTGGTTCTGGGTACTTTTATCTGTCCCCTCCACCCCACAGTGGGG CCACTAGGGACAGGATTGGTGACAGAAAAGCCCCATCCTTAGGCCTCCTCCTTCCTAGTCTCCTGATATTGGGTCTAACCCCCACCTCCTGTTAGGCAGATTCCTTATCTGGTGACACACCCCCATTTCCTGGA
[0043] The PCR primers for amplifying the targeted region in the human genome are as follows: AAVS1-R: CTCGGCATTCCTGCTGAACCGCTCTTCCGATCTacaggaggtgggggttagac AAVS1-F.1: ACACTCTTTCCCTACACGACGCTCTTCCGATCTCGTGATtatattcccagggccggtta AAVS1-F.2: ACACTCTTTCCCTACACGACGCTCTTCCGATCTACATCGtatattcccagggccggtta AAVS1-F.3: ACACTCTTTCCCTACACGACGCTCTTCCGATCTGCCTAAtatattcccagggccggtta AAVS1-F.4: ACACTCTTTCCCTACACGACGCTCTTCCGATCTTGGTCAtatattcccagggccggtta AAVS1-F.5: ACACTCTTTCCCTACACGACGCTCTTCCGATCTCACTGTtatattcccagggccggtta AAVS1-F.6: ACACTCTTTCCCTACACGACGCTCTTCCGATCTATTGGCtatattcccagggccggtta AAVS1-F.7: ACACTCTTTCCCTACACGACGCTCTTCCGATCTGATCTGtatattcccagggccggtta AAVS1-F.8: ACACTCTTTCCCTACACGACGCTCTTCCGATCTTCAAGTtatattcccagggccggtta AAVS1-F.9: ACACTCTTTCCCTACACGACGCTCTTCCGATCTCTGATCtatattcccagggccggtta AAVS1-F.10: ACACTCTTTCCCTACACGACGCTCTTCCGATCTAAGCTAtatattcccagggccggtta AAVS1-F.11: ACACTCTTTCCCTACACGACGCTCTTCCGATCTGTAGCCtatattcccagggccggtta AAVS 1-F.12: ACACTCTTTCCCTACACGACGCTCTTCCGATCTTACAAGtatattcccagggccggtta
[0044] To analyze the HR event using the DNA donor in Figure 2C, the following primers were used: HR_AAVS1-F CTGCCGTCTCTCTCCTGAGT HR_Puro-R GTGGGCTTGTACTCGGTCAT
[0045] Example 5 A bioinformatics approach to calculating CRISPR targets for human exons and a methodology for their multiplex synthesis A set of gRNA gene sequences that maximizes targeting of specific sites in human exons while minimizing targeting of other sites in the genome was determined as follows. According to one embodiment, maximal gRNA targeting efficiency is achieved with a 23-nucleotide (nt) sequence, of which the 5'-most 20 nucleotides must be exactly complementary to the desired site and the 3'-most 3 bases must be in the form NGG. Furthermore, the 5'-most nucleotide must be G to establish the Pol-III transcription start site. However, according to reference (2), mispairing of the 5'-most 6 nucleotides of a 20-bp gRNA to a genomic target does not inhibit Cas9-mediated cleavage as long as the last 14 nucleotides are properly paired, whereas mispairing of the 5'-most 8 nucleotides and pairing of the last 12 nucleotides inhibits cleavage. Mispairing of the 5'-most 7 nucleotides and pairing of the 3'-most 13 nucleotides have not been tested. One condition for being conservative with respect to off-target effects is that cleavage is permitted in the seven most 5'-nucleotide mismatches, as well as in the six-nucleotide mismatches, so that 13 most 3'-nucleotide mismatches are sufficient for cleavage. To identify CRISPR target sites within human exons that are capable of cleavage without off-target cleavage, we examined all 23-bp sequences of the form 5'-GBBBB BBBBB BBBBB BBBBB NGG-3' (first form). Here, B represents the base of the exonic site, and for this sequence, no other sequences of the form 5'-NNNNN NNBBB BBBBB BBBBB NGG-3' (second form) exist anywhere else in the human genome. Specifically, we (i) downloaded BED files for the coding regions of all RefSeq genes in the GRCh37 / hg19 human genome from the UCSC Genome Browser (references 15-17). The coding exon sites in this BED file consisted of a set of 346,089 mappings of RefSeq mRNA accessions to the hg19 genome.However, some RefSeq mRNA accessions mapped to multiple genomic sites (likely gene duplications), and many accessions mapped to subsets of the same set of exon sites (multiple isoforms of the same gene). To distinguish clearly duplicated gene instances and consolidate multiple references to the same genomic exon instance by multiple RefSeq isoform accessions, we (ii) assigned unique numerical suffixes to the 705 RefSeq accession numbers with multiple genomic sites, and (iii) merged the duplicated exon sites into a merged exon region using the mergeBed function in BEDTools (18) (v2.16.2-zip-87e3926). These steps reduced the initial set of 346,089 RefSeq exon sites to 192,783 distinct genomic regions. Using the UCSC Table Browser, we downloaded the hg19 sequences of all merged exon regions and added 20 bp of padding to each end. (iv) Custom Perl code was used to identify 1,657,793 instances of the first form within this exon sequence. (v) These sequences were then filtered for the presence of second form off-targets. Specifically, for each merged exon's first form target, we extracted the 3′-most 13-bp specific (B) "core" sequence. For each core, we created four 16-bp sequences: 5′-BBB BBBBB BBBBB NGG-3′ (N = A, C, G, and T). We then searched the entire hg19 genome for exact matches to these 6,631,172 sequences using Bowtie version 0.12.8 (19) with the following parameters: -l 16 -v 0 -k 2. Any exon target site with more than one match was discarded. Because any specific 13 bp core sequence followed by the sequence NGG only confers 15 bp of specificity, there should be an average of about 5.6 matches to the extended core sequence in a random sequence of about 3 Gb (both strands).Therefore, most of the 1,657,793 initially identified targets were discarded, but 189,864 sequences passed this filter. These comprise the set of CRISPR-targetable exonic sites in the human genome. These 189,864 sequences target sites in 78,028 merged exonic regions (approximately 40.5% of the total 192,783 merged human exonic regions) at a multiplicity of approximately 2.4 sites per targeted exonic region. To assess targeting at the gene level, RefSeq mRNA mappings were clustered such that any two RefSeq accessions overlapping a merged exonic region (including the gene duplicates identified in (ii)) were counted as a single gene cluster. The 189,864 exon-specific CRISPR sites target 17,104 of the 18,872 gene clusters (approximately 90.6% of all gene clusters) at a multiplicity of approximately 11.1 per targeted gene cluster. (Note that these gene clusters collapse RefSeq mRNA accessions representing multiple isoforms of a single transcribed gene into a single entity, but also collapse overlapping distinct genes and genes with antisense transcripts.) At the original RefSeq accession level, 189,864 sequences targeted exonic regions in 30,563 of the total 43,726 mapped RefSeq accessions (including distinct overlapping genes) (approximately 69.9%), with a multiplicity of approximately 6.2 sites per targeted mapped RefSeq accession.
[0046] In certain embodiments, the database can be refined by correlating performance with factors such as the base composition and secondary structure of both the gRNA and the genomic target (References 20, 21), as well as the epigenetic state of these targets in human cell lines where information on epigenetic state is available (Reference 22).
[0047] Example 6 multiplex synthesis The target sequences were engineered into a 200-bp format compatible with multiplex synthesis on DNA arrays (References 23, 24). In one embodiment, this method allows targeted recovery of specific gRNA sequences or pools of gRNA sequences from DNA array-based oligonucleotide pools and rapid cloning into a general expression vector (Figure 13A). Specifically, a 12k oligonucleotide pool was synthesized by CustomArray Inc. Furthermore, optimal gRNAs were successfully recovered from this library (Figure 13B). The error rate was approximately 4 mutations per 1000 bp of synthesized DNA.
[0048] Example 7 RNA-guided genome editing requires both Cas9 and guide RNA for successful targeting Using the GFP reporter assay described in Figure 1B, all possible combinations of repair DNA donor, Cas9 protein, and gRNA were tested for their ability to successfully perform HR (in 293T cells). As shown in Figure 4, GFP+ cells were observed only when all three components were present, confirming that these CRISPR components are essential for RNA-guided genome editing. Data are means ± SEM (N=3).
[0049] Example 8 Analysis of gRNA- and Gas9-mediated genome editing We investigated the CRISPR-mediated genome editing process using either the previously described GFP reporter assay, the results of which are shown in Figure 5A, or (B) deep sequencing of the target site (in 293T cells), the results of which are shown in Figure 5B. For comparison, we tested the D10A mutant of Cas9, which has previously been shown to function as a nickase in in vitro assays. As shown in Figure 5, both Cas9 and Cas9D10A were able to successfully perform HR at approximately the same rate. However, deep sequencing confirmed that while Cas9 exhibits robust NHEJ at the target site (as expected from its presumed ability to simply nick DNA), the D10A mutant exhibited a significantly reduced NHEJ rate. Furthermore, consistent with the known biochemistry of the Cas9 protein, the NHEJ data confirmed that the majority of base pair deletions or insertions occurred at the 3' end of the target sequence, with a peak approximately 3-4 bases upstream of the PAM site and a median deletion frequency of approximately 9-10 bp. Data are means ± SEM (N=3).
[0050] Example 9 RNA-guided genome editing is target sequence specific Similar to the GFP reporter assay described in Figure 1B, we developed three 293T stable cell lines, each harboring a distinct GFP reporter construct. These were distinguished by the sequence of the AAVS1 fragment insert (as shown in Figure 6). One cell line had the wild-type fragment, while the other two cell lines had a 6-bp mutation (highlighted in red). Each cell line was then targeted with one of four reagents: a GFP-ZFN combination, which could target all cell types because its target sequence was present in all cell lines due to being in the adjacent GFP fragment; an AAVS1 TALEN, which could potentially target only the wt-AAVS1 fragment because mutations in the other two cell lines likely prevented the left TALEN from binding to their site; a T1 gRNA, which could also potentially target only the wt-AAVS1 fragment because this target site was also disrupted in the two mutant cell lines; and a T2 gRNA, which, unlike T1 gRNA, was likely capable of targeting all three cell lines because its target site was unchanged in the three cell lines. The ZFNs modified all three cell types, the AAVS1 TALEN and T1gRNA targeted only the wt-AAVS1 cell type, and the T2gRNA successfully targeted all three cell types. Together, these results confirm that guide RNA-mediated editing is target sequence specific. Data are means ± SEM (N = 3).
[0051] Example 10 Guide RNAs targeting GFP sequences enable robust genome editing In addition to the two gRNAs targeting the AAVS1 insert, we tested two new gRNAs targeting the flanking GFP sequence of the reporter described in Figure 1B (in 293T). As shown in Figure 7, these gRNAs were also able to confer robust HR at this modified locus. Data are means ± SEM (N = 3).
[0052] Example 11 RNA-guided genome editing is target sequence specific and exhibits targeting efficiency similar to that of ZFNs or TALENs Similar to the GFP reporter assay described in Figure 1B, we developed two 293T stable cell lines, each carrying a distinct GFP reporter construct. These were distinguished by the sequence of the fragment insert (as shown in Figure 8). One cell line carried a 58-bp fragment from the DNMT3a gene, while the other carried a homologous 58-bp fragment from the DNMT3b gene. The sequence differences are highlighted in red. Each cell line was then targeted with one of six reagents: a GFP-ZFN combination, which is capable of targeting all cell types because its target sequence is present in all cell lines due to its location in the adjacent GFP fragment; a TALEN combination potentially targeting either the DNMT3a or DNMT3b fragment; a gRNA combination potentially targeting only the DNMT3a fragment; and a gRNA potentially targeting only the DNMT3b fragment. As shown in Figure 8, the ZFN modified all three cell types, while the TALEN and gRNA modified only their respective targets. Furthermore, the targeting efficiency was comparable across the six targeting reagents. Taken together, these results confirm that RNA-guided editing is target sequence specific and exhibits targeting efficiency similar to that of ZFNs or TALENs. Data are means ± SEM (N=3).
[0053] Example 12 RNA-induced NHEJ in human iPS cells Human iPS cells (PGP1) were nucleofected with the constructs shown in the left panel of Figure 9. Four days after nucleofection, the NHEJ rate was measured by assessing the genomic deletion and insertion rates at DNA double-strand breaks (DSBs) by deep sequencing. Panel 1: Deletion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of deletions at each nucleotide position is plotted on the black line, and the deletion rate was calculated as the percentage of reads with deletions. Panel 2: Insertion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of insertions at the genomic position where the first insertion junction was detected is plotted on the black line, and the insertion rate was calculated as the percentage of reads with insertions. Panel 3: Distribution of deletion sizes. The frequency of deletions of different sizes among the entire NHEJ population is plotted. Panel 4: Distribution of insertion sizes. The frequencies of insertions of different sizes across the total NHEJ population were plotted. iPS targeting by both gRNAs was efficient (2–4%) and sequence-specific (as indicated by the shift in the position of the NHEJ deletion distribution), reaffirming the results in Figure 4. NGS-based analysis also indicates that both the Cas9 protein and gRNA are essential for NHEJ events at the target locus.
[0054] Example 13 RNA-induced NHEJ in K562 cells The constructs shown in the left panel of Figure 10 were nucleated into K562 cells. Four days after nucleofection, the NHEJ rate was measured by assessing the genomic deletion and insertion rates at DSBs by deep sequencing. Panel 1: Deletion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of deletions at each nucleotide position is plotted on the black line, and the deletion rate was calculated as the percentage of reads with deletions. Panel 2: Insertion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of insertions at the genomic position where the first insertion junction was detected is plotted on the black line, and the insertion rate was calculated as the percentage of reads with insertions. Panel 3: Deletion size distribution. The frequency of deletions of different sizes across all NHEJ populations is plotted. Panel 4: Insertion size distribution. The frequency of insertions of different sizes across all NHEJ populations is plotted. Targeting of K562 by both gRNAs was efficient (13–38%) and sequence-specific (as indicated by a shift in the position of the NHEJ deletion distribution). Importantly, co-transfection of the T1 and T2 guide RNAs resulted in highly efficient deletion of the 19-bp fragment located between them, as evidenced by a peak in the histogram of observed deletion size frequencies, demonstrating the feasibility of multiplexed editing of genomic loci using this approach.
[0055] Example 14 RNA-induced NHEJ in 293T cells The constructs shown in the left panel of Figure 11 were transfected into 293T cells. Four days after nucleofection, the NHEJ rate was measured by assessing the genomic deletion and insertion rates at DSBs by deep sequencing. Panel 1: Deletion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of deletions at each nucleotide position is plotted on the black line, and the deletion rate was calculated as the percentage of reads with deletions. Panel 2: Insertion rate detected in the target region. Red dashed line: Boundary of the T1 RNA target site; green dashed line: Boundary of the T2 RNA target site. The occurrence of insertions at the genomic position where the first insertion junction was detected is plotted on the black line, and the insertion rate was calculated as the percentage of reads with insertions. Panel 3: Distribution of deletion sizes. The frequency of deletions of different sizes across all NHEJ populations is plotted. Panel 4: Distribution of insertion sizes. The frequency of insertions of different sizes across all NHEJ populations is plotted. 293T targeting by both gRNAs is efficient (10–24%) and sequence-specific (as indicated by a shift in the position of the NHEJ deletion distribution).
[0056] Example 15 HR at the endogenous AAVS1 locus using either dsDNA donors or short oligonucleotide donors As shown in Figure 12A, PCR screening (see Figure 2C) confirmed successful targeting in 21 / 24 randomly selected 293T clones. As shown in Figure 12B, similar PCR screening confirmed successful targeting in 3 / 7 randomly selected PGP1-iPS clones. As shown in Figure 12C, a short 90-mer oligo was also capable of robust targeting at the endogenous AAVS1 locus (shown here for K562 cells).
[0057] Example 16 A method for multiplex synthesis, recovery, and cloning of U6 expression vectors of guide RNAs targeting genes in the human genome We created a resource of approximately 190k bioinformatically calculated unique gRNA sites targeting approximately 40.5% of all exons in genes in the human genome. As shown in Figure 13A, the gRNA target sites were assembled into a 200-bp format compatible with multiplex synthesis on DNA arrays. Specifically, this design enables (i) targeted recovery of specific gRNA targets or pools of gRNA targets from DNA array oligonucleotide pools (by three consecutive rounds of nested PCR, as shown in the schematic); and (ii) rapid cloning into generic expression vectors, which serve as recipients for incorporation of gRNA insert fragments via Gibson assembly after linearization with AflII. As shown in Figure 13B, using this method, we achieved targeted recovery of 10 unique gRNAs from a 12k oligonucleotide pool synthesized by CustomArray.
[0058] Example 17 CRISPR-mediated RNA-guided transcriptional activation The CRISPR-Cas system is a bacterial adaptive immune defense system that functions to "cut" invading nucleic acids. In one embodiment, the CRISPR-CAS system was modified to function in human cells and "cut" genomic DNA. This is achieved by using a short guide RNA that guides the Cas9 protein (which has nuclease function) to a target sequence complementary to a spacer in the guide RNA. The ability to "cut" DNA enables many applications related to genome editing and targeted genome regulation. To this end, the Cas9 protein was mutated to be nuclease-deficient by introducing mutations that are predicted to prevent binding to Mg2+ (known to be important for the nuclease function of the RuvC-like and HNH-like domains). Specifically, a combination of D10A, D839A, H840A, and N863A mutations was introduced. The resulting Cas9 nuclease-deficient protein (hereafter referred to as Cas9R-H-), whose ability to not cleave DNA was confirmed by sequencing analysis, was then coupled to a transcriptional activation domain (VP64) to enable the CRISPR-cas system to function as an RNA-guided transcriptional regulator (see Figure 14). The Cas9R-H-+VP64 fusion conjugate enables RNA-guided transcriptional activation in the two reporters shown. Specifically, both FACS analysis and immunofluorescence imaging demonstrate that this protein enables gRNA sequence-specific targeting of the corresponding reporters, and furthermore, the resulting transcriptional activation, as assayed by dTomato fluorescent protein expression, was at a level similar to that induced by a conventional TALE-VP64 fusion protein.
[0059] Example 18 Flexibility of gRNA sequences and their applications We systematically assayed various random sequence insertions in the 5', middle, and 3' portions of the gRNA to determine the flexibility of the gRNA scaffold sequence for designer sequence insertion. Specifically, we engineered 1-bp, 5-bp, 10-bp, 20-bp, and 40-bp inserts into the gRNA sequence at the 5', middle, and 3' ends of the gRNA (the exact locations of the insertions are highlighted in red in Figure 15). The functionality of these gRNAs was then tested by their ability to induce HR in a GFP reporter assay (as described herein). It is clear that gRNAs are flexible to sequence insertions at the 5' and 3' ends (as measured by retained HR-induction activity). Thus, aspects of the present disclosure relate to tagging small molecule-responsive RNA aptamers or visualizing gRNAs, which can trigger the initiation of gRNA activity. Additionally, embodiments of the present disclosure relate to tethering the ssDNA donor to the gRNA by hybridization, which allows coupling of genomic target cleavage with immediate physical localization of the repair template, thereby facilitating homologous recombination rates compared to error-prone non-homologous end-joining.
[0060] The following references, identified by number in the Examples section, are incorporated by reference in their entirety for all purposes.
[0061] References 1. KS Makarova et al, Evolution and classification of the CRISPR-Cas systems. Nature reviews. Microbiology 9, 467 (Jun, 2011). 2. M. Jinek et al., A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816 (Aug 17, 2012). 3. P. Horvath, R. Barrangou, CRISPR / Cas, the immune system of bacteria and archaea. Science 327, 167 (Jan 8, 2010). 4. H. Deveau et al., Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. Journal of bacteriology 190, 1390 (Feb, 2008). 5. J. R. van der Ploeg, Analysis of CRISPR in Streptococcus mutans suggests frequent occurrence of acquired immunity against infection by M102-like bacteriophages. Microbiology 155, 1966 (Jun, 2009). 6. M. Rho, Y. W. Wu, H. Tang, T. G. Doak, Y. Ye, Diverse CRISPRs evolving in human microbiomes. PLoS genetics 8, e1002441 (2012). 7. D. T. Pride et al, Analysis of streptococcal CRISPRs from human saliva reveals substantial sequence diversity within and between subjects over time. Genome research 21, 126 (Jan, 2011). 8. G. Gasiunas, R. Barrangou, P. Horvath, V. Siksnys, Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria. Proceedings of the National Academy of Sciences of the United States of America 109, E2579 (Sep 25, 2012). 9. R. Sapranauskas et al, The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli. Nucleic acids research 39, 9275 (Nov, 2011). 10. J. E. Garneau et al, The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA. Nature 468, 67 (Nov 4, 2010). 11. K. M. Esvelt, J. C. Carlson, D. R. Liu, A system for the continuous directed evolution of biomolecules. Nature 472, 499 (Apr 28, 2011). 12. R. Barrangou, P. Horvath, CRISPR: new horizons in phage resistance and strain identification. Annual review of food science and technology 3, 143 (2012). 13. B. Wiedenheft, S. H. Sternberg, J. A. Doudna, RNA-guided genetic silencing systems in bacteria and archaea. Nature 482, 331 (Feb 16, 2012). 14. N. E. Sanjana et al, A transcription activator-like effector toolbox for genome engineering. Nature protocols 7, 171 (Jan, 2012). 15. W. J. Kent et al, The human genome browser at UCSC. Genome Res 12, 996 (Jun, 2002). 16. T. R. Dreszer et al., The UCSC Genome Browser database: extensions and updates 2011. Nucleic Acids Res 40, D918 (Jan, 2012). 17. D. Karolchik et ah, The UCSC Table Browser data retrieval tool. Nucleic Acids Res 32, D493 (Jan 1, 2004). 18. A. R. Quinlan, I. M. Hall, BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841 (Mar 15, 2010). 19. B. Langmead, C. Trapnell, M. Pop, S. L. Salzberg, Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009). 20. R. Lorenz et al., ViennaRNA Package 2.0. Algorithms for molecular biology : AMB 6, 26 (2011). 21. DH Mathews, J. Sabina, M. Zuker, DH Turner, Expanded sequence dependence of thermodynamic parameters improves prediction of RNA secondary structure. Journal of molecular biology 288, 911 (May 21, 1999). 22. RE Thurman et al., The accessible chromatin landscape of the human genome. Nature 489, 75 (Sep 6, 2012). 23. S. Kosuri et al., Scalable gene synthesis by selective amplification of DNA pools from high-fidelity microchips. Nature biotechnology 28, 1295 (Dec, 2010). 24. Q. Xu, MR Schlabach, GJ Hannon, SJ Elledge, Design of 240,000 orthogonal 25mer DNA barcode probes. Proceedings of the National Academy of Sciences of the United States of America 106, 2289 (Feb 17, 2009).
[0062] Aspects of the present disclosure include the following aspects. <1> transfecting a eukaryotic cell with a nucleic acid encoding RNA complementary to the genomic DNA of the eukaryotic cell; and transfecting the eukaryotic cell with a nucleic acid encoding an enzyme that interacts with the RNA and site-specifically cleaves the genomic DNA; Including, A method of modifying a eukaryotic cell, wherein the cell expresses the RNA and the enzyme, the RNA binds to complementary genomic DNA, and the enzyme site-specifically cleaves the genomic DNA. <2> the enzyme is Cas9; <1> The method described below. <3> The eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. <1> The method described below. <4> the RNA comprises about 10 nucleotides to about 250 nucleotides; <1> The method described below. <5> The RNA comprises about 20 nucleotides to about 100 nucleotides. <1> The method described below. <6> transfecting a human cell with a nucleic acid encoding an RNA complementary to the genomic DNA of the eukaryotic cell; and transfecting said human cells with a nucleic acid encoding an enzyme that interacts with said RNA and site-specifically cleaves said genomic DNA; Including, A method of modifying a human cell, wherein the human cell expresses the RNA and the enzyme, the RNA binds to complementary genomic DNA, and the enzyme site-specifically cleaves the genomic DNA. <7> the enzyme is Cas9; <6> The method described below. <8> the RNA comprises about 10 nucleotides to about 250 nucleotides; <6> The method described below. <9> The RNA comprises about 20 nucleotides to about 100 nucleotides. <6> The method described below. <10> transfecting a eukaryotic cell with a plurality of nucleic acids encoding RNAs complementary to different sites on the genomic DNA of the eukaryotic cell; and transfecting said eukaryotic cell with a nucleic acid encoding an enzyme that interacts with said RNA and site-specifically cleaves said genomic DNA; Including, A method of modifying a eukaryotic cell at multiple genomic DNA sites, wherein the cell expresses the RNA and the enzyme, the RNA binds to complementary genomic DNA, and the enzyme cleaves the genomic DNA in a site-specific manner. <11> the enzyme is Cas9; <10> The method described below. <12> The eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. <10> The method described below. <13> the RNA comprises about 10 nucleotides to about 250 nucleotides; <10> The method described below. <14> The RNA comprises about 20 nucleotides to about 100 nucleotides. <10> The method described below.
Claims
1. (1) a guide RNA sequence or a first nucleic acid sequence encoding the guide RNA sequence, and (2) a Cas enzyme of a type II CRISPR system that forms a complex with the guide RNA sequence, or a second nucleic acid sequence encoding the Cas enzyme of the type II CRISPR system; Including, The guide RNA sequence comprises a spacer sequence and a scaffold sequence complementary to a target nucleic acid sequence in a eukaryotic cell, and the guide RNA sequence is a crRNA-tracrRNA fusion transcript of 100 to 250 nucleotides. An RNA-guided genome editing system for use in eukaryotic cells.
2. The scaffold sequence The RNA-guided genome editing system of claim 1, comprising GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 45).
3. The scaffold sequence GUUUUAGAGCUAGAAAAUAGCAAGUUAAAAAAAAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGCU The RNA-guided genome editing system of claim 1, comprising UUU (sequence number 46).
4. The RNA-guided genome editing system described in claim 1, wherein the eukaryotic cell is a yeast cell, a plant cell, a mammalian cell or a human cell.
5. The RNA-induced genome editing system described in claim 1, wherein the eukaryotic cell is a stem cell.
6. The RNA-induced genome editing system described in claim 1, wherein the eukaryotic cell is a human iPS cell.
7. (1) The RNA-guided genome editing system described in claim 1, comprising the first nucleic acid sequence encoding the guide RNA sequence, and further comprising a control element operable in a eukaryotic cell operably linked to the first nucleic acid sequence encoding the guide RNA sequence.
8. (2) The RNA-guided genome editing system described in claim 1, comprising the second nucleic acid sequence encoding a Cas enzyme of the type II CRISPR system, wherein the second nucleic acid sequence encoding the Cas enzyme of the type II CRISPR system is a human codon-optimized nucleic acid and encodes a nuclear localization signal.
9. (2) The RNA-guided genome editing system described in claim 1, comprising the second nucleic acid sequence encoding a Cas enzyme of the type II CRISPR system, wherein the second nucleic acid sequence encoding the Cas enzyme of the type II CRISPR system is a human codon-optimized nucleic acid comprising a control element operable in a eukaryotic cell.
10. The RNA-guided genome editing system described in claim 1, wherein the Cas enzyme of the type II CRISPR system is Cas9.
11. A nucleic acid sequence comprising a spacer sequence complementary to a target nucleic acid sequence in a eukaryotic cell, and a scaffold sequence, The scaffold sequence is A guide RNA for use in eukaryotic cells comprising: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 45).
12. A nucleic acid sequence comprising a spacer sequence complementary to a target nucleic acid sequence in a eukaryotic cell, and a scaffold sequence, The scaffold sequence is GUUUUAGAGCUAGAAAAUAGCAAGUUAAAAAAAAAGGCUAGUCCGUUAUCAACUUGAAAAAAGUGGCACCGAGUCGGUGCU A guide RNA for use in eukaryotic cells comprising UUU (SEQ ID NO: 46).
13. A method for ex vivo modification of a eukaryotic cell containing a target nucleic acid, comprising: (a) a guide RNA or a nucleic acid encoding said guide RNA, wherein said guide RNA comprises (i) a spacer sequence complementary to a sequence in an AAV integration site in said target nucleic acid, and (ii) a scaffold sequence; (b) a Cas9 protein or a nucleic acid encoding the Cas9 protein, wherein the Cas9 protein is configured to interact with the guide RNA and site-specifically cleave the target nucleic acid; and (c) a single-stranded DNA oligonucleotide comprising a donor nucleic acid sequence; to said eukaryotic cell, the guide RNA binds to the sequence in the AAV integration site, the Cas9 protein site-specifically cleaves the target nucleic acid, and the donor nucleic acid sequence is inserted into the AAV integration site of the target nucleic acid, thereby modifying the eukaryotic cell. The method.
14. The method described in claim 13, wherein the guide RNA has a length of 10 to 250 nucleotides, or a length of 20 to 100 nucleotides.
15. The method described in claim 13, wherein the guide RNA comprises the sequence of SEQ ID NO:
45.
16. The method described in claim 13, comprising providing the guide RNA to the eukaryotic cell.
17. The method described in claim 13, comprising providing the nucleic acid encoding the guide RNA to the eukaryotic cell.
18. The method described in claim 17, wherein providing the nucleic acid encoding the guide RNA comprises a viral delivery method.
19. The method described in claim 18, wherein the viral delivery method includes an adeno-associated virus.
20. The method described in claim 17, wherein providing the nucleic acid encoding the guide RNA comprises a non-viral delivery method.
21. The method of claim 13, comprising providing the Cas9 protein to the eukaryotic cell.
22. The method described in claim 13, comprising providing the nucleic acid encoding the Cas9 protein to the eukaryotic cell.
23. The method described in claim 22, wherein providing the nucleic acid encoding the Cas9 protein comprises a viral delivery method.
24. The method of claim 23, wherein the viral delivery method comprises an adeno-associated virus.
25. The method of claim 13, comprising introducing the Cas9 protein or the nucleic acid encoding the Cas9 protein into the eukaryotic cell by a non-viral delivery method.
26. The method described in claim 13, wherein the Cas9 protein is Cas9 nickase.
27. The method described in claim 13, wherein the Cas9 protein includes a nuclear localization signal.
28. The method described in claim 13, wherein the eukaryotic cell is a non-human cell.
29. The method of claim 13, wherein the donor nucleic acid sequence is inserted into the AAV integration site by homologous recombination.
30. (a) a guide RNA or a nucleic acid encoding said guide RNA, wherein said guide RNA comprises (i) a spacer sequence complementary to a sequence in an AAV integration site in a target nucleic acid in a eukaryotic cell, and (ii) a scaffold sequence; (b) a Cas9 protein or a nucleic acid encoding the Cas9 protein, wherein the Cas9 protein is configured to interact with the guide RNA and site-specifically cleave the target nucleic acid; and (c) a single-stranded DNA oligonucleotide comprising a donor nucleic acid sequence; a eukaryotic cell produced by a method comprising providing to the eukaryotic cell the guide RNA binds to the sequence in the AAV integration site, the Cas9 protein site-specifically cleaves the target nucleic acid, and the donor nucleic acid sequence is inserted into the AAV integration site of the target nucleic acid, thereby modifying the eukaryotic cell. The eukaryotic cell.
31. (a) a guide RNA or a nucleic acid encoding said guide RNA, wherein said guide RNA comprises (i) a spacer sequence complementary to a sequence in an AAV integration site in a target nucleic acid in a eukaryotic cell, and (ii) a scaffold sequence; (b) a Cas9 protein or a nucleic acid encoding the Cas9 protein, wherein the Cas9 protein is configured to interact with the guide RNA and site-specifically cleave the target nucleic acid; and (c) a single-stranded DNA oligonucleotide comprising a donor nucleic acid sequence; to a eukaryotic cell, the guide RNA binds to the sequence in the AAV integration site, the Cas9 protein site-specifically cleaves the target nucleic acid, and the donor nucleic acid sequence is inserted into the AAV integration site of the target nucleic acid, thereby modifying the eukaryotic cell. The organ.