RNA-guided human genome engineering
By using the CRISPR-Cas9 system to guide the Cas9 enzyme to cut genomic DNA in human cells with short RNA, the efficiency and accuracy problems of genome editing in existing technologies have been solved, enabling efficient, low-toxicity genome editing and multisite editing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2013-12-16
- Publication Date
- 2026-03-27
AI Technical Summary
Existing technologies struggle to achieve efficient and specific genome editing in eukaryotic cells, especially in human cells, and conventional methods may have toxicity and off-target effects.
Using the CRISPR-Cas9 system, short RNA (gRNA) guides the Cas9 enzyme to specifically cut genomic DNA in human cells, and genome editing is achieved by expressing RNA and enzymes complementary to the genomic DNA.
It enables efficient and specific genome editing in human cells, reduces toxicity, and improves the accuracy and efficiency of editing, supporting the editing and integration of multiple genomic sites.
Smart Images

Figure CN121737181A_ABST
Abstract
Description
[0001] Relevant application data
[0002] This application claims priority to U.S. Provisional Application No. 61 / 779,169, filed March 13, 2013, and U.S. Provisional Application No. 61 / 738,355, filed December 17, 2012, the entire contents of which are incorporated herein by reference for all purposes.
[0003] Government Rights Statement
[0004] This invention was completed with government support under P50 HG005550 granted by the National Institutes of Health. The government holds certain rights to this invention. Background Technology
[0005] The CRISPR systems of bacteria and archaea rely on crRNA complexed with Cas proteins to directly degrade complementary sequences within invading viral and plasmid DNA (1-3). *Streptococcus pyogenes* (… S. pyogenes Recent in vitro reconstruction of the type II CRISPR system has shown that the crRNA fused with the normally trans-encoded tracrRNA is sufficient to guide the Cas9 protein sequence to specifically cleave the target DNA sequence matching the crRNA. 4 ). Summary of the Invention
[0006] This disclosure uses numerical references, which are listed at the end of this disclosure. References corresponding to numbers are incorporated herein by reference as supporting citations for the entirety of that number.
[0007] According to one aspect of this disclosure, eukaryotic cells are transfected with a two-component system comprising RNA complementary to genomic DNA and an enzyme interacting with the RNA. The RNA and the enzyme are expressed by the cells. The RNA / enzyme complex then binds to the complementary genomic DNA. The enzyme then performs a function, such as cleavage of the genomic DNA. The RNA comprises between about 10 nucleotides and about 250 nucleotides. The enzyme comprises between about 20 nucleotides and about 100 nucleotides. According to some aspects, the enzyme can perform any desired function in a site-specific manner, for which the enzyme has been engineered. According to one aspect, the eukaryotic cell is a yeast cell, plant cell, or mammalian cell. According to one aspect, the enzyme cleaves a genomic sequence targeted by the RNA sequence (see References). 4-6 This process produces eukaryotic cells with altered genomes.
[0008] According to one aspect, this disclosure provides a method for genetically altering human cells by including nucleic acids encoding RNA complementary to genomic DNA into the cell's genome and nucleic acids encoding enzymes on the genomic DNA that perform desired functions into the cell's genome. According to one aspect, RNA and enzymes are expressed. According to one aspect, RNA hybridizes with complementary genomic DNA. According to one aspect, when RNA hybridizes with complementary genomic DNA, the enzyme is activated to perform a desired function, such as cleavage, in a site-specific manner. According to one aspect, RNA and enzymes are components of a bacterial type II CRISPR system.
[0009] According to one aspect, a method for altering eukaryotic cells is provided, comprising transfecting eukaryotic cells with a nucleic acid encoding RNA complementary to the genomic DNA of the eukaryotic cell, and transfecting the eukaryotic cells with a nucleic acid encoding an enzyme that interacts with the RNA and cleaves the genomic DNA in a site-specific manner, wherein the cells express the RNA and the enzyme, the RNA binding to the complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner. According to one aspect, the enzyme is Cas9 or a modified Cas9 or a homolog of Cas9. According to one aspect, the eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. According to one aspect, the RNA comprises about 10 to about 250 nucleotides. According to one aspect, the RNA comprises about 20 to about 100 nucleotides.
[0010] According to one aspect, a method for altering human cells is provided, comprising transfecting human cells with a nucleic acid encoding RNA complementary to the genomic DNA of eukaryotic cells, and transfecting the human cells with a nucleic acid encoding an enzyme that interacts with the RNA and cleaves the genomic DNA in a site-specific manner, wherein the human cells express the RNA and the enzyme, the RNA binding to the complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner. According to one aspect, the enzyme is Cas9 or a modified Cas9 or a homolog of Cas9. According to one aspect, the RNA comprises about 10 to about 250 nucleotides. According to one aspect, the RNA comprises about 20 to about 100 nucleotides.
[0011] According to one aspect, a method is provided for altering eukaryotic cells at multiple genomic DNA sites, comprising transfecting eukaryotic cells with multiple nucleic acids encoding RNA complementary to different sites on the eukaryotic cell's genomic DNA, and transfecting eukaryotic cells with nucleic acids encoding an enzyme that interacts with the RNA and cleaves the genomic DNA in a site-specific manner, wherein the cells express the RNA and the enzyme, the RNA binding to the complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner. According to one aspect, the enzyme is Cas9. According to one aspect, the eukaryotic cell is a yeast cell, a plant cell, or a mammalian cell. According to one aspect, the RNA comprises about 10 to about 250 nucleotides. According to one aspect, the RNA comprises about 20 to about 100 nucleotides. Attached Figure Description
[0012] The present invention will be further described below with reference to the accompanying drawings, which are only for illustrating the embodiments of the present invention and are not intended to limit the scope of the present invention.
[0013] Figure 1A Genome editing in human cells was performed using engineered type II CRISPR systems (SEQ ID NO:17 and 45-46). Figure 1B The GFP coding sequence integrated into the genome was disrupted by the insertion of a stop codon and a 68 bp genomic fragment at the AAVS1 locus (SEQ ID NO:18). Figure 1C-1 The bar chart describes the efficiency of homologous recombination (HR) induced by T1, T2, and TALEN nucleases at target loci, as measured by flow cytometry (FACS). Figure 1C-2 FACS images and microscopic imaging of targeted cells.
[0014] Figure 2A RNA-guided genome editing of the natural AAVS1 locus in multiple cell types (SEQ ID NO:19). Figure 2B-1 , 2B-2 Total number and location distribution of deletions caused by non-homologous end joining (NHEJ) in 2B-3:293T, K562 and PGP1 iPS cells. Figure 2C : DNA donor structure for homologous recombination at the AAVS1 locus, and the location of sequencing primers used to detect successful targeting events (indicated by arrows). Figure 2D PCR test results three days after transfection. Figure 2E Homologous recombination was confirmed by Sanger sequencing of PCR amplicon (SEQ ID NO:20-21). Figure 2F Imaging of 293T cell-targeted clones.
[0015] Figure 3A-1 and3A-2 Human codon-optimized Cas9 protein and the full sequence of the Cas9 gene insertion (SEQ ID NO:22). Figure 3B : Guide RNA (gRNA) expression protocol based on U6 promoter and predicted RNA transcript secondary structure (SEQ ID NO:23). Figure 3C : Targeting the GN20GG genomic locus and the selected gRNA (SEQ ID NO:24-31).
[0016] Figure 4 The effect of all possible combinations of DNA repair donors, Cas9 protein, and gRNA on homologous recombination efficiency was tested in 293T cells.
[0017] Figures 5A-5B Analysis of gRNA and Cas9-mediated genome editing (SEQ ID NO:19).
[0018] Figures 6A-6B : 293T stable cell lines with different GFP reporter gene constructs (SEQ ID NO:32-34).
[0019] Figure 7 Targeted Figure 1B The reporter was side-linked with a GFP sequence gRNA (in 293T cells).
[0020] Figures 8A-8B : 293T stable cell lines with different GFP reporter gene constructs (SEQ ID NO:35-36).
[0021] Figures 9A-9C Human iPS cells (PGP1) transfected with the construct (SEQ ID NO:19).
[0022] Figures 10A-10B RNA-guided non-homologous end joining (NHEJ) in K562 cells (SEQ ID NO:19).
[0023] Figure 11A-11B RNA-guided non-homologous end joining (NHEJ) in 293T cells (SEQ ID NO:19).
[0024] Figures 12A-12C Homologous recombination was achieved at the endogenous AAVS1 locus using dsDNA donors or short oligonucleotide donors (SEQ ID NO:37-38).
[0025] Figure 13A-1 and 13A-2 and Figure 13BMethods for multiple synthesis, recovery and U6 expression vector cloning of guide gRNAs targeting genes in the human genome (SEQ ID NO:39-41).
[0026] Figures 14A-14D CRISPR-mediated RNA-guided transcriptional activation (SEQ ID NO:42-43).
[0027] Figures 15A-15B gRNA sequence flexibility and its applications (SEQ ID NO:44-46). Detailed Implementation
[0028] According to one aspect, a human codon-optimized version of the Cas9 protein with a C-terminal SV40 nuclear localization signal was synthesized and cloned into a mammalian expression system. Figure 1A and Figure 3A Therefore, Figure 1 illustrates genome editing in human cells using an engineered type II CRISPR system. Figure 1A As shown, RNA-guided gene targeting in human cells involves the co-expression of a Cas9 protein with a C-terminal SV40 nuclear localization signal and one or more guide RNAs (gRNAs) expressed by the human U6 polymerase III promoter. Upon recognition of the target sequence by the gRNA, Cas9 unwinds the DNA double helix and cleaves the double strand, but only if the correct protospacer adjacent motif (PAM) is present at the 3' end. Form GN 20 In principle, any genomic sequence of GG can be targeted. For example... Figure 1B As shown, the GFP coding sequence of genome integration was disrupted by inserting a stop codon and a 68 bp genomic fragment from the AAVS1 locus. The GFP sequence was then restored via homologous recombination (HR) using a suitable donor to produce GFP that could be quantified by FACS. + Cells. The T1 and T2 gRNA target sequences are located within the AAVS1 fragment. Underlined sections indicate the binding sites for the two halves of the TAL effector nuclease heterodimer (TALEN). Figure 1C As shown, the bar charts depict the HR efficiency induced by T1 and T2, and the TALEN-mediated nuclease activity at the target site as measured by FACS. Representative FACS plots and microscopic images of the targeted cells are described below (scale bar at 100 μm). Data are mean ± SEM (N=3).
[0029] According to one approach, to guide Cas9 to cleave sequences of interest, a crRNA-tracrRNA fusion transcript, hereinafter referred to as guide RNA (gRNA), is expressed by the human U6 polymerase III promoter. According to another approach, gRNA is transcribed directly by the cell. This approach advantageously avoids the remodeling of RNA processing mechanisms employed by bacterial CRISPR systems. Figure 1A and Figure 3B (See references (4, 7-9)). According to one aspect, a method for altering genomic DNA is provided, using U6 transcription initiated by a G and PAM (protospacer adjacent motif) sequence – NGG – followed by a 20 bp crRNA target. According to this aspect, the target genomic site is GN… 20 GG form (see Figure 3C ).
[0030] Based on one aspect, a GFP reporter analysis in 293T cells similar to that described above (see reference (10)) was developed. Figure 1B This was used to test the functionality of the genome engineering methods described herein. According to one aspect, a stable cell line was established with a genome-integrated GFP coding sequence interrupted by the insertion of a stop codon and a 68 bp genomic fragment from the AAVS1 locus, resulting in an fluorescence-free expressed protein fragment. Homologous recombination (HR) using an appropriate repair donor restored the normal GFP sequence, allowing for quantification of the generated GFP via flow cytometry-activated cell sorting (FACS). + cell.
[0031] According to one aspect, a method for homologous recombination (HR) is provided. Two gRNAs, T1 and T2, were constructed targeting the inserted AAVS1 fragment (Fig. 1b). Their activities were compared with those of previously described TAL effector nuclease heterodimers (TALENs) targeting the same region (see reference (11)). Successful HR events were observed using all three targeting reagents, along with gene correction rates of approximately 3% and 8% respectively using T1 and T2 gRNAs. Figure 1C This RNA-mediated editing process is particularly rapid, in which the first detectable GFP... + Cells appeared approximately 20 hours post-transfection, compared to approximately 40 hours for AAVS1TALEN. HR was observed only after the simultaneous introduction of the repair donor, Cas9 protein, and gRNA, confirming that all components were required for genome editing. Figure 4Although no obvious toxicity was noted associated with Cas9 / crRNA expression, work with ZFN and TALEN has shown that nicking only one strand further reduces toxicity. Therefore, the Cas9D10A mutant (known to act as an in vitro nicking enzyme) was tested, producing a similar HR but a lower rate of non-homologous end joining (NHEJ) (Figure 5) (see references (4, 5)). Consistent with (4) which shows that the associated Cas9 protein cleaves the two upstream strands of PAM by 6 bp, the NHEJ data confirm that most deletions or insertions occur at the 3' end of the target sequence ( Figure 5B It also confirmed that mutations at targeted genomic sites prevented gRNAs from affecting the HR at that locus, demonstrating that CRISPR-mediated genome editing is sequence-specific. Figure 6 It has been shown that two gRNA targeting sites in the GFP gene, and three additional gRNA targeting fragments from homologous regions of the DNA methyltransferase 3a (DNMT3a) and DNMT3b genes, can sequence-specifically induce significant HR in engineered reporter cell lines. Figure 7 , Figure 8 These results collectively confirm that RNA-guided genome targeting in human cells induces strong HR across multiple target sites.
[0032] In some respects, the natural locus was modified. The gRNA was used to target the AAVS1 locus in the PPP1R12C gene located on chromosome 19, which spans most tissues in 293T, K562, and PGP1 human iPS cells. Figure 2A The gene is widely expressed (see reference (12)) and the results were analyzed by next-generation sequencing of the targeted loci. Therefore, Figure 2 illustrates RNA-guided genome editing of the native AAVS1 locus in multiple cell types. Figure 2A As shown, the T1 (red) and T2 (green) gRNAs target the introns of the PPP1R12C gene at the AAVS1 locus on chromosome 19. Figure 2B As shown, this provides the total number and location of NHEJ-induced deletions in 293T, K562, and PGP1 iPS cells after expression of Cas9 and T1 or T2 gRNA, quantified by next-generation sequencing. Red and green dashed lines demarcate the boundaries of T1 and T2 gRNA target sites. The NHEJ frequencies of T1 and T2 gRNAs were 10% and 25% in 293T cells, 13% and 38% in K562 cells, and 2% and 4% in PGP1 iPS cells, respectively. Figure 2CThe diagram depicts the DNA donor architecture for HR at the AAVS1 locus, and the positions of the sequencing primers (arrows) used to detect successful targeting events. Figure 2D As shown, PCR analysis 3 days after transfection demonstrated that only cells expressing donor, Cas9, and T2gRNA exhibited successful HR events. Figure 2E As shown, successful HR was confirmed by Sanger sequencing of the PCR amplicons, which demonstrated the presence of the expected DNA bases at the donor and insert boundaries. Figure 2F As shown, successful targeted clones of 293T cells selected with puromycin persisted for 2 weeks. Two typical GFP strains are shown. + Microscopic image of the clone (scale bar is 100 micrometers).
[0033] Consistent with the results of GFP reporter analysis, high numbers of NHEJ events were observed at endogenous loci for all three cell types. The two gRNAs, T1 and T2, achieved NHEJ rates of 10% and 25% in 293T cells, respectively, 13% and 38% in K562 cells, and 2% and 4% in PGP1-iPS cells, respectively. Figure 2B No significant toxicity was observed in inducing NHEJ expression of Cas9 and crRNA in any of these cell types. Figure 9 As expected, for both T1 and T2, the NHEJ-mediated deletions were centered on the target site, further validating the sequence specificity of this targeting process. Figure 9 , Figure 10 , Figure 11 The simultaneous introduction of both T1 and T2 gRNAs resulted in a highly efficient deletion of the 19 bp interference fragment. Figure 10 This demonstrates that multiplex editing of genomic loci is feasible using this method.
[0034] According to one aspect, HR is used to integrate dsDNA donor constructs (see reference (13)) or oligomeric donors into the natural AAVS1 locus. Figure 2C Figure 12). Using PCR ( Figure 2D Figure 12) and Sanger sequencing ( Figure 2E Two methods were used to confirm HR-mediated integration. Using puromycin selection, 293T or iPS clones were readily obtained from modified cell pools after more than two weeks. Figure 2F (Figure 12). These results demonstrate that Cas9 can efficiently integrate exogenous DNA at endogenous loci in human cells. Therefore, one aspect of this disclosure includes a method for integrating exogenous DNA into the genome of a cell using homologous recombination and Cas9.
[0035] According to one aspect, an RNA-guided genome editing system is provided that can be readily adapted to modify other genomic sites by simply modifying the sequence of a gRNA expression vector to match compatible sequences at loci of interest. According to this aspect, 190,000 specific gRNA-targetable sequences targeting approximately 40.5% of the exons of genes in the human genome are generated. These target sequences are incorporated into a 200 bp format compatible with multiplex synthesis of DNA arrays (see reference (14)) (Figure 13). According to this aspect, prepared whole-genome references for potential target sites in the human genome and methodologies for multiplex gRNA synthesis are provided.
[0036] According to one aspect, methods for multiple genomic alterations in cells are provided, said methods using one or more RNA / enzyme systems described herein to alter the cell's genome at multiple sites. According to one aspect, the target site perfectly matches the PAM sequence NGG and the 8-12 base "seed sequence" at the 3' end of the gRNA. According to some aspects, a perfect match of the remaining 8-12 bases is not required. According to some aspects, Cas9 will have a single mismatch at the 5' end. According to some aspects, the underlying chromatin structure and epigenetic state of the target locus can affect the efficiency of Cas9 function. According to some aspects, Cas9 homologues with higher specificity are included as useful enzymes. Those skilled in the art will be able to identify or engineer suitable Cas9 homologues. According to one aspect, CRISPR-targetable sequences include those with different PAM requirements (see Reference (9)) or those with directed evolution. According to one aspect, inactivation of one of the Cas9 nuclease domains increases the HR to NHEJ ratio and can reduce toxicity ( Figure 3A (Figure 5)(4, 5), and inactivating the two domains allows Cas9 to act as a retargetable DNA-binding protein. Embodiments of this disclosure have broad applications in synthetic biology (see references (21, 22)), direct and multiple perturbations of gene networks (see references (13, 23)), and targeted ex vivo (see references (24-26)) and in vivo gene therapy (see reference (27)).
[0037] According to some aspects, a “re-engineerable organism” is provided as a model system for biological discovery and in vivo screening. According to one aspect, a “re-engineerable mouse” with an inducible Cas9 transgene is provided, and local delivery of gRNA libraries targeting multiple genes or regulatory elements (e.g., using adeno-associated virus) allows screening for mutations that lead to tumor development in target tissue types. The use of Cas9 homologues or nuclease-free variants with effector domains (e.g., activators) allows for multiple activation or repression of genes in vivo. According to this aspect, factors that enable phenotypes such as tissue regeneration, transdifferentiation, etc., can be screened. According to some aspects, (a) the use of DNA arrays enables multiple synthesis of defined gRNA libraries (see Figure 13); and (b) the use of numerous non-viral or viral delivery methods to package and deliver small-sized gRNAs (see Figure 3b).
[0038] According to one aspect, the lower toxicity observed with "nicking enzymes" for genome engineering applications is achieved by inactivating one of the Cas9 nuclease domains: either by creating a nick in the DNA strand paired with RNA bases or by creating a nick in its complementary sequence. Inactivating both domains allows Cas9 to function as a retargetable DNA-binding protein. According to one aspect, Cas9-retargetable DNA-binding proteins are attached to...
[0039] (a) Transcriptional activation or repression domains used to regulate the expression of target genes, including but not limited to chromatin remodeling, histone modification, silencing, insulation, and direct interaction with transcriptional mechanisms; (b) Nuclease domains such as FokI activate “highly specific” genome editing, depending on the dimerization of adjacent gRNA-Cas9 complexes; (c) Fluorescent proteins used to visualize genomic loci and chromosome dynamics; or (d) Other fluorescent molecules, such as protein or nucleic acid bound organic fluorophores, quantum dots, molecular beacons and echo probes or molecular beacon replacements; (e) Enables programmable manipulation of the whole genome 3D architecture of multivalent ligand-binding protein domains.
[0040] According to one aspect, transcriptional activation and repression components can employ natural or synthetic vertical (orthogonal) CRISPR systems, allowing gRNAs to bind only to the activator or inhibitor class of Cas. This allows a large set of gRNAs to be tuned to multiple targets.
[0041] In some respects, the use of RNA offers more capabilities than mRNA, partly due to its smaller size—100 versus 2000 nucleotides in length, respectively. This is particularly valuable when nucleic acid delivery is size-constrained (e.g., in viral packaging). This enables multiple instances of cleavage, nicking, activation, or inhibition—or combinations thereof. The ability to easily target multiple regulatory targets allows for coarse or fine tuning or regulation of networks without being constrained to the downstream of the natural regulatory loops of specific regulators (e.g., the four mRNAs used in reprogramming fibroblasts into IPSCs). Examples of multiple applications include: 1. Establish (primary and secondary) histocompatibility alleles, haplotypes, and genotypes for human (or animal) tissue / organ transplantation. This involves generating, for example, a set of gRNAs in HLA-homozygous cell lines or humanized animal breeds—or capable of superimposing such HLA alleles onto another desired cell line or breed.
[0042] 2. Mutations in multiple cis-regulatory elements (CREs = signals used for transcription, splicing, translation, RNA and protein folding, degradation, etc.) within a single cell (or cell set) can be used to efficiently study the regulatory interactions of complex sets, which can occur in specific episodes of normal development or pathology, synthesis or pharmacology. In one respect, CREs are (or can be made) somewhat perpendicular to each other (i.e., low crosstalk), allowing many to be tested in a single setup, for example, in expensive animal embryonic time series. An exemplary application is RNA fluorescence in situ sequencing (FISSeq).
[0043] 3. Multiple combinations of CRE mutations and / or epigenetic activation or inhibition of CREs can be used to alter or adapt iPSCs or ESCs or other stem cell or non-stem cell types or combinations thereof for use in organ-on-a-chip, or for other cell and organ cultures for purposes such as testing drugs (small molecules, proteins, RNA, cells, animal, plant or microbial cells, aerosols and other delivery methods), transplantation strategies, personalization strategies, etc.
[0044] 4. Preparation of multiple mutant human cells for use in diagnostic tests (and / or DNA sequencing) in medical genetics. In relation to the chromosomal location and structure of human genomic alleles (or epigenetic markers) that can affect the accuracy of clinical genetic diagnosis, it is important to have alleles present in the correct location on the reference genome, rather than in ectopic (aka transgenic) locations or in individual fragments of synthetic DNA. One implementation is a series of separate cell lines, each diagnosing a human SNP, or structural variant. Alternatively, one implementation includes multiple recombinant alleles in the same cells. In some cases, multiple variations in one gene (or multiple genes) would be ideal under the assumption of individual detection. In other cases, specific haplotype combinations of alleles allow testing sequencing (genotyping) methods that precisely establish haplotype phases (i.e., whether one or two copies of a gene in an individual or somatic cell type are affected).
[0045] 5. Engineered Cas+gRNA systems can be used in microbial, plant, animal, or human cells to target repetitive elements or endogenous viral elements to reduce harmful transposons or aid in sequencing or other analytical genomic / transcriptomic / proteomic / diagnostic tools (where nearly identical copies may be problematic).
[0046] The following references, identified by number in the foregoing sections, are incorporated herein by reference in their entirety.
[0047] 1. B. Wiedenheft, SH Sternberg, JA Doudna, Nature 482, 331 (Feb16, 2012).
[0048] 2. D. Bhaya, M. Davison, R. Barrangou, Annual review of genetics 45,273 (2011).
[0049] 3. MP Terns, RM Terns, Current opinion in microbiology 14, 321 (Jun, 2011).
[0050] 4. M. Jinek et al. , Science 337, 816 (Aug 17, 2012).
[0051] 5. G. Gasiunas, R. Barrangou, P. Horvath, V. Siksnys, Proceedings of the National Academy of Sciences of the United States of America109, E2579(September 25, 2012).
[0052] 6. R. Sapranauskas et al. , 1999 . Nucleic acids research 39, 9275 (Nov,2011).
[0053] 7. TR Brummelkamp, R Bernards, R Agami, Science 296, 550 (Apr19, 2002).
[0054] 8. M. Miyagishi, K. Taira, Nature biotechnology 20, 497 (May, 2002).
[0055] 9. E. Deltcheva et al. , 1999 . Nature 471, 602 (Mar. 31, 2011).
[0056] [ PMC free article ] [ PubMed ] 10. Zou J, Mali P, Huang X, Dowey SN, Cheng L, Blood 118, 4599(Oct 27, 2011).
[0057] 11. AND Sanjana et al. , 1999 . Nature protocols 7, 171 (January, 2012).
[0058] 12. JH Lee et al. , 1999 . PLoS Genet 5, e1000718 (Nov, 2009).
[0059] [ PubMed ] 13. D. Hockemeyer et al. , 1999 . Nature biotechnology 27, 851 (September, 2009).
[0060] 14. S. Kosuri et al. , 1999 . Nature biotechnology 28, 1295 (Dec, 2010).
[0061] [ PubMed ] 15. Pattanayak V, Ramirez CL, Joung JK, Liu DR, Nature methods8, 765 (Sep, 2011).
[0062] 16. NM King, O. Cohen-Haguenauer, Molecular therapy : the journal of the American Society of Gene Therapy 16, 432 (Mar, 2008).
[0063] 17. YG Kim, J. Cha, S. Chandrasegaran, Proceedings of the National Academy of Sciences of the United States of America 93, 1156 (Feb 6, 1996).
[0064] 18. EJ Rebar, CO Turkey, Science 263, 671 (Feb 4, 1994).
[0065] 19. J. Boch et al. , Science 326, 1509 (Dec 11, 2009).
[0066] 20. MJ Moscow, AJ Bogdanove, Science 326, 1501 (Dec 11, 2009).
[0067] 21. AS Khalil, JJ Collins, Nature reviews. Genetics 11, 367(May, 2010).
[0068] 22. PE Purnick, R. Weiss, Nature reviews. Molecular cell biology 10, 410 (June, 2009).
[0069] 23. J. Zou et al. , Cell stem cell 5, 97 (Jul 2, 2009).
[0070] 24. N. Holt et al. , Nature biotechnology 28, 839 (Aug, 2010).
[0071] 25. FD Urnov et al. , Nature 435, 646 (June 2, 2005).
[0072] 26. A. Lombardo et al. , Nature biotechnology 25, 1298 (Nov, 2007).
[0073] 27. H. Li et al. , Nature 475, 217 (Jul 14, 2011).
[0074] The following embodiments are given as representative examples of this disclosure. These embodiments should not be construed as limiting the scope of this disclosure, as these and other equivalent embodiments will be apparent from the present disclosure, the accompanying drawings, and the appended claims.
[0075] Example I
[0076] Type II CRISPR-Cas system
[0077] According to one aspect, embodiments of this disclosure utilize short RNA to recognize exogenous nucleic acids for nuclease activity in eukaryotic cells. According to certain aspects of this disclosure, eukaryotic cells are modified to include nucleic acids encoding one or more short RNAs and one or more nucleases within their genome, the nucleases being inactivated by the binding of the short RNA to a target DNA sequence. According to certain aspects, exemplary short RNA / enzyme systems, such as CRISPR / CRISPR-associated (Cas) systems using short RNA to guide the degradation of exogenous nucleic acids, can be recognized within bacteria or archaea. CRISPR (“regularly clustered short palindromic repeats”) defense involves: acquiring a new target “spacer” from invading viral or plasmid DNA and integrating it into a CRISPR locus; expressing and processing a short guiding CRISPR RNA (crRNA) composed of spacer repeat units; and cleaving nucleic acids (most commonly DNA) complementary to the spacer.
[0078] Three classes of CRISPR systems are generally known and referred to as type I, type II, or type III. According to one aspect, the enzyme particularly useful for cleaving dsDNA according to this disclosure is the single-effects enzyme Cas9, typically type II (see reference (1)). In bacteria, the type II effector system consists of a long pre-crRNA transcribed from a CRISPR locus containing spacers, a multifunctional Cas9 protein, and tracrRNA, which is important for gRNA processing. The tracrRNA hybridizes to a repeating region that separates the pre-crRNA spacers, initiates dsRNA cleavage via endogenous RNAe III, followed by a second cleavage event via Cas9 within each spacer, producing mature crRNA that remains associated with both tracrRNA and Cas9. According to one aspect, eukaryotic cell engineering of this disclosure is generally performed to avoid the use of RNAe III and crRNA processing. See reference (2).
[0079] According to one aspect, the enzymes disclosed herein, such as Cas9, unwind the DNA double helix and search for a matching crRNA sequence to cleave. Target recognition occurs after complementarity detection between the "protospacer" sequence in the target DNA and the remaining spacer sequences in the crRNA. Importantly, Cas9 cleaves the DNA only if the correct protospacer adjacent motif (PAM) is also present at the 3' end. According to some aspects, different protospacer adjacent motifs can be used. For example, *Streptococcus pyogenes* (…). S. pyogenes The system requires an NGG sequence, where N can be any nucleotide. Streptococcus thermophilus (… S. thermophilus Type II systems require NGNG (see reference (3)) and NNAGAAW (see figure (4)), respectively, while different Streptococcus mutans ( S. mutans The system is tolerant of NGG or NAAR (see reference (5)). Bioinformatics analysis has generated a large database of CRISPR loci in various bacteria, which can be used to identify additional useful PAMs and expand the group of CRISPR-targetable sequences (see references (6, 7)). In Streptococcus thermophilus, Cas9 produces blunt-ended double-strand breaks 3 bp before the 3' end of the protospacer sequence (see reference (8)), by two catalytic domains in the Cas9 protein: the HNH domain that cuts the complementary strand of DNA and the RuvC-like domain that cuts the non-complementary strand (see reference (8)). Figure 1A The process mediated by (and Figure 3). Although the Streptococcus pyogenes system was not characterized to the same level of precision, DSB formation also occurred toward the 3' end of the original spacer sequence. If one of the two nuclease domains is inactivated, Cas9 will act as a nicking enzyme in vitro (see reference (2)) and in human cells (see Figure 5).
[0080] According to one aspect, in eukaryotic cells, the specificity of gRNA-guided Cas9 cleavage serves as a mechanism for genome engineering. According to another aspect, gRNA hybridization does not need to be 100% for the enzyme to recognize the gRNA / DNA hybrid and affect cleavage. Some off-target activities can occur. For example, the *Streptococcus pyogenes* system tolerates the first 6 base mismatches in the 20 bp mature spacer sequence in vitro. According to another aspect, higher tightness may be advantageous in vivo when a potential off-target site match (the last 14 bp) NGG is present within the human reference genome of the gRNA. The effects of mismatches and enzyme activity are generally described in references (9), (2), (10), and (4).
[0081] Specificity can be improved in several ways. AT-rich target sequences can have fewer off-target sites when the interference is sensitive to the melting temperature of the gRNA-DNA hybrid. Careful selection of target sites to avoid spurious sites with at least 14 bp matching sequences elsewhere in the genome can improve specificity. Using Cas9 variants that require longer PAM sequences can reduce the frequency of off-target sites. Directed evolution can improve Cas9 specificity to a level sufficient to completely eliminate off-target activity, ideally requiring a perfect 20 bp gRNA that matches the minimum PAM. Therefore, modification of the Cas9 protein is a representative implementation of this disclosure. In this way, the new method allows for multiple rounds of evolution in a very short time (see reference (11)) and predictably. CRISPR systems that can be used in this disclosure are described in references (12, 13).
[0082] Example II
[0083] plasmid construction
[0084] The Cas9 gene sequence was codon-optimized and assembled from nine 500 bp g-blocks ordered from IDT via hierarchical fusion PCR assembly. This is for a type II CRISPR system engineered for human cells. Figure 3AThe expression form and full sequence of the Cas9 gene insert are shown. The RuvC-like and HNH motifs, as well as the C-terminal SV40 NLS, are highlighted in blue, brown, and orange, respectively. Cas9_D10A was constructed using a similar method. The resulting full-length product was cloned into the pcDNA3.3-TOPO vector (Invitrogen). The target gRNA expression construct was purchased directly from IDT as individual 455 bp gBlocks and either cloned into the pCR-BluntII-TOPO vector (Invitrogen) or amplified by PCR. Figure 3B An expression protocol for guide RNA based on the U6 promoter and the predicted RNA transcript secondary structure are illustrated. The use of the U6 promoter constrains the first position of the RNA transcript to "G", thus enabling the targeting of the GN form. 20 All genomic loci of GG. Figure 3C The seven gRNAs used are shown.
[0085] By using a stop codon and a 68 bp AAVS1 fragment (or its mutants; see Figure 6 ), or a 58 bp fragment from the DNMT3a and DNMT3b genomic loci (see Figure 8 Fusion PCR assembly of the GFP sequence (assembled in an EGIP lentiviral vector (plasmid #26777) from Addgene) was used to construct vectors involving fragmented GFP for HR reporter assays. These lentiviral vectors were then used to establish stable GFP reporter lines. TALENs used in this study were constructed using the protocol described in (14). All DNA reagents developed in this study were obtained from Addgene.
[0086] Example III
[0087] Cell culture
[0088] PGP1 iPS cells were maintained in Matrigel (BD Biosciences)-coated plates using mTeSR1 (Stemcell Technologies). Cultures were passaged every 5–7 days using TrypLE Express (Invitrogen). K562 cells were grown and maintained in RPMI (Invitrogen) containing 15% FBS. HEK 293T cells were cultured in Dulbecco's modified Eagle's medium (DMEM, Invitrogen) supplemented with high glucose and 10% fetal bovine serum (FBS, Invitrogen), penicillin / streptomycin (pen / strep, Invitrogen), and non-essential amino acids (NEAA, Invitrogen). All cells were maintained in a humidified incubator at 37°C and 5% CO2.
[0089] Example IV
[0090] Gene targeting of PGP1 iPS, K562, and 293T
[0091] Prior to nuclear transfection, PGP1 iPS cells were cultured for 2 h in a Rho kinase (ROCK) inhibitor (Calbiochem). Cells were harvested using TrypLE Express (Invitrogen) and resuspended 2 × 10⁶ cells in P3 reagent (Lonza) containing 1 μg Cas9 plasmid, 1 μg gRNA, and / or 1 μg DNA donor plasmid. 6 Cells were transfected according to the manufacturer's instructions (Lonza). Cells were then plated on mTeSR1-spread plates in mTeSR1 medium supplemented with ROCK inhibitors for the first 24 hours. For K562, cells were resuspended 2 × 10⁶ cells in SF reagent (Lonza) containing 1 μg Cas9 plasmid, 1 μg gRNA, and / or 1 μg DNA donor plasmid. 6 Cells were transfected according to the manufacturer's instructions (Lonza). For 293T cells, Lipofectamine 2000 was used to transfect 0.1 × 10⁶ cells with 1 μg Cas9 plasmid, 1 μg gRNA, and / or 1 μg DNA donor plasmid, following the manufacturer's protocol. 6 Each cell. The DNA donor used for endogenous AAVS1 targeting is dsDNA donor ( Figure 2C () or 90mer oligonucleotides. The former has flanking short homologous arms and an SA-2A-puromycin-CaGGS-eGFP box to enrich successfully targeted cells.
[0092] The following evaluation assesses targeting efficiency. Cells were harvested 3 days after nuclear transfection, and ~1×10⁶ cells were extracted using prepGEM (ZyGEM). 6 Genomic DNA from individual cells was used. Guided PCR was employed to amplify target regions containing cellular genomic DNA, and amplicon was deep sequenced using a MiSeq Personal Sequencer (Illumina) with coverage >200,000 reads. Sequencing data were analyzed to assess NHEJ efficiency. The reference AAVS1 sequence for analysis was: CACTTCAGGACAGCATGTTTGCTGCCTCCAGGGATCCTGTGTCCCCGAGCTGGGACCACCTTATATTCCCAGGGCCGGTTAATGTGGCTCTGGTTCTGGGTACTTTTATCTGTCCCCTCCACCCCACAGTGGGG CCACTAGGGACAGGATTGGTGACAGAAAAGCCCCATCCTTAGGCCTCCTCCTTCCTAGTCTCCTGATATTGGGTCTAACCCCCACCTCCTGTTAGGCAGATTCCTTATCTGGTGACACACCCCCATTTCCTGGA The PCR primers used to amplify target regions of the human genome are: AAVS1-R CTCGGCATTCCTGCTGAACCGCTCTTCCGATCTacaggaggtgggggttagac AAVS1-F.1 ACACTCTTTCCCTACACGACGCTCTTCCGATCTCGTGATtatattcccagggccggtta AAVS1-F.2 ACACTCTTTCCCTACACGACGCTCTTCCGATCTACATCGtatattcccagggccggtta AAVS1-F.3 ACACTCTTTCCCTACACGACGCTCTTCCGATCTGCCTAAtatattcccagggccggtta AAVS1-F.4 ACACTCTTTCCCTACACGACGCTCTTCCGATCTTGGTCAtatattcccagggccggtta AAVS1-F.5 ACACTCTTTCCCTACACGACGCTCTTCCGATCTCACTGTtatattcccagggccggtta AAVS1-F.6 ACACTCTTTCCCTACACGACGCTCTTCCGATCTATTGGCtatattcccagggccggtta AAVS1-F.7 ACACTCTTTCCCTACACGACGCTCTTCCGATCTGATCTGtatattcccagggccggtta AAVS1-F.8 ACACTCTTTCCCTACACGACGCTCTTCCGATCTTCAAGTtatattcccagggccggtta AAVS1-F.9 ACACTCTTTCCCTACACGACGCTCTTCCGATCTCTGATCtatattcccagggccggtta AAVS1-F.10 ACACTCTTTCCCTACACGACGCTCTTCCGATCTAAGCTAtatattcccagggccggtta AAVS1-F.11 ACACTCTTTCCCTACACGACGCTCTTCCGATCTGTAGCCtatattcccagggccggtta AAVS1-F.12 ACACTCTTTCCCTACACGACGCTCTTCCGATCTTACAAGtatattcccagggccggtta In order to analyze in Figure 2C The primers used in the HR event using a DNA donor were: HR_AAVS1-F CTGCCGTCTCTCTCCTGAGT HR_Puro-R GTGGGCTTGTACTCGGTCAT Example V Bioinformatics methods for calculating human exon CRISPR targets and methods for their multiple synthesis The following is a set of gRNA gene sequences that are maximally targeted at specific locations in human exons but minimally targeted at other locations in the genome. According to one aspect, maximally effective targeting by gRNA is achieved through a 23 nt sequence (the 20 nt at the 5' end is precisely complementary to the required position, while the three 3' ends must be in the form of NGG). Additionally, the 5' end nt must be G to establish the pol-III transcription start site. However, according to (2), mismatches of the six 5' end nts of the 20 bp gRNA targeting the genomic target do not eliminate Cas9-mediated cleavage as long as the last 14 nts are properly paired, but mismatches of the eight 5' end nts paired along the last 12 nts eliminate Cas9-mediated cleavage, while seven 5' end mismatches and 13 3' pairs were not tested. One condition for conservatism regarding off-target effects is that the case of 7 5' end mismatches (similar to the case of 6) allows cleavage, thus the pairing of the 13 nts at the 3' end is sufficient for cleavage. To identify CRISPR target sites within human exons that should be cleavable and without off-target nicks, all 23 bp sequences of the form 5'-GBBBB BBBBBBBBBB BBBBB NGG-3' (form 1) were examined, where B represents the base at the exon location. For this, no sequence of the form 5'-NNNNN NNBBB BBBBB BBBBB NGG-3' (form 2) was present at any other location in the human genome. Specifically, (i) BED files of coding region locations for all RefSeq genes in the GRCh37 / hg19 human genome were downloaded from the UCSC Genome Browser (15-17). The coding exon locations in this BED file included a set of 346089 mappings of RefSeq mRNA accession numbers in the hg19 genome. However, some RefSeq mRNA accession numbers mapped to multiple genomic locations (possible gene duplication), and many accession numbers mapped to subsets of the same set of exon locations (multiple isotypes of the same gene). To clearly distinguish duplicate gene instances and consolidate multiple references for the same genomic exon instances via multiple RefSeq isotype accessions, (ii) a unique numerical suffix was added to RefSeq accessions with multiple genomic locations, and (iii) the mergeBed function of BEDTools (18) (v2.16.2-zip-87e3926) was used to consolidate overlapping exon locations into merged exon regions. These steps reduced the initially set 346,089 RefSeq exon locations to 192,783 different genomic fragments. The hg19 sequence of all merged exon regions was downloaded using the UCSC Table Browser, and 20 bp padding was added to each end.(iv) Using custom Perl code, 1,657,793 instances of form 1 were identified within these exon sequences. (v) These sequences were then filtered for off-target occurrences of form 2: For each merged exon form 1 target, a 13 bp specific (B) "core" sequence was extracted from the 3' end, and for each core, four 16 bp sequences 5'-BBBBBBBBBBBNGG-3' (N=A, C, G, and T) were generated, and the entire hg19 genome was searched for exact matches to these 6,631,172 sequences using Bowtie version 0.12.8 (19) with parameters -1 16 –v 0 –k 2. Any exon target site with more than one match was rejected. Note that since any specific 13 bp core sequence immediately following the NGG sequence only confers an average of 15 bp specificity, there should be an average of ~5.6 matches to the extended core sequence in random ~3 Gb sequences (two strands). Therefore, the majority of the 1,657,793 initially identified targets were rejected; however, 189,864 sequences passed this filter. These were groups of CRISPR-targetable exon sites in the human genome. Of the 189,864 target sequences in 78,028 merged exon regions (representing approximately 40.5% of the total 192,783 merged human exon regions), each exhibited a multiplicity of approximately 2.4 sites per target exon region. To assess targeting at the gene level, clustered Ref Seq mRNA mapping was performed such that any two overlapping merged exon regions' RefSeq accession numbers (including gene repeats distinguished in (ii)) were counted as single gene clusters, with 17,104 of the 189,864 exon-specific CRISPR sites targeting 18,872 gene clusters (approximately 90.6% of all gene clusters), each exhibiting a multiplicity of approximately 11.1. (Note that although these gene clusters will fold multiple isotypes of RefSeq mRNA accessions representing a single transcribed gene into a single entity, they will also fold overlapping different genes as well as genes with antisense transcripts.) At the level of the original RefSeq accession, 189,864 of the 30,563 (~69.9%) of the 43,726 mapped RefSeq accessions (including distinguishable gene repetitions) targeted exon regions with a multiplicity of ~6.2 sites per targeted mapped RefSeq accession.
[0093] According to one aspect, the database can be refined by correlating performance with factors such as base composition and the secondary structure of the two gRNAs and genomic targets (20, 21), as well as the epigenetic status of these targets in human cell lines (for which this information is available) (22).
[0094] Example VI
[0095] Multiple synthesis
[0096] The target sequence is merged into a 200bp format, which is compatible with multiplex synthesis on DNA arrays (23, 24). According to one aspect, this method allows for the retrieval of specific gRNA sequences or pools of gRNA sequences targeted from DNA array-based oligonucleotide pools, and their rapid cloning into co-expression vectors. Figure 13A Specifically, a 12k oligonucleotide pool from CustomArray Inc. was synthesized. Furthermore, gRNA selection from this library was successfully retrieved. Figure 13B We observed an error rate of approximately 4 mutations per 1000 bp of synthesized DNA.
[0097] Example VII
[0098] RNA-guided genome editing requires Cas9 and guide RNA for successful targeting.
[0099] Used in the description Figure 1B The GFP reporter assay detects all possible combinations of DNA donors, Cas9 protein, and gRNA to assess the ability to successfully complete HR (in 293T). Figure 4 As shown, GFP+ cells were observed only when all three components were present, confirming that these CRISPR components are essential for RNA-guided genome editing. Data are mean + / - SEM (N=3).
[0100] Example VIII
[0101] Analysis of gRNA and Cas9-mediated genome editing
[0102] Using (A) as described above, the results are shown in Figure 5A (A) GFP reporter assay and (B) deep sequencing of targeted loci (in 293T) examined the CRISPR-mediated genome editing process, and the results were presented in Figure 5BAs shown in Figure 5, the D10A mutation of Cas9 was tested, which has been shown in earlier reports to act as a nicking enzyme in in vitro assays. Both Cas9 and Cas9D10A successfully completed HR at almost the same rate. However, deep sequencing confirmed that while Cas9 exhibited strong NHEJ at the targeted locus, the D10A mutation had a significantly reduced NHEJ rate (as expected from its presumed ability to only nick the DNA). Furthermore, consistent with known Cas9 protein biochemistry, the NHEJ data confirmed that most base pair deletions or insertions occurred near the 3' end of the target sequence: the peak was ~3-4 bases upstream of the PAM site, with a median deletion frequency of ~9-10 bp. Data are mean ± SEM (N=3).
[0103] Example IX
[0104] RNA-guided genome editing is target sequence specific.
[0105] Similar to Figure 1B The GFP reporter assay described in [the original text] was used to develop three stable 293T lines, each with a different GFP reporter construct. These were distinguished by the sequence of the AAVS1 fragment insertion (e.g., [the original text]). Figure 6 (As shown in the diagram). One cell line carried the wild-type fragment, while the other two lines had mutations at 6 bp (highlighted in red). Each line was then targeted with one of four reagents: a GFP-ZFN pair, which could target all cell types because its target sequence was in the flanking GFP fragment and therefore present along with the cell line; an AAVS1 TALEN, which could only effectively target the wt-AAVS1 fragment because the mutations in the other two lines should prevent the left TALEN from binding to their sites; a T1 gRNA, which could also effectively target only the wt-AAVS1 fragment because its target site was also disrupted in the two mutant lines; and finally, a T2 gRNA, which should be able to target all three cell lines because, unlike the T1 gRNA, its target site was unchanged across the three lines. ZFN modified all three cell types, AAVS1 TALEN and T1 gRNA targeted only the wt-AAVS1 cell type, while T2 gRNA successfully targeted all three cell types. These results together confirm that guide RNA-mediated editing is target sequence specific. The data is the mean plus / - SEM (N=3).
[0106] Example X
[0107] Guide RNAs targeting GFP sequences enable strong genome editing.
[0108] In addition to the two gRNAs targeting AAVS1 insertion, tests were conducted on... Figure 1B Two additional gRNAs targeting the flanking GFP sequence of the reporter (described in 293T). For example... Figure 7 As shown, these gRNAs also achieve strong HR at this engineered locus. Data are mean ± SEM (N=3).
[0109] Example XI
[0110] RNA-guided genome editing is target sequence specific. It also showed similar targeting efficiency to ZFN or TALEN. Similar to Figure 1B The GFP reporter assay described in [the original text] was used to develop two stable 293T lines, each with a different GFP receptor construct. These were distinguished by fragment insertion sequences (e.g., [the original text]). Figure 8 (As shown). One cell line carries a 58 bp fragment from the DNMT3a gene, while another carries a homologous 58 bp fragment from the DNMT3b gene. Sequence differences are highlighted in red. Each cell line was then targeted with one of the following six reagents: a GFP-ZFN pair, which can target all cell types because its target sequence is in the flanking GFP fragment and is therefore present with the cell line; a TALEN pair, which potentially targets either the DNMT3a or DNMT3b fragment; a gRNA pair, which can potentially target only the DNMT3a fragment; and finally, a gRNA, which should potentially target only the DNMT3b fragment. Figure 8 As shown, ZFN modified all three cell types, while TALEN and gRNA targeted only their respective targets. Furthermore, targeting efficiencies were comparable across six targeting agents. These results together confirm that RNA-guided editing is target sequence specific and exhibits targeting efficiencies similar to ZFN or TALEN. Data are mean ± SEM (N=3).
[0111] Example XII
[0112] RNA-guided NHEJ in human iPS cells
[0113] Used in Figure 9The left image shows the construct nuclear transfection of human iPS cells (PGP1). Four days after nuclear transfection, the NHEJ rate was measured by deep sequencing, assessing genomic deletion and insertion rates at double-strand breaks (DSBs). Image 1: Deletion rate detected at target regions. Red dashed line: Boundary of T1 RNA target site; Green dashed line: Boundary of T2 RNA target site. Deletion occurrence at each nucleotide position is plotted in black, and the deletion rate is calculated as the percentage of reads carrying the deletion. Image 2: Insertion rate detected at target regions. Red dashed line: Boundary of T1 RNA target site; Green dashed line: Boundary of T2 RNA target site. Insertion occurrence at genomic positions where the first insertion link was detected is plotted in black, and the insertion rate is calculated as the percentage of reads carrying the insertion. Image 3: Deletion size distribution. The frequency of different deletion sizes is plotted across the entire NHEJ population. Image 4: Insertion size distribution. The frequency of different insertion sizes is plotted across the entire NHEJ population. iPS targeting via two gRNAs is effective (2-4%), sequence-specific (as shown by shifting at the NHEJ deletion assignment site), and this further confirms... Figure 4 The results, based on NGS analysis, also showed that Cas9 protein and gRNA are required at the target locus for NHEJ events.
[0114] Example XIII
[0115] RNA-guided NHEJ in K562 cells
[0116] Used in Figure 10The left image shows the construct nuclear transfection of K562 cells. Four days after nuclear transfection, the NHEJ rate was measured by deep sequencing, assessing genomic deletion and insertion rates at the DSB. Image 1: Deletion rate detected at the target region. Red dashed line: Boundary of the T1 RNA target site; Green dashed line: Boundary of the T2 RNA target site. Deletion occurrence at each nucleotide position is plotted in black, and the deletion rate is calculated as the percentage of reads carrying the deletion. Image 2: Insertion rate detected at the target region. Red dashed line: Boundary of the T1 RNA target site; Green dashed line: Boundary of the T2 RNA target site. Insertion occurrence at the genomic position where the first insertion link was detected is plotted in black, and the insertion rate is calculated as the percentage of reads carrying the insertion. Image 3: Deletion size distribution. The frequency of different deletion sizes is plotted across the entire NHEJ population. Image 4: Insertion size distribution. The frequency of different insertion sizes is plotted across the entire NHEJ population. Targeting K562 via two gRNAs is effective (13–38%) and sequence-specific (as shown by shifting at the NHEJ deletion assignment site). Importantly, as shown by the peaks in the observed deletion size frequency histogram, the simultaneous introduction of T1 and T2 guide RNAs resulted in highly efficient deletion of a disruptive 19 bp fragment, indicating that multiple editing of genomic loci is also convenient using this method.
[0117] Example XIV
[0118] RNA-guided NHEJ in 293T cells
[0119] Used in Figure 11 The left image shows the construct transfected into 293T cells. Four days after nuclear transfection, the NHEJ rate was measured by deep sequencing, assessing genomic deletion and insertion rates at the DSB. Image 1: Deletion rate detected at the target region. Red dashed line: Boundary of the T1 RNA target site; Green dashed line: Boundary of the T2 RNA target site. Deletion occurrence at each nucleotide position is plotted in black, and the deletion rate is calculated as the percentage of reads carrying the deletion. Image 2: Insertion rate detected at the target region. Red dashed line: Boundary of the T1 RNA target site; Green dashed line: Boundary of the T2 RNA target site. Insertion occurrence at the genomic position where the first insertion link was detected is plotted in black, and the insertion rate is calculated as the percentage of reads carrying the insertion. Image 3: Deletion size distribution. The frequency of different deletion sizes is plotted across the entire NHEJ population. Image 4: Insertion size distribution. The frequency of different insertion sizes is plotted across the entire NHEJ population. Targeting 293T via two gRNAs is effective (10-24%) and sequence-specific (as shown by shifting at the NHEJ deletion assignment site).
[0120] Example XV
[0121] Use dsDNA donors or short oligonucleotide donors
[0122] HR at the endogenous AAVS1 locus
[0123] like Figure 12A As shown, PCR screening (refer to...) Figure 2C It was confirmed that 21 / 24 randomly selected 293T clones were successfully targeted. Figure 12B As shown, similar PCR screening confirmed that 3 / 7 randomly selected PGP1-iPS clones were successfully targeted. Figure 12C As shown, short 90-mer oligonucleotides can also be strongly targeted at the endogenous AAVS1 locus (K562 cells are shown here).
[0124] Example XVI
[0125] Multiple synthesis of guide RNA-targeting genes in the human genome Methods for retrieving and cloning the U6 expression vector A resource that generates unique gRNA sites targeting approximately 190k bioinformatics from all exons of ~40.5% of the genes in the human genome. For example... Figure 13A As shown, gRNA target sites are incorporated into a 200bp format, which is compatible with multiplex synthesis on a DNA array. Specifically, this design allows for (i) targeted retrieval of specific gRNA targets or pools of gRNA targets from the oligonucleotide pool of the DNA array (via three consecutive rounds of nested PCR schematically shown in the figure); and (ii) rapid cloning into a common expression vector, which, after linearization of the incorporated receiver mediated by the Gibson component using the AflII site as the gRNA insert fragment, is shown below. Figure 13B As shown, this method is used to target and retrieve 10 unique gRNAs from a 12k oligonucleotide library synthesized by CustomArray Inc.
[0126] Example XVII
[0127] CRISPR-mediated RNA-guided transcriptional activation
[0128] The CRISPR-Cas system possesses the adaptive immune defense system and the ability to "cut" invading nucleic acids in bacteria. On one hand, the CRISPR-Cas system has been engineered to function in human cells and to "cut" genomic DNA. This is achieved by guiding the Cas9 protein (which has nuclease function) to a short guide RNA that is complementary to a target sequence in the guide RNA's spacer. The ability to "cut" DNA enables numerous applications related to genome editing, as well as targeted genome regulation. For this purpose, the Cas9 protein is mutated to render it nuclease-free by introducing a mutation, which is predicted to cancel the interaction with Mg... 2+ (It is known to be important for the function of RuvC-like and HNH-like nuclease domains) binding: specifically, introducing a combination of D10A, D839A, H840A, and N863A mutations. The resulting Cas9 nuclease-free protein (as confirmed by sequencing analysis through its inability to cleave DNA) and referred to below as Cas9R-H-, is then bound to the transcription activation domain (here, VP64), enabling the CRISPR-Cas system to function as an RNA-guided transcription factor (see...). Figure 14 The Cas9R-H-+VP64 fusion enables RNA-guided transcriptional activation at the two reporter sites shown. Specifically, both FACS analysis and immunofluorescence imaging demonstrated that this protein enables the gRNA to sequence-specifically target the corresponding reporter, and furthermore, as assessed by expression of dTomato fluorescent protein, the resulting transcriptional activation is at a level similar to that induced by conventional TALE-VP64 fusion proteins.
[0129] Example XVIII
[0130] gRNA sequence flexibility and its applications
[0131] The flexibility of the gRNA backbone sequence to the designed sequence insertion was determined by systematically measuring the range of random sequence insertions at the 5′, middle, and 3′ ends of the gRNA sequence. Specifically, 1bp, 5bp, 10bp, 20bp, and 40bp inserts were created at the 5′, middle, and 3′ ends of the gRNA sequence (the exact insertion location is determined by...). Figure 15(Highlighted in red). The gRNA was then functionally tested (as described herein) by its ability to induce HR in a GFP reporter assay. Clearly, the gRNA's sequence insertions at the 5' and 3' ends are flexible (as demonstrated by the activity assay using retained HR). Therefore, aspects of this disclosure relate to linking small reactive RNA aptamers that can initiate gRNA activity, or gRNA visualization. Additionally, aspects of this disclosure relate to binding ssDNA donors to gRNA via hybridization, thus enabling genomic target cleavage and immediate, natural template localization for repair, which can promote homologous recombination rates at error-prone non-homologous end junctions.
[0132] The following references, which are identified by numbers in the Embodiments section, are incorporated herein by reference in their entirety for all purposes.
[0133] References
[0134] 1. KS Makarova et al. , Evolution and classification of the CRISPR-Cas systems. Nature reviews. Microbiology 9, 467 (Jun, 2011).
[0135] 2. M. Jinek et al. , A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science 337, 816 (Aug 17, 2012).
[0136] 3. P. Horvath, R. Barrangou, CRISPR / Cas, the immune system of bacteria and archaea. Science 327, 167 (Jan 8, 2010).
[0137] 4. H. Deveau et al. , Phage response to CRISPR-encoded resistance in Streptococcus thermophilus. Journal of bacteriology 190, 1390 (Feb, 2008).
[0138] 5. J. R. van der Ploeg, Analysis of CRISPR in Streptococcus mutanssuggests frequent occurrence of acquired immunity against infection by M102-like bacteriophages. Microbiology 155, 1966 (Jun, 2009).
[0139] 6. M. Rho, Y. W. Wu, H. Tang, T. G. Doak, Y. Ye, Diverse CRISPRsevolving in human microbiomes. PLoS genetics 8, e1002441 (2012).
[0140] 7. D. T. Pride et al. , Analysis of streptococcal CRISPRs from humansaliva reveals substantial sequence diversity within and between subjectsover time. Genome research 21, 126 (Jan, 2011).
[0141] 8. G. Gasiunas, R. Barrangou, P. Horvath, V. Siksnys, Cas9-crRNAribonucleoprotein complex mediates specific DNA cleavage for adaptiveimmunity in bacteria. Proceedings of the National Academy of Sciences of the United States of America 109, E2579 (Sep 25, 2012).
[0142] 9. R. Sapranauskas et al. , The Streptococcus thermophilus CRISPR / Cassystem provides immunity in Escherichia coli. Nucleic acids research 39, 9275(Nov, 2011).
[0143] 10. J. E. Garneau et al. , The CRISPR / Cas bacterial immune systemcleaves bacteriophage and plasmid DNA. Nature 468, 67 (Nov 4, 2010).
[0144] 11. K. M. Esvelt, J. C. Carlson, D. R. Liu, A system for thecontinuous directed evolution of biomolecules. Nature 472, 499 (Apr 28,2011).
[0145] 12. R. Barrangou, P. Horvath, CRISPR: new horizons in phageresistance and strain identification. Annual review of food science and technology 3, 143 (2012).
[0146] 13. B. Wiedenheft, S. H. Sternberg, J. A. Doudna, RNA-guided geneticsilencing systems in bacteria and archaea. Nature 482, 331 (Feb 16, 2012).
[0147] 14. N. E. Sanjana et al. , A transcription activator-like effectortoolbox for genome engineering. Nature protocols 7, 171 (Jan, 2012).
[0148] 15. W. J. Kent et al. , The human genome browser at UCSC. Genome Res 12, 996 (Jun, 2002).
[0149] 16. T. R. Dreszer et al., The UCSC Genome Browser database:extensions and updates 2011. Nucleic Acids Res 40, D918 (Jan, 2012).
[0150] 17. D. Karolchik et al. , The UCSC Table Browser data retrieval tool. Nucleic Acids Res 32, D493 (Jan 1, 2004).
[0151] 18. A. R. Quinlan, I. M. Hall, BEDTools: a flexible suite ofutilities for comparing genomic features. Bioinformatics 26, 841 (Mar 15,2010).
[0152] 19. B. Langmead, C. Trapnell, M. Pop, S. L. Salzberg, Ultrafast andmemory-efficient alignment of short DNA sequences to the human genome. Genome Biol 10, R25 (2009).
[0153] 20. R. Lorenz et al. , ViennaRNA Package 2.0. Algorithms for molecular biology : AMB 6, 26 (2011).
[0154] 21. D. H. Mathews, J. Sabina, M. Zuker, D. H. Turner, Expandedsequence dependence of thermodynamic parameters improves prediction of RNAsecondary structure. Journal of molecular biology 288, 911 (May 21, 1999).
[0155] 22. R. E. Thurman et al., The accessible chromatin landscape of thehuman genome. Nature 489, 75 (Sep 6, 2012).
[0156] 23. S. Kosuri et al. , Scalable gene synthesis by selectiveamplification of DNA pools from high-fidelity microchips. Nature biotechnology 28, 1295 (Dec, 2010).
[0157] 24.Q. Xu, M. R. Schlabach, G. J. Hannon, S. J. Elledge, Design of240,000 orthogonal 25mer DNA barcode probes. Proceedings of the National Academy of Sciences of the United States of America 106, 2289 (Feb 17, 2009)。
Claims
1. A method for altering eukaryotic cells, comprising: The eukaryotic cells were transfected with nucleic acids encoding RNA complementary to the genomic DNA of the eukaryotic cells. The eukaryotic cells were transfected with a nucleic acid that encodes an enzyme that interacts with the RNA and cuts the genomic DNA in a site-specific manner. in, The cell expresses the RNA and the enzyme, the RNA binding to complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner.
2. The method according to claim 1, wherein, The enzyme is Cas9.
3. The method according to claim 1, wherein, The eukaryotic cells are yeast cells, plant cells, or mammalian cells.
4. The method according to claim 1, wherein, The RNA consists of approximately 10 to approximately 250 nucleotides.
5. The method according to claim 1, wherein, The RNA consists of approximately 20 to approximately 100 nucleotides.
6. A method for altering human cells, including The human cells were transfected with nucleic acids encoding RNA complementary to the genomic DNA of eukaryotic cells. The human cells were transfected with nucleic acids encoding an enzyme that interacts with the RNA and cuts the genomic DNA in a site-specific manner. in, The human cells express the RNA and the enzyme, the RNA binding to complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner.
7. The method according to claim 6, wherein, The enzyme is Cas9.
8. The method according to claim 6, wherein, The RNA consists of approximately 10 to approximately 250 nucleotides.
9. The method according to claim 6, wherein, The RNA consists of approximately 20 to approximately 100 nucleotides.
10. A method for altering eukaryotic cells at multiple genomic DNA sites, including The eukaryotic cells were transfected with multiple nucleic acids encoding RNA complementary to different sites on the genomic DNA of the eukaryotic cells. The eukaryotic cells were transfected with a nucleic acid that encodes an enzyme that interacts with the RNA and cuts the genomic DNA in a site-specific manner. in, The cell expresses the RNA and the enzyme, the RNA binding to complementary genomic DNA and the enzyme cleaving the genomic DNA in a site-specific manner.