Genome Engineering
Modified TALENs with reduced repetitive sequences and RNA-guided DNA-binding proteins improve genome editing in human iPS cells by enhancing editing efficiency and sensitivity of NHEJ and HDR evaluation, addressing synthesis and delivery challenges.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PRESIDENT & FELLOWS OF HARVARD COLLEGE
- Filing Date
- 2024-01-10
- Publication Date
- 2026-07-22
AI Technical Summary
Current genome editing tools, such as TALENs, face challenges in human induced pluripotent stem cell engineering due to repetitive sequences complicating DNA construct synthesis and hindering lentiviral delivery, while existing assays for non-homologous end joining (NHEJ) and homology repair (HDR) evaluation lack sensitivity and accuracy, particularly in human iPS cells.
Modified TALENs lacking repetitive sequences are used to cleave target DNA, combined with donor nucleic acids for insertion, and RNA-guided DNA-binding proteins for precise genetic modification, enabling efficient and sensitive evaluation of NHEJ and HDR through methods like homologous recombination.
This approach enhances genome editing efficiency in human iPS cells, allowing for accurate and sensitive detection of NHEJ and HDR frequencies, and enables multiple genetic modifications with high precision and scalability.
Smart Images

Figure 0007893426000038 
Figure 0007893426000039 
Figure 0007893426000040
Abstract
Description
Technical Field
[0004] , , ,
[0001] Data of related applications This application claims priority to U.S. Provisional Patent Application No. 61 / 858,866, filed Jul. 26, 2013, the entire contents of which are hereby incorporated by reference herein for all purposes.
[0002] Description of government interests This invention was made with government support under grant number P50 HG003170 from the National Human Genome Research Institute Genome Science Center of the United States. The United States government has certain rights in this invention.
Background Art
[0003] Genome editing via sequence-specific nucleases is known. See References 1, 2, and 3, which are hereby incorporated by reference in their entirety herein. Nuclease-mediated double-stranded DNA (dsDNA) breaks in the genome can be repaired by two major mechanisms: non-homologous end joining (NHEJ), which often introduces non-specific insertions or deletions (indels), or homology-directed repair (HDR), which incorporates a homologous strand as a repair template. See Reference 4, which is hereby incorporated by reference in its entirety herein. When a sequence-specific nuclease is delivered with a homologous donor DNA construct containing a desired mutation, the gene targeting efficiency increases 1000-fold compared to the case of the donor construct alone. See Reference 5, which is hereby incorporated by reference in its entirety herein. It has been reported that single-stranded oligodeoxyribonucleotides (“ssODNs”) are used as DNA donors. See References 21 and 22, which are hereby incorporated by reference in their entirety herein.
[0004] Despite significant advances in gene editing tools, many challenges and questions remain regarding the use of custom-designed nucleases in human induced pluripotent stem cell (human iPS cell: "hiPSC") engineering. Firstly, activator-like effector nucleases (TALENs), despite their simple design, target specific DNA sequences through tandem copying of repeating variable duodecimal (RVD) domains. See Reference 6, incorporated herein by reference. While the modularity of RVDs simplifies TALEN design, the repeating sequences of RVDs complicate the synthesis of their DNA constructs (see References 2, 9, and 15-19, incorporated herein by reference) and hinder their use with lentiviral gene delivery vehicles. See Reference 13, incorporated herein by reference.
[0005] In current practice, non-homologous end joining (NHEJ) and homology repair (HDR) are often evaluated using separate assays. Mismatch-sensitive endonuclease assays (see reference 14, which is incorporated herein by reference) are frequently used to evaluate non-homologous end joining (NHEJ), but the quantitative accuracy of this method can vary, and its sensitivity is limited to NHEJ frequencies above approximately 3%. See reference 15, which is incorporated herein by reference. Homologous repair (HDR) is often evaluated by cloning and sequencing, which are entirely different and often cumbersome methods. While high editing frequencies of around 50% are frequently reported for some cell types such as U2OS and K562 (see references 12 and 14, which are incorporated herein by reference), the frequency is generally lower in human iPS cells, so sensitivity remains a concern. See reference 10, which is incorporated herein by reference. Recently, high editing frequencies have been reported in human iPS cells and human ES cells using TALENs (see Reference 9, which is incorporated herein by reference in its entirety), and even higher frequencies have been reported using the CRISPR Cas9-gRNA system (see References 16-19, which are incorporated herein by reference in its entirety). However, editing rates appear to vary considerably at different sites (see Reference 17, which is incorporated herein by reference in its entirety), and editing may not be detectable at all in some sites (see Reference 20, which is incorporated herein by reference in its entirety).
[0006] The bacterial and archaeal CRISPR-Cas system relies on short-chain guide RNA, which is complexed with the Cas protein, to degrade complementary sequences present in invasive foreign nucleic acids. Deltcheva, E. et al., CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III. Nature, Vol. 471, pp. 602-607 (2011); Gasiunas, G., Barrangou, R., Horvath, P., and Siksnys, V., Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria. Proceedings of the National Academy of Sciences of the United States of America, Vol. 109, pp. E2579-2586 (2012); Jinek, M. et al., A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science, Vol. 337, pp. 816-821 (2012); Sapranauskas, R. et al., The Streptococcus thermophilus CRISPR / Cas system provides immunity See also: in Escherichia coli. Nucleic Acids Research, Vol. 39, pp. 9275-9282 (2011); and Bhaya, D., Davison, M., and Barrangou, R., CRISPR-Cas systems in bacteria and archaea: versatile small RNAs for adaptive defense and regulation. Annual Review of Genetics, Vol. 45, pp. 273-297 (2011).Recent in vitro rearrangements of the type II CRISPR system in Streptococcus pyogenes (S. pyogenes) have shown that crRNA ("CRISPR RNA") fused to tracrRNA ("trans-activated CRISPR RNA"), which is normally trans-encoded, is sufficient to cause the Cas9 protein to sequence-specifically cleave a target DNA sequence matching the crRNA. The expression of gRNA (guide RNA) homologous to the target site triggers Cas9 recruitment and degradation of the target DNA. See H. Deveau et al., "Phage response to CRISPR-encoded resistance in Streptococcus thermophilus," Journal of Bacteriology, Vol. 190, p. 1390 (February 2008). [Overview of the project] [Problems that the invention aims to solve]
[0007] Some aspects of this disclosure relate to the use of modified transcription activator-like effector nucleases (TALENs) to genetically modify cells (such as somatic cells or stem cells). TALENs are known to contain repetitive sequences. Some aspects of this disclosure relate to a method for modifying target DNA in a cell, comprising introducing a TALEN lacking 100 bp or more of repetitive sequences into the cell, wherein the TALEN cleaves the target DNA, non-homologous end joining occurs in the cell, and modified DNA is produced in the cell. According to one aspect, a repetitive sequence of a desired length is removed from the TALEN. According to one aspect, the TALEN lacks a repetitive sequence of a desired length. According to one aspect, a TALEN from which a repetitive sequence of a desired length has been removed is provided. According to one aspect, the TALEN is modified to remove a repetitive sequence of a desired length. According to one aspect, the TALEN is designed to remove a repetitive sequence of a desired length. [Means for solving the problem]
[0008] Some aspects of this disclosure include a method for modifying target DNA in a cell, comprising combining a TALEN lacking 100 bp or more of repetitive sequences with a donor nucleic acid sequence, wherein the TALEN cleaves the target DNA and the donor nucleic acid sequence is inserted into the DNA in the cell. Some aspects of this disclosure relate to a virus comprising a nucleic acid sequence encoding a TALEN lacking 100 bp or more of repetitive sequences. Some aspects of this disclosure relate to a cell comprising a nucleic acid sequence encoding a TALEN lacking 100 bp or more of repetitive sequences. According to some aspects described herein, the TALEN lacks 100 bp or more, 90 bp or more, 80 bp or more, 70 bp or more, 60 bp or more, 50 bp or more, 40 bp or more, 30 bp or more, 20 bp or more, 19 bp or more, 18 bp or more, 17 bp or more, 16 bp or more, 15 bp or more, 14 bp or more, 13 bp or more, 12 bp or more, 11 bp or more, or 10 bp or more of repetitive sequences.
[0009] Some aspects of this disclosure relate to the production of a TALE, including combining an endonuclease, a DNA polymerase, a DNA ligase, an exonuclease, a plurality of nucleic acid dimer blocks encoding repeat variable two-residue domains, and an endonuclease cleavage site; activating the endonuclease to cleave the TALE-N / TF backbone vector at the endonuclease cleavage site to produce a first and second end; activating the exonuclease to create 3' and 5' overhangs on the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks to anneal the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks in a desired order; and activating the DNA polymerase and the DNA ligase to link the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks. Those skilled in the art will be able to readily identify suitable endonucleases, DNA polymerases, DNA ligases, exonucleases, nucleic acid dimer blocks encoding repeating variable two-residue domains, and TALE-N / TF backbone vectors based on this disclosure.
[0010] Some aspects of the present disclosure relate to a method for modifying target DNA in a stem cell, which expresses an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA, the method comprising: (a) introducing a first exogenous nucleic acid into a stem cell, which encodes RNA complementary to the target DNA and guides the enzyme to the target DNA, wherein the RNA and the enzyme are members of a co-localization complex to the target DNA; and introducing a second exogenous nucleic acid into the stem cell, which encodes a donor nucleic acid sequence, wherein the RNA and the donor nucleic acid sequence are expressed, the RNA and the enzyme co-localize to the target DNA, the enzyme cleaves the target DNA, and the donor nucleic acid is inserted into the target DNA to produce modified DNA in the stem cell.
[0011] Some aspects of this disclosure relate to stem cells comprising a first exogenous nucleic acid that encodes an enzyme that forms a colocalization complex with RNA complementary to target DNA and site-specifically cleaves the target DNA.
[0012] Some aspects of this disclosure relate to a cell comprising a first exogenous nucleic acid encoding an enzyme that forms a colocalization complex with RNA complementary to target DNA and site-specifically cleaves the target DNA, and comprising an inducible promoter for promoting the expression of the enzyme. Expression can be controlled in this way, for example, by initiating expression and by stopping expression.
[0013] Some aspects of the present disclosure relate to a cell comprising a first exogenous nucleic acid encoding an enzyme that forms a colocalization complex with RNA complementary to a target DNA and site-specifically cleaves the target DNA, wherein the first exogenous nucleic acid is removable from the cell's genomic DNA using a removal enzyme such as a transposase.
[0014] Some aspects of the present disclosure relate to a method for modifying target DNA in a cell expressing an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA, the method comprising (a) introducing a first exogenous nucleic acid encoding a donor nucleic acid sequence into the cell, and introducing RNA complementary to the target DNA and guiding the enzyme to the target DNA from a medium surrounding the cell into the cell, wherein the RNA and the enzyme are members of a co-localization complex to the target DNA, the donor nucleic acid sequence is expressed, the RNA and the enzyme co-localize to the target DNA, the enzyme cleaves the target DNA, the donor nucleic acid is inserted into the target DNA, and modified DNA is produced in the cell.
[0015] Several aspects of this disclosure relate to the use of RNA-guided DNA-binding proteins for genetically modifying stem cells. In one aspect, the stem cells are genetically modified to contain a nucleic acid encoding the RNA-guided DNA-binding protein, and the stem cells express the RNA-guided DNA-binding protein. According to one aspect, a donor nucleic acid for introducing a specific mutation is optimized for genome editing using the modified TALEN or the RNA-guided DNA-binding protein.
[0016] Some aspects of this disclosure relate to DNA modification (e.g., multiple DNA modification) in stem cells, using one or more guide RNAs (ribonucleic acid) that guide an enzyme having nuclease activity (e.g., a DNA-binding protein having nuclease activity) expressed by the stem cell to a target site on the DNA (deoxyribonucleic acid), wherein the enzyme cleaves the DNA and an exogenous donor nucleic acid is inserted into the DNA, for example, by homologous recombination. Some aspects of this disclosure include creating stem cells having multiple DNA modifications within the cell by cycling or repeating the DNA modification process in the stem cell. The modification may include the insertion of an exogenous donor nucleic acid.
[0017] Insertion of multiple exogenous nucleic acids can be achieved by introducing multiple RNAs and nucleic acids encoding multiple exogenous donor nucleic acids into stem cells expressing the enzyme in a single step (for example, by co-transformation). In this configuration, the multiple RNAs are expressed, each RNA in the multiple RNAs guides the enzyme to a specific site in the DNA, the enzyme cleaves the DNA, and one of the multiple exogenous nucleic acids is inserted into the DNA at the cleavage site. According to this configuration, a large number of changes or modifications to the DNA within the cell are produced in a single cycle.
[0018] Multiple insertions of foreign nucleic acids can be achieved within a cell by repeating a process or cycle of introducing one or more RNAs or multiple RNAs and one or more foreign nucleic acids or one or more nucleic acids encoding multiple foreign nucleic acids into the stem cell expressing the enzyme. In this process, the RNA is expressed to guide the enzyme to a specific site in the DNA, the enzyme cleaves the DNA, and the foreign nucleic acid is inserted into the DNA at the cleavage site. As a result, cells are produced that have multiple modifications or multiple insertions of foreign DNA in the DNA of the stem cell. According to one embodiment, the stem cell expressing the enzyme is genetically modified to express the enzyme, for example, by introducing nucleic acids that encode the enzyme and can be expressed by the stem cell into the cell. Accordingly, some aspects of this disclosure include a cycle of introducing RNA into a stem cell expressing the enzyme, introducing an exogenous donor nucleic acid into the stem cell, expressing the RNA, forming a colocalization complex between the RNA, the enzyme, and the DNA, enzymatically cleaving the DNA with the enzyme, and inserting the donor nucleic acid into the DNA. By cycling or repeating the above steps, multiple genetic modifications of the stem cell occur at multiple loci. That is, stem cells with multiple genetic modifications are produced.
[0019] In one embodiment, a DNA-binding protein or enzyme within the scope of this disclosure includes a protein that forms a complex with a guide RNA, wherein the guide RNA guides the complex to a double-stranded DNA sequence, and the complex binds to the DNA sequence. In one embodiment, the enzyme may be an RNA-guided DNA-binding protein, for example, an RNA-guided DNA-binding protein of a type II CRISPR system that binds to the DNA and is guided by the RNA. In one embodiment, the RNA-guided DNA-binding protein is a Cas9 protein.
[0020] This aspect of the present disclosure may be referred to as colocalization of the RNA and DNA-binding protein to or with double-stranded DNA. Thus, the DNA-binding protein-guide RNA complex may be used to cleave multiple sites of double-stranded DNA to produce stem cells having multiple genetic modifications, such as multiple insertions of exogenous donor DNA.
[0021] According to one embodiment, a method is provided for making multiple modifications to target DNA in a stem cell that expresses an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA, the method comprising: (a) introducing a first exogenous nucleic acid into the stem cell that encodes one or more RNAs complementary to the target DNA and guide the enzyme to the target DNA, wherein the one or more RNAs and the enzyme are members of a co-localization complex to the target DNA; introducing a second exogenous nucleic acid into the stem cell that encodes one or more donor nucleic acid sequences, wherein the one or more RNAs and the one or more donor nucleic acid sequences are expressed, the one or more RNAs and the enzyme co-localize to the target DNA, the enzyme cleaves the target DNA, the donor nucleic acid is inserted into the target DNA to produce modified DNA in the stem cell; and repeating step (a) multiple times to cause multiple modifications to the DNA in the stem cell.
[0022] According to one embodiment, the RNA is between approximately 10 nucleotides and approximately 500 nucleotides. According to another embodiment, the RNA is between approximately 20 nucleotides and approximately 100 nucleotides.
[0023] According to one embodiment, the one or more RNAs are guide RNAs. According to one embodiment, the one or more RNAs are tracrRNA-crRNA fusions.
[0024] According to one embodiment, the DNA is genomic DNA, mitochondrial DNA, viral DNA, or exogenous DNA.
[0025] According to one aspect, cells can be genetically modified to reversibly contain a nucleic acid encoding a DNA-binding enzyme using a vector that can be easily removed using an enzyme. Useful vectors and methods are known to those skilled in the art and include lentiviruses, adeno-associated viruses, nuclease- and integrase-mediated target insertion methods, and transposon-mediated insertion methods. According to one aspect, a nucleic acid encoding a DNA-binding enzyme, such as a nucleic acid encoding an added DNA-binding enzyme using a cassette or vector, can be removed in its entirety - for example, in genomic DNA, for example, without leaving a portion of the nucleic acid, cassette, or vector, the cassette and vector can be removed in their entirety. Such removal is called "scarless" removal in the art because the genome is the same as before the addition of the nucleic acid, cassette, or vector. One exemplary embodiment of insertion and scarless removal is the PiggyBac vector commercially available from System Biosciences.
[0026] Further features and advantages of specific embodiments of the present invention will become more fully apparent from the following description of the embodiments and their drawings, as well as from the claims.
[0027] The foregoing and other features and advantages of the present embodiment will be more fully understood from the following detailed description of the exemplary embodiments taken in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0028] [Figure 1A]Figure 1 shows a diagram illustrating the functional testing of re-TALEN in human somatic cells and stem cells. (a) Schematic illustration of an experimental design to test genome targeting efficiency. The GFP coding sequence integrated into the genome is disrupted by insertion of a 68 bp genomic fragment derived from a stop codon and the AAVS1 locus (bottom). Restoration of the GFP sequence by nuclease-mediated homologous recombination with a tGFP donor (top) results in GFP-positive cells that can be quantified by FACS. re-TALEN and TALEN target the same sequence within the AAVS1 fragment. [Figure 1B] (b) Bar graphs (N=3, error bars = SD (standard deviation)) showing the percentage of GFP-positive cells induced at target loci by tGFP donor alone, tGFP donor combined with TALEN, and tGFP donor combined with re-TALEN, as measured by FACS. A representative FACS plot is shown below. [Figure 1C] (c) A schematic overview illustrating a targeting strategy for the unmodified AAVS1 locus. A donor plasmid containing splicing acceptor (SA)-2A (self-cleaving peptide), puromycin resistance gene (PURO), and GFP is described (see reference 10, which is incorporated herein by reference in its entirety). The locations of the PCR primers used to detect successful editing events are indicated as blue arrows. [Figure 1D] (d) Clones successfully targeted from PGP1 human iPS cells were selected with puromycin (0.5 ug / mL) for two weeks. Microscopic images of three representative GFP-positive clones are shown. Cells were also stained for the pluripotency marker TRA-1-60. Scale bar: 200 μm. [Figure 1E] (e) PCR assays performed on these monoclonal GFP-positive human iPS cell clones demonstrated successful insertion of the donor cassette at the AAVS1 site (lanes 1, 2, and 3), whereas untreated human iPS cells showed no evidence of successful insertion (lane C). [Figure 1F]This figure shows the functional testing of re-TALEN in human somatic cells and stem cells. [Figure 2A] Figure 2 shows a comparison of the genome targeting efficiency of reTALEN and Cas9-gRNA against CCR5 in iPS cells. (a) Schematic illustration of the genome engineering experimental design. A 90-mer single-stranded oligodeoxyribonucleotide with a 2 bp mismatch to genomic DNA at the targeting site of the re-TALEN pair or Cas9-gRNA was delivered to PGP1 human iPS cells along with the reTALEN construct or Cas9-gRNA construct. The cleavage sites of their nucleases are indicated by red arrows in the figure. [Figure 2B] (b) Deep sequencing analysis of homology repair (HDR) efficiency and non-homologous end joining (NHEJ) efficiency of re-TALEN pairs (CCR5 3) and single-stranded oligodeoxyribonucleotides, or Cas9-gRNA and single-stranded oligodeoxyribonucleotides. Changes in the genome of human iPS cells were analyzed from high-throughput sequence data using the Genome Editing Assessment System (GEAS). Top): Homologous repair (HDR) was quantified from the percentage of reads containing a 2bp point mutation (blue) occurring in the center of a single-stranded oligodeoxyribonucleotide, and non-homologous end joining (NHEJ) activity was quantified from the percentage of deletions (gray) / insertions (red) at each specific site in the genome. In the graphs of reTALEN and single-stranded oligodeoxyribonucleotides, a green dashed line is plotted to indicate the outer boundary of the re-TALEN pair binding site, which is located at -26bp and +26bp relative to the center of the two re-TALEN binding sites. In the graphs of Cas9-gRNA and single-stranded oligodeoxyribonucleotides, the green dashed line indicates the outer boundary of the gRNA targeting site, which is located at -20 bp and -1 bp relative to the PAM sequence. Below: Deletion / insertion size distribution in human iPS cells analyzed from the entire non-homologous end joining (NHEJ) population after performing the processing described above. [Figure 2C](c) Genome editing efficiency of re-TALEN and Cas9-gRNA targeting CCR5 in PGP1 human iPS cells. Top): Schematic illustration of the targeted genome editing sites in CCR5. The blue arrows below illustrate 15 targeting sites. For each site, a pair of re-TALENs and their corresponding single-stranded oligodeoxyribonucleotide donors with a 2bp mismatch to the genomic DNA were cotransplanted into the cells. Genome editing efficiency was assayed 6 days after transtransplantation. Similarly, 15 different Cas9-gRNAs, along with their corresponding single-stranded oligodeoxyribonucleotides, were individually transplanted into PGP1-human iPS cells, targeting the same 15 sites, and the efficiency was analyzed 6 days after transtransplantation. Bottom): Genome editing efficiency of re-TALEN and Cas9-gRNA targeting CCR5 in PGP1 human iPS cells. Panels 1 and 2 show reTALEN-mediated non-homologous end joining (NHEJ) and homology repair (HDR) efficiencies. Panels 3 and 4 show Cas9-gRNA-mediated non-homologous end joining (NHEJ) and homology repair (HDR) efficiencies. The non-homologous end joining (NHEJ) rate was calculated by the frequency of genomic alleles carrying deletions or insertions in the targeting region, and the homology repair (HDR) rate was calculated by the frequency of genomic alleles with a 2bp mismatch. Panel 5, DNase I high-sensitivity site (HS) profiles of human iPS cell lines from the ENCODE database (Duke DNase HS, iPS NIHi7 DS). Note that the scales of the different panels are different. (f) Sanger sequencing of PCR amplicons derived from three targeted human iPS cell colonies confirmed the presence of the expected DNA bases at the genomic insertion boundaries. M: DNA ladder. C: Control using untreated human iPS cell genomic DNA. [Figure 3A]Figure 3 shows the testing of functional parameters governing single-strand oligodeoxyribonucleotide-mediated homology repair (HDR) by re-TALEN or Cas9-gRNA in PGP1 human iPS cells. (a) re-TALEN pair (number 3) and single-strand oligodeoxyribonucleotides of various lengths (50nt, 70nt, 90nt, 110nt, 130nt, 150nt, 170nt) were cotransferred into PGP1 human iPS cells. All single-strand oligodeoxyribonucleotides had the same 2bp mismatch to genomic DNA in the middle of their sequences. Optimal homology repair (HDR) was achieved in the targeted genome by the 90mer single-strand oligodeoxyribonucleotide. The efficiency of homology repair (HDR), non-homologous end joining (NHEJ)-induced deletions and insertions is evaluated as described herein. [Figure 3B] (b) The effect of homology deviation along the single-stranded oligodeoxyribonucleotide was examined using a 90-mer single-stranded oligodeoxyribonucleotide corresponding to re-TALEN pair 3, each containing a 2 bp mismatch in the center (A) and an additional 2 bp mismatch at different positions off-center from A (B), with the deviations ranging from -30 bp to 30 bp. The genome editing efficiency of each single-stranded oligodeoxyribonucleotide was evaluated in PGP1 human iPS cells. The bar graph below shows the incorporation frequencies of A only, B only, and A and B in the targeted genome. The homology repair (HDR) rate decreases as the distance of homology deviation from the center increases. [Figure 3C](c) Single-stranded oligodeoxyribonucleotides targeting sites at varying distances from the target site of the third re-TALEN pair (-620 bp to 480 bp) were tested to evaluate the maximum distance at which a single-stranded oligodeoxyribonucleotide could be positioned to introduce a mutation. All single-stranded oligodeoxyribonucleotides had a 2 bp mismatch in the middle of their sequence. The lowest homology repair (HDR) efficiency (≤0.06%) was observed when the mismatch of the single-stranded oligodeoxyribonucleotide was located 40 bp away from the middle of the re-TALEN pair binding site. [Figure 3D] (d) Cas9-gRNA (AAVS1) and single-stranded oligodeoxyribonucleotides with different orientations (Oc: complementary to gRNA; On: non-complementary to gRNA) and varying lengths (30nt, 50nt, 70nt, 90nt, 110nt) were cotransplanted into PGP1 human iPS cells. All single-stranded oligodeoxyribonucleotides had the same 2bp mismatch to genomic DNA in the middle of their sequence. Optimal homology repair (HDR) was achieved in the targeted genome with the 70mer Oc. [Figure 4A] Figure 4 shows the acquisition of monoclonal genome-edited human iPS cells without selection using re-TALEN and single-stranded oligodeoxyribonucleotides. (a) Experimental timeline. [Figure 4B] (b) Genomic engineering efficiency of re-TALEN pairs and single-stranded oligodeoxyribonucleotides (3) as evaluated by the NGS (next-generation sequencing) platform shown in Figure 2b. [Figure 4C] (c) Results of Sanger sequencing of monoclonal human iPS cell colonies after genome editing. A 2bp heterogeneous genotype (CT / CT→TA / CT) was successfully introduced into the genomes of the PGP1-iPS-3-11 and PGP1-iPS-3-13 colonies. [Figure 4D](d) Immunofluorescence staining of targeted PGP1-iPS-3-11. Cells were stained for the pluripotency markers Tra-1-60 and SSEA4. [Figure 4E] (e) Hematoxylin and eosin staining of teratoma sections derived from monoclonal PGP1-iPS-3-11 cells. [Figure 5A] Figure 5 shows the reTALE design. (a) Sequence alignment of monomers in re-TALE-16.5 (re-TALE-M1 → re-TALE-M17) with the original TALE RVD monomer. Nucleotide changes from the original sequence are highlighted in gray. [Figure 5B] (b) Examination of the repeatability of re-TALE by PCR. The upper panel shows the structure of re-TALE / TALE and the position of the primers in the PCR reaction. The lower panel shows the PCR bands under the conditions shown below. In the original TALE template (right lane), the PCR ladder is visible. [Figure 6A] Figure 6 shows the design and implementation of a TALE single-incubation assembly (TASA). (a) Schematic illustration of the library of re-TALE dimer blocks for TASA assembly. There is a library consisting of 10 types of re-TALE dimer blocks encoding two RVDs. All 16 dimers in each block share the same DNA sequence except for the RVD coding sequence. Dimers in different blocks have different sequences, but are designed to have a 32 bp overlap with adjacent blocks. The DNA sequence and amino acid sequence of one of the dimers (block 6_AC) are shown on the right. [Figure 6B](b) Schematic illustration of TASA assembly. The left panel shows the TASA assembly method, in which a one-pot incubation reaction is performed using an enzyme mixture / re-TALE block / re-TALE-N / TF backbone vector. The reaction product can be used directly for bacterial transformation. The right panel shows the mechanism of TASA. The destination vector is linearized with an endonuclease at 37°C to excise the ccdB counterselection cassette, and the ends of each block and the ends of the linearized vector are processed by an exonuclease to expose single-stranded DNA overhangs at the ends of each fragment, thereby annealing the block group and vector backbone in the specified order. When the temperature rises to 50°C, polymerase and ligase work together to fill the gaps, thereby creating the final construct ready for transformation. [Figure 6C] (c) TASA assembly efficiency of re-TALE with different monomer lengths. The blocks used for assembly are shown on the left, and the assembly efficiency is shown on the right. [Figure 7A] This figure shows the functionality and alignment completeness of the wrench-reTALE. [Figure 7B] This figure shows the functionality and alignment completeness of the wrench-reTALE. [Figure 7C] This figure shows the functionality and alignment completeness of the wrench-reTALE. [Figure 7D] This figure shows the functionality and alignment completeness of the wrench-reTALE. [Figure 8A]Figure 8 shows the sensitivity and reproducibility of the Genome Editing Assessment System (GEAS). (A) Information-based analysis of the Homologous Repair (HDR) detection limit. In a given re-TALEN (number 10) / single-stranded oligodeoxyribonucleotide dataset, reads containing expected edits (HDR) were identified, and these Homologous Repair (HDR) reads were systematically removed to create various artificial datasets with "diluted" editing signals. Datasets were created with 100%, 99.8%, 99.9%, 98.9%, 97.8%, 89.2%, 78.4%, 64.9%, 21.6%, 10.8%, 2.2%, 1.1%, 0.2%, 0.1%, 0.02%, and 0% removal of Homologous Repair (HDR) reads to create artificial datasets with homologous recombination (HR) efficiencies ranging from 0 to 0.67%. For each individual dataset, the mutual information (MI) between the background signal (purple) and the signal obtained at the targeting region (green) was evaluated. When the homology repair (HDR) efficiency exceeds 0.0014%, the mutual information (MI) at the targeting region is significantly higher than that at the background. The limit of homology repair (HDR) detection was estimated to be between 0.0014% and 0.0071%. The calculation of mutual information (MI) is described herein. [Figure 8B](B) Examination of the reproducibility of the genome editing evaluation system. The results of the evaluation of homology repair (HDR) and non-homologous end joining (NHEJ) in two replication experiments using the re-TALEN pairs and cell types shown above are shown in paired plots (top and bottom). Nucleofection, target genome amplification, deep sequencing, and data analysis were performed independently for each experiment. The variance of the genome editing evaluation in the replication experiments was calculated as √2(|Homology Repair (HDR)1-HDR2|) / (HDR+HDR2) / 2)=ΔHDR / HDR and √2(|Non-homologous End Joining (NHEJ)1-NHEJ2|) / ((NHEJ1+NHEJ2) / 2)=ΔNHEJ / NHEJ. The results of the variance are shown below the plot. The mean variance of the system was (19%+11%+4%+9%+10%+35%) / 6=15%. Factors that may contribute to this dispersion include the state of cells during nucleofection, nucleofection efficiency, and sequencing coverage and quality. [Figure 9A] Figure 9 shows a statistical analysis of the efficiency of non-homologous end joining (NHEJ) and homology repair (HDR) mediated by reTALEN and Cas9-gRNA on CCR5. (a) Correlation between reTALEN-mediated homologous recombination (HR) efficiency and non-homologous end joining (NHEJ) efficiency at the same site in iPS cells (r=0.91, P<1×10⁻⁵). [Figure 9B] (b) Correlation between the efficiency of Cas9-gRNA-mediated homologous recombination (HR) and non-homologous end joining (NHEJ) at the same site in iPS cells (r=0.74, P=0.002). [Figure 9C] (c) Correlation between Cas9-gRNA-mediated non-homologous end joining (NHEJ) efficiency and Tm temperature of gRNA targeting sites in iPS cells (r=0.52, P=0.04). [Figure 10]This figure shows the correlation analysis between genome editing efficiency and epigenetic status. Pearson correlation was used to analyze possible relationships between DNase I sensitivity and genome engineering efficiency (homologous recombination (HR), non-homologous end joining (NHEJ)). Observed correlations were compared to a randomized set (N=100,000). Observed correlations higher than the 95th percentile or lower than the 5th percentile of the simulated distribution were considered possible relationships. No significant correlation was observed between DNase I sensitivity and non-homologous end joining (NHEJ) / homologous recombination (HR) efficiency. [Figure 11A] Figure 11 shows the effect of homology pairing in single-strand oligodeoxyribonucleotide-mediated genome editing. (a) In the experiment described in Figure 3b, global homology repair (HDR), measured by the rate of integration of the central 2b mismatch (A), decreased as the distance of the secondary mismatch B from A increased (the relative position of B to A varied from -30 bp to 30 bp). The higher integration rate when B was only 10 bp away from A (-10 bp and +10 b) may reflect a lower need for single-strand oligodeoxyribonucleotide pairing with genomic DNA proximal to double-strand DNA breaks. [Figure 11B](b) Distribution of gene conversion lengths along a single-stranded oligodeoxyribonucleotide. At each distance from A to B, a certain proportion of homology repair (HDR) events incorporate only A, and another proportion incorporate both A and B. These two types of events can be explained in terms of gene conversion tracts (Elliott et al., 1998), according to which A+B events represent longer conversion tracts extending beyond B, and A-only events represent shorter conversion tracts that do not reach B. Under this interpretation, the distribution of gene conversion lengths in both directions along the oligo can be evaluated (the center of the single-stranded oligodeoxyribonucleotide is defined as 0, the conversion track toward the 5' end of the single-stranded oligodeoxyribonucleotide is defined as the - direction, and the conversion track toward the 3' end is defined as the + direction). As the length of the gene conversion tract increases, the incidence of the gene conversion gradually decreases. This result is very similar to the gene conversion tract distribution observed for double-stranded DNA donors, but in the case of single-stranded DNA oligos, the distance scale is much more compressed, ranging from tens of base pairs to hundreds of base pairs in the case of double-stranded DNA donors. [Figure 11C](c) Assay of a gene conversion tract measuring sequential integration using a single single-stranded oligodeoxyribonucleotide containing a series of mutations. A single-stranded oligodeoxyribonucleotide donor (top) was used, which had three pairs of 2bp mismatches (orange) spaced 10nt apart on either side of a central 2bp mismatch. Of the more than 300,000 reads sequenced in this region, only a few genome sequencing reads retained one or more mismatches defined by the single-stranded oligodeoxyribonucleotide (see reference 62, the entirety of which is incorporated herein by reference). All of these reads were plotted (bottom), and the read sequences were color-coded. Orange: defined mismatch; Green: wild-type sequence. This genome editing with single-stranded oligodeoxyribonucleotide resulted in a pattern where only the central mutation was integrated in 85% (53 / 62) of cases, and multiple B mismatches were integrated in the others. The number of events incorporating B was too small to estimate the distribution of tract lengths exceeding 10 bp, but it is clear that short tract regions of -10 to 10 bp are dominant. [Figure 12] This figure shows the genome editing efficiency using Cas9-gRNA nuclease and nickes. PGP1 iPS cells were cotransplanted with either nuclease (C2) (Cas9-gRNA) or nickes (Cc) (Cas9D10A-gRNA) and combinations of single-stranded oligodeoxyribonucleotides (Oc and On) with different orientations. All single-stranded oligodeoxyribonucleotides have an identical 2 bp mismatch with genomic DNA at the center of their sequence. Homology repair (HDR) evaluation is described herein. [Figure 13]This figure shows the design and optimization of the re-TALE sequence. The re-TALE sequence was evolved by removing repeats in several design cycles. In each cycle, synonymous sequences derived from each repeat were evaluated. The sequence with the largest Hamming distance to the evolving DNA was selected. The final sequence has cai = 0.59 and ΔG = -9.8 kcal / mol. An R package was provided to implement this overall framework for synthetic protein design. [Figure 14] This gel image shows PCR confirmation of Cas9 genomic insertion within PGP1 cells. Lines 3, 6, 9, and 12 are PCR products from a normal (plain) PGP1 cell line. [Figure 15] This graph shows the mRNA expression levels of Cas9 mRNA under induction. [Figure 16] This graph shows the genome targeting efficiency using various RNA designs. [Figure 17] This graph shows the genome targeting efficiency of homologous recombination achieved by guide RNA-donor DNA fusions, which is 44%. [Figure 18] This diagram shows the genotypes of the PGP1 cell line isogenically produced by the system described herein. PGP1-iPS-BTHH has a single nucleotide deletion phenotype as a BTHH patient. PGP1-NHEJ has a 4 bp deletion that causes a frameshift mutation in a different way. [Figure 19] This graph shows that cardiomyocytes derived from the PGP1 iPS cell lineage reproduced the same ATP production defects and F1F0 ATPase-specific activity defects observed in patient-specific cells. [Figure 20A] The sequences of the re-TALEN and re-TALE-TF (transcription factor) backbone are shown. [Figure 20B] The sequences of the re-TALEN and re-TALE-TF backbones are shown. [Figure 20C] The sequences of the re-TALEN and re-TALE-TF backbones are shown. [Figure 20D] The sequences of the re-TALEN and re-TALE-TF backbones are shown. [Figure 20E] The sequences of the re-TALEN and re-TALE-TF backbones are shown. [Figure 20F] The sequences of the re-TALEN and re-TALE-TF backbones are shown. [Modes for carrying out the invention]
[0029] Some aspects of the present invention relate to the use of TALENs lacking a particular repeat sequence for nucleic acid engineering, which is performed, for example, by cleaving a double-stranded nucleic acid. The use of the TALEN to cleave a double-stranded nucleic acid may result in non-homologous end joining (NHEJ) or homologous recombination (HR). Some aspects of the present disclosure also envision the use of TALENs lacking a repeat sequence for nucleic acid engineering, for example by cleaving a double-stranded nucleic acid, in the presence of a donor nucleic acid, and also envision the insertion of the donor nucleic acid into the double-stranded nucleic acid, for example by non-homologous end joining (NHEJ) or homologous recombination (HR).
[0030] Transcription activator-like effector nucleases (TALENs) are known in the art and include artificial restriction enzymes created by fusing a TAL effector DNA-binding domain to a DNA-cleaving domain. Restriction enzymes are enzymes that cleave DNA strands at specific sequences. Transcription activator-like effectors (TALEs) can be modified to bind to desired DNA sequences. See, in its entirety, Boch, Jens (February 2011), "TALEs of genome targeting," Nature Biotechnology, Vol. 29 (No. 2): pp. 135-136, which is incorporated here by reference. By combining such a modified TALE with a DNA-cleaving domain (which cleaves DNA strands), TALENs that are restriction enzymes specific to any desired DNA sequence can be created. In one embodiment, the TALEN is introduced into cells for in situ targeted nucleic acid editing, such as in situ genome editing.
[0031] According to one embodiment, a hybrid nuclease active in yeast, plant, and animal cells can be constructed using a nonspecific DNA cleavage domain derived from the terminal of a FokI endonuclease. The FokI domain functions as a dimer, requiring two constructs, each having its own unique DNA-binding domain, each targeting a site in the target genome with appropriate orientation and spacing. Both the number of amino acid residues between the TALE DNA-binding domain and the FokI cleavage domain, and the number of bases between the two individual TALEN binding sites, affect the activity.
[0032] The relationship between the amino acid sequence of the TALE binding domain and DNA recognition allows for protein design. TALE constructs can be designed using software programs such as DNAWorks. Other methods for designing TALE constructs are known to those skilled in the art.The following references are incorporated here by reference: Cermak, T.; Doyle, EL; Christian, M.; Wang, L.; Zhang, Y.; Schmidt, C; Bailer, JA; Somia, NV et al., (2011) "Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting", Nucleic Acids Research. doi:10.1093 / nar / gkr218; Zhang, Feng et al., (February 2011) "Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription", Nature Biotechnology, Vol. 29 (No. 2): pp. 149-153; Morbitzer, R.; Elsaesser, J.; Hausner, J.; Lahaye, T., (2011) "Assembly of custom TALE-type DNA binding domains by modular cloning", Nucleic Acids Research. doi:10'1093 / nar / gkr151; Li, T.; Huang, S.; Zhao, X.; Wright, DA; Carpenter, S.; Spalding, MH; Research magazine.doi''10.1093 / nar / gkr188;.
[0033] [ka]
[0034] See Weber, E.; Gruetzner, R.; Werner, S.; Engler, C; Marillonnet, S., (2011), "Assembly of Designer TAL Effectors by Golden Gate Cloning," in Bendahmane, Mohammed, PLoS ONE, Vol. 6 (No. 5): e19722.
[0035] In exemplary embodiments, once the TALEN gene is produced, in certain embodiments, the TALEN gene may be inserted into a plasmid; and the plasmid may be used to translocate a target cell, in which the gene product is expressed, enters the cell nucleus, and reaches the genome. In exemplary embodiments, the TALENs described herein can be used to edit a target nucleic acid, such as the genome, by inducing double-strand breaks (DSBs) and having the cell respond to the double-strand breaks with repair mechanisms. Exemplary repair mechanisms include non-homologous end joining (NHEJ), which rejoins the DNA on both sides of a double-strand break with very little or no sequence overlap for annealing. These repair mechanisms may induce errors in the genome by insertion or deletion (indel) or induce chromosomal rearrangement, and such errors may render the gene product encoded at that site nonfunctional. The entire work is incorporated here by reference to Miller, Jeffrey et al., (February 2011). "A TALE nuclease architecture for efficient genome editing." Nature Biotechnology, Vol. 29 (No. 2): pp. 143-148. Since this activity can vary depending on the species, cell type, target gene, and nuclease used, the activity can be monitored by using a heterodouble-strand break assay that detects any differences between two alleles amplified by PCR. The cleavage products can be visualized using a simple agarose gel or slab gel system.
[0036] Alternatively, DNA can be introduced into the genome by non-homologous end joining (NHEJ) in the presence of exogenous double-stranded DNA fragments. Homologous recombination repair can also introduce exogenous DNA at double-strand breaks (DSBs) when the translocated double-stranded sequence is used as a template for repair enzymes. In one embodiment, stable modified human embryonic stem cell clones and induced pluripotent stem cell (iPS cell) clones can be produced using the TALENs described herein. In another embodiment, knockout species such as the nematode *C. elegans*, knockout rats, knockout mice, or knockout zebrafish can be produced using the TALENs described herein.
[0037] According to one aspect of this disclosure, embodiments relate to the use of exogenous DNA, a nuclease enzyme such as a DNA-binding protein, and a guide RNA for co-localizing with DNA in stem cells and for digesting or cleaving DNA to insert the exogenous DNA. Such DNA-binding proteins that bind to DNA for various purposes will be readily known to those skilled in the art. Such DNA-binding proteins may be native. DNA-binding proteins included within the scope of this disclosure include those that can be guided by RNA, referred to herein as guide RNA. According to this aspect, the guide RNA and the RNA-guided DNA-binding protein form a co-localization complex at the DNA. Such DNA-binding proteins having nuclease activity are known to those skilled in the art and include native DNA-binding proteins having nuclease activity, such as the Cas9 protein, which is present in the type II CRISPR system. Such Cas9 proteins and type II CRISPR systems are well described in the literature in the art. For the entire work, including all supplementary information, please refer to Makarova et al., Nature Reviews, Microbiology, Vol. 9, June 2011, pp. 467-477.
[0038] Exemplary DNA-binding proteins possessing nuclease activity function to introduce nicks into double-stranded DNA or to cleave double-stranded DNA. Such nuclease activity can arise from DNA-binding proteins having one or more polypeptide sequences exhibiting nuclease activity. Such exemplary DNA-binding proteins may have two separate nuclease domains, each responsible for cleaving or nick formation on a specific strand of double-stranded DNA, respectively. Exemplary polypeptide sequences possessing nuclease activity known to those skilled in the art include McrA-HNH nuclease-associated domains and RuvC-like nuclease domains. Thus, an exemplary DNA-binding protein is a DNA-binding protein that essentially contains one or more of the McrA-HNH nuclease-associated domains and RuvC-like nuclease domains.
[0039] The exemplary DNA-binding protein is the RNA-inducible DNA-binding protein of the type II CRISPR system. The exemplary DNA-binding protein is the Cas9 protein.
[0040] In Streptococcus pyogenes (S. pyrogenes), Cas9 forms a blunt-ended double-strand break 3 bp upstream of the protospacer-adjacent motif (PAM). This break is mediated by two catalytic domains in the protein: the HNH domain, which cleaves the complementary DNA strand, and the RuvC-like domain, which cleaves the non-complementary strand. For the complete process, please refer to Jinke et al., Science, Vol. 337, pp. 816-821 (2012), which is incorporated here by reference. The Cas9 protein is known to be present in numerous type II CRISPR systems, including the following, as identified in the supplementary information in Makarova et al., Nature Reviews, Microbiology, Vol. 9, June 2011, pp. 467-477: Methanococcus malipaldis C7; Corynebacterium diphtheriae; Corynebacterium eficiens YS-314; Corynebacterium glutamicum ATCC13032 Kitasato; Corynebacterium glutamicum ATCC13032 Bielefeld; Corynebacterium glutamicum R; Corynebacterium clopenstedii DSM44385; Mycobacterium abscesses ATCC19977; Nocardia farsinica IFM10152; Rhodococcus erythropolis PR4; Rhodococcus josti RHA1; Rhodococcus opacus B4 uid36573; Acidothermus ceruloriticus 11B; Arthrobacter chlorophenoricus A6; Crybella flavida DSM17836 uid43465; Thermospora carbata DSM43183; Bifidobacterium denthium Bd1; Bifidobacterium longum DJO10A; Slacchia heliotrini reducens DSM20476; Persephonera marina EX H1; Bacteroides fragilis NCTC9434; Capnocytophaga ochracea DSM7271; Flavobacterium cyclophyllum JIP02 86; Ackermansia muciniphylla ATCC BAA835; Roseiflexus castenholzii DSM13941; Roseiflexus RSI; Synechocystis PCC6803; Elusimicrobium minutum Pei191;Non-cultured termite group 1 bacterial phylogenetic type Rs D17; Fibrobacter succinogenes S85; Bacillus cereus ATCC10987; Listeria inocure; Lactobacillus casei; Lactobacillus rhamnosus GG; Lactobacillus salivarius UCC118; Streptococcus agalactiae A909; Streptococcus agalactiae NEM316; Streptococcus agalactiae 2603; Streptococcus disgalactiae equisimilis GGS124; Streptococcus equi zuepidemicus MGCS10565; Streptococcus galloreticus UCN34 uid46061; Streptococcus goldonii Challis subst CH1; Streptococcus mutans NN2025 uid46353; Streptococcus mutans; Streptococcus pyogenes M1 GAS; Streptococcus pyogenes MGAS5005; Streptococcus pyogenes MGAS2096; Streptococcus pyogenes MGAS9429; Streptococcus pyogenes MGAS10270; Streptococcus pyogenes MGAS6180; Streptococcus pyogenes MGAS315; Streptococcus pyogenes SSI-1; Streptococcus pyogenes MGAS10750; Streptococcus pyogenes NZ131; Streptococcus thermophilus Streptococcus thermophiles CNRZ1066; Streptococcus thermophiles LMD-9; Streptococcus thermophilus LMG18311; Clostridium botulinum A3 Loch Maree; Clostridium botulinum B Eklund 17B; Clostridium botulinum Ba4 657; Clostridium botulinum F Langeland; Clostridium ceruloticum H10; Finegordia magna ATCC29328; Eubacterium rectore ATCC33656; Mycoplasma galliseptum; Mycoplasma mobile 163K; Mycoplasma penetrans; Mycoplasma sinobiae 53; Streptobacillus moniliformis DSM12112;Bradyrisobium BTAil; Nitrobacter hamburgensis X14; Rhodopseudomonas palustris BisB18; Rhodopseudomonas palustris BisB5; Barbibacrum labmentivorance DS-1; Dinoroseobacter shibae DFL12; Gluconacetobacter diazotrophicus Pal 5 FAPERJ; Gluconacetobacter diazotrophicus Pal 5 JGI; Azospirillum B510 uid46085; Rhodospirillum rubrum ATCC11170; Diaphorobacter TPSY uid29975; Fermineforobacter eiseniae EF01-2; Neisseria meningitides 053442; Neisseria meningitides α14; Neisseria meningitides (Neisseria meningitides)Z2491; Desulfovibrio salexigens DSM2638; Campylobacter jejuni doirei 269 97; Campylobacter jejuni 81116; Campylobacter jejuni; Campylobacter lari RM2100; Helicobacter hepaticus; Worinella succinogenes; Tormonas auensis DSM9187; Pseudoalteromonas atlantica T6c; Shewanella pealeana ATCC700345; Legionella pneumophila Paris; Actinobacillus succinogenes 130Z; Pasteurella multocida; Francisella tularensis nobicida U112; Francisella tularensis horalkutica; Francisella tularensis FSC198; Francisella tularensis tularensis; Francisella tularensis WY96-3418; and Treponema denticola ATCC35405. Therefore, some aspects of this disclosure relate to the Cas9 protein present in the type II CRISPR system.
[0041] The Cas9 protein is sometimes referred to as Csn1 by those skilled in the art in the literature. The Cas9 protein of Streptococcus pyogenes (S. pyrogenes) is shown below. See Deltcheva et al., Nature, Vol. 471, pp. 602-607 (2011), which is incorporated here by reference.
[0042]
number
[0043]
number
[0044] According to one embodiment, the RNA-guided DNA-binding protein includes homologs and orthologues of Cas9 that bind to DNA, are guided by RNA, and retain the ability of the protein to cleave DNA. According to one embodiment, the Cas9 protein includes a sequence shown for native Cas9 from Streptococcus pyogenes (S. pyrogenes), and a protein sequence that has at least 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% homology to said sequence and is a DNA-binding protein (e.g., an RNA-guided DNA-binding protein).
[0045] According to one embodiment, a modified Cas9-gRNA system is provided that optionally enables site-specific RNA-induced genome cleavage and modification of the stem cell genome by insertion of exogenous donor nucleic acids within the stem cell. The guide RNA is complementary to a target site or target locus on the DNA. The guide RNA may be a crRNA-tracrRNA chimera. The guide RNA can be introduced from a medium surrounding the cell. In this way, a method is provided for continuously modifying cells, insofar as various guide RNAs are supplied to the surrounding medium, the guide RNA is taken up by the cell, and additional guide RNA is replenished in the medium. The replenishment may be continuous. The Cas9 binds to the target genomic DNA or its vicinity. One or more guide RNAs bind to the target genomic DNA or its vicinity. The Cas9 cleaves the target genomic DNA, and exogenous donor DNA is inserted into the DNA at the cleavage site.
[0046] Therefore, one method relates to using a guide RNA together with the Cas9 protein and exogenous donor nucleic acid to multiplex insert the exogenous donor nucleic acid into the DNA in a Cas9-expressing stem cell by cycling through the insertion of a nucleic acid encoding a guide RNA (or provision of RNA from an ambient medium), insertion of an exogenous donor nucleic acid, expression of the RNA (or uptake of the RNA), colocalization of the RNA, Cas9 and the DNA in a manner that cleaves the DNA, and insertion of the exogenous donor nucleic acid. The steps of the method can be cycled any number of times to produce any desired number of DNA modifications. Therefore, some methods of this disclosure relate to editing a target gene using the Cas9 protein and guide RNA described herein, and to providing multiplex genetic engineering and epigenetic engineering of stem cells.
[0047] Further aspects of this disclosure relate to the use of DNA-binding proteins or DNA-binding systems in general (such as modified TALENS or Cas9 as described herein) for the multiple insertion of exogenous donor nucleic acids into the DNA (e.g., genomic DNA) of stem cells (e.g., human stem cells). Those skilled in the art will readily identify exemplary DNA-binding systems based on this disclosure.
[0048] Cells relating to this disclosure include any cells that can be introduced and expressed with exogenous nucleic acids as described herein, unless otherwise expressly stated. It should be understood that the basic concepts of this disclosure as described herein are not limited by cell type. Cells relating to this disclosure include somatic cells, stem cells, eukaryotic cells, prokaryotic cells, animal cells, plant cells, fungal cells, archaeal cells, and bacterial cells. Cells include eukaryotic cells such as yeast cells, plant cells, and animal cells. Specific cells include mammalian cells such as human cells. Furthermore, cells include any cells in which DNA modification is beneficial or desirable.
[0049] The target nucleic acid includes any nucleic acid sequence in which a TALEN or RNA-induced DNA-binding protein having the nuclease activity described herein may be useful for nicking or cleaving. The target nucleic acid includes any nucleic acid sequence in which a co-localization complex described herein may be useful for nicking or cleaving. The target nucleic acid includes genes. For the purposes of this disclosure, DNA such as double-stranded DNA may contain the target nucleic acid, and the co-localization complex may bind to the DNA or co-localize with the DNA in a different manner, or the TALEN may bind to the DNA, at the location of the target nucleic acid, adjacent to the target nucleic acid, or in the vicinity of the target nucleic acid, so that the co-localization complex or TALEN may have the desired effect on the target nucleic acid. Such target nucleic acids may include endogenous (or native) nucleic acids and exogenous (or foreign) nucleic acids. Those skilled in the art will be able to readily identify or design, based on this disclosure, guide RNA and Cas9 proteins that co-localize to DNA containing the target nucleic acid, or TALENs that bind to DNA containing the target nucleic acid. Those skilled in the art may also identify transcription regulatory protein or domain (e.g., transcription activators or transcription repressors) that co-localize similarly with the DNA containing the target nucleic acid. The DNA may include genomic DNA, mitochondrial DNA, viral DNA, or exogenous DNA. In one aspect, materials and methods useful for carrying out the disclosure include those described in Di Carlo et al., Nucleic Acids Research, 2013, Vol. 41, No. 7, pp. 4336-4343, which are incorporated herein by reference in their entirety for all purposes, including exemplary strains and media, plasmid construction, plasmid transformation, transient gRNA cassette and donor nucleic acid electroporation, transformation of donor DNA and gRNA plasmid into Cas9-expressing cells, galactose induction of Cas9, and identification of CRISPR-Cas targets in the yeast genome.Additional references, each containing information, materials, and methods useful to those skilled in the art for carrying out the present invention, are incorporated herein by reference in their entirety for all purposes: Mali, P., Yang, L., Esvelt, KM, Aach, J., Guell, M., DiCarlo, JE, Norville, JE, and Church, GM (2013), RNA-Guided human genome engineering via Cas9. Science, 10.1126fscience.1232033; Storici, F., Durham, CL, Gordenin, DA, and Resnick, MA (2003), Chromosomal site-specific double-strand breaks are efficiently targeted for repair by oligonucleotides in yeast. PNAS, Vol. 100, pp. 14994-14999; and Jinek, M., Chylinski, K., Fonfara, L., Hauer, M., Doudna, JA, and Charpentier, E. (2012), A programmable dual-RNA-Guided DNA endonuclease in adaptive bacterial immunity. Available in Science, Vol. 337, pp. 816-821.
[0050] The introduction of exogenous nucleic acids (i.e., nucleic acids that are not part of the cell's natural nucleic acid composition) into cells may be carried out using any method known to those skilled in the art for such introduction. Such methods include transtransfer, transduction, viral transduction, microinjection, lipofection, nucleofection, nanoparticle bombardment, transformation, and conjugation. Those skilled in the art will readily understand and adapt such methods using readily identifiable literature.
[0051] The donor nucleic acid includes any nucleic acid to be inserted into the nucleic acid sequence described herein.
[0052] The embodiments described below are shown as representative of the present disclosure. These embodiments and other equivalent embodiments are evident in light of the present disclosure, the drawings and the appended claims, and should not be construed as limiting the scope of the present disclosure.
[0053] Example 1 Guide RNA assembly A 19bp portion of the selected target sequence (i.e., 5'-N19 of 5'-N19-NGG-3') was incorporated into two complementary 100-mer oligonucleotides (TTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCGN19GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCC). Each 100-mer oligonucleotide was suspended in water at a concentration of 100 mM, mixed in equal volumes, and annealed using a thermocycle machine (95°C, 5 mins, gradient to 4°C, 0.1°C / sec). To prepare the destination vector, the gRNA cloning vector (Addgene plasmid number 41824) was linearized using AfIII and purified. A gRNA assembly reaction (10 µl) was performed at 50°C for 30 minutes using 10 ng of annealed 100 bp fragments, 100 ng of destination backbone, and 1 × Gibson assembly reaction mix (New England Biolabs). The reaction product can be directly processed for bacterial transformation to colonize individual assemblies.
[0054] Example 2 Design and assembly of recoded TALEs To facilitate assembly and improve expression, re-TALE was optimized at various levels. First, the re-TALE DNA sequence was co-optimized for human codon usage frequency and low mRNA folding energy at the 5' end (GeneGA, Bioconductor). The resulting sequence was evolved through several cycles to remove repeats longer than 11 bp (forward or reverse) (see Figure 12). In each cycle, synonymous sequences for each repeat were evaluated. The sequence with the maximum Hamming distance to the evolving DNA was selected. One of the re-TALE sequences with 16.5 monomers is as follows:
[0055]
number
[0056]
number
[0057] In one embodiment, a TALE having at least 80% sequence identity, at least 85% sequence identity, at least 90% sequence identity, at least 95% sequence identity, at least 98% sequence identity, or at least 99% sequence identity with respect to the above sequence can be used. Those skilled in the art will readily understand which parts of the above sequence can be altered while maintaining the DNA-binding activity of the TALE.
[0058] Two re-TALE dimer blocks encoding two RVDs (see Figure 6A) were constructed by two rounds of PCR under standard Kapa HIFI (KPAP) PCR conditions. The first round of PCR introduced the RVD coding sequence, and the second round of PCR created the entire dimer block with a 36 bp overlap with the adjacent block. The PCR products were purified using the QIAquick 96 PCR purification kit (QIAGEN), and concentrations were measured by nanodrop. The primer and template sequences are listed in Tables 1 and 2 below.
[0059] [Table 1-1]
[0060] [Table 1-2]
[0061] [Table 1-3]
[0062] [Table 2-1]
[0063] [Table 2-2]
[0064] Re-TALEN and re-TALE-TF destination vectors were constructed by modifying the TALE-TF and TALEN cloning backbones (see Reference 24, which is incorporated herein by reference in its entirety). The 0.5RVD region on these vectors was re-edited, and a SapI cleavage site was incorporated into the designated re-TALE cloning site. The sequences of the re-TALEN and re-TALE-TF backbones are provided in Figure 20. Plasmids can be pre-treated with SapI (New England Biolabs) under manufacturer-recommended conditions, and these plasmids can be purified using the QIAquick PCR purification kit (QIAGEN).
[0065] A (10 μl) one-pot TASA assembly reaction was performed using 200 ng of each block, 500 ng of destination backbone, a 1×TASA enzyme mixture (2 U SapI, 100 U ampligase (Epicentre), 10 mU T5 exonuclease (Epicentre), 2.5 U Phusion DNA polymerase (New England Biolabs)), and the previously described 1× isothermal assembly reaction buffer (see Reference 25, which is incorporated herein by reference in its entirety) (5% PEG-8000, 100 mM Tris-HCl pH 7.5, 10 mM MgCl2, 10 mM DTT, 0.2 mM each of four types of dNTPs, and 1 mM NAD). Incubation was performed at 37°C for 5 minutes and at 50°C for 30 minutes. The TASA assembly reaction product can be directly processed for bacterial transformation to colonize individual assemblies. The efficiency of obtaining full-length constructs with this approach is approximately 20%. Alternatively, an efficiency of over 90% can be achieved with a three-step assembly. First, 10 µl of re-TALE assembly reaction is performed using 200 ng of each block, a 1× re-TALE enzyme mixture (100 U ampligase, 12.5 mU T5 exonuclease, 2.5 U Phusion DNA polymerase), and 1× isothermal assembly buffer at 50°C for 30 minutes. This is followed by a standardized Kapa HIFI PCR reaction, agarose gel electrophoresis, and QIAquick gel extraction (Qiagen) to concentrate the full-length re-TALE. Next, 200 ng of re-TALE amplicon can be mixed with 500 ng of Sap1 pre-treated destination backbone, 1× re-TALE assembly mixture, and 1× isothermal assembly reaction buffer, and incubated at 50°C for 30 minutes. The re-TALE final assembly reaction product can be directly treated for bacterial transformation to colonize individual assemblies. Those skilled in the art will be able to easily select endonucleases, exonucleases, polymerases, and ligases from known ones to carry out the method described herein.For example, type II endonucleases such as FokI, BtsI, EarI, and SapI can be used. Titralable exonucleases such as lambda exonuclease, T5 exonuclease, and exonuclease III can be used. Non-hot-start polymerases such as Phusion DNA polymerase, Taq DNA polymerase, and VentR DNA polymerase can be used. Thermostable ligases such as Amprigase, Pfu DNA ligase, and Taq DNA ligase can be used in this reaction. In addition, depending on the specific species used, various reaction conditions can be used to activate such endonucleases, exonucleases, polymerases, and ligases.
[0066] Example 3 Cell lines and cell cultures PGP1 iPS cells were maintained on Matrigel (BD Biosciences) coated plates using mTeSR1 (Stemcell Technologies). Cultures were subculturified every 5-7 days using TrypLE Express (Invitrogen). 293T and 293FT cells were cultured and maintained in high-glucose Dulbecco's modified Eagle medium (DMEM, Invitrogen) supplemented with 10% fetal bovine serum (FBS, Invitrogen), penicillin / streptomycin (pen / strep, Invitrogen), and non-essential amino acids (NEAA, Invitrogen). K562 cells were cultured and maintained in RPMI (Invitrogen) supplemented with 10% fetal bovine serum (FBS, Invitrogen, 15%) and penicillin / streptomycin (pen / strep, Invitrogen). All cells were maintained in a humidified incubator at 37°C and 5% CO2.
[0067] Stable 293T cell lines for detecting homology repair (HDR) efficiency were constructed, as described in their entirety in Reference 26, which is incorporated herein by reference. Specifically, these reporter cell lines hold a GFP coding sequence integrated onto a genome disrupted by the insertion of a 68 bp genomic fragment derived from a stop codon and the AAVS1 locus.
[0068] Example 4 Testing of re-TALEN activity 293T reporter cells were seeded at a density of 2 × 10⁵ cells per well in a 24-well plate, and 1 μg of each re-TALEN plasmid and 2 μg of DNA donor plasmid were transfused into these cells using Lipofectamine 2000 according to the manufacturer's protocol. Approximately 18 hours after transfusion, the cells were harvested using TrypLE Express (Invitrogen) and resuspended in 200 μl of medium for flow cytometry analysis using an LSR Fortessa cell analyzer (BD Biosciences). Flow cytometry data were analyzed using FlowJo (FlowJo). At least 25,000 events were analyzed for each transfused sample. For endogenous AAVS1 locus targeting experiments in 293T, the transfusion method was the same as described above, and puromycin selection was performed at a drug concentration of 3 μg / ml one week after transfusion.
[0069] Example 5 Functional lentiviral generation evaluation Lentiviral vectors were constructed using standard PCR and cloning techniques. Lentiviruses were generated by transposing the lentiviral plasmid into 293FT cultured cells (Invitrogen) using lipofectamine 2000 with a lentiviral packaging mix (Invitrogen). The supernatant was collected 48 and 72 hours after transposition, sterile filtered, and 100 µl of the filtered supernatant was added to 5 × 10⁵ new 293T cells along with polyblen. The lentiviral infectivity titer was calculated based on the following formula: Viral infectivity titer = (Percentage of GFP-positive 293T cells × Initial number of cells at transduction) / (Volume of original viral supernatant used in the transduction experiment). To test the functionality of the lentivirus, three days after transduction, a plasmid containing 30 ng of mCherry reporter and a plasmid containing 500 ng of pUC19 plasmid were transfused into lentivirus-transfused 293T cells using lipofectamine 2000 (Invitrogen). Eighteen hours after transduction, the cell morphology was analyzed using an Axio Observer Z.1 (Zeiss), the cells were harvested using TrypLE Express (Invitrogen), and resuspended in 200 μl of medium for flow cytometry analysis using an LSR Fortessa cell analyzer (BD Biosciences). Flow cytometry data were analyzed using a BD FACS Diva (BD Biosciences).
[0070] Example 6 Testing of genome editing efficiency of re-TALEN and Cas9-gRNA Two hours prior to nucleofection, PGP1 iPS cells were cultured in the Rho kinase (ROCK) inhibitor Y-27632 (Calbiochem). Translocation was performed using the P3 Primary Cell 4D Nucleofector X Kit (Lonza). Specifically, cells were harvested using TrypLE Express (Invitrogen) and 2 × 10⁶ cells were resuspended in a 20 μl nucleofection mixture containing 16.4 μl of P3 nucleofector solution, 3.6 μl of supplements, 1 μg of each re-TALEN plasmid or 1 ug of Cas9 and 1 ug of gRNA constructs, and 2 μl of 100 μm of single-stranded oligodeoxyribonucleotides. The mixture was then transferred to a 20 μl nucleocuvette strip, and nucleofection was performed using the CB150 program. For the first 24 hours, cells were seeded in mTeSR1 medium supplemented with a ROCK inhibitor on Matrigel-coated plates. For endogenous AAVS1 locus targeting experiments using double-stranded DNA donors, the same procedure was followed, except that 2 μg of double-stranded DNA donor was used and puromycin was added to mTeSR1 medium at a concentration of 0.5 ug / mL one week after translocation.
[0071] Information on the reTALEN, gRNA, and single-stranded oligodeoxyribonucleotides used in this example is provided in Tables 3 and 4 below.
[0072] [Table 3-1]
[0073] [Table 3-2]
[0074] [Table 4-1]
[0075] [Table 4-2]
[0076] [Table 4-3]
[0077] [Table 4-4]
[0078] [Table 4-5]
[0079] [Table 4-6]
[0080] [Table 4-7]
[0081] [Table 4-8]
[0082] [Table 4-9]
[0083] Example 7 Preparation of an amplicon library for targeting regions Cells were harvested 6 days after nucleofection, and 0.1 μl of prepGEM tissue protease enzyme (ZyGEM) and 1 μl of prepGEM gold buffer (ZyGEM) were added to 2–5 × 10⁵ cells in 8.9 μl of culture medium. Then, 1 μl of the reaction product was added to 9 μl of PCR mix containing 5 μl of 2 × KAPA HIFI hot start ready mix (KAPA Biosystems) and the corresponding 100 nm amplification primer pair. The reaction product was incubated at 95°C for 5 minutes, followed by 15 cycles of 98°C for 20 seconds, 65°C for 20 seconds, and 72°C for 20 seconds. To add the Illumina sequencing adapter, 5 μl of reaction product was added to 20 μl of PCR mix containing 12.5 μl of 2×KAPA HIFI hot-start ready mix (KAPA Biosystems) and a primer with a 200 nM Illumina sequencing adapter. The reaction product was incubated at 95°C for 5 minutes, followed by 25 cycles of 98°C for 20 seconds, 65°C for 20 seconds, and 72°C for 20 seconds. The PCR product was purified using the QIAquick PCR purification kit, mixed at approximately the same concentration, and sequenced using a MiSeq personal sequencer. The PCR primers are listed in Table 5 below.
[0084] [Table 5-1]
[0085] [Table 5-2]
[0086] [Table 5-3]
[0087] [Table 5-4]
[0088] [Table 5-5]
[0089] Example 8 Genome Editing Assessment System (GEAS) Next-generation sequencing is being used to detect rare genomic alterations. See references 27–30, which are incorporated herein by reference for their entirety. To enable the widespread use of this approach for rapidly evaluating homology repair (HDR) and non-homologous end joining (NHEJ) efficiency in human iPS cells, we have created software called a "pipeline" for analyzing genome engineering data. This pipeline is integrated into a single Unix module, which utilizes various tools such as R, BLAT, and the FASTX toolkit.
[0090] Barcode Splitting: A group of samples were combined and sequenced using MiSeq 150bp paired-end (PE150) (Illumina Next Generation Sequencing), followed by DNA barcode-based separation using the FASTX toolkit.
[0091] Quality filtering: Nucleotides with low sequence quality (phred score < 20) were removed. After removal, reads shorter than 80 nucleotides were excluded.
[0092] Mapping: Paired reads were independently mapped to a reference genome using BLAT, and a .psl file was created as output.
[0093] Indel calling: Indels were defined as full-length reads containing two block matches during alignment. Only reads following this pattern were considered for both paired-end reads. For quality control, these indel reads were required to have a minimum 70 nt match with the reference genome, and both blocks were required to be at least 20 nt in length. The size and location of the indels were calculated based on the position of each block relative to the reference genome. Non-homologous end joining (NHEJ) was evaluated as a percentage of reads containing indels (see Equation 1 below). The majority of non-homologous end joining (NHEJ) events were detected near the targeting site.
[0094] Homology repair (HDR) efficiency: Using pattern matching (grep) within a 12 bp window centered on double-strand breaks (DSBs), we counted specific signatures corresponding to reads containing the reference sequence, reads containing modifications to the reference sequence (2 bp intended mismatch), and reads containing only 1 bp mutation within a 2 bp intended mismatch (see Equation 1 below).
[0095] Equation 1. Evaluation of non-homologous end joining (NHEJ) and homology repair (HDR). A = Same lead as the standard: XXXXXABXXXXX B = Reads containing a 2bp mismatch programmed by single-stranded oligodeoxyribonucleotide:XXXXXabXXXXX C = Reads containing only 1 bp of mutation within the target site: e.g., XXXXXaBXXXXX or XXXXXAbXXXXX D = Leads containing the indels listed above.
[0096]
number
[0097]
number
[0098] Example 9 Genotyping screening of colonized human iPS cells Prior to FACS sorting, human iPS cells in feeder-free cultures were pretreated for at least 2 hours in mTesr-1 medium supplemented with SMC4 (5 μM thiazovibin, 1 μM CHIR99021, 0.4 μM PD0325901, 2 μM SB431542) (see Reference 23, which is incorporated herein by reference in its entirety). The cultures were dissociated using accutase (Millipore) and resuspended at a concentration of 1–2 × 10⁷ cells / mL in mTesr-1 medium supplemented with SMC4 and the viability dye ToPro-3 (Invitrogen). Under sterile conditions, viable human iPS cells were single-cell sorted into 96-well plates (Global Stem) coated with irradiated CF-1 mouse embryonic fibroblasts using a BD FACS Aria II SORP UV (BD Biosciences) with a 100 μm nozzle. Each well contained human ES cell medium (see Reference 31, which is incorporated herein by reference) containing 100 ng / ml recombinant human basic fibroblast growth factor (bFGF) (Millipore) supplemented with SMC4 and 5 μg / ml fibronectin (Sigma). After sorting, the plates were centrifuged at 70 × g for 3 minutes. Colony formation was observed 4 days after sorting, and the medium was replaced with human ES cell medium supplemented with SMC4. SMC4 could be removed from the human ES cell medium 8 days after sorting.
[0099] Eight days after fluorescence-activated cell sorting (FACS), several thousand cells were harvested, and 0.1 μl of prepGEM tissue protease enzyme (ZyGEM) and 1 μl of prepGEM gold buffer (ZyGEM) were added to the cells in 8.9 μl of culture medium. The reaction product was then added to 40 μl of PCR mix containing 35.5 ml of platinum 1.1 × supermix (Invitrogen), 250 nM dNTPs, and 400 nM primers. The reaction product was incubated at 95°C for 3 minutes, followed by 30 cycles of 95°C for 20 seconds, 65°C for 30 seconds, and 72°C for 20 seconds. The product was Sanger sequenced using one of the PCR primers listed in Table 5, and the sequences were analyzed using DNASTAR (DNASTAR).
[0100] Example 10 Immunostaining and teratoma assay of human iPS cells Cells were incubated in KnockOut DMEM / F-12 medium at 37°C for 60 minutes using the following antibodies: anti-SSEA-4 PE (Millipore) (diluted 1:500); Tra-1-60 (BD Pharmingen) (diluted 1:100). After incubation, cells were washed three times with KnockOut DMEM / F-12 and images were obtained using an Axio Observer Z.1 (ZIESS).
[0101] To analyze teratoma formation, human iPS cells were harvested using type IV collagenase (Invitrogen), resuspended in 200 μl of Matrigel, and intramuscularly injected into the hind limbs of Rag2γ knockout mice. Teratomas were isolated 4–8 weeks after injection and fixed in formalin. The teratomas were then analyzed by hematoxylin-eosin staining.
[0102] Example 11 Targeting genomic loci in human somatic cells and human stem cells using reTALENS In one embodiment, a TALE known to those skilled in the art is modified or re-edited to remove repetitive sequences. Such TALEs suitable for modification and use in viral delivery vehicles and genome editing methods in various cell lines and organisms described herein are disclosed in their entirety in References 2, 7-12, which are incorporated herein by reference. Several strategies have been developed for assembling repetitive TALE RVD array sequences (see References 14 and 32-34, which are incorporated herein by reference in their entirety). However, even after assembly, the TALE sequence repeats remain unstable, thus limiting the broad utility of this tool, particularly its utility in viral gene delivery vehicles (see References 13 and 35, which are incorporated herein by reference in their entirety). Therefore, one embodiment of this disclosure relates to a TALE lacking repeats, for example, one completely lacking repeats. Such re-edited TALEs have the advantage of enabling faster and simpler synthesis of extended TALE RVD arrays.
[0103] To eliminate repeats, the nucleotide sequence of the TALE RVD array was computationally evolved to minimize the number of sequence repeats while maintaining the amino acid composition. The re-edited TALE (Re-TALE) encoding 16 tandem RVD DNA recognition monomers and the last half of the RVD repeats completely lacks the 12 bp repeat (see Figure 5a). Notably, this level of re-editing is sufficient to enable PCR amplification of any specific monomer or subunit derived from the full-length re-TALE construct (see Figure 5b). The improved re-TALE design can be synthesized using standard DNA synthesis techniques without incurring the additional costs or methods associated with highly repetitive sequences (see Reference 36, which is incorporated herein by reference in its entirety). Furthermore, the re-edited sequence design enables efficient assembly of the re-TALE construct using the modified isothermal assembly reaction described in the methods herein and with reference to Figure 6.
[0104] We performed statistical analysis on genome editing NGS (next-generation sequencing) data as follows: For homology repair (HDR) specificity analysis, we used an exact binomial test to calculate the probability of observing various numbers of sequence reads, including 2bp mismatches. Based on the sequencing results of 10bp windows before and after the targeting site, we evaluated the maximum base change rates in two windows (P1 and P2). Using the null hypothesis that the changes in each of the two target base pairs are independent, we calculated the expected probability of accidentally observing a 2bp mismatch at the targeting site as the product of these two probabilities (P1 × P2). Based on a dataset containing N total reads and n homology repair (HDR) reads, we calculated the p-value of the observed homology repair (HDR) efficiency. For homology repair (HDR) sensitivity analysis, the single-stranded oligodeoxyribonucleotide DNA donor contained a 2bp mismatch with the targeting genome, making it highly probable that the integration of this single-stranded oligodeoxyribonucleotide into the targeting genome would result in the coexistence of base changes in the two target base pairs. The probability of other unexpected observed sequence changes occurring simultaneously was low. Therefore, the interdependence of unexpected changes was much lower. Based on these assumptions, the interdependence of simultaneous changes in the two base pairs at all other paired sites was measured using mutual information (MI), and the homology repair (HDR) detection limit was evaluated as the minimum homology repair (HDR) when the MI at the target 2bp site was greater than the MI of all other pairs combined. For a given experiment, a set of fastq files with diluted homology repair (HDR) efficiency was simulated by identifying homology repair (HDR) reads with a predetermined 2bp mismatch from the original fastq file and systematically removing varying numbers of homology repair (HDR) reads from the original dataset. Mutual information (MI) was calculated between all pairs of positions within a 20bp window centered on the targeting site. In these calculations, the mutual information of base composition between any two positions was calculated.Unlike the homology repair (HDR) specificity measurement described above, this measurement does not assess the tendency of positional pairs to change to specific target base pairs, but only the tendency of those positional pairs to change simultaneously (see Figure 8A). Table 6 shows the efficiencies of homology repair (HDR) and non-homologous end joining (NHEJ) of re-TALEN / single-stranded oligodeoxyribonucleotides targeting CCR5, as well as the NHEL efficiency of Cas9-gRNA. We coded our analysis in R and calculated the MI using the infotheo package.
[0105] [Table 6]
[0106] The correlation between genome editing efficiency and epigenetic status was addressed as follows: Pearson correlation coefficients were calculated to test possible relationships between epigenetic parameters (DNaseI HS (highly sensitive sites) or nucleosome occupancy) and genome engineering efficiency (homologous repair (HDR), non-homologous end joining (NHEJ)). A dataset of DNAaseI hypersensitivity was downloaded from the UCSC Genome Browser: hiPSCs DNase I HS: / gbdb / hg19 / bbi / wgEncodeOpenChromDnaseIpsnihi7Sig.bigWig
[0107] To calculate the p-value, observed correlations were compared to simulated distributions constructed by randomizing the positions of epigenetic parameters (N=100,000). Observed correlations higher than the 95th percentile or lower than the 5th percentile were considered latent relationships.
[0108] The function of reTALEN in human cells was determined compared to the corresponding non-re-edited TALEN. The HEK293 cell line, containing a GFP reporter cassette with a frameshift insertion, was used as described in reference 37, which is incorporated herein by reference. See also Figure 1a. Delivery of a TALEN or reTALEN targeting the insertion sequence, along with a promoter-less GFP donor construct, induced double-strand break (DSB)-induced homology repair (HDR) of the GFP cassette, allowing for evaluation of nuclease cleavage efficiency using GFP repair efficiency. See reference 38, which is incorporated herein by reference. reTALEN induced GFP repair in 1.4% of transfected cells, similar to that achieved by TALEN (1.2%) (see Figure 1b). We examined the activity of reTALEN at the AAVS1 locus in PGP1 human iPS cells (see Figure 1c), and successfully recovered cell clones containing specific insertions (see Figures 1d and e), confirming that reTALEN is active in both human somatic cells and human pluripotent cells.
[0109] Removal of repeats enabled the generation of functional lentiviruses using re-TALE cargo. Specifically, lentiviral particles encoding re-TALE-2A-GFP were packaged, and the activity of re-TALE-TF encoded by the lentiviral particles was examined by transfecting a pool of lenti-reTALE-2A-GFP-infected 293T cells with an mCherry reporter. 293T cells transfected with lenti-re-TALE-TF showed 36-fold increased reporter expression activation compared to negative results with reporter alone (see Figures 7a, b, and c). Sequence integrity of re-TALE-TF in lentivirus-infected cells was examined, and full-length reTALE was detected in all 10 clones examined (see Figure 7d).
[0110] Example 12 Comparison of ReTALE efficiency and Cas9-gRNA efficiency in human iPS cells using the Genome Editing Assessment System (GEAS). To compare the editing efficiency of re-TALEN and Cas9-gRNA in human iPS cells, a next-generation sequencing platform (genome editing evaluation system) was developed to identify and quantify both non-homologous end joining (NHEJ) and homology repair (HDR) gene editing events. Both re-TALEN pairs and Cas9-gRNAs (re-TALEN and Cas9-gRNA pair number 3 in Table 3), targeting the upstream region of CCR5, were designed and constructed with 90nt single-strand oligodeoxyribonucleotide donors identical to the target site except for a 2bp mismatch (see Figure 2a). These nuclease constructs and donor single-strand oligodeoxyribonucleotides were transfused into human iPS cells. To quantify gene editing efficiency, paired-end deep sequencing of the target genomic region was performed 3 days after transfusion. Homologous repair (HDR) efficiency was measured by the percentage of reads containing the exact 2bp mismatch. Non-homologous end joining (NHEJ) efficiency was measured by the percentage of reads containing indels.
[0111] When single-stranded oligodeoxyribonucleotides alone were delivered to human iPS cells, the lowest homology repair (HDR) and non-homologous end joining (NHEJ) rates were obtained. However, when re-TALEN and single-stranded oligodeoxyribonucleotides were delivered, a homology repair (HDR) efficiency of 1.7% and a non-homologous end joining (NHEJ) efficiency of 1.2% were obtained (see Figure 2b). When Cas9-gRNA was introduced together with single-stranded oligodeoxyribonucleotides, a homology repair (HDR) efficiency of 1.2% and a non-homologous end joining (NHEJ) efficiency of 3.4% were obtained. Notably, the rates of genomic deletion and insertion peaked in the middle of the spacer region between the two reTALEN binding sites, while peaks were reached 3-4 bp upstream of the protospacer adjacent motif (PAM) sequence of the Cas9-gRNA targeting site (see Figure 2b). This result was expected, as double-strand breaks occur in these regions. A median genomic deletion size of 6 bp and an insertion size of 3 bp were observed with re-TALEN, and a median deletion size of 7 bp and an insertion size of 1 bp were observed with Cas9-gRNA (see Figure 2b), which was consistent with the DNA damage pattern typically caused by non-homologous end joining (NHEJ) (see Reference 4, which is incorporated herein by reference in its entirety). Several analyses of the next-generation sequencing platform revealed that the genome editing evaluation system (GEAS) could detect homology repair (HDR) at a low rate of 0.007%, which was highly reproducible (coefficient of variation between replicas = ±15% × measured efficiency) and 400 times more sensitive than the most commonly used mismatch-sensitive endonuclease assays (see Figure 8).
[0112] We constructed re-TALEN pairs and Cas9-gRNA targeting 15 sites in the CCR5 genomic locus and measured their editing efficiency (see Figure 2c and Table 3). These sites were selected to represent a broad range of DNase I sensitivity (see reference 39, which is incorporated herein by reference in its entirety). The nuclease constructs, along with their corresponding single-strand oligodeoxyribonucleotide donors (see Table 3), were transfused into PGP1 human iPS cells. Six days after transfusion, we created genome editing efficiency profiles for these sites (Table 6). Non-homologous end joining (NHEJ) and homology repair (HDR) were detected above the statistical detection threshold for 13 of the 15 re-TALEN pairs + single-strand oligodeoxyribonucleotide donors, with an average non-homologous end joining (NHEJ) efficiency of 0.4% and an average homology repair (HDR) efficiency of 0.6% (see Figure 2c). In addition, a statistically significant positive correlation (r²=0.81) was found between homologous recombination (HR) efficiency and non-homologous end joining (NHEJ) efficiency at the same targeting locus (P<1×10⁻⁴) (see Figure 9a), suggesting that double-strand break (DSB) generation, an upstream process common to both homology repair (HDR) and non-homologous end joining (NHEJ), is the rate-limiting step in reTALEN-mediated genome editing.
[0113] In contrast, all 15 Cas9-gRNA pairs showed significant levels of non-homologous end joining (NHEJ) and homologous recombination (HR), with an average NHEJ efficiency of 3% and an average homology repair (HDR) efficiency of 1.0% (see Figure 2c). In addition, a positive correlation was detected between the efficiency of NHEJ and homology repair induced by Cas9-gRNA (see Figure 9b) (r²=0.52, p=0.003), which was consistent with the observations for reTALEN. The NHEJ efficiency achieved by Cas9-gRNA was significantly higher than that achieved by reTALEN (t-test, paired-end, P=0.02). A moderate but statistically significant correlation was observed between non-homologous end joining (NHEJ) efficiency and the melting point of the gRNA targeting sequence (see Figure 9c) (r²=0.28, p=0.04), suggesting that the 28% variability in Cas9-gRNA-mediated double-strand break (DSB) generation efficiency can be explained by the strength of base pairing between the gRNA and its genomic target. Despite Cas9-gRNA generating an average of 7 times higher non-homologous end joining (NHEJ) level than its corresponding reTALEN, Cas9-gRNA only achieved a homology repair (HDR) level (mean = 1.0%) similar to that of its corresponding reTALEN (mean = 0.6%). While we do not wish to be constrained by scientific theory, these results suggest that the concentration of single-strand oligodeoxyribonucleotides in double-strand breaks (DSBs) is the rate-limiting factor in homology repair (HDR), or that the genome break structure created by Cas9-gRNA is unfavorable for effective homology repair (HDR). No correlation was observed between DNaseI HS (highly sensitive sites) and genome targeting efficiency in either method (see Figure 10).
[0114] Example 13 Optimization of single-strand oligodeoxyribonucleotide donor design for homology repair (HDR) High-performance single-stranded oligodeoxyribonucleotides for human iPS cells were designed as follows: A set of single-stranded oligodeoxyribonucleotide donors of different lengths (50-170 nt) was designed, each having the same 2 bp mismatch in the center of the spacer region of the target site of the third CCR5 re-TALEN pair. Homologous repair (HDR) efficiency was observed to change with the length of the single-stranded oligodeoxyribonucleotide, with an optimal homologous repair (HDR) efficiency of approximately 1.8% observed with 90 nt single-stranded oligodeoxyribonucleotides, while the homologous repair (HDR) efficiency decreased with longer single-stranded oligodeoxyribonucleotides (see Figure 3a). When double-stranded DNA donors are used with nucleases, the longer the homologous region, the better the homology repair (HDR) rate (see reference 40, which is incorporated herein by reference in its entirety). Possible reasons for this result are that, compared to double-stranded DNA donors, single-stranded oligodeoxyribonucleotides are used in alternative genome repair processes, longer single-stranded oligodeoxyribonucleotides are less available to the genome repair machinery, or longer single-stranded oligodeoxyribonucleotides have a negative effect that offsets any improvements gained from longer homology (see reference 41, which is incorporated herein by reference in its entirety). However, if either of the first two reasons is true, since single-stranded oligodeoxyribonucleotide donors are not involved in non-homologous end joining (NHEJ) repair, the non-homologous end joining (NHEJ) rate should be unaffected or increase with increasing single-stranded oligodeoxyribonucleotide length. However, since the rate of non-homologous end joining (NHEJ) was observed to decrease with homology repair (HDR) (see Figure 3a), it was suggested that longer single-stranded oligodeoxyribonucleotides have a counteracting effect.One possible hypothesis is that longer single-stranded oligodeoxyribonucleotides are more toxic to cells (see Reference 42, which is incorporated herein by reference in its entirety), or that translocation of longer single-stranded oligodeoxyribonucleotides saturates the DNA processing machinery, reducing molar DNA uptake and impairing the cell's ability to take up or express re-TALEN plasmids.
[0115] We investigated how the mismatch integration rate in single-stranded oligodeoxyribonucleotide donors changes with the distance from the mismatch to the double-strand break ("DSB"). A series of 90nt single-stranded oligodeoxyribonucleotides were designed, each containing an identical 2bp mismatch (A) in the center of the spacer region of re-TALEN pair 3. Each single-stranded oligodeoxyribonucleotide contained a second 2bp mismatch (B) at various distances from the center (see Figure 3b). Single-stranded oligodeoxyribonucleotides containing only the central 2bp mismatch were used as controls. Each of these single-stranded oligodeoxyribonucleotides was individually introduced with re-TALEN pair 3, and the results were analyzed using the Genome Editing Assessment System (GEAS). We found that overall homology repair (HDR), measured by the rate at which A mismatches are incorporated (A only, or A+B), decreases as the B mismatch moves further away from the center (see Figure 3b and Figure 11a). The higher overall homology repair (HDR) rate observed when B is only 10 bp away from A may reflect a lower need for single-stranded oligodeoxyribonucleotides to anneal to genomic DNA near the double-strand break.
[0116] For each distance from A to B, a certain proportion of homology repair (HDR) events incorporated only A mismatches, while another proportion incorporated both A and B mismatches (see Figure 3b (A only, and A+B)). These two results may be due to gene conversion tracts along the length of the single-stranded DNA oligo (see Reference 43, which is incorporated herein by reference in its entirety), thereby resulting in A+B mismatch incorporation from longer conversion tracts extending beyond the B mismatch, and A-only incorporation from shorter tracts that do not reach B. Under this interpretation, we evaluated the bidirectional distribution of gene conversion lengths along single-stranded oligodeoxyribonucleotides (see Figure 11b). The evaluated distribution showed a gradual decrease in their frequency as the length of the gene conversion tract increased, a result very similar to the gene conversion tract distribution observed for double-stranded DNA donors, but on a much more compressed distance scale of tens of bases for single-stranded DNA donors compared to hundreds of bases for double-stranded DNA donors. Consistent with this result, experiments using single-stranded oligodeoxyribonucleotide donors containing three pairs of 2bp mismatches spaced 10nt apart on both sides of the central 2bp mismatch "A" showed a pattern where A alone was incorporated in 86% of cases, while multiple B mismatches were incorporated in other cases (see Figure 11c). The number of B-only incorporation events was too small to evaluate the distribution of tract lengths less than 10bp, but it is clear that short tract regions of less than 10bp at nuclease sites were dominant (see Figure 11b). Finally, in all experiments with a single B mismatch, a small percentage of B-only inclusion events were observed (0.04% to 0.12%), and this was roughly constant regardless of the distance of B from A.
[0117] Furthermore, we analyzed how far away single-stranded oligodeoxyribonucleotide donors could be placed from re-TALEN-induced double-stranded DNA breaks while observing integration. We targeted a range greater than -600 bp to +400 bp away from the re-TALEN-induced double-stranded DNA break site and examined a set of 90 nt single-stranded oligodeoxyribonucleotides with a central 2 bp mismatch. When the single-stranded oligodeoxyribonucleotide matched more than 40 bp away, we observed a low homology repair (HDR) efficiency of less than 1 / 30th compared to a control single-stranded oligodeoxyribonucleotide located on the break region and centrally (see Figure 3c). The observed low levels of integration may be due to processes unrelated to the double-stranded DNA break, as can be seen in experiments where the genome is modified by single-stranded DNA donors alone (see reference 42, which is incorporated herein by reference in its entirety). On the other hand, the low level of homology repair (HDR) present when single-stranded oligodeoxyribonucleotides are separated by approximately 40 bp may be due to a combination of reduced homology on the mismatch-containing side of the double-stranded DNA break and insufficient single-stranded oligodeoxyribonucleotide oligo length on the opposite side of the double-stranded DNA break.
[0118] Single-stranded oligodeoxyribonucleotides for Cas9-gRNA-mediated targeting DNA donor designs were examined. Cas9-gRNA (C2) targeting the AAVS1 locus was constructed, and single-strand oligodeoxyribonucleotide donors with varying orientations (Oc: complementary to gRNA, and On: non-complementary to gRNA) and lengths (30, 50, 70, 90, 110 nt) were designed. Oc achieved superior efficiency compared to On, with a 70-mer Oc achieving an optimal homology repair (HDR) rate of 1.5% (see Figure 3d). Despite the fact that homology repair (HDR) efficiency mediated by Cas9-derived nicks with single-strand oligodeoxyribonucleotides (Cc: Cas9_D10A) was significantly lower than that of C2 (t-test, paired-end, P=0.02), the same single-strand oligodeoxyribonucleotide bias was detected even when using Cc (see Figure 12).
[0119] Example 14 Isolation of modified human iPS cell clones The Genome Editing Assessment System (GEAS) revealed that re-TALEN pair 3 achieves precise genome editing in human iPS cells with an efficiency of approximately 1%, a level at which precisely edited cells can typically be isolated by clonal screening. As single cells, human iPS cells have low viability. Using the optimized protocol described in Reference 23, which is incorporated herein by reference, along with a single-cell FACS sorting method, we constructed a robust platform for sorting and maintaining single human iPS cells, enabling the recovery of human iPS cell clones with a viability rate of over 25%. Combining this method with a rapid, high-efficiency genotyping system, we enabled large-scale genotyping of edited human iPS cells by performing chromosomal DNA extraction and targeted genome amplification in a 1-hour single-in-vitro reaction. In summary, these methods include a pipeline for robustly obtaining genome-edited human iPS cells without selection.
[0120] To demonstrate this system (see Figure 4a), a pair of re-TALENs targeting CCR5 and a single-stranded oligodeoxyribonucleotide were transfused into PGP1 human iPS cells at site 3 (see Table 3). A genome editing evaluation system (GEAS) was performed on a portion of the transfused cells, and a homology repair (HDR) frequency of 1.7% was found (see Figure 4b). This information, along with a 25% recovery rate of the selected single-cell clones, allows for an assessment that at least one accurately edited clone can be obtained from five 96-well plates with a 98% Poisson probability (assuming μ = 0.017 × 0.25 × φθ × 5 × 2). Human iPS cells were facs-selected six days after transfusion, and 100 human iPS cell clones were screened eight days after selection. Sanger sequencing revealed that 2 out of every 100 unselected human iPS cell colonies contained heterozygous genotypes with a 2bp mutation introduced by a single-strand oligodeoxyribonucleotide donor (see Figure 4c). The targeting efficiency of 1% (1% = 2 / 2 × 100, 2 monoallelic modified clones out of 100 screened cells) was consistent with next-generation sequencing analysis (1.7%) (see Figure 4b). The pluripotency of the resulting human iPS cells was confirmed by immunostaining for SSEA4 and TRA-1-60 (see Figure 4d). The successfully targeted human iPS cell clones were able to form mature teratomas possessing characteristics of all three germ layers (see Figure 4e).
[0121] Example 15 Methods for continuous cell genome editing In one embodiment, a method is provided for genome editing in cells (such as human cells, e.g., human stem cells) by genetically modifying cells to include a nucleic acid encoding an enzyme that forms a co-localization complex with RNA complementary to target DNA and site-specifically cleaves the target DNA. Such enzymes include RNA-induced DNA-binding proteins, e.g., RNA-induced DNA-binding proteins of the type II CRISPR system. One exemplary enzyme is Cas9. In this embodiment, the cells express the enzyme, and guide RNA is provided to the cells from a surrounding medium. The guide RNA and the enzyme form a co-localization complex at the target DNA, where the enzyme cleaves the DNA. Optionally, a donor nucleic acid may be present at the cleavage site for insertion into the DNA, the insertion being, for example, by non-homologous end joining or homologous recombination. According to one embodiment, a nucleic acid encoding an enzyme (e.g., Cas9) that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA is under the influence of a promoter, which can, for example, be activated and silenced. Such promoters are well known to those skilled in the art. One exemplary promoter is the dox (doxycycline) inducible promoter. According to one embodiment, the cell is genetically modified by reversibly inserting a nucleic acid encoding an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA into the cell's genome. Once inserted, the nucleic acid can be removed using a reagent such as a transposase. Thus, the nucleic acid can be easily removed after use.
[0122] According to one embodiment, a serial genome editing system in human induced pluripotent stem cells (hiPSCs) using a CRISPR system is provided. According to an exemplary embodiment, the method includes the use of a human iPS cell line having Cas9 reversibly inserted into the genome (Cas9-hiPSC), and a gRNA modified from its native form to allow migration from the surrounding medium into the cell for use with the Cas9. Such gRNA is treated with a phosphatase to remove a phosphate group. Genome editing in the cells is performed using Cas9 by adding the phosphatase-treated gRNA to tissue culture medium. This approach enables traceless genome editing in human iPS cells with an efficiency of up to 50% in a single day of processing, which is 2 to 10 times more efficient than the best efficiency reported to date. Furthermore, the method is easy to use and significantly less cytotoxic. Embodiments of this disclosure include single editing of human iPS cells for biological research and therapeutic applications, multiple editing of human iPS cells for biological research and therapeutic applications, directed human iPS cell evolution, and phenotypic screening of human iPS cells and their derived cells.
[0123] In some embodiments, in addition to stem cells, other cell lines and organisms described herein can also be used. For example, the method described herein can be used on animal cells such as mouse cells or rat cells, thereby enabling the creation of mouse and rat cells in which Cas9 is stably incorporated, and tissue-specific genome editing can be performed by locally introducing phosphatase-treated gRNA from a medium surrounding the cells. Furthermore, other Cas9 derivatives can be inserted into numerous cell lines and organisms, enabling targeted genome modifications such as sequence-specific nick formation, gene activation, repression, and epigenetic modification.
[0124] Some aspects of this disclosure relate to the creation of stable human iPS cells using Cas9 inserted into the genome. Some aspects of this disclosure relate to the modification of RNA so that the RNA can enter the cell through the cell wall and co-localize with Cas9 while evading the cell's immune response. Such modified guide RNA can achieve optimal translocation efficiency with minimal toxicity. Some aspects of this disclosure relate to optimized genome editing in Cas9-human iPS cells using phosphatase-treated gRNA. Some aspects of this disclosure involve the removal of Cas9 from human iPS cells to achieve traceless genome editing, in which case the nucleic acid encoding Cas9 is reversibly positioned within the cell genome. Some aspects of this disclosure involve biomedical engineering to form desired genetic mutations using human iPS cells having Cas9 inserted into the genome. Such modified human iPS cells can maintain pluripotency and successfully differentiate into various cell types (such as cardiomyocytes) that perfectly reproduce the phenotype of the patient cell line.
[0125] Some aspects of this disclosure include libraries of phosphatase-treated gRNAs for multiple genome editing. Some aspects of this disclosure include constructing libraries of PGP cell lines, each having one to several specified mutations in its genome, which can serve as resources for drug screening. Some aspects of this disclosure include constructing PGP1 cell lines in which all retrotrans elements are barcoded with various sequences in order to track the location and activity of the retrotrans elements.
[0126] Example 16 Generation of stable human iPS cells containing Cas9 inserted into the genome. Cas9 was encoded under a dox-inducible promoter, and the construct was placed in a Piggybac vector that could be inserted into or removed from the genome with the help of a Piggybac transposase. Stable insertion of the vector was confirmed by PCR (see Figure 14). Inducible Cas9 expression was determined by RT-QPCR. Cas9 mRNA levels increased 1000-fold 8 hours after the addition of 1 ug / mL of DOX to the culture medium, and Cas9 mRNA levels decreased to normal levels approximately 20 hours after the removal of DOX (see Figure 15).
[0127] According to one embodiment, the Cas9-human iPS cell system-based genome editing bypasses the Cas9 plasmid / RNA translocation procedure, which is a large construct with translocation efficiency typically less than 1% in human iPS cells. This Cas9-human iPS cell system can serve as a platform for performing highly efficient genome engineering in human stem cells. In addition, the Cas9 cassette introduced into human iPS cells using the Piggybac system can be easily removed from the genome by transposase introduction.
[0128] Example 17 Phosphatase-treated guide RNA To enable serial genome editing in Cas9-human iPS cells, a series of modified RNAs encoding gRNAs were created and added to Cas9-iPS culture medium in conjugate form with liposomes. Phosphatase-treated unmodified RNA without any capping achieved an optimal homology repair (HDR) efficiency of 13%, which is 30 times higher than that of previously reported 5' cap-added modified RNAs (see Figure 16).
[0129] In one embodiment, the guide RNA is physically bound to the donor DNA. This provides a method for linking Cas9-mediated genomic breaks with single-strand oligodeoxyribonucleotide-mediated homology repair (HDR), thereby promoting sequence-specific genome editing. Optimal concentrations of DNA and gRNA bound to single-strand oligodeoxyribonucleotide donors achieved 44% homology repair (HDR) and 2% nonspecific non-homologous end joining (NHEJ) (see Figure 17). It is noteworthy that this method does not result in the visible toxicity observed with nucleofection or electroporation.
[0130] In one aspect, the disclosure provides an in vitro modified RNA structure that encodes a gRNA and works in cooperation with Cas9 inserted into the genome to achieve high translocation efficiency and genome editing efficiency. In addition, the disclosure provides a gRNA-DNA chimeric construct that links genome cleavage events with homology-directed recombination reactions.
[0131] Example 18 Reversible removal of Cas9 from human iPS cells to achieve traceless genome editing. According to one embodiment, a Cas9 cassette is inserted into the genome of human iPS cells using a reversible vector. Thus, a Cas9 cassette was reversibly inserted into the genome of human iPS cells using a PiggyBac vector. The Cas9 cassette was removed from the genome-edited human iPS cells by transposation of the cells with a plasmid encoding a transposase. Therefore, some embodiments of this disclosure involve the use of a reversible vector known to those skilled in the art. A reversible vector is, for example, a vector that can be inserted into the genome and subsequently removed by a corresponding vector removal enzyme. Such vectors and corresponding vector removal enzymes are known to those skilled in the art. Colonized iPS cells were screened, and colonies lacking the Cas9 cassette were collected, as confirmed by a PCR reaction. Therefore, this disclosure provides a genome editing method that does not affect the rest of the genome by having a permanent Cas9 cassette present in the cells.
[0132] Example 19 Genome editing in iPGP1 cells Research into the pathogenesis of cardiomyopathy has historically been hampered by the lack of suitable model systems. The differentiation of patient-derived induced pluripotent stem cells (iPS cells) into cardiomyocytes offers a promising pathway to overcome this barrier, and reports of iPS cell modeling of cardiomyopathy are beginning to emerge. However, realizing this prospect requires approaches to overcome the genetic heterogeneity of patient-derived iPS cell lines.
[0133] Using the Cas9-iPGP1 cell line and phosphatase-treated guide RNA bound to DNA, three isogenic iPS cell lines were generated, excluding the sequence of TAZ exon 6, which was identified as carrying a single nucleotide deletion in patients with Baas syndrome. A homology repair (HDR) efficiency of approximately 30% was achieved with a single RNA transfusion. Modified Cas9-iPGP1 cells with the desired mutation were colonized (see Figure 18), and the cell lines were differentiated into cardiomyocytes. Cardiomyocytes derived from the modified Cas9-iPGP1 adequately reproduced the cardiolipin, mitochondrial, and ATP defects observed in patient-derived iPS cells and neonatal rat TAZ knockdown models (see Figure 19). Thus, a method is provided for correcting disease-causing mutations in pluripotent cells and subsequently differentiating those cells into desired cell types.
[0134] Example 20 material and method 1. Construction of stable human iPS / ES strains induced by PiggyBac Cas9 dox. 1. After the cells reach 70% density, pretreat the culture overnight with a final concentration of 10 μM ROCK inhibitor Y27632. 2. The following day, prepare the nucleofection solution by mixing 82 μl of human stem cell nucleofector solution and 18 μl of additive 1 in a 1.5 ml sterile Eppendorf tube. Mix thoroughly. Incubate the solution at 37°C for 5 minutes. 3. Aspirate the mTeSR1 and gently wash the cells with 2 mL of DPBS per well in a 6-well plate. 4. Aspirate the DPBS, add 2 mL / well of Versene, and return the culture to a 37°C incubator until the cells become rounded and loosely adhered, but not detached. This should take 3–7 minutes. 5. Gently aspirate Versene and add mTeSR1. Add 1 ml of mTeSR1 and gently flow the mTeSR1 over the cells using a 1,000 uL micropipette to detach the cells. 6. The detached cells are collected, gently triturated into a single-cell suspension, quantified using a hemocytometer, and the cell density is adjusted to 1 million cells per ml. 7. Add 1 ml of cell suspension to a 1.5 ml Eppendorf tube and centrifuge in a benchtop centrifuge at 1100 RPM for 5 minutes. 8. Resuspend the cells in 100 μl of human stem cell nucleofector solution from step 2. 9. Transfer the cells to a nucleofector cuvette using a 1 ml pipette tip. Add 1 μg of plasmid transposonase and 5 μg of PB Cas9 plasmid to the cell suspension in the cuvette. Gently swirl to mix the cells and DNA. 10. Place the cuvette into the nucleofector. Program B-016 is selected. Press button X to nucleofect the cells. 11. After nucleofection, 500 µl of mTeSR1 medium containing a ROCK inhibitor is added to the cuvette. 12. Using the provided Pasteur plastic pipette, aspirate the nucleofected cells from the cuvette. Then, drop the cells into the Matrigel-coated wells of a 6-well plate of mTeSR1 medium containing the ROCK inhibitor. Incubate the cells overnight at 37°C. 13. The following day and 72 hours after translocation, the culture medium is changed to mTesr1. Puromycin is added at a final concentration of 1 ug / ml. The strain is then established within 7 days.
[0135] 2. RNA preparation 1. Prepare a DNA template with a T7 promoter upstream of the gRNA coding sequence. 2. Purify the DNA using MegaClear Purification and normalize its concentration. 3. Prepare custom-made NTP mixtures for various gRNA productions.
[0136] [Table 7-1]
[0137] [Table 7-2]
[0138] [Table 7-3]
[0139] [Table 7-4]
[0140] 4. Prepare the in vitro transfer mix at room temperature.
[0141] [Table 8]
[0142] 5. Incubate at 37°C (thermocycler) for 4 hours (3-6 hours is also acceptable). 6. Add 2 μl of Turbo DNAse (Ambion's MEGAscript kit) to each sample. Mix gently and incubate at 37°C for 15 minutes. 7. Purify the DNAse-treated reaction product using Ambion's Megaclear according to the manufacturer's instructions. 8. Purify the RNA using MegaClear (the purified RNA can be stored at -80°C for several months). 9. Remove the phosphate group to avoid the host cell's Toll2 immune response.
[0143] [Table 9]
[0144] 3. RNA translocation 1. Seed 10-20K cells per 48 wells without the use of antibiotics. For translocation, the cells should be 30-50% confluent. 2. At least two hours before translocation, change the cell medium to one containing B18R (200 ng / ml), DOX (1 ug / ml), and puromycin (2 ug / ml). 3. Prepare a translocation reagent containing gRNA (0.5 ug to 2 ug), donor DNA (0.5 ug to 2 ug), and RNAiMax, incubate the mixture at room temperature for 15 minutes, and then transfer it into the cells.
[0145] 4. Single human iPS cell seeding and single clone pickup 1. Four days after dox induction and one day after dox removal, aspirate the culture medium and gently wash with DPBS. Add 2 mL / well of Versene and return the culture to a 37°C incubator until the cells become rounded and loosely adhered but not detached. This takes 3–7 minutes. 2. Gently aspirate Versene and add mTeSR1. Add 1 ml of mTeSR1 and gently flow the mTeSR1 over the cells using a 1,000 uL micropipette to detach the cells. 3. Collect the detached cells, gently grind them into a single-cell suspension, quantify them using a hemocytometer, and adjust the cell density to 100K cells per ml. 4. Seed the cells in 10cm dishes coated with Matrigel containing mTeSR1 and a ROCK inhibitor at cell densities of 50K, 100K, and 400K per 10cm dish.
[0146] 5. Screening of single-cell-forming clones 1. Culture the clones in a 10cm dish for 12 days until they are large enough to be identified with the naked eye, then label the clones with colon markers. Ensure that the clones do not grow too large and adhere to each other. 2. Place the 10cm dish in the culture hood and use a P20 pipette with a filter tip (set to 10ul) to aspirate 10ul of medium into one well of the 24-well plate. Pick up the clones by scraping them into smaller pieces and transfer them to one well of the 24-well plate. Change the filter tip for each clone. 3. After 4-5 days, the clone in one well of the 24-well plate will be large enough to split. 4. Aspirate the culture medium and wash with 2 mL / well of DPBS. 5. Aspirate the DPBS and replace it with 250 ul / well of dispase (0.1 U / mL), then incubate the cells in the dispase at 37°C for 7 minutes. 6. Replace the dispase with 2 ml of DPBS. 7. Add 250 µl of mTeSR1. Use a cell scraper to detach the cells and collect them. 8. Transfer 125 µl of cell suspension into the wells of a Matrigel-coated 24-well plate. 9. Transfer 125 µl of cell suspension to a 1.5 ml Eppendorf tube for genomic DNA extraction.
[0147] 6. Clone screening 1. Centrifuge the tube from step 7.7. 2. Aspirate the culture medium and add 250 μl of lysis buffer per well (10 mM Tris pH 7.5 (or 8.0), 10 mM EDTA, 10 mM). 3. NaCl + 10% SDS + 40 ug / mL of proteinase K (add fresh before using the buffer). 4. Incubate at 55°C overnight. 5. Add 250 µl of isopropanol to precipitate the DNA. 6. Centrifuge at high speed for 30 minutes. Wash with 70% ethanol. 7. Gently remove the ethanol. Allow to air dry for 5 minutes. 8. Resuspend the gDNA in 100-200 µl of dH₂O. 9. PCR amplification of a target genomic region using specific primers. 10. Sanger sequencing of PCR products using each primer. 11. Analysis of Sanger sequence data and propagation of target clones.
[0148] 7. Removal of the Piggybac vector 1. Repeat steps 2.1 to 2.9. 2. Transfer the cells to a nucleofector cuvette using a 1 ml pipette tip. Add 2 μg of transposonase plasmid to the cell suspension in the cuvette. Gently swirl to mix the cells and DNA. 3. Repeat steps 2.10 to 2.11. 4. Using the provided Pasteur plastic pipette, aspirate the nucleofected cells from the cuvette. Then, drop the cells into a 10 cm dish of Matrigel-coated wells containing mTeSR1 medium and a ROCK inhibitor. Incubate the cells overnight at 37°C. 5. The following day, change the culture medium to mTesr1 and change the medium daily for the next four days. 6. Once the clones have grown sufficiently, select 20-50 clones and sow them in 24 wells. 7. Determine the genotype of the clones using PB Cas9 PiggyBac vector primers and growth-negative clones.
[0149] References References are designated throughout this specification by the following numbers and are incorporated herein as if they were fully described herein. Each of the subsequent references is incorporated herein by reference in its entirety. 1. Carroll, D. (2011) Genome engineering with zinc-finger nucleases. Genetics, 188, 773-82. 2. Wood,A.J., Lo,T.-W., Zeitler,B., Pickle,C.S., Ralston,E.J., Lee,A.H., Amora,R., Miller,J.C., Leung,E., Meng,X., et al. (2011) Targeted genome editing across species using ZFNs and TALENs. Science (New York, N.Y.), 333, 307. 3. Perez-Pinera,P., Ousterout,D.G. and Gersbach,C.A. (2012) Advances in targeted genome editing. Current opinion in chemical biology, 16, 268-77. 4. Symington,L.S. and Gautier,J. (2011) Double-strand break end resection and repair pathway choice. Annual review of genetics, 45, 247-71.
[0150] 5. Urnov,F.D., Miller,J.C., Lee,Y.-L., Beausejour,C.M., Rock,J.M., Augustus,S., Jamieson,A.C., Porteus,M.H., Gregory,P.D. and Holmes,M.C. (2005) Highly efficient endogenous human gene correction using designed zinc-finger nucleases. Nature, 435, 646-51. 6. Boch,J., Scholze,H., Schornack,S., Landgraf,A., Hahn,S., Kay,S., Lahaye,T., Nickstadt,A. and Bonas,U. (2009) Breaking the code of DNA binding specificity of TAL-type III effectors. Science (New York, N.Y.), 326, 1509-12. 7. Cell,P., Replacement,K.S., Talens,A., Type,A., Collection,C, Ccl-,A. and Quickextract,E. Genetic engineering of human pluripotent cells using TALE nucleases. 8. Mussolino,C., Morbitzer,R., Lutge,F., Dannemann,N., Lahaye,T. and Cathomen,T. (2011) A novel TALE nuclease scaffold enables high genome editing activity in combination with low toxicity. Nucleic acids research, 39, 9283-93.
[0151] 9. Ding,Q., Lee,Y., Schaefer,E.A.K., Peters,D.T., Veres,A., Kim,K., Kuperwasser,N., Motola,D.L., Meissner,T.B., Hendriks,W.T., et al. (2013) Resource A TALEN Genome-Editing System for Generating Human Stem Cell-Based Disease Models. 10. Hockemeyer, D., Wang, H., Kiani, S., Lai, C. S., Gao, Q., Cassady, J. P., Cost, G. J., Zhang, L., Santiago, Y., Miller, J. C., et al. (2011) Genetic engineering of human pluripotent cells using TALE nucleases. Nature biotechnology, 29, 731-4. 11. Bedell, V. M., Wang, Y., Campbell, J. M., Poshusta, T. L., Starker, C. G., Krug Ii, R. G., Tan, W., Penheiter, S. G., Ma, A. C., Leung, A. Y. H., et al. (2012) In vivo genome editing using a high-efficiency TALEN system. Nature, 490, 114-118. 12. Miller, J. C., Tan, S., Qiao, G., Barlow, K. a, Wang, J., Xia, D. F., Meng, X., Paschon, D. E., Leung, E., Hinkley, S. J., et al. (2011) A TALE nuclease architecture for efficient genome editing. Nature biotechnology, 29, 143-8.
[0152] [Chemical formula]
[0153] 14. Reyon,D., Tsai,S.Q., Khayter,C., Foden,J. a, Sander,J.D. and Joung,J.K. (2012) FLASH assembly of TALENs for high-throughput genome editing. Nature Biotechnology, 30, 460-465. 15. Qiu,P., Shandilya,H., D'Alessio,J.M., 0'Connor,K., Durocher,J. and Gerard,G.F. (2004) Mutation detection using Surveyor nuclease. BioTechniques, 36, 702-7. 16. Mali,P., Yang,L., Esvelt,K.M., Aach,J., Guell,M., DiCarlo,J.E., Norville,J.E. and Church,G.M.(2013) RNA-guided human genome engineering via Cas9. Science (New York, N.Y.), 339, 823-6.
[0154] 17. Ding,Q., Regan,S.N., Xia,Y., Oostrom,L.A., Cowan,C.A. and Musunuru,K. (2013) Enhanced Efficiency of Human Pluripotent Stem Cell Genome Editing through Replacing TALENs with CRISPRs. Cell Stem Cell, 12, 393-394. 18. Cong,L., Ran,F.A., Cox,D., Lin,S., Barretto,R., Habib,N., Hsu,P.D., Wu,X., Jiang,W., Marraffini,L. a, et al. (2013) Multiplex genome engineering using CRISPR / Cas systems. Science (New York, N.Y.), 339, 819-23. 19. Cho,S.W., Kim,S., Kim,J.M. and Kim,J.-S. (2013) Targeted genome engineering in human cells with the Cas9 RNA-guided endonuclease. Nature biotechnology, 31, 230-232. 20. Hwang, W.Y., Fu,Y., Reyon,D., Maeder,M.L., Tsai,S.Q., Sander,J.D., Peterson,R.T., Yeh,J.- R.J. and Joung,J.K. (2013) Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology, 31, 227-229.
[0155] 21. Chen,F., Pruett-Miller,S.M., Huang,Y., Gjoka,M., Duda,K., Taunton,J., Collingwood,T.N., Frodin,M. and Davis,G.D. (2011) High-frequency genome editing using ssDNA oligonucleotides with zinc-finger nucleases. Nature methods, 8, 753-5.
[0156] [Chemical formula]
[0157] 23. Valamehr,B., Abujarour,R., Robinson,M., Le,T., Robbins,D., Shoemaker,D. and Flynn,P.(2012) A novel platform to enable the high-throughput derivation and characterization of feeder- free human iPSCs. Scientific reports, 2, 213. 24. Sanjana,N.E., Cong,L., Zhou,Y., Cunniff,M.M., Feng,G. and Zhang,F. (2012) A transcription activator- like effector toolbox for genome engineering. Nature protocols, 7, 171-92.
[0158] 25. Gibson,D.G., Young,L., Chuang,R., Venter, J.C., Iii,C.A.H., Smith,H.O. and America,N. (2009) Enzymatic assembly of DNA molecules up to several hundred kilobases. 6, 12-16. 26. Zou,J., Maeder,M.L., Mali,P., Pruett-Miller,S.M., Thibodeau-Beganny,S., Chou,B.-K., Chen,G., Ye,Z., Park,I.-H., Daley,G.Q., et al. (2009) Gene targeting of a disease-related gene in human induced pluripotent stem and embryonic stem cells. Cell stem cell, 5, 97-110. 27. Perez,E.E., Wang,J., Miller, J. C, Jouvenot,Y., Kim,K. a, Liu,O., Wang,N., Lee,G., Bartsevich,V. V, Lee,Y.-L., et al. (2008) Establishment of HIV- 1 resistance in CD4+ T cells by genome editing using zinc-finger nucleases. Nature biotechnology, 26, 808-16. 28. Bhakta,M.S., Henry,I.M., Ousterout,D.G., Das,K.T., Lockwood,S.H., Meckler,J.F., Wallen,M.C, Zykovich,A., Yu,Y., Leo,H., et al. (2013) Highly active zinc-finger nucleases by extended modular assembly. Genome research, 10.1101 / gr.143693.112.
[0159] 29. Kim,E., Kim,S., Kim,D.H., Choi,B.-S., Choi,I.-Y. and Kim,J.-S. (2012) Precision genome engineering with programmable DNA-nicking enzymes. Genome research, 22, 1327-33. 30. Gupta,A., Meng,X., Zhu,L.J., Lawson,N.D. and Wolfe,S. a (2011) Zinc finger protein-dependent and -independent contributions to the in vivo off-target activity of zinc finger nucleases. Nucleic acids research, 39, 381-92. 31. Park,I.-H., Lerou,P.H., Zhao,R., Huo,H. and Daley,G.Q. (2008) Generation of human-induced pluripotent stem cells. Nature protocols, 3, 1180-6. 32. Cermak,T., Doyle,E.L., Christian,M., Wang,L., Zhang,Y., Schmidt,C., Baller,J.A., Somia,N. V, Bogdanove,A.J. and Voytas,D.F. (2011) Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic acids research, 39, e82.
[0160] 33. Briggs,A.W., Rios,X., Chari,R., Yang,L., Zhang,F., Mali,P. and Church,G.M. (2012) Iterative capped assembly: rapid and scalable synthesis of repeat-module DNA such as TAL effectors from individual monomers. Nucleic acids research, 10.1093 / nar / gks624. 34. Zhang,F., Cong,L., Lodato,S., Kosuri,S., Church,G.M. and Arlotta,P. (2011) LETTErs Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. 29, 149-154. 35. Pathak,V.K. and Temin,H.M. (1990) Broad spectrum of in vivo forward mutations, hypermutations, and mutational hotspots in a retroviral shuttle vector after a single replication cycle: substitutions, frameshifts, and hypermutations. Proceedings of the National Academy of Sciences of the United States of America, 87,6019-23. 36. Tian,J., Ma,K. and Saaem,I. (2009) Advancing high-throughput gene synthesis technology. Molecular bioSystems, 5, 714-22.
[0161] 37. Zou,J., Mali,P., Huang,X., Dowey,S.N. and Cheng,L. (2011) Site-specific gene correction of a point mutation in human iPS cells derived from an adult patient with sickle cell disease. Blood, 118, 4599-608. 38. Mali,P., Yang,L., Esvelt,K.M., Aach,J., Guell,M., Dicarlo,J.E., Norville,J.E. and Church,G.M.(2013) RNA-Guided Human Genome. 39. Boyle,A.P., Davis,S., Shulha,H.P., Meltzer,P., Margulies,E.H., Weng,Z., Furey,T.S. and Crawford,G.E. (2008) High-resolution mapping and characterization of open chromatin across the genome. Cell, 132, 311-22. 40. Orlando,S.J., Santiago,Y., DeKelver,R.C., Freyvert,Y., Boydston,E. a, Moehle,E. a, Choi,V.M., Gopalan,S.M., Lou,J.F., Li,J., et al. (2010) Zinc-finger nuclease-driven targeted integration into mammalian genomes using donors with limited chromosomal homology. Nucleic acids research, 38, e152.
[0162] 41. Wang,Z., Zhou,Z.-J., Liu,D.-P. and Huang,J.-D. (2008) Double-stranded break can be repaired by single-stranded oligonucleotides via the ATM / ATR pathway in mammalian cells. Oligonucleotides, 18, 21-32. 42. Rios,X., Briggs,A.W., Christodoulou,D., Gorham,J.M., Seidman,J.G. and Church,G.M. (2012) Stable gene targeting in human cells using single-strand oligonucleotides with modified bases. PloS one, 7, e36697. 43. Elliott,B., Richardson,C, Winderbaum,J., Jac,A., Jasin,M. and Nickoloff,J.A.C.A. (1998) Gene Conversion Tracts from Double-Strand Break Repair in Mammalian Cells Gene Conversion Tracts from Double-Strand Break Repair in Mammalian Cells. 18. 44. Lombardo,A., Genovese,P., Beausejour,C.M., Colleoni,S., Lee,Y.-L., Kim,K. a,Ando,D., Urnov,F.D., Galli,C, Gregory,P.D., et al. (2007) Gene editing in human stem cells using zinc finger nucleases and integrase-defective lentiviral vector delivery. Nature biotechnology, 25, 1298-306.
[0163] 45. Jinek,M., Chylinski,K., Fonfara,I., Hauer,M., Doudna,J. a and Charpentier,E. (2012) A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science (New York, N.Y.), 337, 816-21. 46. Shrivastav,M., De Haro,L.P. and Nickoloff,J. . a (2008) Regulation of DNA double-strand break repair pathway choice. Cell research, 18, 134-47. 47. Kim,YG, Cha,J. and Chandrasegaran,S. (1996) Hybrid restriction enzymes: zinc finger fusions to Fok I cleavage domain. Proceedings of the National Academy of Sciences of the United States of America, 93, 1156-60. 48. Mimitou,EP and Symington,LS (2008) Sae2, Exo1 and Sgs1 collaborate in DNA double-strand break processing. Nature, 455, 770-4. 49. Doyon,Y., Choi,VM, Xia,DF, Vo,TD, Gregory,PD and Holmes,MC (2010) Transient cold shock enhances zinc-finger nuclease-mediated gene disruption. Nature methods, 7, 459-60. Embodiments of the present invention include the following: <1> A method for modifying target DNA in a cell, comprising introducing a TALEN lacking 100 bp or more of repetitive sequences into a cell, wherein the TALEN cleaves the target DNA, and the cell undergoes non-homologous end joining to produce modified DNA within the cell. <2> The aforementioned TALEN lacks a repeating sequence of 90 bp or more. <1> Methods used. <3> The aforementioned TALEN lacks a repeating sequence of 80 bp or more. <1> Methods used. <4> The aforementioned TALEN lacks a repeating sequence of 70 bp or more. <1> Methods used. <5> The aforementioned TALEN lacks a repeating sequence of 60 bp or more. <1> Methods used. <6> The method according to <1>, wherein the TALEN lacks a repeat sequence of 50 bp or more. <7> The method according to <1>, wherein the TALEN lacks a repeat sequence of 40 bp or more. <8> The method according to <1>, wherein the TALEN lacks a repeat sequence of 30 bp or more. <9> The method according to <1>, wherein the TALEN lacks a repeat sequence of 20 bp or more. <10> The method according to <1>, wherein the TALEN lacks a repeat sequence of 19 bp or more. <11> The method according to <1>, wherein the TALEN lacks a repeat sequence of 18 bp or more. <12> The method according to <1>, wherein the TALEN lacks a repeat sequence of 17 bp or more. <13> The method according to <1>, wherein the TALEN lacks a repeat sequence of 16 bp or more. <14> The method according to <1>, wherein the TALEN lacks a repeat sequence of 15 bp or more. <15> The method according to <1>, wherein the TALEN lacks a repeat sequence of 14 bp or more. <16> The method according to <1>, wherein the TALEN lacks a repeat sequence of 13 bp or more. <17> The method according to <1>, wherein the TALEN lacks a repeat sequence of 12 bp or more. <18> The method according to <1>, wherein the TALEN lacks a repeat sequence of 11 bp or more. <19> The method according to <1>, wherein the TALEN lacks a repeat sequence of 10 bp or more. <20> The method according to <1>, wherein the cell is a eukaryotic cell. <二十一> The method according to <1>, wherein the cell is a yeast cell, a plant cell or an animal cell. <22> The aforementioned cells are somatic cells. <1> Methods used. <23> The aforementioned cells are stem cells. <1> Methods used. <24> The aforementioned cells are human stem cells. <1> Methods used. <25> This includes introducing a first exogenous nucleic acid encoding the TALEN into the cells, The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <1> Methods used. <26> The process includes introducing a virus containing a first exogenous nucleic acid encoding the TALEN into the cells, The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <1> Methods used. <27> The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <1> Methods used. <28> The method according to <1>, wherein the TALEN is expressed and the TALEN cleaves the target DNA to produce modified DNA in the cell. <29> A method for modifying a target DNA in a cell, comprising combining a TALEN lacking a repetitive sequence of 100 bp or more in the cell with a donor nucleic acid sequence, wherein the TALEN cleaves the target DNA and the donor nucleic acid sequence is inserted into the DNA in the cell. <30> The method according to <29>, wherein the cell undergoes non-homologous end joining to produce modified DNA in the cell. <31> The method according to <29>, wherein the cell undergoes homologous recombination to produce modified DNA in the cell. <32> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 90 bp or more. <33> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 80 bp or more. <34> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 70 bp or more. <35> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 60 bp or more. <36> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 50 bp or more. <37> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 40 bp or more. <38> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 30 bp or more. <39> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 20 bp or more. <40> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 19 bp or more. <41> The method according to <29>, wherein the TALEN lacks a repetitive sequence of 18 bp or more. <42> The aforementioned TALEN lacks a repeating sequence of 17 bp or more. <29> Methods used. <43> The aforementioned TALEN lacks a repeating sequence of 16 bp or more. <29> Methods used. <44> The aforementioned TALEN lacks a repeating sequence of 15 bp or more. <29> Methods used. <45> The aforementioned TALEN lacks a repeating sequence of 14 bp or more. <29> Methods used. <46> The aforementioned TALEN lacks a repeating sequence of 13 bp or more. <29> Methods used. <47> The aforementioned TALEN lacks a repeating sequence of 12 bp or more. <29> Methods used. <48> The aforementioned TALEN lacks a repeating sequence of 11 bp or more. <29> Methods used. <49> The aforementioned TALEN lacks a repeating sequence of 10 bp or more. <29> Methods used. <50> The aforementioned cells are eukaryotic cells. <29> Methods used. <51> The aforementioned cells are yeast cells, plant cells, or animal cells. <29> Methods used. <52> The aforementioned cells are somatic cells. <29> Methods used. <53> The aforementioned cells are stem cells. <29> Methods used. <54> The aforementioned cells are human stem cells. <29> Methods used. <55> This includes introducing a first exogenous nucleic acid encoding the TALEN into the cells, The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <29> Methods used. <56> The process includes introducing a virus containing a first exogenous nucleic acid encoding the TALEN into the cells, The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <29> Methods used. <57> The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <29> Methods used. <58> The TALEN is expressed, and the TALEN cleaves the target DNA to produce modified DNA within the cell. <29> Methods used. <59> A virus containing a nucleic acid sequence that encodes a TALEN lacking more than 100 bp of repetitive sequences. <60> The aforementioned TALEN lacks a repeating sequence of 90 bp or more. <59> The virus described above. <61> The aforementioned TALEN lacks a repeating sequence of 80 bp or more. <59> The virus described above. <62> The aforementioned TALEN lacks a repeating sequence of 70 bp or more. <59> The virus described above. <63> The aforementioned TALEN lacks a repeating sequence of 60 bp or more. <59> The virus described above. <64> The aforementioned TALEN lacks a repeating sequence of 50 bp or more. <59> The virus described above. <65> The aforementioned TALEN lacks a repeating sequence of 40 bp or more. <59> The virus described above. <66> The aforementioned TALEN lacks a repeating sequence of 30 bp or more. <59> The virus described above. <67> The aforementioned TALEN lacks a repeating sequence of 20 bp or more. <59> The virus described above. <68> The aforementioned TALEN lacks a repeating sequence of 19 bp or more. <59> The virus described above. <69> The aforementioned TALEN lacks a repeating sequence of 18 bp or more. <59> The virus described above. <70> The aforementioned TALEN lacks a repeating sequence of 17 bp or more. <59> The virus described above. <71> The aforementioned TALEN lacks a repeating sequence of 16 bp or more. <59> The virus described above. <72> The aforementioned TALEN lacks a repeating sequence of 15 bp or more. <59> The virus described above. <73> The aforementioned TALEN lacks a repeating sequence of 14 bp or more. <59> The virus described above. <74> The aforementioned TALEN lacks a repeating sequence of 13 bp or more. <59> The virus described above. <75> The aforementioned TALEN lacks a repeating sequence of 12 bp or more. <59> The virus described above. <76> The aforementioned TALEN lacks a repeating sequence of 11 bp or more. <59> The virus described above. <77> The aforementioned TALEN lacks a repeating sequence of 10 bp or more. <59> The virus described above. <78> It is a lentivirus. <59> The virus described above. <79> Cells containing nucleic acid sequences encoding TALENs that lack repetitive sequences of 100 bp or more. <80> The aforementioned TALEN lacks a repeating sequence of 90 bp or more. <70> The cells described. <81> The aforementioned TALEN lacks a repeating sequence of 80 bp or more. <70> The cells described. <82> The aforementioned TALEN lacks a repeating sequence of 70 bp or more. <70> The cells described. <83> The aforementioned TALEN lacks a repeating sequence of 60 bp or more. <70> The cells described. <84> The aforementioned TALEN lacks a repeating sequence of 50 bp or more. <70> The cells described. <85> The aforementioned TALEN lacks a repeating sequence of 40 bp or more. <70> The cells described. <86> The aforementioned TALEN lacks a repeating sequence of 30 bp or more. <70> The cells described. <87> The aforementioned TALEN lacks a repeating sequence of 20 bp or more. <70> The cells described. <88> The aforementioned TALEN lacks a repeating sequence of 19 bp or more. <70> The cells described. <89> The aforementioned TALEN lacks a repeating sequence of 18 bp or more. <70> The cells described. <90> The aforementioned TALEN lacks a repeating sequence of 17 bp or more. <70> The cells described. <91> The aforementioned TALEN lacks a repeating sequence of 16 bp or more. <70> The cells described. <92> The aforementioned TALEN lacks a repeating sequence of 15 bp or more. <70> The cells described. <93> The aforementioned TALEN lacks a repeating sequence of 14 bp or more. <70> The cells described. <94> The aforementioned TALEN lacks a repeating sequence of 13 bp or more. <70> The cells described. <95> The aforementioned TALEN lacks a repeating sequence of 12 bp or more. <70> The cells described. <96> The aforementioned TALEN lacks a repeating sequence of 11 bp or more. <70> The cells described. <97> The aforementioned TALEN lacks a repeating sequence of 10 bp or more. <70> The cells described. <98> Eukaryotic cells <70> The cells described. <99> These are yeast cells, plant cells, or animal cells. <70> The cells described. <100> somatic cells <70> The cells described. <101> stem cells <70> The cells described. <102> Human stem cells <70> The cells described. <103> By combining TALE-N / TF backbone vectors containing endonucleases, DNA polymerases, DNA ligases, exonucleases, multiple nucleic acid dimer blocks encoding repeating variable two-residue domains, and endonuclease cleavage sites, The first and second ends are produced by activating the endonuclease and cleaving the TALE-N / TF backbone vector at the endonuclease cleavage site. The exonuclease is activated to create 3' overhangs and 5' overhangs in the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks, and the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks are annealed in a desired order. A method for producing a TALE, comprising activating the DNA polymerase and the DNA ligase to link the TALE-N / TF backbone vector and the plurality of nucleic acid dimer blocks. <104> A method for modifying target DNA in stem cells expressing an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA, wherein the method is: (a) Introducing into the stem cell a first exogenous nucleic acid that encodes RNA complementary to the target DNA and which guides the enzyme to the target DNA; where the RNA and the enzyme are members of a colocalization complex to the target DNA. Introducing a second exogenous nucleic acid encoding the donor nucleic acid sequence into the stem cells, Includes, The method comprising the expression of the RNA and the donor nucleic acid sequence, the colocalization of the RNA and the enzyme to the target DNA, the cleavage of the target DNA by the enzyme, and the insertion of the donor nucleic acid into the target DNA to produce modified DNA within the stem cell. <105> The enzyme is an RNA-induced DNA-binding protein. <104> Methods used. <106> The enzyme is an RNA-inducible DNA-binding protein of the type II CRISPR system. <104> Methods used. <107> The enzyme is Cas9. <104> Methods used. <108> The RNA is between approximately 10 nucleotides and approximately 500 nucleotides. <104> Methods used. <109> The RNA is between approximately 20 and 100 nucleotides. <104> Methods used. <110> The aforementioned RNA is a guide RNA. <104> Methods used. <111> The RNA is a tracrRNA-crRNA fusion. <104> Methods used. <112> The DNA in question is genomic DNA, mitochondrial DNA, viral DNA, or exogenous DNA. <104> Methods used. <113> The donor nucleic acid sequence is inserted by recombination. <104> Methods used. <114> The donor nucleic acid sequence is inserted by homologous recombination. <104> Methods used. <115> The donor nucleic acid sequence is inserted by non-homologous end joining. <104> Methods used. <116> The RNA and the donor nucleic acid sequence are present on one or more plasmids. <104> Methods used. <117> The process further includes repeating step (a) multiple times to cause multiple modifications to the DNA within the cell. <104> Methods used. <118> After creating modified DNA within the stem cell, the nucleic acid encoding the enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA is removed from the stem cell genome. <104> Methods used. <119> The RNA and the donor nucleic acid sequence are expressed as a nucleic acid sequence to which the RNA and the donor nucleic acid sequence are bound. <104> Methods used. <120> A stem cell comprising a first exogenous nucleic acid that forms a co-localization complex with RNA complementary to the target DNA and encodes an enzyme that site-specifically cleaves the target DNA. <121> The present invention further comprises a second exogenous nucleic acid that encodes RNA complementary to the target DNA and guides the enzyme to the target DNA, wherein the RNA and the enzyme are members of a colocalization complex with respect to the target DNA. <120> Stem cells as described above. <122> Further comprising a third exogenous nucleic acid encoding a donor nucleic acid sequence, <121> Stem cells as described above. <123> The inductive promoter for promoting the expression of the enzyme further comprises <120> Stem cells as described above. <124> The first exogenous nucleic acid can be removed from the genomic DNA of the cell using a transpose. <120> Stem cells as described above. <125> The enzyme is an RNA-induced DNA-binding protein. <120> Stem cells as described above. <126> The enzyme is an RNA-inducible DNA-binding protein of the type II CRISPR system. <120> Stem cells as described above. <127> The enzyme is Cas9. <120> Stem cells as described above. <128> The RNA is between approximately 10 nucleotides and approximately 500 nucleotides. <120> Stem cells as described above. <129> The RNA is between approximately 20 and 100 nucleotides. <120> Stem cells as described above. <130> The aforementioned RNA is a guide RNA. <120> Stem cells as described above. <131> The RNA is a tracrRNA-crRNA fusion. <120> Stem cells as described above. <132> The target DNA is genomic DNA, mitochondrial DNA, viral DNA, or exogenous DNA. <120> Stem cells as described above. <133> A cell comprising a first exogenous nucleic acid that comprises an enzyme that forms a co-localization complex with RNA complementary to target DNA and site-specifically cleaves the target DNA, and an inducible promoter for promoting the expression of the enzyme. <134> The present invention further comprises a second exogenous nucleic acid that encodes RNA complementary to the target DNA and guides the enzyme to the target DNA, wherein the RNA and the enzyme are members of a colocalization complex with respect to the target DNA. <133> The cells described. <135> Further comprising a third exogenous nucleic acid encoding a donor nucleic acid sequence, <134> Stem cells as described above. <136> A cell comprising a first exogenous nucleic acid that forms a colocalization complex with RNA complementary to target DNA and encodes an enzyme that site-specifically cleaves the target DNA, wherein the first exogenous nucleic acid can be removed from the cell's genomic DNA using a transposase. <137> The present invention further comprises a second exogenous nucleic acid that encodes RNA complementary to the target DNA and guides the enzyme to the target DNA, wherein the RNA and the enzyme are members of a colocalization complex with respect to the target DNA. <136> The cells described. <138> Further comprising a third exogenous nucleic acid encoding a donor nucleic acid sequence, <137> Stem cells as described above. <139> A cell comprising a first exogenous nucleic acid that forms a colocalization complex with RNA complementary to target DNA and encodes an enzyme that site-specifically cleaves the target DNA, wherein the first exogenous nucleic acid is reversibly inserted into the cell's genomic DNA. <140> The present invention further comprises a second exogenous nucleic acid that encodes RNA complementary to the target DNA and guides the enzyme to the target DNA, wherein the RNA and the enzyme are members of a colocalization complex with respect to the target DNA. <139> The cells described. <141> Further comprising a third exogenous nucleic acid encoding a donor nucleic acid sequence, <140> Stem cells as described above. <142> A method for modifying target DNA in a cell expressing an enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA, wherein the method is: (a) Introducing a first exogenous nucleic acid encoding a donor nucleic acid sequence into the cells, Introducing RNA complementary to the target DNA and guiding the enzyme to the target DNA from the culture medium surrounding the cell into the cell; where the RNA and the enzyme are members of a colocalization complex for the target DNA. Includes, The donor nucleic acid sequence is expressed, The method comprising the RNA and the enzyme co-localizing to the target DNA, the enzyme cleaving the target DNA, and the donor nucleic acid being inserted into the target DNA to produce modified DNA within the cell. <143> The RNA includes a 5' cap structure. <142> Methods used. <144> The RNA lacks a phosphate group. <142> Methods used. <145> The enzyme is an RNA-induced DNA-binding protein. <142> Methods used. <146> The enzyme is an RNA-inducible DNA-binding protein of the type II CRISPR system. <142> Methods used. <147> The enzyme is Cas9. <142> Methods used. <148> The RNA is between approximately 10 nucleotides and approximately 500 nucleotides. <142> Methods used. <149> The RNA is between approximately 20 and 100 nucleotides. <142> Methods used. <150> The aforementioned RNA is a guide RNA. <142> Methods used. <151> The RNA is a tracrRNA-crRNA fusion. <142> Methods used. <152> The DNA in question is genomic DNA, mitochondrial DNA, viral DNA, or exogenous DNA. <142> Methods used. <153> The donor nucleic acid sequence is inserted by recombination. <142> Methods used. <154> The donor nucleic acid sequence is inserted by homologous recombination. <142> Methods used. <155> The donor nucleic acid sequence is inserted by non-homologous end joining. <142> Methods used. <156> The process further includes repeating step (a) multiple times to cause multiple modifications to the DNA within the cell. <142> Methods used. <157> After creating modified DNA within the cell, the nucleic acid encoding the enzyme that forms a co-localization complex with RNA complementary to the target DNA and site-specifically cleaves the target DNA is removed from the cell's genome. <142> Methods used. <158> The RNA and the donor nucleic acid sequence are expressed as a nucleic acid sequence to which the RNA and the donor nucleic acid sequence are bound. <142> Methods used.
Claims
1. (a) Introducing a single-stranded oligodeoxyribonucleotide (ssODN) of length 50 to 110 nt into a eukaryotic cell, wherein the ssODN comprises a donor nucleic acid sequence, the donor nucleic acid sequence comprises a mismatch and an additional mismatch to the target DNA, and the additional mismatch in the donor nucleic acid sequence is located 30 nt or less away from the mismatch, and (b) To provide the cells of the eukaryotes with an RNA-inducible DNA-binding protein of a type II CRISPR system that site-specifically cleaves target DNA at a target site, Includes, The donor nucleic acid sequence is inserted into the target DNA to produce modified DNA within the eukaryotic cell. A method for modifying target DNA in vitro within eukaryotic cells, When the RNA-inducible DNA-binding protein of the type II CRISPR system is provided, the method further comprises providing the eukaryotic cell with a guide RNA containing a sequence complementary to the target sequence of the target DNA, The guide RNA and the RNA-inducible DNA-binding protein of the type II CRISPR system form a colocalization complex that binds to the target sequence of the target DNA. The aforementioned method.
2. The method according to claim 1, wherein the RNA-inducible DNA-binding protein of the type II CRISPR system comprises Cas9 nuclease.
3. The method according to claim 1, wherein the length of the ssODN is 50 to 90 nt.
4. The method according to claim 1, wherein, after insertion of the donor nucleic acid sequence into the target DNA, the mismatch in the modified DNA is located 40 nt or less away from the target site.
5. The method according to claim 1, wherein the donor nucleic acid sequence is inserted by homologous recombination.
6. The method according to claim 1, wherein the donor nucleic acid sequence is inserted by non-homologous end joining.
7. The method according to claim 1, wherein the eukaryotic cell is an animal cell.
8. The method according to claim 1, wherein the RNA-inducible DNA-binding protein of the type II CRISPR system and the ssODN are simultaneously introduced into the eukaryotic cell.
9. The method according to claim 1, further comprising repeating steps (a) and (b) to bring about multiple foreign nucleic acid insertions into the cells.