DNA knock-in system
By combining sequence-specific nucleases and exonucleases in DNA knock-in technology, and utilizing homologous recombination and specific cutting of exonucleases, the problem of low homologous recombination efficiency in existing technologies is solved, and efficient and low-cost exogenous gene insertion is achieved.
Patent Information
- Application Number
- CN201580041157.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2014-08-14
- Filing Date
- 2015-08-13
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2035-12-23
AI Technical Summary
In existing DNA knock-in technologies, the efficiency of homologous recombination is low, making it difficult to efficiently insert exogenous genes into specific loci. In particular, its application in various species is limited and costly.
The method of sequence-specific nuclease combined with exonuclease and donor construct is used to insert the donor sequence on the chromosome through homologous recombination. The homology of the 5' and 3' homology arms to the target site is utilized, and the insertion efficiency is improved by combining 5' to 3' exonuclease.
It significantly improves the insertion efficiency of exogenous sequences at target sites, is applicable to a variety of cells and animals, reduces operating costs, and expands the scope of application.
Smart Images

Figure CN107429241B_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 037,551, filed August 14, 2014, the contents of which are incorporated herein by reference in their entirety.
[0003] Submission of Sequence Listing in ASCII Text File
[0004] The contents of the following submission on ASCII text file are incorporated herein by reference in their entirety: a computer readable form (CRF) of the Sequence Listing (file name: 735782000140SeqList.txt, date recorded: August 11, 2015, size: 24 KB). TECHNICAL FIELD
[0005] The present invention relates to an enhanced DNA knock-in (EKI) system and uses thereof. BACKGROUND
[0006] Recent developments in DNA editing allow for the efficient introduction of exogenous genes into the chromosome of a host cell (gene knock-in). This allows for the insertion of a cDNA sequence encoding a protein into a specific locus in the genome of an organism, e.g., insertion of a mutation or an exogenous gene at a specific locus on a chromosome. For example, a point mutation can be introduced into a target gene by knock-in to mimic a human genetic disease. Furthermore, an exogenous gene such as a reporter gene (EGFP, mRFP, mCherry, tdTomato, etc.) can be introduced into a specific locus of a target gene by homologous recombination, and the exogenous gene can be used to track the expression of the target gene and study its expression profile by expression of the reporter gene.
[0007] In many cases, gene knock-in involves the homologous recombination machinery within an organism. Under natural conditions, the probability of homologous recombination between an exogenous targeting vector and the cell genome is very low, about 1 / 10 5 to 1 / 10 6 Spontaneous gene targeting usually occurs in mammalian cells at a very low frequency, with an efficiency of one in a million cells. The presence of a double-strand break often generates recombination and can increase the frequency of homologous recombination by several thousand-fold. See Jasin, 1996, "Genetic manipulation of genomes with rare-cutting endonucleases," Trends in genetics : TIG 12(6): 224-228. In plants, it is known that the creation of a double-strand break in DNA can increase the frequency of homologous recombination from about 10-3 up to 10 -4 approximately 100-fold. See Hanin et al., 2001, "Gene targeting in Arabidopsis," Plant J. 28:671-77. The establishment of mouse embryonic stem cell lines has made it possible to generate genetically modified mice through homologous recombination. For example, a targeting vector can be constructed using a bacterial artificial chromosome (BAC) and introduced into mouse embryonic stem cells by transfection (e.g., electroporation). Positive embryonic stem cell clones are selected and injected into mouse blastocysts microcell mass, which are then implanted into a surrogate mouse to generate genetically engineered chimeric mice. However, methods for establishing embryonic stem cell lines from other species have not been successful and widely used.
[0008] Recently developed technologies, including ZFN (zinc finger nuclease), TALEN (transcription activator-like effector nuclease), CRISPR / Cas9, and other site-specific nuclease technologies, make it possible to create double-stranded DNA breaks at desired genetic locus sites. These controlled double-stranded breaks can facilitate homologous recombination at these particular genetic locus sites. This process relies on targeting a specific sequence of a nucleic acid molecule (e.g., a chromosome) with an endonuclease that recognizes and binds to such a sequence and induces a double-stranded break in the nucleic acid molecule. The double-stranded break is repaired by error-prone nonhomologous end-joining (NHEJ) or by homologous recombination (HR).
[0009] Homologous recombination occurring during DNA repair tends to result in non-crossover products, effectively repairing the damaged DNA molecule as it existed prior to the double-strand break, or creating a recombinant molecule by incorporating the sequence of the template. The latter has been used for gene targeting, protein engineering, and gene therapy. If a template for homologous recombination is provided in trans (e.g., by introducing an exogenous template into a cell), the provided template can be used to repair a double-strand break in the cell. In gene targeting, the initial double-strand break increases the efficiency of targeting by orders of magnitude compared to conventional homologous recombination-based gene targeting. In principle, this method can be used to insert any sequence at the repair site, as long as it is flanked by appropriate regions homologous to sequences near the double-strand break. While this method has been successful in a number of species (e.g., mouse and rat), the efficiency and success rate of homologous recombination remains low, which hinders the widespread use of this method. For example, the method remains expensive and technically challenging for scientific and commercial uses.
[0010] The disclosures of all publications, patents, patent applications and published patent applications referred to in this document are hereby incorporated by reference in their entirety. SUMMARY
[0012] In one embodiment, disclosed herein is a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, comprising (a) introducing into the cell a sequence-specific nuclease that cleaves the chromosome at the insertion site; (b) introducing into the cell a donor construct; and (c) introducing into the cell an exonuclease. In one aspect, the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid. In another aspect, the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm. In certain aspects, the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and the 3' homology arm is homologous to a sequence downstream of the cleavage site on the chromosome. In some embodiments, the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid. In other embodiments, the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0013] In some embodiments, the sequence-specific nuclease used herein is a zinc finger nuclease (ZFN). In some embodiments, the sequence-specific nuclease used herein is a transcription activator-like effector nuclease (TALEN). In some embodiments, the sequence-specific nuclease used herein is an RNA-guided nuclease. In one aspect, the RNA-guided nuclease is a Cas, such as Cas9.
[0014] In any of the foregoing embodiments involving an RNA-guided nuclease, the method can further comprise introducing into the cell a guide RNA (gRNA) that recognizes the insertion site.
[0015] In any of the foregoing embodiments, the sequence-specific nuclease can be introduced into the cell as a protein, mRNA, or cDNA.
[0016] In some embodiments, the sequence homology between the 5' homology arm and the sequence 5' to the insertion site is at least about 80%. In some embodiments, the sequence homology between the 3' homology arm and the sequence 3' to the insertion site is at least about 80%. In any of the foregoing embodiments, the 5' and 3' homology arms can be at least about 200 base pairs (bp).
[0017] In any of the foregoing embodiments, the exonuclease is a 5' to 3' exonuclease. In one aspect, the exonuclease is a herpes simplex virus type 1 (HSV-1) exonuclease. In one embodiment, the exonuclease is UL12.
[0018] In any of the foregoing embodiments, the donor construct can be a linear nucleic acid. In some embodiments, the donor construct is circular when introduced into the cell and is cleaved within the cell to produce a linear nucleic acid. In one aspect, the donor construct further comprises a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm. In one embodiment, the 5' or 3' flanking sequence is from about 1 to about 500 bp.
[0019] In some embodiments, the methods disclosed herein further comprise introducing into the cell a second sequence-specific nuclease that cleaves the donor construct at one or both flanking sequences, thereby producing a linear nucleic acid. In certain embodiments, the sequence-specific nuclease is an RNA-guided nuclease and the method further comprises introducing into the cell a second guide RNA that recognizes one or both of the flanking sequences.
[0020] In any of the foregoing embodiments, the eukaryotic cell can be a mammalian cell. In some embodiments, the mammalian cell is a zygote or a pluripotent stem cell.
[0021] In some aspects, disclosed herein are methods of generating a genetically modified animal comprising an insertion of a donor sequence at a predetermined insertion site on a chromosome of the animal. In one embodiment, the method comprises (a) introducing a sequence-specific nuclease that cleaves the chromosome at the insertion site into the cell; (b) introducing a donor construct into the cell; (c) introducing an exonuclease into the cell; and (d) introducing the cell into a carrier animal to generate the genetically modified animal. In one embodiment, the donor construct is a linear nucleic acid or can be cleaved within the cell to generate a linear nucleic acid. In one aspect, the linear nucleic acid comprises a 5' homology arm, a donor sequence, and a 3' homology arm. In one aspect, the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome. In one aspect, the 3' homology arm is homologous to a sequence downstream of the cleavage site on the chromosome. In some embodiments, the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid. In some embodiments, the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0022] In one aspect, the genetically modified animal is a rodent. In other aspects, the cell is a zygote or a pluripotent stem cell.
[0023] In some aspects, provided herein is a genetically modified animal produced by the method of any one of the above methods.
[0024] In other aspects, provided herein is a kit for inserting a donor sequence at an insertion site on a chromosome in a eukaryotic cell, comprising: (a) a sequence-specific nuclease that cleaves the chromosome at the insertion site; (b) a donor construct; and (c) an exonuclease. In one embodiment, the donor construct is a linear nucleic acid or can be cleaved within the cell to generate a linear nucleic acid. In one aspect, the linear nucleic acid comprises a 5' homology arm, a donor sequence, and a 3' homology arm. In one aspect, the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome. In one aspect, the 3' homology arm is homologous to a sequence downstream of the cleavage site on the chromosome. In one aspect, the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid. In some embodiments, the sequence-specific nuclease is an RNA-guided nuclease. In other embodiments, the kit further comprises a guide RNA (gRNA) that recognizes the insertion site.
[0025] In some embodiments, the donor construct is circular. In one aspect, the donor construct further comprises a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm. In another aspect, the 5' flanking sequence or the 3' flanking sequence is from about 1 to about 500 bp. In any of the foregoing embodiments, the sequence-specific nuclease can be an RNA-guided nuclease, and the kit can further comprise a second guide RNA that recognizes one or both of the flanking sequences. In any of the foregoing embodiments, the exonuclease can be a 5' to 3' exonuclease. In one aspect, the exonuclease is a herpes simplex virus type 1 (HSV-1) exonuclease. In another aspect, the exonuclease is UL12. In one embodiment, the UL12 is fused to a nuclear localization sequence (NLS).
[0026] In some embodiments, the disclosure provides a method of producing a linear nucleic acid in a eukaryotic cell, wherein the linear nucleic acid comprises a 5' homology arm, a donor sequence, and a 3' homology arm; the 5' homology arm is homologous to a sequence upstream of a nuclease cleavage site on a chromosome, and the 3' homology arm is homologous to a sequence downstream of the cleavage site on the chromosome; and the 5' homology arm and the 3' homology arm are proximal to the 5' and 3' ends, respectively, of the linear nucleic acid. In one aspect, the method comprises: (a) introducing into the cell a circular donor construct comprising the linear nucleic acid and further comprising a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm; and (b) introducing into the cell a sequence-specific nuclease, wherein the sequence-specific nuclease cleaves the circular donor construct at the 5' flanking sequence and the 3' flanking sequence, thereby producing the linear nucleic acid. In one aspect, the sequence-specific nuclease is a ZFN. In another aspect, the sequence-specific nuclease is a TALEN. In another aspect, the sequence-specific nuclease is an RNA-guided nuclease. In one embodiment, the RNA-guided nuclease is a Cas. In some embodiments, the RNA-guided nuclease is a Cas9. In any of the foregoing embodiments, the method can further comprise introducing into the cell a guide RNA that recognizes the 5' flanking sequence and / or the 3' flanking sequence. BRIEF DESCRIPTION OF DRAWINGS
[0028] Figure 1 A targeting scheme to knock in EGFP-ACTB in the U2OS cell line is shown.
[0029] Figure 2 are fluorescent microscopy images showing expression of EGFP-ACTB fusion protein in U2OS cells after knock-in using the high efficiency knock-in ("EKI") system.
[0030] Figure 3Figures 3A, 3B and 3C depict flow cytometry results showing that knockin efficiency of EGFP-ACTB in U20S cells is significantly improved using the EKI system compared to conventional CRISPR / Cas9-mediated knockin. Figure 3 A is the result of a control experiment, Figure 3 B is the result of conventional CRISPR / Cas9-mediated knockin, Figure 3 C is the result of EKI system-mediated knockin.
[0031] Figure 4 Figure 4A, 4B and 4C depict flow cytometry results showing that knockin efficiency of EGFP-LMNB1 in C6 cells is significantly improved using the EKI system compared to conventional CRISPR / Cas9-mediated knockin.
[0032] Figure 5 Figure 5 is a fluorescent microscopy image showing expression of EGFP-LMNB1 fusion protein after knockin in C6 cells using the EKI system.
[0033] Figure 6 Figures 6A, 6B and 6C are flow cytometry results showing that knockin efficiency of EGFP-LMNB1 in C6 cells is significantly improved using the EKI system compared to conventional CRISPR / Cas9-mediated knockin. Figure 6 A is the result of a control experiment, Figure 6 B is the result of conventional CRISPR / Cas9-mediated knockin, Figure 6 C is the result of EKI system-mediated knockin.
[0034] Figure 7 Figures 7A and 7B show targeting schemes for double knockin of EGFP-ACTB and mCherry-LMNB1 in U20S cell line. Figure 7 A shows the targeting scheme for knockin of EGFP-ACTB, Figure 7 B shows the targeting scheme for knockin of mCherry-LMNB1.
[0035] Figure 8 Figure 8 is a fluorescent microscopy image showing expression of EGFP-ACTB and mCherry-LMNB1 fusion proteins after successful double knockin in the same U20S cells mediated by the EKI system.
[0036] Figure 9 Figure 9 shows a targeting scheme for making CD4-2A-dsRed knockin rats.
[0037] Figure 10 Figure 10 is a 5’ end PCR reaction genotype detection result of F0 generation CD4-2A-dsRed knockin rats.
[0038] Figure 11is the result of 3' end PCR reaction genotype detection of F0 generation CD4-2A-dsRed knock-in rats.
[0039] Figure 12 A and 12B are the results of southern blot hybridization of CD4-2A-dsRed knock-in rats. Figure 12 A is the southern blot hybridization strategy, Figure 12 B is the result of southern blot hybridization using 5' probe, 3' probe and dsRed probe. Two F1 generation rats (No. 19 and No. 21) used for southern blot hybridization test are Figure 10 and Figure 11 Offspring of No. 22 F0 rat in
[0040] Figure 13 A targeting strategy for knocking in TH-GFP in H9 cell line is shown.
[0041] Figure 14 is a fluorescent microscopy image showing the expression of TH-GFP fusion protein in H9 cells after knock-in using EKI system.
[0042] Figure 15 The position of genotype detection primers for TH-GFP knock-in in H9 cells is shown. LA: left homology arm. RA: right homology arm.
[0043] Figure 16 A and 16B are the results of genotype detection of TH-GFP knock-in in H9 cells. Figure 16 A is the result of 5' end PCR reaction, Figure 16 B is the result of 3' end PCR reaction. WT: wild type H9 cells.
[0044] Figure 17 is Figure 16 the sequencing result of PCR product of No. 1 H9-TH-GFP cell line shown in
[0045] Figure 18 is Figure 16 the karyotype of No. 1 H9-TH-GFP cell line shown in
[0046] Figure 19 A targeting strategy for knocking in OCT4-EGFP in H9 cell line is shown.
[0047] Figure 20 is a fluorescent microscopic image showing expression of OCT4-EGFP fusion protein in H9 cells after knock-in using the EKI system.
[0048] Figure 21 shows the position of genotyping primers for OCT4-EGFP knock-in in H9 cells. LA: left homology arm. RA: right homology arm.
[0049] Figure 22 A, 22B and 22C are the results of genotyping of OCT4-EGFP knock-in in H9 cells. Figure 22 A is the result of 5' end PCR reaction, Figure 22 B is the result of 3' end PCR reaction, Figure 22 C is the result of full-length PCR reaction. wt: wild type H9 cells. c-: H2O as a negative control.
[0050] Figure 23 is Figure 22 is the result of sequencing of PCR product of No. 6 H9-OCT4-EGFP cell line shown in
[0051] Figure 24 is DETAILED DESCRIPTION is the karyotype of No. 6 H9-OCT4-EGFP cell line shown in Figure 1
[0053] The present application provides novel DNA knock-in methods that allow for the introduction of one or more exogenous sequences into a specific target site on a cell chromosome with significantly higher efficiency compared to traditional DNA knock-in methods using sequence-specific nucleases, such as CRISPR / Cas9 or TALEN-based gene knock-in systems. In addition to the use of sequence-specific nucleases, the methods of the present application also utilize an exonuclease (e.g., a 5' to 3' exonuclease, such as UL12) in conjunction with a donor construct that is a linear nucleic acid or that can be cleaved within the cell to produce a linear nucleic acid. This DNA knock-in system allows for the insertion of donor sequences into any desired target site with high efficiency, making it useful for many purposes, such as the generation of transgenic animals expressing exogenous genes, the modification (e.g., mutation) of genomic loci, and gene editing, by, for example, adding exogenous non-coding sequences (e.g., sequence tags or regulatory elements) to the genome. Cells and animals generated using the methods provided herein can have a variety of applications, such as cell therapy, disease models, research tools, and as humanized animals for a variety of purposes.
[0054] Accordingly, in one aspect, the present application provides a method of inserting a donor sequence into a predetermined insertion site on a chromosome of a eukaryotic cell.
[0055] In another aspect, a method of producing a linear nucleic acid within a cell is provided.
[0056] In another aspect, a kit for use in any of the methods described herein is provided.
[0057] In another aspect, a method of producing a genetically modified animal by using the gene knock-in system described herein is provided.
[0058] References in this document to values or parameters "of about" include (and describe) variations in the value or parameter itself. For example, descriptions of "about X" include descriptions of "X".
[0059] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0060] The compositions and methods of the present application can comprise, consist of, or consist essentially of the essential elements and limitations of the application described herein, in addition to any additional or optional ingredients, components, or limitations described herein, or otherwise useful in nutritional or pharmaceutical applications.
[0061] Method of the invention
[0062] The present invention provides methods of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell. In some embodiments, the method comprises: a) introducing a sequence-specific nuclease that cleaves the chromosome at the insertion site into the cell; b) introducing a donor construct into the cell; and c) introducing an exonuclease into the cell; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0063] In some embodiments, the sequence-specific nuclease, the exonuclease, and the donor construct are introduced into the cell simultaneously. In some embodiments, at least one of the three components is introduced into the cell at a different time than the other components. For example, the donor construct can be introduced into the cell first, followed by the sequence-specific nuclease and the exonuclease. In some embodiments, the sequence-specific nuclease is introduced into the cell first, followed by the donor construct and the exonuclease. In some embodiments, all three components are introduced at different points in time relative to each other. For example, the three components can be administered sequentially one after the other in a particular order.
[0064] In some embodiments, the sequence-specific nuclease and / or the exonuclease is introduced into the cell as a cDNA. In some embodiments, the sequence-specific nuclease and / or the exonuclease is introduced into the cell as an mRNA. In some embodiments, the sequence-specific nuclease and / or the exonuclease is introduced into the cell as a protein.
[0065] For example, in some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a nucleic acid sequence encoding a sequence-specific nuclease that cleaves the chromosome at the insertion site; b) introducing into the cell a donor construct; and c) introducing into the cell a nucleic acid sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is mRNA. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is cDNA. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is introduced into the cell by transfection, including, for example, by transfection by electroporation. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is introduced into the cell by injection.
[0066] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease that cleaves the chromosome at the insertion site; b) introducing into the cell a donor construct; and c) introducing into the cell a vector comprising a nucleic acid sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the vector comprising the nucleic acid encoding the sequence-specific nuclease and / or the vector comprising the nucleic acid encoding the exonuclease is introduced into the cell by transfection, including, for example, by transfection by electroporation.
[0067] In some embodiments, a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell is provided, the method comprising: a) introducing a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease that cleaves a chromosome at an insertion site and a nucleic acid sequence encoding an exonuclease into a cell; b) introducing a donor construct into the cell; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the vector comprising a nucleic acid encoding a sequence-specific nuclease and a nucleic acid encoding an exonuclease is introduced into a cell by transfection, including, e.g., by transfection by electroporation.
[0068] In some embodiments, a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell (e.g., a zygote) is provided, the method comprising: a) introducing (e.g., injecting) into a cell an mRNA sequence encoding a sequence-specific nuclease that cleaves a chromosome at an insertion site; b) introducing (e.g., injecting) into a cell a donor construct; and c) introducing (e.g., injecting) into a cell an mRNA sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the introducing (e.g., injecting) is performed in vitro. In some embodiments, the introducing (e.g., injecting) is performed in vivo. In some embodiments, the method further comprises transcribing a nucleic acid encoding a sequence-specific nuclease into an mRNA in vitro. In some embodiments, the method further comprises transcribing a nucleic acid encoding an exonuclease into an mRNA in vitro.
[0069] For example, in some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in an immune cell (e.g., a T cell), the method comprising: a) introducing into the cell a nucleic acid sequence encoding a sequence-specific nuclease that cleaves the chromosome at the insertion site; b) introducing into the cell a donor construct; and c) introducing into the cell a nucleic acid sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5’ homology arm, the donor sequence, and a 3’ homology arm, wherein the 5’ homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3’ homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5’ homology arm and the 3’ homology arm are adjacent to the 5’ and 3’ ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is mRNA. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is cDNA. In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is introduced into the cell by transfection (including, e.g., by electroporation). In some embodiments, the nucleic acid encoding the sequence-specific nuclease and / or the nucleic acid encoding the exonuclease is introduced into the cell by injection.
[0070] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in an immune cell (e.g., a T cell), the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease that cleaves the chromosome at the insertion site; b) introducing into the cell a donor construct; and c) introducing into the cell a vector comprising a nucleic acid sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5’ homology arm, the donor sequence, and a 3’ homology arm, wherein the 5’ homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3’ homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5’ homology arm and the 3’ homology arm are adjacent to the 5’ and 3’ ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the vector comprising a nucleic acid encoding the sequence-specific nuclease and / or the vector comprising a nucleic acid encoding the exonuclease is introduced into the cell by transfection (including, e.g., by electroporation).
[0071] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in an immune cell (e.g., a T cell), the method comprising: a) introducing a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease that cleaves a chromosome at an insertion site and a nucleic acid sequence encoding an exonuclease into the cell; b) introducing a donor construct into the cell; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the vector comprising a nucleic acid encoding a sequence-specific nuclease and a nucleic acid encoding an exonuclease is introduced into the cell by transfection (including, e.g., by electroporation).
[0072] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in an immune cell (e.g., a T cell), the method comprising: a) introducing a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease that cleaves a chromosome at an insertion site and a nucleic acid sequence encoding an exonuclease into the cell; b) introducing a donor construct into the cell; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the vector comprising a nucleic acid encoding a sequence-specific nuclease and a nucleic acid encoding an exonuclease is introduced into the cell by transfection (including, e.g., by electroporation).
[0073] The cells described herein can be any eukaryotic cell, for example, an isolated cell of an animal, for example, a totipotent, pluripotent, or somatic stem cell, a zygote, or a somatic cell. In some embodiments, the cells are from a primary cell culture. In some embodiments, the cells used in the methods are human cells. In some embodiments, the cells used in the methods are yeast cells. In some embodiments, the cells are from a domesticated animal (e.g., a cow, a sheep, a cat, a dog, and a horse). In some embodiments, the cells are from a primate (e.g., a non-human primate such as a monkey). In some embodiments, the cells are from a rabbit. In some embodiments, the cells are from a fish (e.g., a zebrafish). In some embodiments, the cells are from a rodent (e.g., a mouse, a rat, a hamster, a guinea pig). In some embodiments, the cells are from an invertebrate animal (e.g., Drosophila melanogaster and Caenorhabditis elegans).
[0074] The cells described herein can be immune cells, including, but not limited to, granulocytes (e.g., basophils, eosinophils, and neutrophils), mast cells, monocytes, dendritic cells (DCs), natural killer (NK) cells, B cells, and T cells (e.g., CD8+ T cells and CD4+ T cells). In some embodiments, the cells are CD4+ T cells (e.g., T helper cells TH1, TH2, TH17, and regulatory T cells).
[0075] The methods and compositions described herein can be used to insert donor sequences into one or more genomic loci in a cell. In certain embodiments, the methods and compositions described herein can be used to target more than one genomic locus within a cell, for example, for dual DNA knockin. In certain embodiments, the methods and compositions described herein are used to target two, three, four, five, six, seven, eight, nine, ten, or more than ten genomic loci within a cell. In some aspects, the dual or multiple knockin can be performed simultaneously or sequentially. For example, reagents for targeting two or more genomic loci within the same cell are mixed and introduced into the cell at substantially the same time. In other embodiments, reagents targeting each of the multiple genomic loci can be introduced into the cell sequentially, one after the other, in a particular order. In other embodiments, a first genomic locus is targeted, and the cells that successfully knock in are selected, enriched, and / or isolated, and targeted to a second genomic locus.
[0076] Double or multiple knock-ins can be achieved, for example, by using two or more different sequence-specific nucleases, each of which recognizes a sequence at one of the predetermined insertion sites. These sequence-specific nucleases can be introduced into the cell simultaneously or sequentially. Thus, for example, in some embodiments, a method is provided for inserting two or more donor sequences, each at a predetermined insertion site on a chromosome in a eukaryotic cell, comprising: a) introducing one or more sequence-specific nucleases that cleave the chromosome at the predetermined insertion site into the cell; b) introducing two or more donor constructs into the cell; and c) introducing an exonuclease into the cell, wherein each of the donor constructs is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the corresponding nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the corresponding nuclease cleavage site on the chromosome, wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends of the linear nucleic acid, respectively, and wherein the two or more donor sequences are inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, a method is provided for inserting two donor sequences, each at a predetermined insertion site on a chromosome in a eukaryotic cell, comprising: a) introducing a first sequence-specific nuclease into the cell, the first sequence-specific nuclease cleaving the chromosome at the first predetermined insertion site; b) introducing a first donor construct into the cell; c) introducing a second sequence-specific nuclease into the cell, the second sequence-specific nuclease cleaving the chromosome at the second predetermined insertion site; and d) introducing a second donor construct into the cell; and e) introducing an exonuclease into the cell, wherein each of the donor constructs is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the corresponding nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the corresponding nuclease cleavage site on the chromosome, wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends of the linear nucleic acid, respectively, and wherein the two donor sequences are inserted into the chromosome at the insertion site by homologous recombination.
[0077] A sequence-specific nuclease (and exonuclease as described herein) can be introduced into a cell in the form of a protein or in the form of a nucleic acid, e.g., mRNA or cDNA, encoding the sequence-specific nuclease (and exonuclease as described herein). The nucleic acid can be delivered as part of a larger construct, e.g., a plasmid or viral vector, or directly, e.g., by electroporation, liposomes, viral transport proteins, microinjection, and biolistic techniques. For example, a sequence-specific nuclease (and exonuclease as described herein) can be introduced into a cell by a variety of methods known in the art, including transfection, calcium phosphate-DNA coprecipitation, DEAE-dextran-mediated transfection, polybrene-mediated transfection, electroporation, microinjection, transduction, cell fusion, liposome fusion, lipofection, protoplast fusion, retroviral infection, use of a gene gun, use of a DNA vector transporter, and biolistic techniques (e.g., particle bombardment) (see, e.g., Wu et al., 1992, J. Biol. Chem., 267:963-967; Wu and Wu, 1988, J. Biol. Chem., 263:14621-14624; and Williams et al., 1991, Proc. Natl. Acad. Sci. USA 88:2726-2730). Receptor-mediated DNA delivery methods can also be used (Curiel et al., 1992, Hum. Gene Ther., 3:147-154; and Wu and Wu, 1987, J. Biol. Chem., 262:4429-4432).
[0078] A donor construct can be introduced into a cell or cleaved within a cell to produce a linear nucleic acid in the form of a linear nucleic acid. It can be delivered by any method suitable for introducing a nucleic acid into a cell. For example, a donor construct can be introduced into a cell by a variety of methods known in the art, including transfection, calcium phosphate-DNA coprecipitation, DEAE-dextran-mediated transfection, polybrene-mediated transfection, electroporation, microinjection, transduction, cell fusion, liposome fusion, lipofection, protoplast fusion, retroviral infection, use of a gene gun, use of a DNA vector transporter, and biolistic techniques (e.g., particle bombardment) (see, e.g., Wu et al., 1992, J. Biol. Chem., 267:963-967; Wu and Wu, 1988, J. Biol. Chem., 263:14621-14624; and Williams et al., 1991, Proc. Natl. Acad. Sci. USA 88:2726-2730). Receptor-mediated DNA delivery methods can also be used (Curiel et al., 1992, Hum. Gene Ther., 3:147-154; and Wu and Wu, 1987, J. Biol. Chem., 262:4429-4432).
[0079] In one embodiment, a target cell can be transfected with a nucleic acid containing a particular gene that results in expression of a gene product in the target cell (e.g., a sequence-specific nuclease or exonuclease as described herein). In another embodiment, a functional protein (e.g., a sequence-specific nuclease or exonuclease) is delivered into a target cell using a membrane-disrupting, pore-forming method or reagent (e.g., microinjection and electroporation), or other reagent (e.g., a liposome that acts as a carrier to deliver the protein across the cell membrane). Introduction of nucleic acids or proteins into a target cell can be confirmed using a variety of assays known in the art, and their effects on cell physiology and / or gene expression can be investigated.
[0080] In some aspects, delivery of the sequence-specific nucleases, donor constructs, and / or exonucleases described herein to a target cell is non-specific, e.g., anything can go into or out of the cell once the membrane is disrupted. In other aspects, delivery of the nucleic acids and / or proteins to a target cell is specific. For example, the sequence-specific nucleases, donor constructs, and / or exonucleases described herein can be delivered into a cell using a protein-transduction domain (PTD) and / or a membrane translocating peptide that mediates the delivery of proteins into cells. These PTDs or signal peptide sequences are naturally occurring polypeptides of 15 to 30 amino acids that generally mediate protein secretion in cells. They consist of a positively charged amino-terminal end, a central hydrophobic core, and a carboxyl-terminal cleavage site recognized by signal peptidases. In certain embodiments, solution-based protein transfection protocols can be used to introduce polypeptides, protein domains, and full-length proteins, including antibodies, into cells. In one aspect, the protein to be introduced into the cell is pre-complexed with a vehicle reagent. In another embodiment, a fusion protein between the protein to be introduced and another moiety is used. For example, the fusion protein comprises a protein of interest (e.g., a sequence-specific nuclease and / or exonuclease) or a domain of a protein of interest covalently fused to a protein or peptide that displays spontaneous cell-penetrating properties. Examples of such membrane transduction peptides include Trojan peptides, human immuodeficiency virus (HIV)-1 transactivator (TAT) protein or functional domain peptides thereof, and other peptides containing protein transduction domains (PTDs) derived from translocation proteins such as the Drosophila homeodomain transcription factor Antennapedia (Antp) and the herpes simplex virus DNA-binding protein VP22. Several commercially available peptides can be used, such as penetratin 1, Pep-1 (Chariot reagent, Active Motif Inc., CA), and the HIV GP41 fragment (519-541).
[0081] In some embodiments, the exonuclease described herein is an alkaline exonuclease. In some embodiments, the exonuclease is a pH-dependent alkaline exonuclease. In some embodiments, the exonuclease interacts with a single-stranded DNA binding protein and facilitates strand exchange. In some embodiments, the exonuclease is also an endonuclease. In some embodiments, the exonuclease is a 5’ to 3’ exonuclease. In some embodiments, the exonuclease is a herpes simplex virus-type 1 (HSV-1) exonuclease. In some embodiments, the exonuclease is UL-12, for example, as described in US 7,135,324 B2, the UL-12 protein (SEQ ID NO: 2, accession number NP_044613.1) encoded by SEQ ID NO: 1 (accession number NC_001806.1, gene ID: 2703382), the disclosure of which is incorporated herein in its entirety for all purposes. In some embodiments, the exonuclease is a UL-12 homolog from Epstaine-Barr virus, bovine herpes virus type 1, pseudorabies virus, and human cytomegalovirus (HCMV).
[0082] In herpes simplex viruses, HSV-1 alkaline nuclease UL-12 and HSV-1 single-stranded DNA binding polypeptide (encoded by the ICP 8 gene, hereinafter “ICP8”; also known in the art as UL-29) act in concert to affect DNA strand exchange. As used herein, UL-12 refers to HSV-1 UL-12 and homologs, orthologs, and paralogs thereof. “Homolog” is a general term used in the art to indicate a polynucleotide or polypeptide sequence that has a high degree of sequence relatedness to a sequence of interest. This relatedness can be quantified by determining the degree of identity and / or similarity between sequences being compared. Falling within the scope of this general term are the terms “ortholog” and “paralog,” the “ortholog” meaning a polynucleotide or polypeptide that is the functional equivalent of a polynucleotide or polypeptide in another species; the “paralog” meaning a functionally similar sequence when considered in the same species. Paralogs present in the same species or orthologs of the UL-12 gene in other species can be readily identified without undue experimentation by molecular biology techniques well known in the art.
[0083] Goldstein and Weller (1998, Virology, 244(2):442-57) examined the region of HSV-1 UL-12 that is highly conserved among herpesvirus homologs and identified seven conserved amino acid regions. The seven regions of homology in herpesviruses were first reported in Martinez et al., 1996, Virology, 215: 152-64. Baculoviruses also encode homologs of this protein (Ahrens et al., 1997, Virology, 229(2):381-99; Ayres et al., 1994, 202:586-605); however, only motifs I-IV are present in these homologs.
[0084] The seven conserved motifs of HSV-1 UL-12 are as follows: Motif I (from amino acid residues 218 to residue 244 of SEQ ID NO: 2), Motif II (from amino acid residues 325 to residue 340 of SEQ ID NO: 2), Motif III (from amino acid residues 362 to residue 377 of SEQ ID NO: 2), Motif IV (from amino acid residues 415 to residue 445 of SEQ ID NO: 2), Motif V (from amino acid residues 455 to residue 465 of SEQ ID NO: 2), Motif VI (from amino acid residues 491 to residue 514 of SEQ ID NO: 2), and Motif VII (from amino acid residues 565 to residue 576 of SEQ ID NO: 2). See Goldstein and Weller, 1998, Virology, 244(2):442-57, the disclosure of which is incorporated herein in its entirety for all purposes. Motif II is one of the most conserved regions. Within this motif, the C-terminal 5 amino acids (336-GASLD-340) represent the most conserved cluster. Asp 340 is an absolutely conserved amino acid, and aspartate residues are required for metal binding in some endo- and exonucleases (Kovall and Matthews, 1997, Science, 277: 1824-7). In Motif II, Gly 336 and Ser 338 are absolutely conserved among the 16 herpesvirus homologs. Goldstein and Weller demonstrated that the D340E mutant and the G336A / S338A mutants of UL-12 lack exonuclease activity, and thus lack in vivo function.
[0085] In some embodiments, the exonuclease has a sequence that is homologous to SEQ ID NO: 1 in any of at least about 70%, 80%, 90%, 95%, 98%, or 99%. In some embodiments, the exonuclease comprises at least 1 (e.g., any of 2, 3, 4, 5, 6, or 7) conserved motifs of UL12.
[0086] In some embodiments, the exonuclease is of eukaryotic or viral origin. In some aspects, the exonuclease is EXOI (eukaryotic) or exo (bacteriophage). In other embodiments, an exonuclease such as ExoIII or bacteriophage T7 gene 6 exonuclease is used. In some embodiments, the exonuclease is Mre11, MRE11A, or MRE11B, e.g., of human origin.
[0087] When introducing an exonuclease and / or donor construct into a cell, a sequence-specific nuclease can be introduced into the cell simultaneously or sequentially. As used herein, the term "sequence-specific endonuclease" or "sequence-specific nuclease" refers to a protein that can recognize and bind to a polynucleotide at a specific nucleic acid sequence and catalyze a single- or double-stranded break in the polynucleotide. In certain embodiments, the sequence-specific nuclease cleaves the chromosome only once, i.e., a single double-stranded break is introduced at the insertion site in the methods described herein.
[0088] Examples of sequence-specific nucleases include zinc finger nucleases (ZFNs). ZFNs are recombinant proteins composed of a DNA-binding zinc finger protein domain and an effector nuclease domain. Zinc finger protein domains are ubiquitous (e.g., transcription factor-associated) protein domains that recognize and bind to specific DNA sequences. A "finger" domain can be composed of about 30 amino acids, which include an invariant set of histidine residues that form a complex with zinc. While more than 10,000 zinc finger sequences have been identified to date, the variety of zinc finger proteins is further expanded by replacing the targeting amino acids in the zinc finger domain to create new zinc finger proteins useful for recognizing specific nucleotide sequences of interest. For example, phage display libraries have been used to screen zinc finger combinatorial libraries to obtain desired sequence specificities (Rebar et al., Science 263:671-673 (1994); Jameson et al., Biochemistry 33:5689-5695 (1994); Choo et al., PNAS 91:11163-11167 (1994), each of which is incorporated herein by reference in its entirety). Zinc finger proteins with the desired sequence specificity can then be linked to an effector nuclease domain, such as Fokl as described in US 6,824,978, as described in PCT Application Publication Nos. WO 1995 / 09233 and WO 1994018313, each of which is incorporated by reference herein in its entirety.
[0089] Another example of a sequence-specific nuclease includes a transcription activator-like effector endonuclease (TALEN) comprising a TAL effector domain that binds to a specific nucleotide sequence and an endonuclease domain that catalyzes a double-stranded break at the target site. Examples of TALENs and methods of making and using are described by PCT Patent Application Publication Nos. WO2011072246 and WO 2013163628 and U.S. Application Publication No. US 20140073015 Al, which are incorporated by reference herein as if in their entirety.
[0090] In one aspect, a transcription activator-like effector (TALE) modulates host gene function by binding to specific sequences within a gene promoter. "Transcription activator-like effector nuclease" or "TALEN" as used interchangeably herein refers to an engineered fusion protein of a catalytic domain of a nuclease (e.g., endonuclease Fokl) and a designed TALE DNA binding domain that can target specific DNA sequences. "TALEN monomer" refers to an engineered fusion protein having a catalytic nuclease domain and a designed TALE DNA binding domain. Two TALEN monomers can be designed for targeting and cleaving a target region. Generally, a TALE includes tandem-like and nearly identical monomers (i.e., repeat domains) flanked by N-terminal and C-terminal sequences. In some embodiments, each monomer contains 34 amino acids, and the sequence of each monomer is highly conserved. Only two amino acids of each repeat (i.e., 12th and 13th residues) are highly variable, also known as repeat variable di-residue (RVD). The RVD determines the nucleotide binding specificity of each TALE repeat domain. An RVD or RVD module generally comprises 33 to 35 amino acids of a TALE DNA binding domain. RVD modules can be combined to create an RVD array. "RVD array length" as used herein refers to the number of RVD modules corresponding to the length of the nucleotide sequence within the target region (i.e., binding region) recognized by the TALEN.
[0091] TALENs can be used to introduce site-specific double-strand breaks at a genomic locus of interest. When two independent TALENs bind to adjacent DNA sequences, a site-specific double-strand break is generated, allowing dimerization of Fokl and cleavage of the target DNA. TALENs have advanced genome editing capabilities due to their high success rate and efficient genetic modification. This DNA cleavage can stimulate natural DNA repair mechanisms, leading to one of two possible repair pathways: homology-directed repair (HDR) or non-homologous end joining (NHEJ) pathways. TALENs can be designed to target any gene, including genes involved in genetic diseases. A TALEN can comprise a nuclease and a TALE DNA-binding domain that binds to a target gene. The target gene can have a mutation, such as a frameshift mutation or a nonsense mutation. If the target gene has a mutation that creates a premature stop codon, a TALEN can be designed to recognize and bind to the nucleotide sequence upstream or downstream of the premature stop codon. In some embodiments, a TALE DNA-binding domain can have an RVD array length of 1 to 30 modules, 1 to 25 modules, 1 to 20 modules, 1 to 15 modules, 5 to 30 modules, 5 to 25 modules, 5 to 20 modules, 5 to 15 modules, 7 to 25 modules, 7 to 23 modules, 7 to 20 modules, 10 to 30 modules, 10 to 25 modules, 10 to 20 modules, 10 to 15 modules, 15 to 30 modules, 15 to 25 modules, 15 to 20 modules, 15 to 19 modules, 16 to 26 modules, 16 to 41 modules, 20 to 30 modules, or 20 to 25 modules. The RVD array length can be any amount of about 5 modules, 8 modules, 10 modules, 11 modules, 12 modules, 13 modules, 14 modules, 15 modules, 16 modules, 17 modules, 18 modules, 19 modules, 20 modules, 22 modules, 25 modules, or 30 modules.
[0092] Another example of a sequence-specific nuclease system that can be used with the methods and compositions described herein includes the Cas / CRISPR system (Wiedenheft, B. et al. Nature 482, 331-338 (2012); Jinek, M. et al. Science 337, 816-821 (2012); Mali, P. et al. Science 339, 823-826 (2013); Cong, L. et al. Science 339, 819-823 (2013)). The Cas / CRISPR (Clustered Regularly Interspaced Short Palindromic Repeat) system utilizes RNA-guided DNA binding and sequence-specific cleavage of target DNA. The guide RNA (gRNA) contains about 20 to 25 (e.g., 20) nucleotides complementary to the target genomic DNA sequence upstream of the genomic PAM (protospacer adjacent motif) site and a constant RNA scaffold region. In certain embodiments, the target sequence is associated with a PAM, which is a short sequence recognized by the CRISPR complex. The precise sequence and length requirements of the PAM vary depending on the CRISPR enzyme used, but the PAM is typically a 2 to 5 bp sequence adjacent to the protospacer sequence (i.e., the target sequence). Examples of PAM sequences are known in the art, and the skilled artisan will be able to identify other PAM sequences for a given CRISPR enzyme. For example, a 5'-N PAM sequence can be identified by searching for a 5'-N PAM sequence on the introduced sequence and the reverse complement of the introduced sequence. x-NGG-3' to identify target sites for Cas9 from S. pyogenes having the PAM sequence NGG. In certain embodiments, the genomic PAM sites used herein are NGG, NNG, NAG, NGGNG, or NNAGAAW. Other PAM sequences and methods for identifying PAM sequences are known in the art, for example as disclosed in U.S. Patent No. 8,697,359, the disclosure of which is incorporated herein by reference for all purposes. In some particular embodiments, Streptococcus pyogenes Cas9 (SpCas9) is used, and the corresponding PAM is NGG. In some aspects, different Cas9 enzymes from different bacterial strains use different PAM sequences. Cas (CRISPR-associated) proteins bind to gRNAs and to the target DNA where the gRNA is bound, and introduce a double-strand break at a defined location upstream of the PAM site. In one aspect, the CRISPR / Cas, Cas / CRISPR, or CRISPR-Cas system (these terms are used interchangeably throughout this application) does not require the production of custom proteins to target specific sequences, but rather a single Cas enzyme can be programmed to recognize a specific DNA target through a short RNA molecule, i.e., the Cas enzyme can be recruited to a specific DNA target using a short RNA molecule.
[0093] In some embodiments, the sequence-specific nuclease is a Type II Cas protein. In some embodiments, the sequence-specific nuclease is Cas9 (also known as Csnl and Csxl2), homologs thereof, or modified versions thereof. In some embodiments, a combination of two or more Cas proteins can be used. In some embodiments, the CRISPR enzyme is Cas9, and can be Cas9 from S. pyogenes or S. pneumoniae. Cas enzymes are known in the art; for example, the amino acid sequence of the S. pyogenes Cas9 protein can be found in the SwissProt database under accession number Q99ZW2.
[0094] In some embodiments, Cas9 is used in the methods described herein. Cas9 carries two separate nuclease domains homologous to HNH and RuvC endonucleases, and by mutating either of these domains, the Cas9 protein can be converted into a nickase that introduces a single-strand break (Cong, L. et al. Science 339, 819-823 (2013)). It is specifically contemplated that the methods and compositions of the application can be used with Cas9 in single- or double-stranded inducing forms, as well as with other RNA-guided DNA nucleases (e.g., other bacterial Cas9-like systems). The sequence-specific nucleases of the methods and compositions described herein can be engineered from organisms, chimeric, or isolated.
[0095] CRISPR, also known as SPIDR (Spacer Interspersed Direct Repeat), constitutes a family of DNA loci that are generally specific to a particular bacterial species. The CRISPR locus contains different classes of interspersed short sequence repeats (SSRs) recognized in E. coli (Ishino et al., 1987, J. Bacteriol., 169: 5429-5433; and Nakata et al., 1989, J. Bacteriol., 171 : 3553-3556) and associated genes. Similar interspersed SSRs have been identified in Haloferax mediterranei, Streptococcus pyogenes, Anabaena, and Mycobacterium tuberculosis (see Groenen et al., 1993, MoI. Microbiol, 10: 1057-1065; Hoe et al., 1999, Emerg. Infect. Dis., 5: 254-263; Masepohl et al., 1996, Biochim. Biophys. Acta 1307: 26-30; and Mojica et al., 1995, MoI. Microbiol, 17: 85-93). The CRISPR locus is generally distinguished from other SSRs by the repetitive structure, which is referred to as short regularly spaced repeat (SRSR) (Janssen et al., 2002, OMICS J. Integ. Biol, 6: 23-33; and Mojica et al., 2000, MoI. Microbiol, 36: 244-246). Generally, the repeats are short elements that occur in clusters, regularly spaced by unique intervening sequences of essentially constant length (Mojica et al., 2000, supra). Although the repeat sequences are highly conserved between strains, the number of interspersed repeats and the sequence of the spacer region generally differ between strains (van Embden et al., 2000, J. Bacteriol, 182: 2393-2401).CRISPR loci have been identified in more than 40 prokaryotic organisms (see, e.g., Jansen et al., 2002, MoI. Microbiol, 43: 1565-1575), including, but not limited to, Aeropyrum, Pyrobaculum, Sulfolobus, Archaeoglobus, Haloarcula, Methanobacterium, Methanococcus, Methanosarcina, Methanopyrus, Pyrococcus, Picrophilus, Thernioplasnia, Corynebacterium, Mycobacterium, Streptomyces, Aquifrx, Porphyromonas, Chlorobium, Thermus, Bacillus, Listeria, Staphylococcus, Clostridium, Thermoanaerobacter, Mycoplasma, Fusobacterium, Azarcus, Chromobacterium, Neisseria, Nitrosomonas, Desulfovibrio, Geohacter, Myrococcus, Campylobacter, Wolinella, Acinetobacter, Erwinia, Escherichia, Legionella, Methylococcus, Pasteurella, Photobacterium, Salmonella, Xanthomonas, Yersinia, Treponema, and Thermotoga.
[0096] A CRISPR system refers to transcripts and other elements collectively involved in expression of a Cas gene or directing activity of a Cas gene, including sequences encoding a Cas gene, a tracr (trans-activating CRISPR) sequence (e.g., a tracrRNA or an active portion of a tracrRNA), a tracr-mate sequence (including a “direct repeat” and a portion of a direct repeat processed by a tracrRNA in the case of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer” in the case of an endogenous CRISPR system), or other sequences and transcripts from a CRISPR locus. In some embodiments, one or more elements of a CRISPR system are from a Type I, Type II, or Type III CRISPR system. In some embodiments, one or more elements of a CRISPR system can be from a particular organism comprising an endogenous CRISPR system, such as Streptococcus pyogenes. In certain embodiments, elements of a CRISPR system contribute to formation of a CRISPR complex at the location of a target sequence (also referred to as a protospacer in the case of an endogenous CRISPR system). In the context of formation of a CRISPR complex, a “target sequence” refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between the target sequence and the guide sequence contributes to formation of a CRISPR complex. It is not necessary for complementarity to be perfect so long as there is sufficient complementarity to cause hybridization and contribute to formation of a CRISPR complex. A target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide. In some embodiments, a target sequence is located in the nucleus or cytoplasm of a cell. In some embodiments, a target sequence can be within an organelle of a eukaryotic cell, such as within a mitochondrion or a chloroplast. A sequence or template that can be used for recombination into a target locus comprising a target sequence is referred to as an “editing template” or “editing polynucleotide” or “editing sequence.” In some aspects of the disclosure, an exogenous template polynucleotide can be referred to as an editing template. In some aspects, recombination is homologous recombination. CRISPR-Cas systems have been used to edit, modulate, and target genomes, for example, as disclosed in Sander and Joung, 2014 Nature Biotechnology 32(4): 347-55, the disclosure of which is incorporated by reference herein for all purposes.
[0097] An exemplary Type II CRISPR system is the Type II CRISPR locus from Streptococcus pyogenes SF370, which contains a cluster of four genes, Cas9, Casl, Cas2, and Csnl, and two non-coding RNA elements, a tracrRNA and an array of repeat sequences (direct repeats) interspaced by short stretches of non-repetitive sequence (spacers, each ~30 bp). In this system, a targeted DNA double-strand break (DSB) is generated in four successive steps. First, two non-coding RNAs (a pre-crRNA array and a tracrRNA) are transcribed from the CRISPR locus. Second, the tracrRNA hybridizes to the direct repeats of the pre-crRNA, which is then processed into mature crRNAs containing individual spacer sequences. Third, the mature crRNA:tracrRNA complex directs Cas9 to a DNA target composed of a protospacer and a corresponding PAM by forming a heteroduplex between the spacer of the crRNA and the protospacer DNA. Finally, Cas9 mediates cleavage of the target DNA upstream of the PAM to create a DSB within the protospacer. Additional descriptions of CRISPR and / or Cas and methods of use can be found in WO 2007025097, US20100093617, US 20130011828, US 13 / 960,796, US 8,546,553, WO 2010011961, US20140093941, US 20100076057, US 20110217739, WO 2010075424, WO 2013142578, WO 2013141680, US 20130326645, WO 2013169802, US 20140068797, WO 2013176772, WO 2013181440, US 20130330778, WO 2013188037, WO 2013188522, WO 2013188638, WO 2013192278, WO 2014018423, CN 103388006, WO 2014022702, US 20140090113, WO 2014039872, WO 2014065596, US 8,697,359, and CN 103725710, the disclosures of which are incorporated by reference in their entireties for all purposes.
[0098] In some embodiments, sequence-specific nuclease is an RNA-guided endonuclease (e.g., Cas / CRIPSR) system. The term "RNA-guided DNA nuclease" or "RNA-guided DNA nuclease" or "RNA-guided endonuclease" as used herein refers to a protein that recognizes and binds to guide RNA and polynucleotides (e.g., target genes) at specific nucleotide sequences and catalyzes single-strand or double-strand breaks in polynucleotides. In some embodiments, guide RNA is an RNA comprising such a 5' region and a 3' region, the 5' region comprising at least one repeat from the CRISPR locus, and the 3' region is complementary to a predetermined insertion site on a chromosome. In certain embodiments, the 5' region comprises a sequence complementary to a predetermined insertion site on a chromosome, and the 3' region comprises at least one repeat from the CRISPR locus. In some aspects, the 3' region of the guide RNA also comprises one or more structural sequences of crRNA and / or trRNA. The 5' region can include, for example, about 1, 2, 3, 4, 5 or more repeats from the CRISPR locus, and can be any length of about 5, 10, 15, 20, 25, 30 or more nucleotides in length. In some embodiments, the 5' region sequence complementary to the predetermined insertion site on the chromosome includes about 17 to about 24 nucleotides. In other embodiments, the 3' region can be, for example, any length of about 5, 10, 15, 20, 25, 30 or more nucleotides in length. In some aspects, the length of the 5' region sequence complementary to the predetermined insertion site on the chromosome can vary, while the length of the 3' region sequence is fixed. In this embodiment where a guide RNA is needed, the step of introducing a sequence-specific nuclease can also include introducing the guide RNA into the cell. The guide RNA can be introduced as, for example, RNA or as a plasmid or other nucleic acid vector encoding the guide RNA. The plasmid or other nucleic acid vector can also include the coding sequence of a sequence-specific nuclease (e.g., Cas) and / or an exonuclease (e.g., UL-12). In some embodiments, the guide RNA includes crRNA and tracrRNA, and the two sections of RNA form a complex by hybridization. In some embodiments, when multiple guide RNAs are used, a single tracrRNA can be used that is paired with different crRNAs.
[0099] Thus, for example, in some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a sequence-specific RNA-guided nuclease (e.g., Cas, e.g., Cas9); b) introducing into the cell a guide RNA that recognizes the insertion site; c) introducing into the cell a donor construct; and d) introducing into the cell a nucleic acid sequence encoding an exonuclease (e.g., UL-12); wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0100] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a nucleic acid sequence encoding a sequence-specific RNA-guided nuclease (e.g., Cas, e.g., Cas9); b) introducing into the cell a nucleic acid sequence encoding a guide RNA that recognizes the insertion site; c) introducing into the cell a donor construct; and d) introducing into the cell a nucleic acid sequence encoding an exonuclease (e.g., UL-12); wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0101] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific RNA-guided nuclease (e.g., Cas, e.g., Cas9); b) introducing into the cell a vector comprising a nucleic acid sequence encoding a guide RNA that recognizes the insertion site; c) introducing into the cell a donor construct; and d) introducing into the cell a DNA vector comprising a nucleic acid sequence encoding an exonuclease (e.g., UL-12); wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0102] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific RNA-guided nuclease (e.g., Cas, e.g., Cas9) and a guide RNA that recognizes the insertion site; b) introducing into the cell a donor construct; and c) introducing into the cell a DNA vector comprising a nucleic acid sequence encoding an exonuclease (e.g., UL-12); wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
[0103] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific RNA-guided nuclease (e.g., Cas, e.g., Cas9) and a nucleic acid sequence encoding an exonuclease (e.g., UL-12); b) introducing into the cell a guide RNA recognizing the insertion site; and c) introducing into the cell a donor construct; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted at the insertion site on the chromosome by homologous recombination.
[0104] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, the method comprising: a) introducing into the cell a vector comprising a nucleic acid sequence encoding a sequence-specific nuclease (e.g., Cas, e.g., Cas9), a guide RNA recognizing the insertion site, and a nucleic acid sequence encoding an exonuclease (e.g., UL-12); b) introducing into the cell a donor construct; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted at the insertion site on the chromosome by homologous recombination.
[0105] In some embodiments, there is provided a method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell (e.g., a zygote), the method comprising: a) injecting into the cell an mRNA sequence encoding a sequence-specific nuclease (e.g., Cas, e.g., Cas9); b) injecting into the cell a guide RNA recognizing the insertion site; c) introducing (e.g., injecting) into the cell a donor construct; and d) injecting into the cell an mRNA sequence encoding an exonuclease; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the injecting is performed in vitro. In some embodiments, the injecting is performed in vivo. In some embodiments, the method further comprises transcribing a nucleic acid encoding a sequence-specific RNA-guided nuclease into mRNA in vitro. In some embodiments, the method further comprises transcribing a nucleic acid encoding a guide RNA recognizing the insertion site into mRNA in vitro. In some embodiments, the method further comprises transcribing a nucleic acid encoding an exonuclease into mRNA in vitro.
[0106] Insertion of the donor sequence can be assessed using any method known in the art. For example, a 5' primer corresponding to the sequence upstream of the 5' homology arm and a corresponding 3' primer corresponding to a region in the donor sequence can be designed to assess the 5' junction of the insertion. Similarly, a 3' primer corresponding to the sequence downstream of the 3' homology arm and a corresponding 5' primer corresponding to a region in the donor sequence can be designed to assess the 3' junction of the insertion. Other methods, such as southern blot hybridization and DNA sequencing techniques, can also be used.
[0107] The insertion site can be at any desired site, provided that a sequence-specific nuclease can be designed to effect cleavage at that site. In some embodiments, the insertion site is at a target locus. In some embodiments, the insertion site is not a locus.
[0108] As used herein, "donor sequence" refers to a nucleic acid to be inserted into a host cell chromosome. In some embodiments, the donor nucleic acid is a sequence that is not present in the host cell. In some embodiments, the donor sequence is an endogenous sequence that is present at a site other than the intended target site. In some embodiments, the donor sequence is a coding sequence. In some embodiments, the donor sequence is a non-coding sequence. In some embodiments, the donor sequence is a mutant locus of a gene.
[0109] The size of the donor sequence can range from about 1 bp to about 100 kb. In certain embodiments, the size of the donor sequence is about 1 bp to about 10 bp, about 10 bp to about 50 bp, about 50 bp to about 100 bp, about 100 bp to about 500 bp, about 500 bp to about 1 kb, about 1 kb to about 10 kb, about 10 kb to about 50 kb, about 50 kb to about 100 kb, or greater than about 100 kb.
[0110] In some embodiments, the donor sequence is an exogenous gene to be inserted into the chromosome. In some embodiments, the donor sequence is a modified sequence that replaces an endogenous sequence at the target site. For example, the donor sequence can be a gene containing a desired mutation and can be used to replace an endogenous gene present on the chromosome. In some embodiments, the donor sequence is a regulatory element. In some embodiments, the donor sequence is a tag or coding sequence that encodes a reporter protein and / or RNA. In some embodiments, the donor sequence can be inserted in-frame to the coding sequence of a target gene, which will allow expression of a fusion protein comprising an exogenous sequence fused to the N- or C-terminus of the target protein.
[0111] The donor constructs described herein are linear nucleic acids or are cleaved within the cell to produce linear nucleic acids. The linear nucleic acids described herein comprise a 5' homology arm, a donor sequence, and a 3' homology arm. The 5' and 3' homology arms are homologous to sequences upstream and downstream of the DNA cleavage site on the target chromosome, allowing homologous recombination to occur.
[0112] The term "homology" or "homologous" as used herein is defined as the percentage of nucleotide residues in a homology arm that are identical with the nucleotide residues in the corresponding sequence on the target chromosome after aligning the sequences and introducing gaps, if necessary, to achieve the maximum percent sequence identity. Alignment for determining percent nucleotide sequence homology can be achieved in various ways that are within the skill in the art, for instance, using publicly available computer software such as BLAST, BLAST-2, ALIGN, ClustalW2, or Megalign (DNASTAR) software. Those skilled in the art can determine appropriate parameters for aligning sequences, including any algorithms needed to achieve maximal alignment over the full length of the sequences being compared. In some embodiments, the homology between the 5' homology arm and the corresponding sequence on the chromosome is at least about any of 80%, 85%, 90%, 95%, 98%, 99%, or 100%. In some embodiments, the homology between the 3' homology arm and the corresponding sequence on the chromosome is at least about any of 80%, 85%, 90%, 95%, 98%, 99%, or 100%.
[0113] In one embodiment, the homology arms are greater than about 30 bp in length, for example greater than about any of 50 bp, 100 bp, 200 bp, 300 bp, 500 bp, 800 bp, 1 kb, 1.5 kb, 2 kb, and 5 kb in length. The 5' and / or 3' homology arms can be homologous to sequences immediately upstream and / or downstream of the DNA cleavage site. Alternatively, the 5' and / or 3' homology arms can be homologous to sequences that are distal to the DNA cleavage site, for example 0 bp from the DNA cleavage site or sequences that partially or completely overlap the DNA cleavage site. In other embodiments, the 5' and / or 3' homology arms can be homologous to sequences that are at least about any of 1, 2, 5, 10, 15, 20, 25, 30, 50, 100, 200, 300, 400, or 500 bp from the DNA cleavage site.
[0114] The 5’ and 3’ homology arms of the linear nucleic acid are proximal to the 5’ and 3’ ends of the linear nucleic acid, i.e., are no more than about 200 bp from the 5’ and 3’ ends of the linear nucleic acid. In some embodiments, the 5’ homology arm is no more than about any of 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 bp from the 5’ end of the linear DNA. In some embodiments, the 3’ homology arm is no more than about any of 1, 2, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 120, 140, 160, 180, or 200 bp from the 3’ end of the linear DNA. In some aspects, the 5’ and / or 3’ homology arms can be directly linked to the 5’ and 3’ ends of the linear DNA, respectively, or partially or completely overlap the 5’ and 3’ ends of the linear DNA, respectively.
[0115] In some embodiments, the donor construct is cleaved within the cell (e.g., by a sequence-specific nuclease that recognizes a cleavage site on the construct) to produce the linear nucleic acid described herein. For example, the donor construct can comprise flanking sequences upstream of the 5’ homology arm and downstream of the 3’ homology arm. In some embodiments, such flanking sequences are not present in the genomic sequence of the host cell, allowing cleavage to occur only on the donor construct. The sequence-specific nuclease can then be designed accordingly to effect cleavage at the flanking sequences, which allows release of the linear nucleic acid without affecting the host sequence. The flanking sequences can be, for example, about 5 to about 500 bp, including any of about 5 to 15, 15 to 30, 30 to 50, 50 to 80, 80 to 100, 100 to 150, 150 to 200, 200 to 300, 300 to 400, or 400 to 500 bp. In some embodiments, the flanking sequences are no more than about any of 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bp. In some embodiments, the portion of the flanking sequence that remains on the linear nucleic acid after sequence-specific cleavage is about any of 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 bp. In some embodiments, the flanking sequences comprise any of the following sequences:
[0116] GGCAGAAATGGCTCCGATCGAGG (SEQ ID NO: 3)
[0117] GGGCGGGATTGATAGCGCGCGGG (SEQ ID NO: 4)
[0118] GGCAGTCGGGAACATCTCGTGGG (SEQ ID NO: 5)
[0119] GGGCGCAGTAATTCTTAGAGCGG (SEQ ID NO: 6)
[0120] GGCTAATAACTTAATCGTGGAGG (SEQ ID NO: 7)
[0121] GGTTAAGCCTTATTGGTGGTCGG (SEQ ID NO: 8)
[0122] GGAGGCCTGCTTGCAAGCATTGG (SEQ ID NO: 9)
[0123] GGTTAGGCCCTAAGCGAATACGG (SEQ ID NO: 10)
[0124] GGAGCCGAGTTGACGGTTAGCGG (SEQ ID NO: 11)
[0125] GGGGTTCCTTCACGAGCGTCCGG (SEQ ID NO: 12)
[0126] GGTACAATGTAACGTTGCGCGGG (SEQ ID NO: 13)
[0127] GGTATTCAAGTCACTAATGTCGG (SEQ ID NO: 14)
[0128] GGAACCCCTTCCGTTCCGTCGGG (SEQ ID NO: 15)
[0129] GGTATTCACTCCTAAAGCGTCGG (SEQ ID NO: 16)
[0130] GGGATGGAACACTAGACTGCGGG (SEQ ID NO: 17)
[0131] GGTTAATCCCTCATGACCGTCGG (SEQ ID NO: 18)
[0132] GGAGCTTCAGTGTCGGTCGTTGG (SEQ ID NO: 19)
[0133] GGTTACGTGCCATATACGTTCGG (SEQ ID NO: 20)
[0134] In some embodiments, the donor construct is a circular DNA construct. The donor construct can also comprise certain sequences that provide structural or functional support, such as sequences of a plasmid or other vector that support amplification of the donor construct (e.g., a pUC19 vector). The donor construct can also optionally comprise certain selectable markers or reporters, some of which can be flanked by recombinase recognition sites for subsequent activation, inactivation, or deletion.
[0135] In some embodiments, when the donor construct is cleaved within the cell to produce a linear nucleic acid, the methods described herein can further comprise introducing a second sequence-specific nuclease into the cell. The second sequence-specific nuclease recognizes a cleavage site on the construct (e.g., a flanking sequence described herein) and causes cleavage of the donor construct within the cell to produce a linear nucleic acid.
[0136] Thus, in some embodiments, there are provided methods of inserting a donor sequence at a predetermined insertion site on a chromosome of a eukaryotic cell, comprising: a) introducing into the cell a first sequence-specific nuclease that cleaves the chromosome at the insertion site; b) introducing into the cell a donor construct (e.g., a circular donor construct), c) introducing into the cell a second sequence-specific nuclease that cleaves the donor construct; and d) introducing into the cell an exonuclease; wherein upon cleavage by the second sequence-specific nuclease, the donor construct produces a linear nucleic acid comprising a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; and wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the first sequence-specific nuclease and the second sequence-specific nuclease are of the same type (e.g., both are ZFNs, TALENs, or CRISPR-based nucleases). In some embodiments, the first sequence-specific nuclease and the second sequence-specific nuclease are of different types. In some embodiments, the first sequence-specific nuclease, the second sequence-specific nuclease, and / or the exonuclease are introduced into the cell simultaneously. In some embodiments, the first sequence-specific nuclease, the second sequence-specific nuclease, and / or the exonuclease are introduced into the cell sequentially. In some embodiments, the first sequence-specific nuclease, the second sequence-specific nuclease, and / or the exonuclease are introduced into the cell as cDNA. In some embodiments, the first sequence-specific nuclease, the second sequence-specific nuclease, and / or the exonuclease are introduced into the cell as mRNA. In some embodiments, the first sequence-specific nuclease, the second sequence-specific nuclease, and / or the exonuclease are introduced into the cell as protein.
[0137] In some embodiments, the first sequence-specific nuclease and the second sequence-specific nuclease are both sequence-specific RNA-guided nucleases. In some such embodiments, a single nuclease can be used with two different guide RNAs (one recognizing the insertion site, the other recognizing the cleavage site on the donor construct). For example, in some embodiments, the method comprises: a) introducing a sequence-specific RNA-guided nuclease into the cell; b) introducing a donor construct (e.g., a circular donor construct) into the cell; c) introducing a first guide RNA recognizing the insertion site into the cell; d) introducing a second guide RNA recognizing the cleavage site on the donor construct into the cell; and e) introducing an exonuclease into the cell; wherein upon cleavage by the sequence-specific nuclease, the donor construct produces a linear nucleic acid comprising a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; and wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the first guide RNA and the second guide RNA are introduced into the cell via a DNA vector (and in some embodiments on the same vector). In some embodiments, the first guide RNA and the second guide RNA are introduced into the cell by injection (e.g., after production by in vitro transcription).
[0138] Accordingly, the present application also provides methods of producing a linear nucleic acid described herein. In some embodiments, a method of producing a linear nucleic acid in a eukaryotic cell is provided, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of a nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of a nuclease cleavage site on the chromosome; and wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, the method comprising: a) introducing into the cell a circular donor construct comprising the linear nucleic acid and further comprising a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm; and b) introducing into the cell a sequence-specific nuclease, wherein the sequence-specific nuclease cleaves the circular nucleic acid construct at the flanking sequences, thereby producing the linear nucleic acid. In some embodiments, the cleavage site is at the 5' flanking sequence and / or the 3' flanking sequence. In some embodiments, the linear DNA is about 200 bp to about 100 kb in length. In certain embodiments, the linear DNA is about 10 bp to about 50 bp, about 50 bp to about 100 bp, about 100 bp to about 150 bp, about 150 bp to about 200 bp, about 200 bp to about 500 bp, about 500 bp to about 1 kb, about 1 kb to about 10 kb, about 10 kb to about 50 kb, about 50 kb to about 100 kb, or longer than about 100 kb in length.
[0139] Uses of the method
[0140] The methods described herein can have a number of uses. For example, the methods described herein can be used to produce genetically modified cells (e.g., immune cells) that can be used in cell therapy.
[0141] In some embodiments, a method of generating a genetically modified animal (e.g., a genetically modified rodent, e.g., a mouse or rat) comprising an inserted donor sequence at a predetermined insertion site on an animal chromosome is provided, the method comprising: a) introducing a sequence-specific nuclease that cleaves the chromosome at the insertion site into a cell of the animal; b) introducing a donor construct into the cell; c) introducing an exonuclease into the cell; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to generate a linear nucleic acid; and d) introducing the cell into a carrier animal to generate the genetically modified animal, wherein the donor construct is a linear nucleic acid or is cleaved within the cell to generate a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. In some embodiments, the cell is an embryonic stem cell. In some embodiments, the cell is a zygote. In some embodiments, the method further comprises breeding the genetically modified animal.
[0142] In some embodiments, the cell is a cell from a blastocyst. After injection of the various components into the cell, a chimeric animal can develop from the injected blastocyst. Heterozygous Fl animals can be obtained by breeding between the chimera and a pure inbred line animal. Homozygous animals can be obtained by crossing between heterozygous animals.
[0143] In some embodiments, the method is used to generate a mutant animal having a particular mutant allele. For example, the donor sequence can comprise a mutant allele, and can be inserted into the genome of the animal (e.g., by replacing a corresponding endogenous locus). The mutant animal can be used for a number of purposes, e.g., as a research tool or disease model. In some embodiments, the animal is modified to have a desired phenotype, e.g., a desired disease phenotype. Animals having various disease phenotypes that can be generated by the methods described herein include, but are not limited to, animals that exhibit a phenotype in a metabolic disease, an immune disease, a neurological disease, a neurodegenerative disease (e.g., Alzheimer's disease), an embryonic development disease, a vascular disease, an inflammatory disease (e.g., asthma and arthritis), an infectious disease, a cancer, a behavioral disease, and a cognitive disease.
[0144] In some embodiments, the methods described herein are used to generate "humanized" animals, e.g., humanized mice or humanized rats. As used herein, "humanized animal" refers to an animal having human-derived donor sequences. The human donor sequences can be inserted at any site on the genome. In some embodiments, the human donor sequences are inserted at the corresponding endogenous loci in the animal cells.
[0145] In some embodiments, humanized rodents can be generated that are capable of producing immunoglobulins comprising human variable domains and / or human constant domains. Rearranged or un-rearranged human immunoglobulin loci containing human immunoglobulin V, D, J, and / or constant gene loci can be placed on the donor constructs described herein and introduced into the genome of the animal cells by homologous recombination. In some embodiments, the human immunoglobulin loci are inserted at the corresponding endogenous immunoglobulin loci in the animal cells. By this manipulation, transgenic animals can be generated that are capable of producing fully human antibodies, chimeric antibodies (e.g., antibodies comprising mouse variable domains and human constant domains), or reverse chimeric antibodies (e.g., antibodies comprising human variable domains and mouse constant regions).
[0146] In some embodiments, the methods described herein are used to generate mouse models for human diseases or disorders. In other aspects, the mouse models reflect or mimic at least one aspect of a human disease or disorder. In other aspects, the compositions and methods disclosed herein are used to knock-in human genes into mice and replace the corresponding mouse genes to generate humanized mice, e.g., mice with humanized TNF-a (TNF-a H-mice) and mice with humanized IL-6 (IL-6 H-mice). Generation of mouse models for human diseases or disorders are disclosed in Wu et al., "Correction of a Genetic Disease in Mouse via Use of CRISPR-Cas9," Cell Stem Cell (2013) 13(6):659-62, and Yang et al., "One-Step Generation of Mice Carrying Reporter and Conditional Alleles by CRISPR / Cas-Mediated Genome Engineering," Cell (2013) 154(6): 1370-9, the disclosures of which are incorporated by reference herein in their entireties for all purposes.
[0147] In some embodiments, transgenic animals having an immune cytokine reporter can be generated by inserting a sequence encoding an immune cytokine reporter at a desired insertion site on an animal cell chromosome via the donor constructs described herein.
[0148] In some embodiments, a transgenic mouse carrying a ROSA26 locus can be produced by inserting a sequence comprising the ROSA26 locus at a desired insertion site on a chromosome of an animal cell by a donor construct described herein.
[0149] Also provided are cells and genetically modified animals produced by any one of the methods described herein.
[0150] Kit
[0151] Also provided herein are kits useful in any one of the methods described herein. For example, in some embodiments, a kit for inserting a donor sequence at an insertion site on a chromosome in a eukaryotic cell is provided, comprising: a) a sequence-specific nuclease that cleaves the chromosome at the insertion site; b) a donor construct, wherein the donor construct comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence 5' to the insertion site on the chromosome, and wherein the 3' homology arm is homologous to a sequence 3' to the insertion site on the chromosome; and c) an exonuclease, wherein the donor construct is a linear nucleic acid or can be cleaved to produce a linear nucleic acid, and wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends of the linear nucleic acid.
[0152] Kits described herein can also comprise packaging that houses the contents of the kit. The packaging optionally provides a sterile, contaminant-free environment, and can be made of any of plastic, paper, foil, glass, etc. In some embodiments, the packaging is a glass vial. In some embodiments, the kit further comprises instructions for performing any one of the methods described herein. Examples
[0153] The following non-limiting examples further illustrate the compositions and methods of the present application. Those skilled in the art will recognize that many embodiments are possible within the scope and spirit of the application. The present application will now be described in greater detail by reference to the following non-limiting examples. The following examples further illustrate the application but, of course, should not be construed as limiting its scope in any way.
[0154] Example 1. EKI system significantly improves knock-in efficiency of EGFP-ACTB in U20S cells
[0155] Figure 2The targeting scheme for expressing the EGFP-ACTB fusion protein is shown in the middle. The targeting vector contains -1 kb homology arms flanking the EGFP sequence, and the sgRNA targets the ACTB allele at a position near the start codon ATG of the ACTB gene. Upon successful homologous recombination, the EGFP sequence will be inserted after the start codon of the ACTB genomic locus for expression of the EGFP-ACTB fusion protein. The sgRNA target sequence for the human ACTB gene is 5'-cgcggcgatatcatcatccatgg-3' (SEQ ID NO: 21).
[0156] The Cas9 sequence (SEQ ID NO: 22) is shown below. The bold underlined sequence is a 3x FLAG tag sequence. The italicized underlined sequence is two SV40 nuclear localization sequences (NLS).
[0157] SEQ ID NO: 22:
[0158] 5'-
[0159] ATG ATGGCC taa-3’
[0160] The UL12 sequence (SEQ ID NO: 23) is shown below. The bold underlined sequence is a 3x FLAG tag sequence. The italicized underlined sequence is an SV40 nuclear localization sequence (NLS).
[0161] SEQ ID NO: 23:
[0162] 5’-
[0163] ATG TGA-3’
[0164] The following plasmids were constructed: Cas9 / sgRNA-hACTB; Cas9 / sgRNA-LS14; targeting vector TV-LS14-hACTB; and pcDNA3.1 Hygro(+)-UL12.
[0165] The constructed plasmids were transfected into U20S cells by electroporation using the Neon (Invitrogen) transfection system.
[0166] The ACTB gene encodes beta-actin, which is a component of the cytoskeleton. Three days after transfection, green fluorescent (GFP) filamentous structures indicative of the cytoskeleton were observed by fluorescence microscopy. Figure 3 ).
[0167] Flow cytometry analysis showed that the efficiency of conventional CRISPR / Cas9-mediated EGFP knock-in was only about 1.91%. In contrast, the EKI system achieved a knock-in efficiency of 15.02% ( Figure 4 ).
[0168] Example 2. The EKI system significantly improves the knock-in efficiency of EGFP-LMNB1 in C6 cells
[0169] Figure 5 A targeting scheme for expressing EGFP-LMNB1 fusion protein is shown in FIG. 2B. The targeting vector contains about 1 kb homology arms flanking the EGFP sequence, and the sgRNA targets the LMNB1 allele at a position near the LMNB1 gene start codon ATG. After successful homologous recombination, the EGFP sequence will be inserted after the start codon of the LMNB1 genomic locus for expression of the EGFP-LMNB1 fusion protein. The sgRNA target sequence for the human LMNB1 gene is 5'-gctgtctccgccgcccgccatgg-3' (SEQ ID NO: 24). The sgRNA target sequence for the rat LMNB1 gene is 5'-gggggtcgcggtcgccatggcgg-3' (SEQ ID NO: 25).
[0170] The following plasmids were constructed: Cas9 / sgRNA-LMNB1; Cas9 / sgRNA-LS14; targeting vector TV-LS14-LMNB1; and pcDNA3.1 Hygro(+)-UL12.
[0171] The constructed plasmids were transfected into C6 cells by electroporation using the Neon (Invitrogen) transfection system.
[0172] The LMNB1 gene encodes lamin BI, which is a component of the nuclear lamina. Three days post-transfection, green fluorescence indicative of the nuclear membrane structure was observed by fluorescence microscopy Figure 6 ).
[0173] Flow cytometry analysis indicated that the efficiency of conventional CRISPR / Cas9-mediated EGFP knock-in was only about 0.19%. In contrast, the EKI system achieved a knock-in efficiency of 3.6% Figure 4 ).
[0174] In another experiment, the targeting scheme shown in Figure 7 was used, with TurboGFP instead of EGFP for expressing TurboGFP-LMNB1 fusion proteins in C6 cells (data not shown).
[0175] Example 3. The EKI system achieves double knock-in of EGFP-ACTB and mCherry-LMNB1 in U20S cells.
[0176] Figure 8 The targeting scheme for expressing EGFP-ACTB fusion proteins and mCherry-LMNB1 fusion proteins after knocking in EGFP and mCherry sequences into the endogenous ACTB and LMNB1 loci, respectively, is shown in
[0177] The following plasmids were constructed: Cas9 / sgRNA-ACTB; Cas9 / sgRNA-LMNB1; Cas9 / sgRNA-LS14; targeting vector TV-LS14-ACTB; targeting vector TV-LS14-LMNB1; and pcDNA3.1 Hygro(+)-UL12.
[0178] The constructed plasmids were transfected into U20S cells by electroporation using the Neon (Invitrogen) transfection system. Three days post-transfection, green fluorescence indicative of filamentous structures representing the cytoskeleton and red fluorescence indicative of the nuclear membrane structure were observed by fluorescence microscopy Figure 9 ).
[0179] Example 4. Efficient production of CD4-2A-dsRed knock-in rats using the EKI system
[0180] GAGGGCCTATTTCCCATGATTCCTTCATATTTGCATATACGATACAAGGCTGTTAGAGAGATAATTGG A targeting scheme for producing CD4-2A-dsRed knock-in rats is shown in FIG. 4. The targeting vector contains -1 kb homology arms flanking the 2A-dsRed sequence, and the sgRNA targets the endogenous rat CD4 allele at a position near the stop codon. Upon successful homologous recombination, the 2A-dsRed sequence inserted near the stop codon of the CD4 genomic locus will be expressed as a CD4-2A-dsRed fusion protein. CD4-positive cells of the knock-in rats will express the 2A-dsRed red fluorescent protein. The sgRNA target sequence for the rat CD4 gene is 5'-gaaaagccacaatctcatatgagg-3' (SEQ ID NO: 26). The sgRNA target sequence for LS14 is 5'-ggtattcactcctaaagcgtcgg-3' (SEQ ID NO: 27).
[0181] An exemplary targeting vector is shown in the 5' to 3' direction: CCGACGCTTTAGGAGTGAATACC (SEQ ID NO: 28, LS14 sequence) - left homology arm - 2A-dsRed - right homology arm - CCGACGCTTTAGGAGTGAATACC (SEQ ID NO: 29, LS14 sequence).
[0182] A U6-sgRNA backbone sequence (SEQ ID NO: 30) can be used. The underlined sequence is the U6 promoter sequence, the bolded sequence is replaced with the target sequence when making the construct, and the italicized sequence is the structural sequence of the sgRNA.
[0183] SEQ ID NO: 30:
[0184] AATTAATTTGACTGTAAACACAAAGATATTAGTACAAAATACGTGACGTAGAAAGTAATAATTTCTTGGGTAGTTT GCAGTTTTAAAATTATGTTTTAAAATGGACTATCATATGCTTACCGTAACTTGAAAGTATTTCGATTTCTTGGCTTTATATA TCTTGTGGAAAGGACGAAACACC TAATACGACTCACTATAGG
[0185] A T7-sgRNA backbone sequence (SEQ ID NO: 31) is shown below. The underlined sequence is the T7 promoter sequence, the bolded sequence is replaced with the target sequence when making the construct, and the italicized sequence is the structural sequence of the sgRNA.
[0186] SEQ ID NO: 31:
[0187] Figure 9
[0188] The following plasmids were constructed: Cas9 / sgRNA-CD4; Cas9 / sgRNA-LS14; targeting vector TV-LS14-CD4; pcDNA3.1Hygro(+)-UL12; T7-Cas9; T7-sgRNA-CD4; and T7-sgRNA-LS14.
[0189] The constructed plasmids were transcribed in vitro to produce UL12 mRNA, Cas9 mRNA, and sgRNAs for CD4 and LS14. UL12 mRNA, Cas9 mRNA, sgRNA-CD4, sgRNA-LS14, and TV-LS14-CD4 were then injected into fertilized rat eggs. These injected eggs were then transplanted into pseudopregnant rats.
[0190] Thirty-three rats of the F0 generation were born and genotyped. 5'-junction PCR reaction was performed using primers LF and INR ( Figure 10 ). The forward primer LF is located at the distal end of the left homology arm, and the reverse primer INR is located within the 2A-dsRed region. Seven of the 33 F0 rats tested were knock-in positive ( Figure 9 ).
[0191] Primers INF and RR were used for 3' end PCR reaction ( Figure 11 ). The forward primer INF is located within the 2A-dsRed region, and the reverse primer RR is located at the distal end of the right homology arm. Seven of the 33 F0 rats tested positive for knock-in ( Figure 12 ).
[0192] Therefore, the 5'-end PCR reaction and the 3'-end PCR reaction produced the same results, and F0 rats No. 6, 7, 14, 17, 22, 23, and 29 were tested as knock-in positive by both PCR assays. The positive rate was 7 / 33 (21.2%).
[0193] Figure 10 The results of southern blot hybridization of knock-in rats are shown, further indicating the insertion of the donor sequence at the intended site. The two F1 rats (No. 19 and No. 21) tested by southern blot hybridization are Figure 11 and Figure 13 Offspring of F0 rat No. 22.
[0194] Example 5. Knock-in of TH-GFP in H9 cells using the EKI system
[0195] This example demonstrates the generation of a TH-GFP knock-in in H9 human embryonic stem cells using the EKI system.
[0196] The human TH gene encodes tyrosine hydroxylase (also known as tyrosine 3-monooxygenase or tyrosinase), which is an enzyme that catalyzes the conversion of the amino acid L-tyrosine to L-3,4-dihydroxyphenylalanine (L-DOPA). L-DOPA is a precursor of dopamine. TH is expressed in the central nervous system, peripheral sympathetic neurons, and adrenal medulla, and is used as a marker for dopaminergic neurons.
[0197] Figure 14 The targeting scheme for generating the H9 cell line of TH-GFP knock-in is shown in FIG. 1. The targeting vector contains -1 kb of homology arms upstream of the F2A-GFP cassette and downstream of the PGK-EM7-Neo-SV40 polyA cassette. The sgRNA targets the TH allele at a position near its stop codon. Upon successful homologous recombination, the F2A-GFP cassette and the PGK-EM7-Neo-SV40 polyA cassette are placed immediately before the stop codon and after the 3' UTR of the TH gene, respectively, for expression of the TH-GFP fusion protein. The sgRNA target sequence of the human TH gene is 5'-ggacgccgtgcacctagccaa tgg-3' (SEQ ID NO: 44).
[0198] The following plasmids were constructed: Cas9 / sgRNA-TH; Cas9 / sgRNA-LS14; targeting vector TV-LS14-TH; and pcDNA3.1 Neo(+)-UL12.
[0199] The constructed plasmids were transfected into H9 cells by electroporation using the Neon (Invitrogen) transfection system. 2 x 10 6 H9 cells, and 2.5 μg each of Cas9 / sgRNA-TH, Cas9 / sgRNA-LS14, TV-LS14-TH, and pcDNA3.1 Neo(+)-UL12. Drug resistant colonies were picked and expanded after 7 to 10 days of selection with G418.
[0200] Green fluorescence from GFP can be used as a marker for TH gene activity, which is observed when H9 cells are differentiated into dopaminergic neurons (see 17.75 μl ). This H9-TH-GFP cell line can be used to support research in many areas, including the differentiation of dopaminergic neurons from human embryonic stem cells.
[0201] The knock-in of TH-GFP in H9 cells was genotyped in three independent cell lines using the primers listed in Table 1, the reaction components listed in Table 2, and the PCR cycle conditions listed in Table 3. Table 1. Primers for H9-TH-GFP genotyping
[0202]
[0203] Table 2. PCR reaction components
[0204] H2O KOD buffer (10x) 3 μl dNTP (2 mM) 3 μl DMSO (0.5%) 1.5 μl Forward primer (10 μM) MgSO4(25 mM) 1.5ul Reverse primer (10 μM) 0.75ul Genomic DNA (100-200 ng / μl) 0.75ul 1 μl KOD-plus 0.75 μl Total volume 30 μl Figure 15
[0205] Table 3. PCR cycle conditions
[0206]
[0207] Primer TH-5'-F and TH-5'-R were used in a 5' end PCR reaction Figure 16 ). Forward primer TH-5'F is located distal to the left homology arm, while reverse primer TH-5'-R is located at the F2A-GFP cassette. All three cell lines tested were knock-in positive Figure 15 A).
[0208] Primer TH-3'-F and TH-3'-R were used in a 3' end PCR reaction Figure 16 ). Forward primer TH-3'-F is located at the PGK-EM7-Neo-SV40 polyA cassette, while reverse primer TH-3'-R is located distal to the right homology arm. All three cell lines tested were knock-in positive Figure 17 B).
[0209] Thus, both the 5' end PCR and the 3' end PCR produced the same results, with cell lines numbered 1, 2, and 3 all tested as knock-in positive by both PCR assays.
[0210] Figure 18 Sequencing results from PCR products from cell line number 1 are shown, which confirm that the cell line correctly targeted the TH locus.
[0211] Figure 19 Cell line number 1 tested is shown to have a normal human karyotype.
[0212] Example 6. Generation of an OCT4-EGFP knock-in in H9 cells using the EKI system
[0213] This example demonstrates the generation of an OCT4-EGFP knock-in in H9 human embryonic stem cells using the EKI system.
[0214] OCT4 (octamer-binding transcription factor 4; also known as POU5F1 : POU domain, class 5, transcription factor 1) encoded by the human POU5F1 gene is a transcription factor that binds an octamer motif (5'-ATTTGCAT-3'). It plays a key role in embryonic development and stem cell self-renewal and pluripotency. OCT4 is expressed in human embryonic stem cells, germ cells, and adult stem cells. Aberrant expression of this gene in adult cells is associated with tumorigenesis.
[0215] Figure 20 A targeting scheme for generating an OCT4-EGFP knock-in H9 cell line is shown in FIG. 1. The targeting vector contains -1 kb homology arms flanking an EGFP-F2A-Puro-SV40-polyA signal sequence cassette. The sgRNA targets the OCT4 allele at a position near its stop codon. Upon successful homologous recombination, the EGFP-F2A-Puro-SV40-polyA signal sequence cassette will be placed immediately before the stop codon of the OCT4 gene for expression of the OCT4-EGFP fusion protein. The sgRNA target sequence for the human OCT4 gene is 5'-tctcccatgcattcaaactgagg-3' (SEQ ID NO: 45).
[0216] The following plasmids were constructed: Cas9 / sgRNA-OCT4; Cas9 / sgRNA-LS14; targeting vector TV-LS14-OCT4; and pcDNA3.1 Puro(+)-UL12.
[0217] The constructed plasmids were transfected into H9 cells by electroporation using the Neon (Invitrogen) transfection system. 2 x 10 6 H9 cells, and 2.5 μg each of Cas9 / sgRNA-OCT4, Cas9 / sgRNA-LS14, TV-LS14-OCT4, and pcDNA3.1 Puro(+)-UL12. Drug resistant colonies were picked and expanded after 7 to 10 days of puromycin selection.
[0218] Green fluorescence from EGFP can be used as a marker for OCT4 gene activity, which is found to be expressed in pluripotent stem cells (see Figure 21 ). This H9-OCT4-EGFP cell line can be used to support research in many areas, including reprogramming and human embryonic stem cell self-renewal and differentiation.
[0219] H9 cells were genotyped for knock-in of OCT4-EGFP in 15 independent cell lines using primers listed in Table 4, reaction components listed in Table 2, and PCR cycling conditions listed in Table 3.
[0220] Table 4. Primers for H9-OCT4-EGFP genotyping
[0221]
[0222] Primer OCT4-5'-F and OCT4-5'-R were used in a 5' end PCR reaction Figure 22 ). Forward primer OCT4-5'-F is located distal to the left homology arm, and reverse primer OCT4-5'-R is located at the EGFP-F2A-Puro-SV40-polyA signal sequence cassette. All 15 cell lines tested were knock-in positive Figure 21 A).
[0223] Primer OCT4-3'-F and OCT4-3'-R were used in a 3' end PCR reaction Figure 22 ). Forward primer OCT4-3'-F is located at the EGFP-F2A-Puro-SV40-polyA signal sequence cassette, and reverse primer OCT4-3'-R is located distal to the right homology arm. Cell lines numbered 1, 2, 4, 6, 7, 10, 11, and 12 tested were knock-in positive Figure 21 B).
[0224] Full-length PCR reactions were also tested using primers OCT4-5'-F and OCT4-3'-R Figure 22 ). Cell lines numbered 1, 3, 4, 5, 6, 7, 10, 11, and 13 tested were knock-in positive Figure 23 C).
[0225] Figure 24 Sequencing results from PCR products of cell line 6 are shown, which confirm that the cell line correctly targeted the OCT4 locus.
[0226] Cell line 6 tested is shown to have a normal human karyotype.
[0227] References
[0228] Iacovitti L, Wei X, Cai J, Kosmicki EW, Lin R, Gorodinsky A, Roman P, Kusek G, Das SS, Dufour A, Martinez TN, Dave KD. 2014. The hTH-GFP reporter rat model for the study of Parkinson's disease. PLoS One 9(12):el 13351. [PubMed: 25462571]
[0229] Hockemeyer D, Wang H, Kiani S, Lai CS, Gao Q, Cassady JP, Cost GJ, Zhang L, Santiago Y, Miller JC, Zeitler B, Cherone JM, Meng X, Hinkley SJ, Rebar EJ, Gregory PD, Urnov FD, Jaenisch R. 2011. Genetic engineering of human pluripotent cells using TALE nucleases. Nat. Biotechnol 29(8):731-4. [PubMed: 21738127]
[0230] Yu J, Vodyanik MA, Smuga-Otto K, Antosiewicz-Bourget J, Frane JL, Tian S, Nie J, Jonsdottir GA, Ruotti V, Stewart R, Slukvin II, Thomson JA. 2007. Induced pluripotent stem cell lines derived from human somatic cells. Science 318(5858): 1917-20. [PubMed: 18029452]
[0231] Boyer LA, Lee TI, Cole MF, Johnstone SE, Levine SS, Zucker JP, Guenther MG, Kumar RM, Murray HL, Jenner RG, Gifford DK, Melton DA, Jaenisch R, Young RA. 2005. Core transcriptional regulatory circuitry in human embryonic stem cells. Cell 122(6):947-56. [PubMed: 16153702] SEQUENCE LISTING <110> Beijing Boao Sai Tujibio Technology Co. Ltd. <120> DNA knock-in system <130> 735782000140 <150> US 62 / 037,551 <151> 2014-08-14 <160> 45 <170> FastSEQ for Windows Version 4.0 <210> 1 <211> 2247 <212> DNA <213> Human herpesvirus type 1 <400> 1 tactgtcgtc ggtggcgctg cctcccgagc ttaagcctct cctggtgctg gtgtcccgcc 60 tgtgtcacac caacccgtgc gcgcggcacg cgctgtcgtg agaatcagcg ttcacccggc 120 ggcgcgctca accaccgctc cccccacgtc gtctcggaaa tggagtccac ggtaggccca 180 gcatgtccgc cgggacgcac cgtgactaag cgtccctggg ccctggccga ggacacccct 240 cgtggccccg acagcccccc caagcgcccc cgccctaaca gtcttccgct gacaaccacc 300 ttccgtcccc tgcccccccc accccagacg acatcagctg tggacccgag ctcccattcg 360 cccgttaacc ccccacgtga tcagcacgcc accgacaccg cagacgaaaa gccccgggcc 420 gcgtcgccgg cactttctga cgcctcaggg cctccgaccc cagacattcc gctatctcct 480 gggggcaccc acgcccgcga cccggacgcc gatcccgact ccccggacct tgactctatg 540 tggtcggcgt cggtgatccc caacgcgctg ccctcccata tactagccga gacgttcgag 600 cgccacctgc gcgggttgct gcgcggcgtc cgcgcccctc tggccatcgg tcccctctgg 660 gcccgcctgg attatctgtg ttccctggcc gtggtcctcg aggaggcggg tatggtggac 720 cgcggactcg gtcggcacct atggcgcctg acgcgccgcg ggcccccggc cgccgcggac 780 gccgtggcgc cccggcccct catggggttt tacgaggcgg ccacgcaaaa ccaggccgac 840 tgccagctat gggccctgct ccggcggggc ctcacgaccg catccaccct ccgctggggc 900 ccccagggtc cgtgtttctc gccccagtgg ctgaagcaca acgccagcct gcggccggat 960 gtacagtctt cggcggtgat gttcgggcgg gtgaacgagc cgacggcccg aagcctgctg 1020 tttcgctact gcgtgggccg cgcggacgac ggcggcgagg ccggcgccga cacgcggcgc 1080 tttatcttcc acgaacccag cgacctcgcc gaagagaacg tgcatacgtg tggggtcctc 1140 atggacggtc acacggggat ggtcggggcg tccctggata ttctcgtctg tcctcgggac 1200 attcacggct acctggcccc agtccccaag acccccctgg ccttttacga ggtcaaatgc 1260 cgggccaagt acgctttcga ccccatggac cccagcgacc ccacggcctc cgcgtacgag 1320 gacttgatgg cacaccggtc cccggaggcg ttccgggcat ttatccggtc gatcccgaag 1380 cccagcgtgc gatacttcgc gcccgggcgc gtccccggcc cggaggaggc tctcgtcacg 1440 caagaccagg cctggtcaga ggcccacgcc tcgggcgaaa aaaggcggtg ctccgccgcg 1500 gatcgggcct tggtggagtt aaatagcggc gttgtctcgg aggtgcttct gtttggcgcc 1560 cccgacctcg gacgccacac catctccccc gtgtcctgga gctccgggga tctggtccgc 1620 cgcgagcccg tcttcgcgaa cccccgtcac ccgaacttta agcagatctt ggtgcagggc 1680 tacgtgctcg acagccactt ccccgactgc cccccccacc cgcatctggt gacgtttatc 1740 ggcaggcacc gcaccagcgc ggaggagggc gtaacgttcc gcctggagga cggcgccggg 1800 gctctcgggg ccgcaggacc cagcaaggcg tccattctcc cgaaccaggc cgttccgatc 1860 gccctgatca ttacccccgt ccgcatcgat ccggagatct ataaggccat ccagcgaagc 1920 agccgcctgg cattcgacga cacgctcgcc gagctatggg cctctcgttc tccggggccc 1980 ggccctgctg ctgccgaaac aacgtcctca tcaccgacga cggggaggtc gtctcgctga 2040 ccgcccacga ctttgacgtc gtggatatcg agtccgaaga ggaaggtaat ttctacgtgc 2100 ccccggatat gcgcggggtt acgcgggccc cggggagaca gcgcctgcgt tcatcggacc 2160 ccccctcgcg ccacactcac cggcggaccc ccggaggcgc ctgccccgcc acccagtttc 2220 caccccccat gtccgatagc gaataaa 2247 <210> 2 <211> 626 <212> PRT <213> 1 type human herpesvirus <400> 2 Met Glu Ser Thr Val Gly Pro Ala Cys Pro Pro Gly Arg Thr Val Thr 1 5 10 15 Lys Arg Pro Trp Ala Leu Ala Glu Asp Thr Pro Arg Gly Pro Asp Ser 20 25 30 Pro Pro Lys Arg Pro Arg Pro Asn Ser Leu Pro Leu Thr Thr Thr Phe 35 40 45 Arg Pro Leu Pro Pro Pro Pro Gln Thr Thr Ser Ala Val Asp Pro Ser 50 55 60 Ser His Ser Pro Val Asn Pro Pro Arg Asp Gln His Ala Thr Asp Thr 65 70 75 80 Ala Asp Glu Lys Pro Arg Ala Ala Ser Pro Ala Leu Ser Asp Ala Ser 85 90 95 Gly Pro Pro Thr Pro Asp Ile Pro Leu Ser Pro Gly Gly Thr His Ala 100 105 110 Arg Asp Pro Asp Ala Asp Pro Asp Ser Pro Asp Leu Asp Ser Met Trp 115 120 125 Ser Ala Ser Val Ile Pro Asn Ala Leu Pro Ser His Ile Leu Ala Glu 130 135 140 Thr Phe Glu Arg His Leu Arg Gly Leu Leu Arg Gly Val Arg Ala Pro 145 150 155 160 Leu Ala Ile Gly Pro Leu Trp Ala Arg Leu Asp Tyr Leu Cys Ser Leu 165 170 175 Ala Val Val Leu Glu Glu Ala Gly Met Val Asp Arg Gly Leu Gly Arg 180 185 190 His Leu Trp Arg Leu Thr Arg Arg Gly Pro Pro Ala Ala Ala Asp Ala 195 200 205 Val Ala Pro Arg Pro Leu Met Gly Phe Tyr Glu Ala Ala Thr Gln Asn 210 215 220 Gln Ala Asp Cys Gln Leu Trp Ala Leu Leu Arg Arg Gly Leu Thr Thr 225 230 235 240 Ala Ser Thr Leu Arg Trp Gly Pro Gln Gly Pro Cys Phe Ser Pro Gln 245 250 255 Trp Leu Lys His Asn Ala Ser Leu Arg Pro Asp Val Gln Ser Ser Ala 260 265 270 Val Met Phe Gly Arg Val Asn Glu Pro Thr Ala Arg Ser Leu Leu Phe 275 280 285 Arg Tyr Cys Val Gly Arg Ala Asp Asp Gly Gly Glu Ala Gly Ala Asp 290 295 300 Thr Arg Arg Phe Ile Phe His Glu Pro Ser Asp Leu Ala Glu Glu Asn 305 310 315 320 Val His Thr Cys Gly Val Leu Met Asp Gly His Thr Gly Met Val Gly 325 330 335 Ala Ser Leu Asp lie Leu Val Cys Pro Arg Asp lie His Gly Tyr Leu 340 345 350 Ala Pro Val Pro Lys Thr Pro Leu Ala Phe Tyr Glu Val Lys Cys Arg 355 360 365 Ala Lys Tyr Ala Phe Asp Pro Met Asp Pro Ser Asp Pro Thr Ala Ser 370 375 380 Ala Tyr Glu Asp Leu Met Ala His Arg Ser Pro Glu Ala Phe Arg Ala 385 390 395 400 Phe lie Arg Ser lie Pro Lys Pro Ser Val Arg Tyr Phe Ala Pro Gly 405 410 415 Arg Val Pro Gly Pro Glu Glu Ala Leu Val Thr Gin Asp Gin Ala Trp 420 425 430 Ser Glu Ala His Ala Ser Gly Glu Lys Arg Arg Cys Ser Ala Ala Asp 435 440 445 Arg Ala Leu Val Glu Leu Asn Ser Gly Val Val Ser Glu Val Leu Leu 450 455 460 Phe Gly Ala Pro Asp Leu Gly Arg His Thr lie Ser Pro Val Ser Trp 465 470 475 480 Ser Ser Gly Asp Leu Val Arg Arg Glu Pro Val Phe Ala Asn Pro Arg 485 490 495 His Pro Asn Phe Lys Gln Ile Leu Val Gin Gly Tyr Val Leu Asp Ser 500 505 510 His Phe Pro Asp Cys Pro Pro His Pro His Leu Val Thr Phe Ile Gly 515 520 525 Arg His Arg Thr Ser Ala Glu Glu Gly Val Thr Phe Arg Leu Glu Asp 530 535 540 Gly Ala Gly Ala Leu Gly Ala Ala Gly Pro Ser Lys Ala Ser Ile Leu 545 550 555 560 Pro Asn Gin Ala Val Pro Ile Ala Leu Ile Ile Thr Pro Val Arg Ile 565 570 575 Asp Pro Glu Ile Tyr Lys Ala Ile Gin Arg Ser Ser Arg Leu Ala Phe 580 585 590 Asp Asp Thr Leu Ala Glu Leu Trp Ala Ser Arg Ser Pro Gly Pro Gly 595 600 605 Pro Ala Ala Ala Glu Thr Thr Ser Ser Ser Pro Thr Thr Gly Arg Ser 610 615 620 Ser Arg 625 <210> 3 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 3 ggcagaaatg gctccgatcg agg 23 <210> 4 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 4 gggcgggatt gatagcgcgc ggg 23 <210> 5 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 5 ggcagtcggg aacatctcgt ggg 23 <210> 6 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 6 gggcgcagta attcttagag cgg 23 <210> 7 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 7 ggctaataac ttaatcgtgg agg 23 <210> 8 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 8 ggttaagcct tattggtggt cgg 23 <210> 9 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 9 ggaggcctgc ttgcaagcat tgg 23 <210> 10 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 10 ggttaggccc taagcgaata cgg 23 <210> 11 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 11 ggagccgagt tgacggttag cgg 23 <210> 12 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 12 ggggttcctt cacgagcgtc cgg 23 <210> 13 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 13 ggtacaatgt aacgttgcgc ggg 23 <210> 14 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 14 ggtattcaag tcactaatgt cgg 23 <210> 15 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 15 ggaacccctt ccgttccgtc ggg 23 <210> 16 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 16 ggtattcact cctaaagcgt cgg 23 <210> 17 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 17 gggatggaac actagactgc ggg 23 <210> 18 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 18 ggttaatccc tcatgaccgt cgg 23 <210> 19 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> synthetic construct <400> 19 ggagcttcag tgtcggtcgt tgg 23 <210> 20 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> synthetic construct <400> 20 ggttacgtgc catatacgtt cgg 23 <210> 21 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> synthetic construct <400> 21 cgcggcgata tcatcatcca tgg 23 <210> 22 <211> 4272 <212> DNA <213> Artificial Sequence <220> <223> synthetic construct <400> 22 atggactata aggaccacga cggagactac aaggatcatg atattgatta caaagacgat 60 gacgataaga tggccccaaa gaagaagcgg aaggtcggta tccacggagt cccagcagcc 120 gacaagaagt acagcatcgg cctggacatc ggcaccaact ctgtgggctg ggccgtgatc 180 accgacgagt acaaggtgcc cagcaagaaa ttcaaggtgc tgggcaacac cgaccggcac 240 agcatcaaga agaacctgat cggagccctg ctgttcgaca gcggcgaaac agccgaggcc 300 acccggctga agagaaccgc cagaagaaga tacaccagac ggaagaaccg gatctgctat 360 ctgcaagaga tcttcagcaa cgagatggcc aaggtggacg acagcttctt ccacagactg 420 gaagagtcct tcctggtgga agaggataag aagcacgagc ggcaccccat cttcggcaac 480 atcgtggacg aggtggccta ccacgagaag taccccacca tctaccacct gagaaagaaa 540 ctggtggaca gcaccgacaa ggccgacctg cggctgatct atctggccct ggcccacatg 600 atcaagttcc ggggccactt cctgatcgag ggcgacctga accccgacaa cagcgacgtg 660 gacaagctgt tcatccagct ggtgcagacc tacaaccagc tgttcgagga aaaccccatc 720 aacgccagcg gcgtggacgc caaggccatc ctgtctgcca gactgagcaa gagcagacgg 780 ctggaaaatc tgatcgccca gctgcccggc gagaagaaga atggcctgtt cggaaacctg 840 attgccctga gcctgggcct gacccccaac ttcaagagca acttcgacct ggccgaggat 900 gccaaactgc agctgagcaa ggacacctac gacgacgacc tggacaacct gctggcccag 960 atcggcgacc agtacgccga cctgtttctg gccgccaaga acctgtccga cgccatcctg 1020 ctgagcgaca tcctgagagt gaacaccgag atcaccaagg cccccctgag cgcctctatg 1080 atcaagagat acgacgagca ccaccaggac ctgaccctgc tgaaagctct cgtgcggcag 1140 cagctgcctg agaagtacaa agagattttc ttcgaccaga gcaagaacgg ctacgccggc 1200 tacattgacg gcggagccag ccaggaagag ttctacaagt tcatcaagcc catcctggaa 1260 aagatggacg gcaccgagga actgctcgtg aagctgaaca gagaggacct gctgcggaag 1320 cagcggacct tcgacaacgg cagcatcccc caccagatcc acctgggaga gctgcacgcc 1380 attctgcggc ggcaggaaga tttttaccca ttcctgaagg acaaccggga aaagatcgag 1440 aagatcctga ccttccgcat cccctactac gtgggccctc tggccagggg aaacagcaga 1500 ttcgcctgga tgaccagaaa gagcgaggaa accatcaccc cctggaactt cgaggaagtg 1560 gtggacaagg gcgcttccgc ccagagcttc atcgagcgga tgaccaactt cgataagaac 1620 ctgcccaacg agaaggtgct gcccaagcac agcctgctgt acgagtactt caccgtgtat 1680 AACGAGCTGACCAAAGTGAAATACGTGACCGAGGGAATGAGAAAGCCC GCC TTCCTGAGC 1740 GGCGAGCAGAAAAGGCCC ATC GTGGACCTGCTGTTCAAGACCAACC GGAAGTGACC GTG 1800 AAGCAGCTGAAAGAGGACTACTTCAAGAAAATCGAGTGCTTCGACTCCGT GGAAATCTCC 1860 GGC GTGGAAGATCGGTTCAACGCCTCCCTGGGCACATACCA CGATCTGCTGAAAATTATC 1920 AAGGACAAGGACTTCCTGGACAATGAGGA AACGAGGACATTCTGGAAGATATCGTGCTG 1980 ACCCTGACACTGT TTGAGGACAGAGAGATGATCGAGGAACGGCTGAAAACCTATGCCCAC 2040 CTGTTCGACGACAAAGTGATGAAGCAGCTGAAGCGGCGGAGATACACCGGCTGGGGC AGG 2100 CTGAGCCGGAA GCTGATCAACGGCATCCGGGACAAGCAGTCCGGCAAGACAATCCTGGAT 2160 TTCCTGAAGTCCGACGGCTTCGCCAACAGAAACTT CATGCAGCTGATCACGACGACAGC 2220 CTGACCTTAAAGAGGACATCCAGAAAGCCCAGGTGTCCGGCCAGGGCGATAGCCTGCAC 2280 GAGCACATTG CCAATCTGGCCGGCAGCCCCGCCATTAAGAAGGGCATCCTGCAGACAGTG 2340 AAGGTGGTGGACGAGCTCGTGAAAGTGATGGGCCGGCACAAGCCCGAGAACATCGTGATC 2400 gaaatggcca gagagaacca gaccacccag aagggacaga agaacagccg cgagagaatg 2460 aagcggatcg aagagggcat caaagagctg ggcagccaga tcctgaaaga acaccccgtg 2520 gaaaacaccc agctgcagaa cgagaagctg tacctgtact acctgcagaa tgggcgggat 2580 atgtacgtgg accaggaact ggacatcaac cggctgtccg actacgatgt ggaccatatc 2640 gtgcctcaga gctttctgaa ggacgactcc atcgacaaca aggtgctgac cagaagcgac 2700 aagaaccggg gcaagagcga caacgtgccc tccgaagagg tcgtgaagaa gatgaagaac 2760 tactggcggc agctgctgaa cgccaagctg attacccaga gaaagttcga caatctgacc 2820 aaggccgaga gaggcggcct gagcgaactg gataaggccg gcttcatcaa gagacagctg 2880 gtggaaaccc ggcagatcac aaagcacgtg gcacagatcc tggactcccg gatgaacact 2940 aagtacgacg agaatgacaa gctgatccgg gaagtgaaag tgatcaccct gaagtccaag 3000 ctggtgtccg atttccggaa ggatttccag ttttacaaag tgcgcgagat caacaactac 3060 caccacgccc acgacgccta cctgaacgcc gtcgtgggaa ccgccctgat caaaaagtac 3120 cctaagctgg aaagcgagtt cgtgtacggc gactacaagg tgtacgacgt gcggaagatg 3180 atcgccaaga gcgagcagga aatcggcaag gctaccgcca agtacttctt ctacagcaac 3240 atcatgaact ttttcaagac cgagattacc ctggccaacg gcgagatccg gaagcggcct 3300 ctgatcgaga caaacggcga aaccggggag atcgtgtggg ataagggccg ggattttgcc 3360 accgtgcgga aagtgctgag catgccccaa gtgaatatcg tgaaaaagac cgaggtgcag 3420 acaggcggct tcagcaaaga gtctatcctg cccaagagga acagcgataa gctgatcgcc 3480 agaaagaagg actgggaccc taagaagtac ggcggcttcg acagccccac cgtggcctat 3540 tctgtgctgg tggtggccaa agtggaaaag ggcaagtcca agaaactgaa gagtgtgaaa 3600 gagctgctgg ggatcaccat catggaaaga agcagcttcg agaagaatcc catcgacttt 3660 ctggaagcca agggctacaa agaagtgaaa aaggacctga tcatcaagct gcctaagtac 3720 tccctgttcg agctggaaaa cggccggaag agaatgctgg cctctgccgg cgaactgcag 3780 aagggaaacg aactggccct gccctccaaa tatgtgaact tcctgtacct ggccagccac 3840 GAGAAGGAGA AGAAGGAGAA GAAGGAGAAG GAGAAGGAGA AGGAG 55 CACAAGCACC TACCTGGAC GAGATCATCG AGCAGATCAG CGAGTTCTCC AAGAGAGTGA TC 3960 CTGGCCGACG CTAATCTGGA CAAAGTGCTG TCCGCCTACA ACAAGCACCG GGATAAGCCC 4020 ATCAGAGAGC AGGCCGAGAA TATCATCCAC CTGTTTACCC TGACCAATCT GGGAGCCCCT 4080 GCCGCCTTCA AGTACTTTGA CACCACCATC GACCAGAAGA GGTACACCAG CACCAAAGAG 4140 GTGCTGGACG CCACCCTGAT CCACCAGAGC ATCACCGGCC TGTCGAGACA CGGATCGAC 4200 CTGTCTCAGC TGGGAGGCGA CAAAAGGCCG GCGGCCACGA AAAAGGCCGG CCAGGCAAAA 4260 AAGAAAAAGT AA 4272 <210> 23 <211> 1968 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 23 ATGCCAAAGA AGAAGCGGAA GGTGAGTCAC GGGAGGCCCA GCAATTCCTG CGGGACGC 60 ACCCTGACTA AGCGTTCCTG GGCCCTGGCC GAGGACACCC CTCGTGGCCC CGACAGCCCC 120 CCCCAAGCGC CCCCGCCTAA CAGTCTTCCG CTGACAACCA CCTTCCGTCC CTGCCCCCCC 180 ccaccccaga cgacgtcagc tgtggaccca agctcccatt cgcccgataa ccccccacgt 240 gatcagcacg ccaccgacac cgcagacgaa aagccccggg ccgcgtcgcc ggcactttct 300 gacgcctcag ggcctccgac cccagacatt ccgctatctc ctgggggcac ccacgcccgc 360 gacccggacg ccgatcccga ctccccggac cttgactcta tgtggtcggc gtcggtgatc 420 cccaacgcgc tgccctccca tatactagcc gagacgttcg agcgccacct gcgcgggttg 480 ctgcgcggcg tccgcgcccc cctggccatc ggtcccctct gggcccgcct ggattatctg 540 tgttccctgg ccgtggtcct cgaggaggcg ggtatggtgg accgcggact cggccggcac 600 ctatggcgcc tgacgcgccg cgggcccccg gccgccgcgg acgccgtggc gccccggccc 660 ctcatggggt tttacgaggc ggccacgcaa aaccaggccg actgccagct atgggccctg 720 ctccggcggg gcctcacgac cgcatccacc ctccgctggg gcccccaggg tccgtgtttc 780 tcgccccagt ggctgaagca caacgccagc ctgcggccgg atgtacagtc ttcggcggtg 840 atgttcgggc gggtgaacga gccgacggcc cgaagcctgc tgtttcgcta ctgcgtgggc 900 cgcgcggacg acggcggcga ggccggcgcc gacacgcggc gctttatctt ccacgaaccc 960 ggcgacctcg ccgaagagaa cgtgcatacg tgtggggtcc tcatggacgg tcacacgggg 1020 atggtcgggg cgtccctgga tattctcgtc tgtcctcggg acactcacgg ctacctggcc 1080 ccagtcccca agacccccct ggccttttac gaggtcaaat gccgggccaa gtacgctttc 1140 gaccccatgg accccagcga ccccacggcc tccgcgtacg aggacttgat ggcacaccgg 1200 tccccggagg cgttccgggc atttatccgg tcgatcccga agcccagcgt gcgatacttc 1260 gcgcccgggc gcgtccccgg cccggaggag gctctcgtca cgcaagacca ggcctggtca 1320 gaggcccacg cctcgggcga aaaaaggcgg tgctccgccg cggatcgggc cttggtggag 1380 ttaaatagcg gcgttgtctc ggaggtgctt ctgtttggcg cccccgacct cggacgccaa 1440 accatctccc ccgtgtcctg gagctccggg gatctggtcc gccgcgagcc cgtcttcgcg 1500 aacccccgtc acccgaactt taagcagatc ttggtgcagg gctacgtgct cgacagccac 1560 ttccccgact gcccccccca cccgcatctg gtgacgttta tcggcaggca ccgcaccagc 1620 gcggaggagg gcgtaacgtt ccgcctggag gacggcgccg gggctctcgg ggccgcagga 1680 cccagcaagg cgtccattct cccgaaccag gccgttccga tcgccctgat cattaccccc 1740 gtccgcatcg atccggagat ctataaggcc atccagcgaa gcagccgcct ggcgttcgac 1800 gacacgctcg ccgagctatg ggcctctcgt tctccggggc ccggccctgc tgctgccgaa 1860 acaacgtcct catcaccgac gacggggagg tcgtctcgcg actataagga ccacgacgga 1920 gactacaagg atcatgatat tgattacaaa gacgatgacg ataagtga 1968 <210> twenty four <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> twenty four gctgtctccg ccgcccgcca tgg 23 <210> 25 <211> twenty three <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 25 gggggtcgcg gtcgccatgg cgg 23 <210> 26 <211> twenty four <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 26 gaaaagccacaatctcatat gagg 24 <210> 27 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 27 ggtattcact cctaaagcgt cgg 23 <210> 28 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 28 ccgacgcttt aggagtgaat acc 23 <210> 29 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 29 ccgacgcttt aggagtgaat acc 23 <210> 30 <211> 341 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <220> <221> misc_feature <222> 250, 251, 252, 253, 254, 255, 256, 257, 258, 259 <223> n = A,T,C or G <400> 30 GAGGGCCTAT TTCCCATGAT TCCTT CATAT TTG CATATA CGATA CAAGGCT GTTAGAGAG 60 ATAAT TGGAA TTAATT TGACT GTA AACACA AAGAT ATTAGT ACAAA ATAC GTGACG TAGA 120 AAGTAATAAT TTCTTGGGTA GTTTGCAGTT TTA AAATTAT GTTTTAA AAT GGACTAT CAT 180 ATGCTTACC GTA ACTTGAAG TATTTCGAT TTCTTGGCTT TATATATCTT GTGGAAAGGA 240 CGAAACACCN NNNNNNNNG TTTTAGAGCT AGAAATAGCA AGT TAAAT AAGGCTAGTCC 300 GTTATCAACT TGAAAAAGTG GCACC GAGTC GGTGCTTTTT 341 <210> 31 <211> 109 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <220> <221> misc_feature <222> 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 <223> n = A,T,C or G <400> 31 TAATACGACT CACTATAGGN NNNNNNNNG TTTTAGAGCT AGAAATAGCA AGT TAAAT AAGGCTAGTCC 60 AGGCTAGTCC GTTATCAACT TGAAAAAGTG GCACC GAGTC GGTGCTTTT 109 <210> 32 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> synthetic construct <400> 32 agtggagtca gtgatgccat tggcctc 27 <210> 33 <211> 26 <212> DNA <213> artificial sequence <220> <223> synthetic construct <400> 33 gcctttggtg ctcttcatct tgttgg 26 <210> 34 <211> 26 <212> DNA <213> artificial sequence <220> <223> synthetic construct <400> 34 tacccgtgat attgctgaag agcttg 26 <210> 35 <211> 25 <212> DNA <213> artificial sequence <220> <223> synthetic construct <400> 35 tttggtagtg ggcaccagct atctg 25 <210> 36 <211> 27 <212> DNA <213> artificial sequence <220> <223> synthetic construct <400> 36 ggtattcagc caaacgacca tctgccg 27 <210> 37 <211> 23 <212> DNA <213> artificial sequence <220> <223> synthetic construct <400> 37 agtcgtgctg cttcatgtgg tcg 23 <210> 38 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 38 tgacacgtgc tacgagattt cgattc 26 <210> 39 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 39 acaggcttca cctgtactgt cagggca 27 <210> 40 <211> 46 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 40 cagacgtacc agtcagtcta cttcgtgtct gagagcttca gtgacg 46 <210> 41 <211> 46 <212> DNA <213> Artificial Sequence <220> <223> Synthetic Construct <400> 41 ctcctctcaa ggaggcaccc atgtcctctc cagctgccgg gcctca 46 <210> 42 <211> 45 <212> DNA <213> Artificial Sequence <220> <223> SYNTHETIC CONSTRUCT <400> 42 gatacccggg gaccttccct ttcttggcct aatttccatt gcttc 45 <210> 43 <211> 46 <212> DNA <213> ARTIFICIAL SEQUENCE <220> <223> SYNTHETIC CONSTRUCT <400> 43 gtgggttaag cggtttgatt cacactgaac caggccagcc cagttg 46 <210> 44 <211> 24 <212> DNA <213> ARTIFICIAL SEQUENCE <220> <223> SYNTHETIC CONSTRUCT <400> 44 ggacgccgtg cacctagcca atgg 24 <210> 45 <211> 23 <212> DNA <213> ARTIFICIAL SEQUENCE <220> <223> SYNTHETIC CONSTRUCT <400> 45 tctcccatgc attcaaactg agg 23
Claims
1. A method of inserting a donor sequence at a predetermined insertion site on a chromosome in a eukaryotic cell, comprising: a) introducing into the cell a sequence-specific nuclease that cleaves the chromosome at the insertion site, wherein the sequence-specific nuclease is introduced into the cell as a protein, mRNA, or cDNA; b) introducing into the cell a donor construct; and c) introducing into the cell an exonuclease, wherein the exonuclease is UL12; wherein the donor construct is a linear nucleic acid or is cleaved within the cell to produce a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
2. The method of claim 1, wherein the sequence-specific nuclease is a zinc finger nuclease (ZFN).
3. The method of claim 1, wherein the sequence-specific nuclease is a transcription activator-like effector nuclease (TALEN).
4. The method of claim 1, wherein the sequence-specific nuclease is an RNA-guided nuclease.
5. The method of claim 4, wherein the RNA-guided nuclease is a Cas.
6. The method of claim 5, wherein the RNA-guided nuclease is Cas9.
7. The method of any one of claims 4 to 6, further comprising introducing into the cell a guide RNA (gRNA) that recognizes the insertion site.
8. The method of any one of claims 1 to 6, wherein the sequence homology between the 5' homology arm and the sequence 5' to the insertion site is at least 80%.
9. The method of any one of claims 1 to 6, wherein the sequence homology between the 3' homology arm and the sequence 3' to the insertion site is at least 80%.
10. The method of any one of claims 1 to 6, wherein the 5' homology arm and the 3' homology arm are at least 200 bp.
11. The method of any one of claims 1 to 6, wherein the exonuclease is a 5' to 3' exonuclease.
12. The method of claim 11, wherein the exonuclease is a herpes simplex virus type 1 (HSV-1) exonuclease.
13. The method of any one of claims 1 to 6 and 12, wherein the donor construct is a linear nucleic acid.
14. The method of any one of claims 1 to 6 and 12, wherein the donor construct is circular when introduced into the cell and is cleaved within the cell to produce a linear nucleic acid.
15. The method of claim 14, wherein the donor construct further comprises a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm.
16. The method of claim 15, wherein the 5' flanking sequence or the 3' flanking sequence is 1 to 500 bp.
17. The method of claim 15 or 16, wherein the method further comprises introducing a second sequence-specific nuclease into the cell that cleaves the donor construct at one or both of the flanking sequences, thereby generating the linear nucleic acid.
18. The method of claim 15 or 16, wherein the sequence-specific nuclease is an RNA- guided nuclease, and wherein the method further comprises introducing a second guide RNA that recognizes one or both of the flanking sequences into the cell.
19. The method of any one of claims 1 to 6, 12, 15, and 16, wherein the eukaryotic cell is a mammalian cell.
20. The method of claim 19, wherein the mammalian cell is a zygote or a pluripotent stem cell.
21. A method of producing a genetically modified animal comprising a donor sequence inserted at a predetermined insertion site on a chromosome of the animal, the method comprising: a) introducing a sequence-specific nuclease that cleaves the chromosome at the insertion site into a cell, wherein the sequence-specific nuclease is introduced into the cell as a protein, mRNA, or cDNA; b) introducing a donor construct into the cell; c) introducing an exonuclease into the cell, wherein the exonuclease is UL12; and d) introducing the cell into a carrier animal to produce the genetically modified animal; wherein the donor construct is a linear nucleic acid or can be cleaved within the cell to generate a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination.
22. The method of claim 21, wherein the genetically modified animal is a rodent.
23. The method of claim 21 or 22, wherein the cell is a zygote or a pluripotent stem cell.
24. A kit for inserting a donor sequence at an insertion site on a chromosome in a eukaryotic cell, comprising: a) a sequence-specific nuclease that cleaves the chromosome at the insertion site, wherein the sequence-specific nuclease is as a protein, mRNA, or cDNA; b) a donor construct; and c) an exonuclease, wherein the exonuclease is UL12; wherein the donor construct is a linear nucleic acid or can be cleaved within the cell to generate a linear nucleic acid, wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid, and wherein the donor sequence is inserted into the chromosome at the insertion site by homologous recombination. wherein the linear nucleic acid comprises a 5' homology arm, the donor sequence, and a 3' homology arm, wherein the 5' homology arm is homologous to a sequence upstream of the nuclease cleavage site on the chromosome, and wherein the 3' homology arm is homologous to a sequence downstream of the nuclease cleavage site on the chromosome; and wherein the 5' homology arm and the 3' homology arm are adjacent to the 5' and 3' ends, respectively, of the linear nucleic acid.
25. The kit of claim 24, wherein the sequence-specific nuclease is an RNA-guided nuclease.
26. The kit of claim 25, wherein the kit further comprises a guide RNA (gRNA) that recognizes the insertion site.
27. The kit of any one of claims 24-26, wherein the donor construct is circular.
28. The kit of claim 27, wherein the donor construct further comprises a 5' flanking sequence upstream of the 5' homology arm and a 3' flanking sequence downstream of the 3' homology arm.
29. The kit of claim 28, wherein the 5' flanking sequence or the 3' flanking sequence is 1-500 bp.
30. The kit of claims 28-29, wherein the sequence-specific nuclease is an RNA-guided nuclease, and wherein the kit further comprises a second guide RNA that recognizes one or both of the flanking sequences.
31. The kit of any one of claims 24-26 and 28-29, wherein the exonuclease is a 5' to 3' exonuclease.
32. The kit of claim 31, wherein the exonuclease is a herpes simplex virus type 1 (HSV-1) exonuclease.
33. The kit of claim 24, wherein the UL12 is fused to a nuclear localization sequence.
34. The method of any one of claims 15, 16, and 20, wherein the 5' flanking sequence and / or the 3' flanking sequence comprises the nucleotide sequence set forth in SEQ ID NO:
28.
35. The kit of any one of claims 28-29 and 32-33, wherein the 5' flanking sequence and / or the 3' flanking sequence comprises the nucleotide sequence set forth in SEQ ID NO: 28.
Citation Information
Patent Citations
TARGET DNA INTERFERENCE WITH crRNA
US20100076057A1
Use
US20100093617A1
Cas6 polypeptides and methods of use
US20110217739A1
Use
US20130011828A1
Methods and compositions for nuclease-mediated targeted integration of transgenes
US20130326645A1