Scarless genome editing by two step homology-directed repair

The two-step HDR method addresses the limitations of existing scarless editing in hPSCs by removing selectable markers, enabling efficient and flexible genome editing at any location with larger edits, thus overcoming the inefficiencies of previous methods.

JP2026012842APending Publication Date: 2026-01-27THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025178375
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2017-07-18
Filing Date
2025-10-23
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing genome editing methods in human pluripotent stem cells (hPSCs) face challenges in achieving scarless editing, including the need for highly active nucleases, prevention of insertions/deletions (INDELs) at non-targeted alleles, and the inefficiency of methods like CORRECT when distance from the cut site to the edited site exceeds 20 base pairs, limiting the range of target sites and requiring cumbersome clone analysis.

Method used

A two-step homology-directed repair (HDR) method involving a first step with a donor polynucleotide and Cas9 to introduce a selectable marker, followed by a second step to remove the marker using a second donor polynucleotide and Cas9, ensuring no residual editing remnants through negative selection.

Benefits of technology

Enables efficient, flexible scarless genome editing at any location, allowing for larger edits without leaving silent mutations or selection markers, reducing the need for clone analysis and increasing editing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026012842000001_ABST
    Figure 2026012842000001_ABST
Patent Text Reader

Abstract

Methods for scarless genome editing are provided, particularly scarless genome modification by using a homology-directed repair (HDR) step to genetically modify a cell and remove an undesired sequence.SOLUTION: Provided is a method for scarless editing of genomic DNA of a cell comprising the following steps. The method comprises the steps of: (a) performing a first Cas9-mediated homologous recombination repair (HDR) step; (b) isolating genetically modified cells based on a positive selection of at least one selectable marker; (c) performing a second Cas9-mediated homologous recombination repair (HDR) step; and (d) isolating genetically modified cells containing the intended editing based on a negative selection of at least one selectable marker, wherein an expression cassette encoding the at least one selectable marker has been deleted.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the fields of genome engineering and scarless genome editing. In particular, the present invention relates to a method of scarless genome editing that utilizes homology-directed repair (HDR) to genetically modify cells to remove unwanted sequences. [Background technology]

[0002] background Precise genome editing has become possible by creating DNA double-strand breaks (DSBs) at specific sites in the genome (Jasin Trends Genet. 12, 224-228 (1996)). DSB creation has been achieved using programmable zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs, Porteus Annu. Rev. Pharmacol. Toxicol. 56, 163-190 (2016)). However, the emergence of the clustered regularly interspaced short palindromic repeats (CRISPR) / Cas9 / guide RNA system (Cas9 / gRNA) has enabled targeted cleavage with unprecedented ease of use, efficiency, and specificity. Each of these nuclease systems stimulates genome editing by initiating sequence-specific double-strand breaks (DSBs) that can then be repaired by either non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR) pathways. The HDR pathway of repair can be divided into two mechanistically distinct strategies: classical gene targeting using the Rad51-dependent homologous recombination pathway (HR) or single-stranded template repair (SSTR) using single-stranded oligonucleotides (Porteus Annu. Rev. Pharmacol. Toxicol. 56, 163-190 (2016) (Non-Patent Document 2); Richardson et al. "CRISPR-Cas9 genome editing in human cells works via the Fanconi Anemia pathway" (2017) (Non-Patent Document 3)). When DSBs are repaired by either NHEJ or MMEJ, loss-of-function substitutions or INDELs can be created. On the other hand, when a homologous donor is provided, HDR allows the introduction of precise insertions, deletions, or substitutions ranging from single nucleotides to large gene cassettes (Hsu et al. Cell 157, 1262-1278 (2014)).

[0003] In cell types traditionally refractory to editing, such as human pluripotent stem cells (hPSCs), selectable markers are typically required to identify and purify genome-edited populations or clones. However, leaving selectable markers in the genome can interfere with transcriptional regulation of the region and would typically be incompatible with clinical translation. Methods exist that do not leave selectable markers (Gonzalez et al. Cell Stem Cell 15, 215-226 (2014) (Non-Patent Document 5); Mitzelfelt et al. Stem Cell Reports 0, 1-9 (2017) (Non-Patent Document 6)), but these rely on the persistent activity of sequence-specific nucleases near the editing site, which has the potential to destroy the original edit (Paquet et al. Nature 533, 1-18 (2016) (Non-Patent Document 7)). To counter this, previous editing methods introduce silent mutations into homology donors to protect edited alleles from persistent nuclease activity.However, the introduction of these silent mutations may still lead to some unintended consequences, such as splicing abnormalities (Yadegari et al. Blood 128, 2144-2152 (2016) (Non-Patent Document 8)), changes in mRNA or protein structure (Duan et al. Human Molecular Genetics 12, 205-216 (2003) (Non-Patent Document 9); Kimchi-Sarfaty et al. Science 315, 525-8 (2007) (Non-Patent Document 10)), or even changes in total protein level due to codon usage bias.Therefore, the ideal genome editing method is "scarless", which still retains the ability to produce HDR-mediated genome editing, but does not leave any remnants of the editing process, either in the form of selection markers or silent mutations.

[0004] Although several scarless editing methods have been reported, all of them have limitations in terms of efficiency and / or versatility in hPSCs. The difficulty of scarless editing is faced with the following challenges: 1) the need for a highly active nuclease to produce a high frequency of desired editing, together with 2) the need to prevent INDELs on non-targeted alleles; and 3) the need to prevent HDR-modified alleles from being re-cut by active nucleases and thus from accumulating INDELs. The solution to the latter is to introduce synonymous changes that prevent re-cutting. However, because Cas9 tolerates small changes, at least one change and sometimes more changes are required to prevent re-cutting, so these synonymous changes correspond to small genetic scars (Fu et al. Nat. Biotechnol. 31, 822-826 (2013)). An alternative approach is to screen hundreds to thousands of clones to find clones with scarless biallelic editing, but there is no structured method to guarantee that such clones exist. Simply improving the frequency of on-target editing does not solve the problem of identifying clones with scarless modifications and may sometimes even make the task more difficult.

[0005] Many previously reported methods rely on target editing that overlaps with the gRNA target, so that Cas9 / gRNA re-cutting is prevented after editing (Gonzalez et al. (Non-Patent Document 5) mentioned above; Yu et al. Cell Stem Cell 16, 142-147 (2015) (Non-Patent Document 12); Liu et al. One-Step Biallelic and Scarless Correction of a β-Thalassemia Mutation in Patient-Specific iPSCs without Drug Selection. Mol Ther Nucleic Acids. 6, 57-67 (2017) (Non-Patent Document 13); Steyer et al. Stem Cell Reports 10, 642-654 (2018) (Non-Patent Document 14)). However, this limits the range of target sites that can be edited without scarring. Paquet et al. (2016) attempted to address this limitation with a two-step editing process using ssODN and gRNA-Cas9, named CORRECT (Paquet et al., supra, Non-Patent Document 7). They introduced mutations to prevent re-cutting in addition to the intended edit, and then repaired the unwanted mutations with a second round of editing. This provided an elegant solution when the distance from the cut site to the edited site was short, but when the distance from the cut site to the target edited site was greater than 20 base pairs, CORRECT was extremely inefficient or not efficient at all (0.0-0.3%). Ultimately, all ssODN-based methods are limited to making small edits, which precludes the possibility of scarless large insertion edits, such as the introduction of cDNA reporters (Yang et al. Nucleic Acids Res. 41, 9049-9061 (2013)). Finally, the CORRECT method remains cumbersome, as it can require the analysis of hundreds of clones in both the first and second stages to identify the correct clone. Although adding a selectable marker without removing it can reduce the number of clones that need to be analyzed, the marker leaves a visible scar in the genome.Therefore, for it to be a scarless system, the integrated marker must be removed. The PiggyBac system is designed to remove the integrated marker using the PiggyBac transposon, excising the marker between the transposon-specific inverted terminal repeats and leaving only the "TTAA" sequence. It is scarless as long as the TTAA is part of the endogenous target (Ye et al. Proc. Natl. Acad. Sci. 111, 9591-9596 (2014)). However, the number of sites where both the gRNA and the adjacent TTAA sequence exist with a short distance from the cleavage to the target is limited. Recently, Kim and colleagues reported that MMEJ-based editing allows scarless removal of the integrated marker if microhomologies are present at both ends of the marker (Kim et al. Nat. Commun. 9, (2018)). In their report, they did not demonstrate true scarless editing because they introduced synonymous mutations during the process. Furthermore, they used a system that integrates a reporter into a gene expressed in pluripotent cells to reduce the frequency of identifying clones with random integrations, a feature that also limits the range of gene targets to which the method can be applied.

[0006] Therefore, more efficient and flexible methods of scarless editing that can modify the genome at any location and allow for edits of larger size are needed to enable genomes to be designed as desired. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] Jasin Trends Genet. 12, 224-228 (1996) [Non-patent document 2] Porteus Annu. Rev. Pharmacol. Toxicol. 56, 163-190 (2016) [Non-licensed document 3] Richardson et al. "CRISPR-Cas9 genome editing in human cells works via the Fanconi Anemia pathway" (2017)

Non-licensed Document 4

Non-licensed Document 5

Non-licensed Document 6

Non-licensed Document 7

Non-licensed literature 9

Non-licensed literature 10

Non-licensed Document 11

Non-licensed Document 12

Non-licensed Document 13

[0008] overview The present invention relates to a method for scarless genome editing. In particular, this method provides scarless genome modification by using HDR step to genetically modify cells and remove unwanted sequences.This method can be used for genome editing, including introducing mutation, deletion or insertion at any position in genome, without leaving silence mutation, selection marker sequence or other additional unwanted sequences in genome.

[0009] In one aspect, the present invention includes a method for scarless editing of genomic DNA of a cell, comprising the steps of: (a) performing a first Cas9-mediated homology-directed repair (HDR) step comprising: (i) introducing into the cell a first donor polynucleotide, wherein the first donor polynucleotide comprises a first left homology arm and a first right homology arm that flank a sequence comprising an intended edit for the target sequence to be modified in the genomic DNA of the cell, and an expression cassette encoding at least one selectable marker; and (ii) introducing into the cell a first guide RNA and Cas9, wherein the first guide RNA comprises a sequence complementary to the genomic target sequence to be modified in the genomic DNA of the cell, such that Cas9 forms a first complex with the first guide RNA, the guide RNA directs the first complex to the genomic target sequence, and the first donor polynucleotide is integrated into the genomic DNA by Cas9-mediated HDR to generate a genetically modified cell; (c) (i) introducing into the genetically modified cells a second donor polynucleotide, the second donor polynucleotide comprising a second left homology arm and a second right homology arm flanking sequences complementary to a target genomic sequence as modified by integration of the first donor polynucleotide into genomic DNA by Cas9-mediated HDR, except that the second donor polynucleotide comprises a deletion of an expression cassette encoding the at least one selectable marker; and (ii) performing a second Cas9-mediated homology-directed repair (HDR) step, comprising: introducing into the cell a second guide RNA and a Cas9 nuclease, wherein Cas9 forms a second complex with the second guide RNA, the second guide RNA directs the second complex to the modified target genomic sequence, Cas9 creates a double-stranded break in the modified genomic DNA, and a second donor polynucleotide is integrated into the genomic DNA by Cas9-mediated HDR, thereby removing an expression cassette encoding at least one selectable marker from the modified genomic DNA by Cas9-mediated HDR;and (d) isolating genetically modified cells containing the intended edit based on negative selection of at least one selectable marker, wherein the expression cassette encoding the at least one selectable marker is deleted;

[0010] In one embodiment, the sequence containing the intended edit is located within or near an expression cassette encoding a selectable marker.

[0011] In certain embodiments, the intended editing introduces mutations, such as insertions, deletions, or substitutions, into the gene in the genomic DNA of the cell.In other embodiments, the intended editing removes mutations from the gene in the genomic DNA of the cell.In another embodiment, the intended editing results in the inactivation of the gene in the genomic DNA of the cell.

[0012] Either one allele or two alleles may be modified in genomic DNA. In certain embodiments, at least one selection marker is a fluorescent marker, and the fluorescence intensity can be measured to determine whether the genetically modified cell contains single-allelic editing or biallelic editing.

[0013] In certain embodiments, the double-stranded break is a non-gene-disrupting double-stranded break. In certain embodiments, the double-stranded break does not affect expression of the at least one selectable marker.

[0014] In certain embodiments, one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, and Cas9 are provided by a vector, for example, a plasmid or a viral vector.In another embodiment, the first guide RNA and Cas9 are provided by a single vector or multiple vectors.In another embodiment, the second donor polynucleotide, the second guide RNA, and Cas9 are provided by a single vector or multiple vectors.

[0015] In certain embodiments, the cells to be genetically modified are derived from eukaryotes, prokaryotes, or archaea. Scarless genome editing as described herein may be performed on cells in vitro or in vivo. Cells may be derived from any type of animal, including vertebrates or invertebrates. For example, cells may be obtained from mammals, such as humans or non-human mammals (e.g., non-human primates, laboratory animals, livestock), or domesticated birds, wild birds, or game birds. In another embodiment, the cells are derived from cell lines. Cells may be immortalized or cancerous. In another embodiment, the cells are selected from the group consisting of K562 cells, embryonic stem cells, and induced pluripotent stem cells.

[0016] In certain embodiments, the at least one selectable marker is selected from the group consisting of a cell surface marker, a drug resistance gene, a reporter gene, and a suicide gene. In other embodiments, the cell surface marker is selected from the group consisting of cleaved CD8, NGFR, cleaved CD19 (tCD19), CCR5, and ABO antigens.

[0017] The genetically modified cells generated from the first HDR step can be isolated by positive selection using a binding agent that specifically binds to the selection marker. In certain embodiments, the binding agent comprises an antibody, antibody mimetic, or aptamer that specifically binds to the selection marker.

[0018] In another embodiment, the binding agent is a monoclonal antibody, a polyclonal antibody, a chimeric antibody, a nanobody, a recombinant fragment of an antibody, a Fab fragment, a Fab' fragment, a F(ab')2 fragment, a F v fragments, and scF v The antibody may be selected from the group consisting of antibody fragments.

[0019] For example, cells genetically modified to express a surface marker may be isolated using a binding agent that specifically binds to that surface marker (e.g., an anti-tCD19 antibody to isolate cells bearing the tCD19 surface marker, an anti-CD8 antibody to isolate cells bearing the CD8 surface marker, or an anti-NGFR antibody to isolate cells bearing the NGFR surface marker).

[0020] In certain embodiments, the binding agent is immobilized on a solid support. Exemplary solid supports include magnetic beads, non-magnetic beads, slides, gels (e.g., agarose or acrylamide), nylon, membranes, glass plates, microtiter plate wells, or metal, glass, or plastic surfaces.

[0021] Any cell separation technique may be used to isolate the genetically modified cells, including but not limited to fluorescence-activated cell sorting (FACS), magnetic-activated cell sorting (MACS), elutriation, immunopurification, or affinity chromatography.

[0022] In another embodiment, the expression cassette encoding at least one selectable marker comprises a UbC promoter, a polynucleotide encoding mCherry, a polynucleotide encoding a T2A peptide, a polynucleotide encoding a truncated CD19 (tCD19), and a polyadenylation sequence.

[0023] In certain embodiments, the first donor polynucleotide further comprises at least one expression cassette that encodes the short hairpin RNA (shRNA) that inhibits the expression of randomly integrated or episomal selection marker.For example, the first donor polynucleotide can comprise a pair of expression cassettes that encodes the short hairpin RNA (shRNA), and the first expression cassette of the pair is located at the 5' of the first left homology arm, and the second expression cassette of the pair is located at the 3' of the first right homology arm.The expression of shRNA reduces the selection of the clone that expresses the selection marker that is randomly integrated into genome.

[0024] In another aspect, the present invention includes a scarless genome editing system comprising: (a) a first donor polynucleotide comprising a first left homology arm and a first right homology arm flanking a sequence comprising an intended edit for a target sequence to be modified in the genomic DNA of a cell, and an expression cassette encoding at least one selectable marker; (b) a second donor polynucleotide comprising a second left homology arm and a second right homology arm flanking a sequence complementary to the sequence of the first donor polynucleotide but comprising a deletion of the expression cassette encoding the at least one selectable marker; (c) a Cas9 nuclease; (d) a first guide RNA capable of forming a complex with the Cas9 nuclease and directing the complex to the target sequence; and (e) a second guide RNA capable of forming a complex with the Cas9 nuclease and directing the complex to the target genomic sequence to be modified by integration of the first donor polynucleotide into the genomic DNA by Cas9-mediated homology-directed repair (HDR).

[0025] In another embodiment, the first donor polynucleotide comprises an expression cassette encoding at least one selectable marker, wherein the expression cassette comprises the sequence of SEQ ID NO:1, or a sequence exhibiting at least about 80-100% sequence identity thereto, including any percent identity thereto within these ranges, e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity thereto; the first donor polynucleotide can be integrated into the genomic DNA of a cell at a target site by Cas9-mediated HDR under conditions suitable for expression of the at least one selectable marker; and cells having the first donor polynucleotide integrated into the genomic DNA of the cell at the target site by Cas9-mediated HDR can be identified by positive selection of the at least one selectable marker encoded by the expression cassette.

[0026] In certain embodiments, one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, and Cas9 in the scarless genome editing system are provided by a vector, for example, a plasmid or a viral vector.In another embodiment, the first guide RNA and Cas9 are provided by a single vector or multiple vectors.In another embodiment, the second donor polynucleotide, the second guide RNA, and Cas9 are provided by a single vector or multiple vectors.

[0027] In another aspect, the invention includes a host cell comprising the scarless genome editing system described herein.

[0028] In another aspect, the present invention includes a kit comprising the scarless genome editing system described herein. Such a kit may comprise a first donor polynucleotide, a second donor polynucleotide, a first guide RNA, a second guide RNA, and a Cas9 nuclease. The kit may further comprise instructions for performing scarless editing of genomic DNA.

[0029] In certain embodiments, the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA and Cas9 in the kit are provided by a vector, for example, a plasmid or a viral vector.In another embodiment, the first guide RNA and Cas9 are provided by a single vector or multiple vectors.In another embodiment, the second donor polynucleotide, the second guide RNA and Cas9 are provided by a single vector or multiple vectors.

[0030] In one embodiment, the first donor polynucleotide comprises a selectable marker expression cassette comprising the sequence of SEQ ID NO:1, or a sequence exhibiting at least about 80-100% sequence identity thereto, including any percent identity thereto within this range, e.g., 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% sequence identity thereto, and cells having the first donor polynucleotide integrated into the genomic DNA of the cell at the target site by Cas9-mediated HDR can be identified by positive selection of the selectable marker encoded by the expression cassette.

[0031] These and other aspects of the present invention will readily occur to those of ordinary skill in the art in view of the disclosure herein. [The present invention 1001] A method for scarless editing of genomic DNA of a cell, comprising the steps of: (a)(i) introducing into a cell a first donor polynucleotide, the first donor polynucleotide comprising a first left homology arm and a first right homology arm flanking a sequence comprising an intended edit to a target sequence to be modified in the genomic DNA of the cell, and an expression cassette encoding at least one selectable marker; and (ii) introducing into a cell a first guide RNA and Cas9, wherein the first guide RNA comprises a sequence complementary to a genomic target sequence to be modified in the genomic DNA of the cell, such that Cas9 forms a first complex with the first guide RNA, the guide RNA directs the first complex to the genomic target sequence, Cas9 creates a double-stranded break in the genomic DNA of the cell, and the first donor polynucleotide is integrated into the genomic DNA by Cas9-mediated homology-directed repair (HDR) to generate a genetically modified cell. performing a first Cas9-mediated homology-directed repair (HDR) step, comprising: (b) isolating the genetically modified cells based on positive selection of at least one selectable marker; (c)(i) introducing into the genetically modified cell a second donor polynucleotide, wherein the second donor polynucleotide comprises a second left homology arm and a second right homology arm flanked by sequences complementary to the target genomic sequence as modified by integration of the first donor polynucleotide into the genomic DNA by Cas9-mediated HDR, except that the second donor polynucleotide comprises a deletion of an expression cassette encoding at least one selectable marker; (ii) introducing a second guide RNA and a Cas9 nuclease into the cell, wherein Cas9 forms a second complex with the second guide RNA, the second guide RNA directs the second complex to the modified target genomic sequence, Cas9 creates a double-stranded break in the modified genomic DNA, and the second donor polynucleotide is integrated into the genomic DNA by Cas9-mediated HDR, thereby removing an expression cassette encoding at least one selectable marker from the modified genomic DNA by Cas9-mediated HDR. performing a second Cas9-mediated homology-directed repair (HDR) step, comprising: (d) isolating genetically modified cells containing the intended edit based on negative selection of at least one selectable marker, wherein the expression cassette encoding the at least one selectable marker is deleted. [The present invention 1002] 1001. The method of claim 1001, wherein the sequence containing the intended edit is located within or near an expression cassette encoding at least one selectable marker. [The present invention 1003] 1001. The method of claim 1001, wherein said intended editing introduces a mutation into a gene in the genomic DNA of the cell. [The present invention 1004] 1004. The method of claim 1003, wherein said mutation is selected from the group consisting of an insertion, a deletion, and a substitution. [The present invention 1005] 1001. The method of claim 1001, wherein said intended editing removes a mutation from a gene in the genomic DNA of the cell. [The present invention 1006] 1001. The method of claim 1001, wherein said intended editing results in the inactivation of a gene in the genomic DNA of the cell. [The present invention 1007] 1001. The method of claim 1001, wherein one allele is modified in the genomic DNA. [The present invention 1008] The method of claim 1001, wherein both alleles are modified in the genomic DNA. [The present invention 1009] 1001. The method of claim 10, wherein said at least one selectable marker is a fluorescent marker. [The present invention 1010] 1009. The method of claim 10, further comprising measuring fluorescence intensity to determine whether said genetically modified cells contain single-allelic edits or biallelic edits. [The present invention 1011] 1001. The method of claim 1001, wherein said cell is from a eukaryotic, prokaryotic, or archaeal organism. [The present invention 1012] 10. The method of claim 1 1 , wherein the cell is mammalian. [The present invention 1013] 1013. The method of claim 1012, wherein the cell is human. [The present invention 1014] 1001. The method of claim 1001, wherein said cells are derived from a cell line. [The present invention 1015] 1001. The method of claim 1001, wherein said cell is an immortalized cell. [The present invention 1016] 1001. The method of claim 1001, wherein said cells are cancerous. [The present invention 1017] 1001. The method of claim 1001, wherein said cell is in vitro or in vivo. [The present invention 1018] 1001. The method of claim 1001, wherein said cells are selected from the group consisting of K562 cells, embryonic stem cells, and induced pluripotent stem cells. [The present invention 1019] 1001. The method of claim 1001, wherein said at least one selectable marker is selected from the group consisting of a cell surface marker, a drug resistance gene, a reporter gene, and a suicide gene. [The present invention 1020] 1020. The method of claim 1019, wherein said cell surface marker is selected from the group consisting of cleaved CD8, NGFR, cleaved CD19 (tCD19), CCR5, and ABO antigens. [The present invention 1021] The method of claim 1019, wherein said cell surface marker is not essential for cell function. [The present invention 1022] The method of claim 1019, wherein the genetically modified cells generated from the first HDR step are isolated by positive selection using a binding agent that specifically binds to the selectable marker. [The present invention 1023] 1023. The method of claim 1022, wherein said binding agent comprises an antibody, antibody mimetic, or aptamer that specifically binds to the selectable marker. [The present invention 1024] The antibody may be a monoclonal antibody, a polyclonal antibody, a chimeric antibody, a nanobody, a recombinant fragment of an antibody, a Fab fragment, a Fab' fragment, a F(ab')2 fragment, a F v fragments, and scF v The method of claim 1023, wherein the fragment is selected from the group consisting of: [The present invention 1025] 1023. The method of claim 1022, wherein said antibody is selected from the group consisting of an anti-tCD19 antibody, an anti-CD8 antibody, and an anti-NGFR antibody. [The present invention 1026] The method of claim 1022, wherein said binding agent is immobilized on a solid support. [The present invention 1027] 1027. The method of claim 1026, wherein said solid support is a magnetic bead, a non-magnetic bead, a slide, a gel, a membrane, or a microtiter plate well. [The present invention 1028] 1001. The method of claim 1001, wherein the expression cassette encoding the at least one selectable marker comprises a UbC promoter, a polynucleotide encoding mCherry, a polynucleotide encoding a T2A peptide, a polynucleotide encoding a truncated CD19 (tCD19), and a polyadenylation sequence. [The present invention 1029] 1001. The method of claim 1001, wherein one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, or Cas9 are provided by a vector. [The present invention 1030] 1001. The method of claim 10, wherein said vector is a plasmid or a viral vector. [The present invention 1031] The method of claim 1029, wherein the first donor polynucleotide, the first guide RNA, and the Cas9 are provided by a single vector or multiple vectors. [The present invention 1032] The method of claim 1029, wherein the second donor polynucleotide, the second guide RNA, and the Cas9 are provided by a single vector or multiple vectors. [The present invention 1033] 1001. The method of claim 1001, wherein the first donor polynucleotide further comprises at least one expression cassette encoding a short hairpin RNA (shRNA) that inhibits expression of a randomly integrated or episomal selectable marker. [The present invention 1034] The method of claim 1033, wherein the first donor polynucleotide comprises a pair of expression cassettes encoding an shRNA, the first expression cassette of the pair being located 5' of the first left homology arm and the second expression cassette of the pair being located 3' of the first right homology arm. [This invention 1035] 1034. The method of claim 1033, wherein said shRNA reduces selection of clones expressing a selectable marker that has been randomly integrated into the genome. [The present invention 1036] 1001. The method of claim 1001, wherein the double-strand break is a non-gene-disrupting double-strand break. [This invention 1037] 1001. The method of claim 1001, wherein said double-stranded break does not affect expression of said at least one selectable marker. [The present invention 1038] Scarless genome editing system, including: (a) a first donor polynucleotide comprising a first left homology arm and a first right homology arm flanking a sequence comprising an intended edit to a target sequence to be modified in the genomic DNA of a cell, and an expression cassette encoding at least one selectable marker; (b) a second donor polynucleotide comprising a second left homology arm and a second right homology arm flanked by sequences complementary to the sequence of the first donor polynucleotide except for the deletion of an expression cassette encoding at least one selectable marker; (c) Cas9 nuclease; (d) a first guide RNA capable of forming a complex with Cas9 nuclease and directing the complex to a target sequence; and (e) A second guide RNA capable of forming a complex with a Cas9 nuclease and directing the complex to a target genomic sequence to be modified by integration of the first donor polynucleotide into genomic DNA by Cas9-mediated homology-directed repair (HDR). [This invention 1039] 1038. The scarless genome editing system of claim 1038, wherein the expression cassette encoding the at least one selectable marker comprises the sequence of SEQ ID NO:1 or a sequence having at least 95% identity to the sequence of SEQ ID NO:1, the first donor polynucleotide can be integrated into the genomic DNA of a cell at a target sequence by Cas9-mediated HDR under conditions suitable for expression of the at least one selectable marker, and cells having the first donor polynucleotide integrated into the genomic DNA of a cell at a target site by Cas9-mediated HDR can be identified by positive selection of the at least one selectable marker encoded by the expression cassette. [The present invention 1040] The scarless genome editing system of the present invention 1038, wherein the second donor polynucleotide comprises a polynucleotide comprising the sequence of SEQ ID NO:2 or a sequence having at least 95% identity to the sequence of SEQ ID NO:2, and the second donor polynucleotide can be integrated into the modified genomic DNA of the cell at the target site by Cas9-mediated HDR to remove an expression cassette encoding at least one selectable marker. [The present invention 1041] The scarless genome editing system of the present invention 1038, wherein one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, or Cas9 are provided by a vector. [The present invention 1042] The scarless genome editing system of the present invention 1038, wherein the vector is a plasmid or a viral vector. [This invention 1043] The scarless genome editing system of the present invention 1038, wherein the first donor polynucleotide, the first guide RNA, and Cas9 are provided by a single vector or multiple vectors. [This invention 1044] The scarless genome editing system of the present invention 1038, wherein the second donor polynucleotide, the second guide RNA, and the Cas9 are provided by a single vector or multiple vectors. [This invention 1045] The scarless genome editing system of the present invention 1038, wherein the first donor polynucleotide further comprises at least one expression cassette encoding a short hairpin RNA (shRNA) that inhibits expression of a randomly integrated or episomal selectable marker. [The present invention 1046] A scarless genome editing system of the present invention 1038, wherein the first donor polynucleotide comprises a pair of expression cassettes encoding shRNA, the first expression cassette of the pair being located 5' of the first left homology arm and the second expression cassette of the pair being located 3' of the first right homology arm. [This invention 1047] A host cell comprising the scarless genome editing system of the present invention. [This invention 1048] A kit comprising the scarless genome editing system of the present invention 1038 and instructions for performing scarless editing of genomic DNA. [This invention 1049] A kit comprising the host cell of the present invention. [Brief explanation of the drawings]

[0032] [Figure 1A] Figure 1A shows the workflow of the two-step HR scarless editing procedure. [Figure 1B]Figure 1B shows a schematic of scarless editing. Figure 1B shows the donor DNA construct and guide RNA for two-step HR. In the case of TBX1 editing, the intended mutation is a G to A substitution at 928 base pairs (bp) of the TBX1 cDNA coding region (TBX1 c.928G>A). The arrow indicates the design of primers for genotyping. [Figure 1C] A schematic of scarless editing is shown. Figure 1C shows the timeline for editing. The total estimated time is 1.5-2 months, depending on the growth rate of the cells. [Figure 2A] Figure 2A shows scarless single-base substitution in human ES cells by two-step HR. Figure 2B shows flow cytometry analysis of transient expression (day 2) and stable expression (day 6). [Figure 2B] Figure 2B shows the scarless single-base substitution in human ES cells after two-step HR. Figure 2B shows the analysis of tCD19 expression by FACS before and after positive selection by magnetic-activated cell sorting (MACS). CD19-positive cells were enriched from 2.1% to 67.8%. [Figure 2C] Figure 2C shows scarless single-base substitution in human ES cells after two-step HR. Fluorescence microscopy of single-cell colonies after MACS positive selection is shown. Some colonies were bright (yellow arrowheads) and others were dim (white arrowheads). The scale bar indicates 200 μm. [Figure 2D] Figure 2D shows scarless single base substitution in human ES cells by two-step HR. Figure 2D shows genotyping of single cell clones after MACS positive selection using the primers shown in Figure 1B. [Figure 2E] Figure 2B shows scarless single-base substitution in human ES cells by two-step HR. Figure 2C shows copy number analysis of the UbC promoter assessed by droplet digital PCR. Data are shown as the mean ± SD from three replicates. [Figure 2F]Figure 2F shows the scarless single-base substitution in human ES cells after two-step HR. Figure 2F shows the CD19-positive cell populations of clones 2 and 13 before and after the second HR and MACS negative selection. CD19-negative cells were enriched from 0.097% to 99.5% in clone 2 and from 0.044% to 94.7% in clone 13. [Figure 2G] Figure 2G shows scarless single-base substitution in human ES cells by two-step HR. Figure 2G shows the genotype of a single-cell clone determined by sequencing. [Figure 2H] Figure 2H shows scarless single-base substitution in human ES cells by two-step HR. Representative sequence data for mono- and bi-allelic edited clones (TBX1 reference sequence (SEQ ID NO: 32) and c.928G>A variant (SEQ ID NO: 33)) are shown. [Figure 3A] Figure 3A shows the scarless reporter gene integration 5' of the RUNX1 stop codon in human iPSCs by two-step editing. The schematic design of the donor DNA and guide RNA for RUNX1 editing is shown. The arrows indicate the primers for genotyping. The size of the PCR amplicon derived from each genotype is shown in the table on the right. [Figure 3B] Figure 3B shows the integration of a scarless reporter gene 5' of the RUNX1 stop codon in human iPSCs by two-step editing. Figure 3B shows PCR genotyping of bright clones after the first edit, followed by MACS positive selection and single-cell cloning. Clone 2 was used for subsequent steps. [Figure 3C] Figure 3C shows the integration of a scarless reporter gene 5' to the RUNX1 stop codon in human iPSCs by two-step editing. Figure 3C shows the genotyping of the second edited clone by PCR, followed by MACS negative selection and single-cell cloning. [Figure 3D]Figure 3D shows scarless reporter gene integration 5' of the RUNX1 stop codon in human iPSCs via two-step editing. A representative image of a RUNX1-mOrange iPSC cyst is shown. Cells were observed 13 days after induction of differentiation. Scale bar = 100 μm. [Figure 3E] Figure 3E shows the integration of a scarless reporter gene 5' to the RUNX1 stop codon in human iPSCs by two-step editing. Figure 3F shows a representative flow cytometry profile of cells 14 days after induction of differentiation. Population A, which is CD34+, CD45 intermediate double-positive, expresses mOrange at the highest intensity among all populations. [Figure 4A] Figures 4A-4D show the integration of biallelic scarless HLA-A*24 cDNA (1.1 kb) immediately before the stop codon of the B2M gene using a donor vector with an shRNA (shGFP) expression cassette in the ES H9 line. Schematic diagram of donor DNA design and genotyping by PCR. Figure 4A shows the first editing without the shGFP expression cassette. [Figure 4B] Figure 4B shows the integration of biallelic scarless HLA-A*24 cDNA (1.1 kb) immediately before the stop codon of the B2M gene using a donor vector with an shRNA (shGFP) expression cassette in the ES H9 line. Schematic diagram of donor DNA design and PCR genotyping. Figure 4B shows the first edit with the shGFP expression cassette. All clones shown expressed bright GFP fluorescence. [Figure 4C] Figure 4C shows the integration of biallelic scarless HLA-A*24 cDNA (1.1 kb) immediately before the stop codon of the B2M gene using a donor vector with an shRNA (shGFP) expression cassette in the ES H9 line. Schematic diagram of donor DNA design and PCR genotyping. Figure 4D shows the second editing using clone number 7. [Figure 4D]Figure 4D shows the integration of biallelic scarless HLA-A*24 cDNA (1.1 kb) immediately before the stop codon of the B2M gene in the ES H9 line using a donor vector with an shRNA (shGFP) expression cassette. Schematic diagram of donor DNA design and PCR genotyping. Figure 4D shows a representative flow cytometry profile of the cells. The class I HLA types of the ES H9 line are (A*02, A*03, B*35, B*44, C2*04, and Cw*07). The iAM9 line is an HLA-A*24-carrying hiPSC line used as a positive control for native HLA-A*24 expression. Cells were treated with 50 ng / mL INFγ for 3 days before FCM analysis. [Figure 5A] This paper examines the practical use of MMEJ-assisted scarless editing (KIM et al., Nature Commun.) in the gRNA-Cas9 system. Figure 5A shows the difference between standard and MMEJ-assisted scarless strategies in the prevention of gRNA-Cas9 re-cutting by marker integration. In standard marker integration, if the marker is integrated at the gRNA target sequence, the edited genome will no longer be cut by gRNA-Cas9. On the other hand, if the marker is integrated with microhomology sequences on both ends of the selection marker, the gRNA target will be retained unless the microhomology is sufficiently short or a mutation is introduced. [Figure 5B]This paper examines the practical use of MMEJ-assisted scarless editing (KIM et al., Nature Commun.) with the gRNA-Cas9 system. Figure 5B shows possible microhomology designs when the gRNA target does not overlap with the intended edit. Cas9 often tolerates a 2-nt 5' mismatch (distal from the PAM), and therefore, a 19-nt microhomology is a practical upper limit for the length to avoid re-cutting. Kim et al. demonstrated that the allelic frequency of MMEJ-based precise excision with 19-nt homology was approximately one-third that with 32-nt homology, meaning that the theoretical frequency of precise biallelic excision with 19-nt homology is approximately 10-fold lower than with 32-nt homology. Assuming that the frequency of precise biallelic excision with 32-nt homology was 10%, the frequency with 19-nt homology would be approximately 1%. [Figure 5C] This figure shows the actual use of MMEJ-assisted scarless editing (KIM et al., Nature Commun.) in the gRNA-Cas9 system. Figure 5C shows a possible microhomology design when the gRNA target overlaps with the intended edit (indicated by "Z"). Homologies longer than 32 nt can be generated, but a single-base mismatch typically does not prevent re-cutting by Cas9 (Fu et al., 2014 Nature Biotech). If a single-base mismatch in the intended edit effectively prevents re-cutting and the gRNA has high HR performance, standard one-step scarless editing is now available with high efficiency. [Figure 6A] The target sites of guide RNA for editing of TBX1 (Figure 6A) are shown. [Figure 6B] The target sites of guide RNA for editing of RUNX1 (Figure 6B) are shown. [Figure 6C] The target sites of guide RNAs for editing of GFI1 (Figure 6C) are shown. [Figure 6D] The target sites of guide RNAs for editing B2M (Figure 6D) are shown. [Figure 7A]Figure 7A shows the second round of HR in TBX1 monoallelic marker-inserted clone 14. Figure 7B shows the sequence data of the marker-free allele of clone 14. There is a single G insertion at the position of the CRISPR / Cas9 cleavage site in the first editing process. [Figure 7B] Figure 7B shows the second round of HR in TBX1 monoallelic marker insertion clone no. 14. Figure 7B shows the genotype of the clone after the second HR in clone no. 14 by sequence analysis. [Figure 7C] Figure 7C shows the second round of HR in TBX1 monoallelic marker insertion clone number 14. Figure 7C shows a pictorial illustration of HR involving donor DNA or chromosomal homologs. [Figure 8] Immunostaining of Oct4, Tra1-60, and Nanog in TBX1-edited ES cells is shown (scale bar = 400 μm). [Figure 9A] Figure 9A shows the scarless marker integration immediately before the stop codon of the GFI1 gene in the RUNX1-mOrange line. Figure 9B shows the schematic design of donor DNA and guide RNA for RUNX1 editing. Arrows indicate primers for genotyping. The size of the PCR amplicon from each genotype is shown in the table on the right. [Figure 9B] Figure 9B shows the scarless marker integration immediately before the stop codon of the GFI1 gene in the RUNX1-mOrange line. Figure 9B shows the PCR genotyping of the bright clones after the first editing step, followed by MACS positive selection and single-cell cloning. The copy number of the UbC promoter was assessed by ddPCR. Clone number 2 was used for subsequent procedures. [Figure 9C]Figure 9C shows the integration of the scarless marker immediately before the stop codon of the GFI1 gene in the RUNX1-mOrange line. Figure 9C shows the genotyping of the second edited clone by PCR, followed by MACS negative selection and single-cell cloning. Because the second donor vector contains the same sequence (EGFP) in addition to the homology arms, unwanted editing events were introduced (numbers 1-3 and 5), as illustrated in Figure 9D. [Figure 9D] See legend to Figure 9C. [Figure 10A] Figure 10A shows the detection of random integration in RUNX1 and GFI1 editing after the first round of editing. Figure 10A shows a schematic diagram of primer design for detecting random integration. When the selection marker is integrated into the genome at the sequences outside the left and right homology arms (derived from the plasmid backbone), respectively, primer sets 1 and 2 can detect random integration. The PCR results are shown for RUNX1 editing (Figure 10B) and for GFI1 editing (Figure 10C). [Figure 10B] See legend to Figure 10A. [Figure 10C] See legend to Figure 10A. [Figure 11A] Figure 11A shows canonical PCR-based copy number analysis of the UbC promoter in B2M editing. Figure 11A shows the GG>AA mutation at positions 314 and 315 upstream from the ATG start codon. [Figure 11B] Figure 11B shows canonical PCR-based copy number analysis of the UbC promoter in B2M editing. Figure 11B shows flow cytometry analysis demonstrating that the WT and GG>AA mutants of the UbC promoter expressed GFP at approximately the same intensity in K562 cells. [Figure 11C] Figure 11C shows a canonical PCR-based copy number analysis of the UbC promoter in B2M editing. Figure 11C shows a schematic diagram of primer design. [Figure 11D]Figure 11D shows canonical PCR-based copy number analysis of the UbC promoter in B2M editing. Figure 11D shows agarose gel electrophoresis of PCR products digested with EcoRI-HF enzymes. DNA was stained with MidoriGreen dye. [Figure 11E] Figure 11E shows a representative image of image analysis by Image J software. [Figure 11F] Figure 11F shows the quantification of fluorescence intensity (FI) in samples derived from standard PCR combined with EcoRI digestion. All samples derived from bright clones showed the same intensity between WT and GG>AA, indicating that two copies of the exogenous UbC promoter had been integrated. The FI of GG>AA in samples derived from dim clones was half that of WT, indicating that one copy of the exogenous UbC promoter had been integrated. [Figure 11G] Figure 11G shows a regular PCR-based copy number analysis of the UbC promoter in B2M editing. Figure 11G shows a copy number analysis of the UbC promoter using a ddPCR-based method. The results from ddPCR were consistent with the results of regular PCR-EcoRI digestion-based experiments. Data are shown as mean ± SD (N = 3). [Figure 12] This shows the range of edits possible in scarless editing in terms of the limit of the distance from cleavage to editing (Paquet et al. 2016) and the tolerance of Cas9 to mismatches (Fu et al., 2013). [Figure 13A]The donor DNA designs for various patterns of scarless editing are shown. We demonstrated that HR-based editing, when combined with appropriate marker selection, can be used in hPSCs to integrate and delete sequences longer than 3 kb with TBX1 editing, and to modify sequences from approximately 3 kb to 1 kb with B2M editing. Combining these techniques theoretically enables additional applications of scarless editing, such as: (Figure 13A) large insertions for optimizing differentiation in pluripotent stem cells or introducing reporter genes, which should be useful for drug screening; (Figure 13B) large deletions for creating disease models or knockout animals; (Figure 13C) large substitutions for creating knockin animals, including humanized models; and (Figure 13D) editing when there are no good target sites for site-specific nucleases (SSNs), including zinc finger nucleases, near the intended mutation site. Using this technique, nearly all limitations on editing sites should be eliminated. [Figure 13B] See legend to Figure 13A. [Figure 13C] See legend to Figure 13A. [Figure 13D] See legend to Figure 13A. [Figure 14] A schematic diagram of a two-step HDR-based scarless genome editing strategy using a combination of ngd-DSB, HDR-driven selectable marker removal, and negative selection is shown. This method allows for 100% enrichment of edited cells. [Figure 15] We demonstrate the application of a selection strategy using surface markers that are not essential for cell function. [Figure 16] 1 shows donor plasmid designs for suppression of GFP expression from randomly integrated or episomally present selectable markers. DETAILED DESCRIPTION OF THE INVENTION

[0033] Detailed Description The practice of the present invention will employ, unless otherwise indicated, conventional methods of genome editing, biochemistry, chemistry, immunology, molecular biology, and recombinant DNA techniques within the skill of the art, and such techniques are fully explained in the literature. For example, Targeted Genome Editing Using Site-Specific Nucleases: ZFNs, TALENs, and the CRISPR / Cas9 System (T. Yamamoto ed., Springer, 2015);Genome Editing: The Next Step in Gene Therapy (Advances in Experimental Medicine and Biology, T. Cathomen, M. Hirsch, and M. Porteus eds., Springer, 2016);Aachen Press Genome Editing (CreateSpace Independent Publishing Platform, 2015); Handbook of Experimental Immunology, Vols. I-IV (DM Weir and CC Blackwell eds., Blackwell Scientific Publications); AL Lehninger, Biochemistry (Worth Publishers, Inc., current addition); Sambrook, et al., Molecular Cloning: A Laboratory Manual (3 rd Edition, 2001); Methods In Enzymology (S. Colowick and N. Kaplan eds., Academic Press, Inc.).

[0034] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety.

[0035] I. Definition In describing the present invention, the following terms will be employed and are intended to be defined as indicated below.

[0036] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a cell" includes a mixture of two or more cells, and the like.

[0037] The term "about," particularly with respect to a given quantity, is meant to encompass a deviation of plus or minus 5 percent.

[0038] As used herein, "cell" refers to any type of cell isolated from a prokaryotic, eukaryotic, or archaeal organism, including bacteria, archaea, fungi, protozoa, plants, and animals, including cells derived from tissues, organs, and biopsies, as well as recombinant cells, cells derived from in vitro cultured cell lines, and cell fragments, cellular components, or organelles containing nucleic acids. The term also encompasses artificial cells, such as nanoparticles, liposomes, polymersomes, or microcapsules that encapsulate nucleic acids. Cells may include fixed cells or living cells. The genome editing methods described herein can be performed on samples containing, for example, single cells or populations of cells. The term also includes genetically modified cells.

[0039] The terms "polypeptide" and "protein" refer to polymers of amino acid residues and are not limited to a minimum length. Thus, peptides, oligopeptides, dimers, multimers, and the like, are encompassed within the definition. Both full-length proteins and fragments thereof are encompassed by the definition. The terms also include post-expression modifications of the polypeptide, such as glycosylation, acetylation, phosphorylation, hydroxylation, and the like. Furthermore, for purposes of the present invention, "polypeptide" refers to proteins containing modifications to the native sequence, such as deletions, additions, and substitutions, so long as the protein maintains the desired activity. These modifications may be deliberate, as by site-directed mutagenesis, or may be accidental, such as due to mutations of hosts producing the protein or errors due to PCR amplification.

[0040] As used herein, the term "Cas9" encompasses type II clustered regularly interspaced short palindromic repeats (CRISPR) system Cas9 endonuclease from any species, including biologically active fragments, variants, analogs, and derivatives thereof that retain Cas9 endonuclease activity (i.e., catalyze site-specific cleavage of DNA to generate double-stranded breaks). Cas9 endonuclease binds to and cleaves DNA at sites containing sequences complementary to its bound guide RNA (gRNA).

[0041] Cas9 polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide, or peptide refers to a molecule derived from any source.The molecule does not necessarily have to be derived from a living organism, but can be synthetic or recombinantly produced.Cas9 sequences from multiple bacterial species are well known in the art and are listed in the National Center for Biotechnology Information (NCBI) database. For example, Streptococcus pyogenes (WP_002989955, WP_038434062, WP_011528583); Campylobacter jejuni (WP_022552435, YP_002344900); Campylobacter coli (WP_060786116); Campylobacter fetus (WP_059434633); Corynebacterium ulcerans (NC_015683, NC_017317); Corynebacterium diphtheriae (Corynebacterium diphtheria (NC_016782, NC_016786); Enterococcus faecalis (WP_033919308); Spiroplasma syrphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanense (NC_021846); Streptococcus iniae (NC_021314); Belliella baltica (NC_018010); Psychroflexus torquis I (NC_018721);Streptococcus thermophilus (YP_820832), Streptococcus mutans (WP_061046374, WP_024786433); Listeria innocua (NP_472073); Listeria monocytogenes (WP_061665472); Legionella pneumophila (WP_062726656); Staphylococcus aureus (WP_001573634); Francisella tularensis See NCBI accessions for Cas9 from Enterococcus tularensis (WP_032729892, WP_014548420), Enterococcus faecalis (WP_033919308); Lactobacillus rhamnosus (WP_048482595, WP_032965177); and Neisseria meningitidis (WP_061704949, YP_002342100);All of these sequences (registered as of the filing date of this application) are incorporated herein by reference. Any of these sequences, or variants thereof, including sequences having at least about 70-100% sequence identity thereto, including any percent identity thereto within this range, for example, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto, can be used for scarless genome editing as described herein, wherein the variant retains biological activity, such as Cas9 site-specific endonuclease activity. For sequence comparisons and discussions of genetic diversity and phylogenetic analysis of Cas9, see also Fonfara et al. (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al. (2015) J. Bacteriol. 198(5):797-807, Shmakov et al. (2015) Mol. Cell. 60(3):385-397, and Chylinski et al. (2014) Nucleic Acids Res. 42(10):6091-6105.

[0042] By "derivative" is intended any suitable modification of a naturally occurring polypeptide of interest, of a fragment of a naturally occurring polypeptide, or of their respective analogs, such as glycosylation, phosphorylation, polymer conjugation (e.g., with polyethylene glycol), or other addition of foreign moieties, so long as the desired biological activity of the naturally occurring polypeptide is retained. Methods for making polypeptide fragments, analogs, and derivatives are generally available in the art.

[0043] By "fragment" is intended a molecule consisting of only a portion of the intact full-length sequence and structure. Fragments can include C-terminal deletions, N-terminal deletions, and / or internal deletions of the polypeptide. Active fragments of a particular protein or polypeptide generally contain at least about 5-10 contiguous amino acid residues of the full-length molecule, preferably at least about 15-25 contiguous amino acid residues of the full-length molecule, and most preferably at least about 20-50 or more contiguous amino acid residues of the full-length molecule, or any integer between 5 amino acids and the full-length sequence, provided that the fragment retains biological activity, such as Cas9 site-specific endonuclease activity.

[0044] "Substantially purified" generally refers to the isolation of a substance (compound, polynucleotide, nucleic acid, protein, polypeptide, polypeptide composition) such that it constitutes the majority percent of the sample to which it belongs. Typically, in a sample, the substantially purified component constitutes 50%, preferably 80%-85%, and more preferably 90%-95% of the sample. Techniques for purifying polynucleotides and polypeptides of interest are well known in the art and include, for example, ion exchange chromatography, affinity chromatography, and sedimentation according to density.

[0045] By "isolated," when referring to a polypeptide, it is meant that the indicated molecule is separate and distinct from the whole organism with which it is naturally found, or exists in the substantial absence of other biological macromolecules of the same type. The term "isolated," with reference to a polynucleotide, refers to a nucleic acid molecule that is devoid of all or part of sequences that are normally associated with it in nature; or a sequence that appears to be naturally occurring but has heterologous sequences associated with it; or a molecule that has been dissociated from a chromosome.

[0046] As used herein, the phrase "heterogeneous population of cells" refers to a mixture of at least two types of cells, one of which is a cell of interest (e.g., has a genomic modification of interest). The heterogeneous population of cells may be derived from any organism.

[0047] As used herein, in the context of selecting a cell or population of cells that have a genomic modification of interest, the terms "isolating" and "isolation" refer to separating a cell or population of cells that have the genomic modification of interest from a heterogeneous population of cells, e.g., by positive or negative selection.

[0048] The term "selectable marker" refers to a marker that can be used to enrich a population of cells from a heterogeneous population of cells, either by positive selection (selecting for cells that express the marker) or by negative selection (excluding cells that express the marker).

[0049] The term "binding agent" refers to any agent that specifically binds to a selectable marker. Examples of binding agents include, but are not limited to, antibodies, antibody fragments, antibody mimetics, and aptamers that specifically bind to a selectable marker.

[0050] The phrase "specifically (or selectively) binds" when referring to a binding agent refers to a binding reaction that determines the presence of cells bearing a particular selection marker in a heterogeneous population of cells and other biologics. Thus, under specified conditions, the binding agent binds to a particular selection marker on cells at least twice the background level and does not substantially bind in significant amounts to other cells present in the sample that do not bear the selection marker. Specific binding to cells bearing a selection marker under such conditions may require a binding agent (e.g., an antibody, antibody mimetic, or aptamer) that has been selected for its specificity for the particular selection marker. Typically, specific or selective binding between a binding agent and a selection marker is at least twice the background signal or noise, and more typically greater than 10-100 times the background signal or noise.

[0051] The term "antibody" refers to polyclonal and monoclonal antibody preparations, as well as preparations including hybrid, altered, chimeric, and humanized antibodies, as well as hybrid (chimeric) antibody molecules (see, e.g., Winter et al. (1991) Nature 349:293-299; and U.S. Pat. No. 4,816,567); F(ab')2 and F(ab)2 fragments; F v molecules (non-covalent heterodimers, see, e.g., Inbar et al. (1972) Proc Natl Acad Sci USA 69:2659-2662; and Ehrlich et al. (1980) Biochem 19:4091-4096); single-chain Fv molecules (sFv) (see, e.g., Huston et al. (1988) Proc Natl Acad Sci USA 85:5879-5883); nanobodies or single-domain antibodies (sdAb) (see, e.g., Wang et al. (2016) Int J Nanomedicine 11:3287-3303, Vincke et al. (2012) Methods Mol Biol 911:15-26); dimeric and trimeric antibody fragment constructs; minibodies (see, e.g., Pack et al. (1992) Biochem 31:1579-1584; Cumber et al. (1992) J Immunology 149B:120-126); humanized antibody molecules (see, e.g., Riechmann et al. (1988) Nature 332:323-327; Verhoeyan et al. (1988) Science 239:1534-1536; and UK Patent Publication No. GB ​​2,276,169, published September 21, 1994); and any functional fragments derived from such molecules, wherein such fragments retain the specific binding properties of the parent antibody molecule.

[0052] As used herein, "solid support" refers to a solid surface such as magnetic beads, latex beads, microtiter plate wells, glass plates, nylon, agarose, acrylamide, and the like.

[0053] The terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" are used herein to include polymeric forms of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. The terms refer only to the primary structure of the molecule. Thus, the terms include triple-, double-, and single-stranded DNA, as well as triple-, double-, and single-stranded RNA. They also include modified, for example, by methylation and / or capping, as well as unmodified forms of polynucleotides. More specifically, the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D-ribose), any other type of polynucleotide that is an N- or C-glycoside of a purine or pyrimidine base, as well as other polymers containing non-nucleotide backbones, such as polyamide (e.g., peptide nucleic acid (PNA)) and polymorpholino (commercially available as Neugene from Anti-Virals, Inc., Corvallis, Oreg.) polymers, and other synthetic sequence-specific nucleic acid polymers, provided that the polymer contains the nucleic acid bases in a configuration that allows for base pairing and base stacking as found in DNA and RNA. There is no intended distinction in length between the terms "polynucleotide," "oligonucleotide," "nucleic acid," and "nucleic acid molecule," and these terms are used interchangeably.Thus, these terms include, for example, 3'-deoxy-2',5'-DNA, oligodeoxyribonucleotide N3' P5' phosphoramidates, 2'-O-alkyl-substituted RNA, double- and single-stranded DNA, and double- and single-stranded RNA, microRNA, DNA:RNA hybrids, and hybrids between PNA and DNA or RNA, and also include known types of modifications, such as labels known in the art, methylation, "caps," substitutions with one or more analogs of naturally occurring nucleotides (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, C5-propynylcytidine, C5-propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine). The term also includes internucleotide modifications, such as those with uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), those with negatively charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), and those with positively charged linkages (e.g., aminoalkyl phosphoramidates, aminoalkyl phosphotriesters), those containing pendant moieties, such as proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those containing intercalators (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, metal oxides, etc.), those containing alkylating agents, those with modified linkages (e.g., α-anomeric nucleic acids, etc.), and unmodified forms of polynucleotides or oligonucleotides. The term also includes locked nucleic acids (e.g., including ribonucleotides with a methylene bridge between the 2'-oxygen atom and the 4'-carbon atom).For example, Kurreck et al. (2002) Nucleic Acids Res. 30: 1911-1918;Elayadi et al. (2001) Curr. Opinion Invest. Drugs 2: 558-561;Orum et al. (2001) Curr. Opinion Mol. Ther. 3: 239-243;Koshkin et al. (1998) Tetrahedron 54: 3607-3630; Obika et al. (1998) Tetrahedron Lett. 39: 5401-5404.

[0054] The terms "hybridize" and "hybridization" refer to the formation of a complex between nucleotide sequences that are sufficiently complementary to form a complex via Watson-Crick base pairing.

[0055] The term "homologous region" refers to a region of nucleic acid that has homology with another nucleic acid region.Therefore, whether a "homologous region" exists in a nucleic acid molecule is determined with respect to another nucleic acid region in the same molecule or different molecules.In addition, since nucleic acids are often double-stranded, the term "homologous region" used herein refers to the ability of nucleic acid molecules to hybridize with each other.For example, a single-stranded nucleic acid molecule can have two homologous regions that can hybridize with each other.Therefore, the term "homologous region" includes nucleic acid segments that have complementary sequences. The homologous region may vary in length, but is typically 4 to 500 nucleotides (e.g., about 4 to about 40, about 40 to about 80, about 80 to about 120, about 120 to about 160, about 160 to about 200, about 200 to about 240, about 240 to about 280, about 280 to about 320, about 320 to about 360, about 360 to about 400, about 400 to about 440, etc.).

[0056] As used herein, the terms "complementary" or "complementarity" refer to polynucleotides that can base pair with each other. Base pairs are typically formed by hydrogen bonding between nucleotide units in an antiparallel orientation between polynucleotide strands. Complementary polynucleotide strands can base pair in a Watson-Crick manner (e.g., AT, AU, CG) or any other manner that allows for the formation of a duplex. As those skilled in the art will recognize, when using RNA as opposed to DNA, uracil (U), rather than thymine (T), is the base considered complementary to adenosine. However, when uracil is referred to in the context of the present invention, its ability to substitute for thymine is implied unless otherwise stated. "Complementarity" can exist between two RNA strands, two DNA strands, or between an RNA strand and a DNA strand. It is generally understood that two or more polynucleotides can be "complementary" and can form a duplex despite having less than perfect or 100% complementarity. Two sequences are "fully complementary" or "100% complementary" if at least a contiguous portion of each polynucleotide sequence, including the region of complementarity, perfectly base-pairs with the other polynucleotide without any mismatches or interruptions within such region. Two or more sequences are considered "fully complementary" or "100% complementary" even if one or both polynucleotides contain additional non-complementary sequences, as long as the contiguous regions of complementarity within each polynucleotide can perfectly hybridize with each other. "Less than perfect" complementarity refers to a situation in which fewer than all contiguous nucleotides within such a region of complementarity can base-pair with each other. Determining the percentage of complementarity between two polynucleotide sequences is a matter of ordinary skill in the art. For purposes of Cas9 targeting, a gRNA may contain a sequence "complementary" to a target sequence (e.g., a major or minor allele) that is capable of sufficient base-pairing to form a duplex (i.e., the gRNA hybridizes with the target sequence).In addition, the gRNA may contain a sequence complementary to the PAM sequence, and the gRNA also hybridizes to the PAM sequence in the target DNA.

[0057] A "target site" or "target sequence" is a nucleic acid sequence recognized (i.e., sufficiently complementary for hybridization) by a homology arm of a guide RNA (gRNA) or donor polynucleotide. Target sites may be allele-specific (e.g., major or minor alleles).

[0058] The term "donor polynucleotide" refers to a polynucleotide that provides the sequence of the intended edit to be integrated into the genome at the target locus by HDR.

[0059] By "homology arm" is meant the portion of the donor polynucleotide responsible for targeting the donor polynucleotide to the genomic sequence to be edited in a cell. The donor polynucleotide typically comprises a 5' homology arm that hybridizes to the 5' genomic target sequence and a 3' homology arm that hybridizes to the 3' genomic target sequence, adjacent to the nucleotide sequence containing the intended edit in the genomic DNA. The homology arms are referred to herein as 5' and 3' (i.e., upstream and downstream) homology arms, which refer to the relative position of the homology arms to the nucleotide sequence containing the intended edit in the donor polynucleotide. The 5' and 3' homology arms hybridize to regions within the target locus in the genomic DNA to be modified, referred to herein as "5' target sequence" and "3' target sequence," respectively. The nucleotide sequence containing the intended edit is integrated into genomic DNA by HDR at the genomic target locus recognized by the 5' and 3' homology arms (i.e., sufficiently complementary for hybridization).

[0060] "Administering" a nucleic acid, such as a donor polynucleotide, guide RNA, or Cas9 expression system, to a cell includes transduction, transfection, electroporation, translocation, fusion, phagocytosis, shooting, or ballistic methods, i.e., any means by which the nucleic acid can be transported across the cell membrane.

[0061] "Selectively binds" in reference to guide RNA means that the guide RNA preferentially binds to the target sequence of interest or binds to the target sequence with higher affinity than other genomic sequences. For example, the gRNA binds to a substantially complementary sequence and does not bind to unrelated sequences. A gRNA that "selectively binds" to a specific allele, such as a specific mutant allele (e.g., an allele containing a substitution, insertion, or deletion), refers to a gRNA that preferentially binds to a specific target allele but binds to a lesser extent to wild-type alleles or other sequences. A gRNA that selectively binds to a specific target DNA sequence preferentially directs the binding of Cas9 to a substantially complementary sequence at the target site, and not to unrelated sequences.

[0062] As used herein, the terms "label" and "detectable label" refer to a molecule capable of being detected, including, but not limited to, a radioisotope, a fluorescer, a chemiluminescer, a chromophore, an enzyme, an enzyme substrate, an enzyme cofactor, an enzyme inhibitor, a semiconductor nanoparticle, a dye, a metal ion, a metal sol, a ligand (e.g., biotin, streptavidin, or a hapten), etc. The term "fluorescer" refers to a substance or portion thereof capable of exhibiting fluorescence in the detectable range. Specific examples of labels that can be used in the practice of the present invention include SYBR green, SYBR gold; CAL Fluor dyes such as CAL Fluor Gold 540, CAL Fluor Orange 560, CAL Fluor Red 590, CAL Fluor Red 610, and CAL Fluor Red 635; Quasar dyes such as Quasar 570, Quasar 670, and Quasar 705; Alexa Fluors such as Alexa Fluor 350, Alexa Fluor 488, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 594, Alexa Fluor 647, and Alexa Fluor 784; Cy cyanine dyes such as Cy3.5, Cy5, Cy5.5, and Cy7; fluorescein, 2',4',5',7'-tetrachloro-4-7-dichlorofluorescein (TET), carboxyfluorescein (FAM), 6-carboxy-4',5'-dichloro-2',7'-dimethoxyfluorescein (JOE), hexachlorofluorescein (HEX), rhodamine, carboxy-X-rhodamine (ROX), tetramethylrhodamine (TAMRA), FITC, dansyl, umbelliferone, dimethyl acridinium ester (DMAE), Texas red, luminol, NADPH, horseradish peroxidase (HRP), and α-β-galactosidase.

[0063] " Homology " refers to the percent identity between two polynucleotides or two polypeptide molecules. Two nucleic acid or two polypeptide sequences are "substantially homologous" to each other if the sequences exhibit at least about 50% sequence identity over the defined length of the molecule, preferably at least about 75% sequence identity, more preferably at least about 80%, 85% sequence identity, more preferably at least about 90% sequence identity, and most preferably at least about 95%, 98% sequence identity. As used herein, "substantially homologous" also refers to a sequence that shows complete identity to a specified sequence.

[0064] Generally, "identity" refers to the exact correspondence of nucleotide to nucleotide or amino acid to amino acid of two polynucleotide or polypeptide sequences, respectively. Percent identity can be determined by direct comparison of the sequence information between two molecules by aligning the sequences, counting the exact number of matches between the two aligned sequences, dividing by the length of the shorter sequence, and multiplying the result by 100. Easily available computer programs, such as ALIGN, Dayhoff, MO, in Atlas of Protein Sequence and Structure, ed., 5 Suppl. 3:353-358, National Biomedical Research Foundation, Washington, DC, which adapts the local homology algorithm of Smith and Waterman Advances in Appl. Math. 2:482-489, 1981, to peptide analysis, can be used to assist in the analysis. The program for determining nucleotide sequence identity is available in Wisconsin Sequence Analysis Package, Version 8 (available from Genetics Computer Group, Madison, WI), for example, BESTFIT, FASTA and GAP program, which also rely on Smith and Waterman algorithm.These programs are easily used with the default parameters recommended by the manufacturer and described in the Wisconsin Sequence Analysis Package mentioned above.For example, the percent identity of a specific nucleotide sequence to a reference sequence can be determined using the Smith and Waterman homology algorithm with default scoring table and gap penalty of 6 nucleotide positions.

[0065] Another method for establishing percent identity in the context of the present invention is to use the programs in the MPSRCH package, copyrighted by the University of Edinburgh, developed by John F. Collins and Shane S. Sturrok, and distributed by IntelliGenetics, Inc. (Mountain View, CA). From this package suite, the Smith Waterman algorithm can be used, and default parameters for the scoring table are used (e.g., a gap open penalty of 12, a gap extension penalty of 1, and a gap of 6). From the resulting data, the "match" value reflects "sequence identity." Other suitable programs for calculating percent identity or similarity between sequences are generally known in the art; for example, another alignment program is BLAST, used with default parameters. For example, BLASTN and BLASTP can be used with the following default parameters: genetic code = standard; filter = none; strand = both; cutoff = 60; expectation = 10; matrix = BLOSUM62; description = 50 sequences; sort by HIGH SCORE; databases = non-redundant GenBank + EMBL + DDBJ + PDB + GenBank CDS translations + Swiss protein + Spupdate + PIR. Details of these programs are readily available.

[0066] Alternatively, homology can be determined by polynucleotide hybridization under conditions that form stable duplexes between homologous regions, followed by digestion with single-strand-specific nucleases and determining the size of the digested fragments.Substantially homologous DNA sequences can be identified, for example, in Southern hybridization experiments under stringent conditions as defined for that particular system.Defining appropriate hybridization conditions is within the skill of the art.See, for example, Sambrook et al., supra; DNA Cloning, supra; Nucleic Acid Hybridization, supra.

[0067] "Recombinant," as used herein to describe a nucleic acid molecule, refers to a polynucleotide of genomic, cDNA, viral, semisynthetic, or synthetic origin that, by virtue of its origin or manipulation, is not associated with all or a portion of the polynucleotide with which it is naturally associated. The term "recombinant," as used with respect to a protein or polypeptide, refers to a polypeptide produced by expression of a recombinant polynucleotide. Generally, a gene of interest is cloned and then expressed in a transformed organism, as described further below. The host organism expresses the foreign gene under expression conditions to produce the protein.

[0068] The term "transformation" refers to the insertion of an exogenous polynucleotide into a host cell, regardless of the method used for the insertion, including, for example, direct uptake, transduction, or f-mating. The exogenous polynucleotide may be maintained as a non-integrated vector, for example, a plasmid, or alternatively, may be integrated into the host genome.

[0069] "Recombinant host cells," "host cells," "cells," "cell lines," "cell cultures," and other such terms referring to microbial or higher eukaryotic cell lines cultured as unicellular entities, refer to cells that can be or have been used as recipients for recombinant vectors or other transferred DNA, and include the original progeny of the original cell that has been transfected.

[0070] A "coding sequence," or a sequence "encoding" a selected polypeptide, is a nucleic acid molecule that is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences (or "control elements"). The boundaries of the coding sequence can be determined by a start codon at the 5' (amino) terminus and a translation stop codon at the 3' (carboxy) terminus. Coding sequences can include, but are not limited to, cDNA from viral, prokaryotic, or eukaryotic mRNA, genomic DNA sequences from viral or prokaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence can also be located 3' to the coding sequence.

[0071] Typical "control elements" include, but are not limited to, transcriptional promoters, transcriptional enhancer elements, transcription termination signals, polyadenylation sequences (located 3' to the translation stop codon), sequences for optimizing translation initiation (located 5' to the coding sequence), and translation termination sequences.

[0072] "Operably linked" refers to an arrangement of elements in which the components so described are configured to perform their normal function. Thus, a given promoter operably linked to a coding sequence is capable of effecting expression of the coding sequence when the proper enzymes are present. A promoter need not be contiguous with the coding sequence, so long as it functions to direct the expression of that sequence. Thus, for example, intervening untranslated but transcribed sequences can be present between the promoter sequence and the coding sequence, and the promoter sequence can still be considered "operably linked" to the coding sequence.

[0073] "Expression cassette" or "expression construct" refers to an assembly capable of directing the expression of a sequence or gene of interest. An expression cassette generally includes control elements as described above, such as a promoter operably linked to the sequence or gene of interest (to direct its transcription), and often also includes a polyadenylation sequence. Within certain embodiments of the present invention, the expression cassettes described herein may be contained within a donor polynucleotide, a plasmid, or a viral vector construct. In addition to the components of the expression cassette, the construct may also include one or more selectable markers, a signal that allows the construct to exist as single-stranded DNA (e.g., an M13 origin of replication), at least one multiple cloning site, and a "mammalian" origin of replication (e.g., an SV40 or adenovirus origin of replication).

[0074] " Purified polynucleotide " refers to a polynucleotide of interest or a fragment thereof that is essentially free of the proteins with which the polynucleotide is naturally associated, for example, less than about 50%, preferably less than about 70%, and more preferably less than about at least 90%. Techniques for purifying a polynucleotide of interest are well known in the art, and include, for example, disrupting cells containing the polynucleotide with a chaotropic agent, and separating polynucleotides and proteins by ion exchange chromatography, affinity chromatography, and sedimentation according to density.

[0075] The term "transfection" is used to refer to the uptake of foreign DNA by a cell. A cell is "transfected" when foreign DNA is introduced inside the cell membrane. Several transfection techniques are generally known in the art. See, for example, Graham et al. (1973) Virology, 52:456; Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3rd edition, Cold Spring Harbor Laboratories, New York; Davis et al. (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw-Hill; and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more foreign DNA moieties into suitable host cells. The term refers to both stable and transient uptake of genetic material, including uptake of peptide- or antibody-conjugated DNA.

[0076] A "vector" can transfer a nucleic acid sequence into a target cell (e.g., a viral vector, a non-viral vector, a particulate carrier, and a liposome). Typically, "vector construct," "expression vector," and "gene transfer vector" refer to any nucleic acid construct that can direct the expression of a nucleic acid of interest and transfer a nucleic acid sequence into a target cell. Thus, the terms include cloning and expression vehicles, as well as plasmids and viral vectors.

[0077] The terms "variant," "analog," and "mutein" refer to biologically active derivatives of a reference molecule that retain a desired activity, such as site-specific Cas9 endonuclease activity. Generally, the terms "variant" and "analog" refer to compounds that have a native polypeptide sequence and structure with one or more amino acid additions, (generally conservative in nature) substitutions, and / or deletions compared to the native molecule, and that are "substantially homologous" to the reference molecule, as defined below, so long as the modifications do not destroy the biological activity. Generally, the amino acid sequence of such an analog has a high degree of sequence homology to the reference sequence when the two sequences are aligned, e.g., greater than 50%, generally greater than 60%-70%, and even more particularly, 80%-85% or more, e.g., at least 90%-95% or more. Often, analogs contain the same number of amino acids but include substitutions, as described herein. The term "mutein" further includes compounds containing only amino and / or imino molecules, polypeptides containing one or more analogs of an amino acid (including, e.g., unnatural amino acids, etc.), polypeptides with substituted linkages, and polypeptides with one or more amino acid-like molecules, including, but not limited to, other modifications known in the art, both natural and unnatural (e.g., synthetic), cyclized, branched molecules, etc. The term also includes molecules containing one or more N-substituted glycine residues ("peptoids") and other synthetic amino acids or peptides. (See, e.g., U.S. Pat. Nos. 5,831,005; 5,877,278; and 5,977,301; Nguyen et al., Chem. Biol. (2000) 7:463-473; and Simon et al., Proc. Natl. Acad. Sci. USA (1992) 89:9367-9371 for a description of peptoids.) Methods for making polypeptide analogs and muteins are known in the art and are further described below.

[0078] As explained above, analogs generally include conservative substitutions, i.e., substitutions that occur within a family of amino acids related in their side chains. Specifically, amino acids are generally divided into four families: (1) acidic—aspartic acid and glutamic acid; (2) basic—lysine, arginine, histidine; (3) nonpolar—alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan; and (4) uncharged polar—glycine, asparagine, glutamine, cysteine, serine, threonine, and tyrosine. Phenylalanine, tryptophan, and tyrosine are sometimes classified as aromatic amino acids. For example, it is reasonably predictable that the simple replacement of leucine with isoleucine or valine, aspartic acid with glutamic acid, or threonine with serine, or similar conservative replacement of amino acids with structurally related amino acids, will not have a significant effect on biological activity. For example, a polypeptide of interest may include up to about 5-10 conservative or non-conservative amino acid substitutions, or even up to about 15-25 conservative or non-conservative amino acid substitutions, or any integer between 5 and 25, so long as the desired function of the molecule remains intact. One of skill in the art can readily determine regions of a molecule of interest that may be tolerant to alteration by reference to Hopp / Woods and Kyte-Doolittle plots, which are well known in the art.

[0079] "Gene transfer" or "gene delivery" refers to a method or system for reliably inserting DNA or RNA of interest into a host cell. Such methods can result in the transient expression of transferred non-integrated DNA, the extrachromosomal replication and expression of transferred replicons (e.g., episomes), or the integration of transferred genetic material into the genomic DNA of the host cell. Gene delivery expression vectors include, but are not limited to, bacterial plasmid vectors, viral vectors, non-viral vectors, adenoviruses, retroviruses, alphaviruses, poxviruses, and vaccinia virus-derived vectors.

[0080] The term "derived from" is used herein to identify the original source of a molecule, but is not meant to limit the manner in which the molecule is made, which can be, for example, by chemical synthesis or recombinant means.

[0081] A polynucleotide "derived from" a specified sequence refers to a polynucleotide sequence that contains a contiguous sequence of approximately at least about 6 nucleotides, preferably at least about 8 nucleotides, more preferably at least about 10-12 nucleotides, and even more preferably at least about 15-20 nucleotides, that corresponds to, i.e., is identical to, or complementary to, a region of the specified nucleotide sequence. A derived polynucleotide is not necessarily physically derived from the nucleotide sequence of interest, but may be produced in any manner, including, but not limited to, chemical synthesis, replication, reverse transcription, or transcription, based on information provided by the sequence of bases in the region from which the polynucleotide is derived. It may therefore represent either the sense or antisense orientation of the original polynucleotide.

[0082] The term "subject" includes both vertebrates and invertebrates, including, but not limited to, humans and non-human mammals, e.g., non-human primates, including chimpanzees and other ape and monkey species; laboratory animals, such as mice, rats, rabbits, hamsters, guinea pigs, and chinchillas; domesticated animals, such as dogs and cats; mammals, including livestock, such as sheep, goats, pigs, horses, and cows; and domesticated, wild, and game birds, including chickens, turkeys, and other gallinaceous birds, ducks, geese, etc. In some cases, the methods of the invention are used in laboratory animals, in veterinary applications, and in the development of animal models for disease, including, but not limited to, rodents, including mice, rats, and hamsters; primates, and transgenic animals.

[0083] II. MODES FOR CARRYING OUT THE INVENTION Before describing the present invention in detail, it is to be understood that this invention is not limited to particular formulations or process parameters, which may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only, and is not intended to be limiting.

[0084] Although a number of methods and materials similar or equivalent to those described herein can be used in the practice of the present invention, the preferred materials and methods are described herein.

[0085] The present invention is based on the development of a scarless genome editing method that uses two-step HDR repair to genetically modify cells and remove unwanted sequences. This method can be used in genome editing to introduce mutations, deletions, or insertions at any position in the genome without leaving silent mutations, selectable marker sequences, or other additional unwanted sequences in the genome. Furthermore, there is virtually no limit to the selection of editing sites. This method can be used not only for small edits but also for large edits of several thousand bases. The major advantages of this genome editing method include its extremely high efficiency and flexibility in various types of cells. The inventors have successfully demonstrated the use of this method for genome editing of human ES / iPS cells (Example 1). The high efficiency of this method makes it possible to identify desired clones, typically without the need to screen more than 10 colonies. In addition, this method can be easily adapted to enable multiple gene edits (see Example 1).

[0086] To advance understanding of the present invention, a more detailed discussion of scarless genome editing using two-step HDR repair is provided below.

[0087] A. Scarless genome editing by two-step HDR repair As described above, the method of the present invention uses two HDR repair steps to achieve scarless genome editing using Cas9 nuclease.In the first HDR step, a donor polynucleotide containing the sequence containing the intended genome editing is used to modify the target genome sequence in a cell, and the donor polynucleotide is integrated into the genome at the target locus by site-specific homologous recombination.The donor polynucleotide can be used to, for example, introduce the intended editing into the genome for the purpose of repairing, modifying, replacing, deleting, attenuating or inactivating the target gene.By including a selection marker expression cassette in the donor polynucleotide, the genetically modified cells generated by the first HDR step can be isolated by positive selection.In the second HDR step, a second donor polynucleotide is used to delete the selection marker expression cassette from the genetically modified cells generated in the first HDR step. The second donor polynucleotide contains a sequence complementary to the target genome sequence that is modified by the integration of the first donor polynucleotide into the genomic DNA in the first HDR step, except that the selection marker expression cassette is deleted. Therefore, integration of the second donor polynucleotide at the target genome locus removes the expression cassette. Genetically modified cells in which the selection marker cassette has been removed can be isolated by negative selection.

[0088] In the donor polynucleotide used in the first HDR step, the sequence containing intended editing is flanked by a pair of homology arms, which are responsible for targeting the donor polynucleotide to the target locus that should be edited in cells.Donor polynucleotide typically comprises a 5' homology arm that hybridizes with the 5' genome target sequence and a 3' homology arm that hybridizes with the 3' genome target sequence.Homology arms are herein referred to as 5' and 3' (i.e., upstream and downstream) homology arms, and this refers to the relative position of the homology arm to the nucleotide sequence that contains intended editing in donor polynucleotide.5' and 3' homology arms hybridize with the region in the target locus in genomic DNA that should be modified, which are herein referred to as "5' target sequence" and "3' target sequence", respectively. The donor polynucleotide used in the second HDR step to remove the selectable marker expression cassette similarly has 5' and 3' homology arms flanked by sequences complementary to the integrated donor polynucleotide from the first HDR step.

[0089] The homology arms must be sufficiently complementary to the target sequence to mediate homologous recombination between the donor polynucleotide and genomic DNA at the target locus. For example, the homology arms may comprise a nucleotide sequence having at least about 80-100% sequence identity to the corresponding genomic target sequence, including any percent identity within this range, such as at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity thereto, such that the nucleotide sequence containing the intended edit is integrated into the genomic DNA by HDR at the genomic target locus recognized by the 5' and 3' homology arms (i.e., sufficiently complementary for hybridization).

[0090] In certain embodiments, the corresponding homologous nucleotide sequences in the genome target sequence (i.e., the "5' target sequence" and the "3' target sequence") flank the specific site for cleavage and / or the specific site for introducing the intended edit. The distance between the specific cleavage site and the homologous nucleotide sequence (e.g., each homology arm) can be several hundred nucleotides. In some embodiments, the distance between the homology arm and the cleavage site is 200 nucleotides or less (e.g., 0, 10, 20, 30, 50, 75, 100, 125, 150, 175, and 200 nucleotides). In most cases, a shorter distance can result in a higher gene targeting rate. In a preferred embodiment, the donor polynucleotide is substantially identical to the target genome sequence over its entire length, except for the sequence change to be introduced into the portion of the genome that encompasses both the specific cleavage site and the portion of the genome target sequence to be modified.

[0091] The homology arms can be any length, e.g., 10 or more nucleotides, 50 or more nucleotides, 100 or more nucleotides, 250 or more nucleotides, 300 or more nucleotides, 350 or more nucleotides, 400 or more nucleotides, 450 or more nucleotides, 500 or more nucleotides, 1000 or more nucleotides (1 kb), 5000 or more nucleotides (5 kb), 10000 or more nucleotides (10 kb), etc. In some examples, the 5' and 3' homology arms are substantially equal in length to each other, e.g., one may be no more than 30% shorter than the other homology arm, no more than 20% shorter than the other homology arm, no more than 10% shorter than the other homology arm, no more than 5% shorter than the other homology arm, no more than 2% shorter than the other homology arm, or only a few nucleotides shorter than the other homology arm. In other examples, the 5' and 3' homology arms may differ substantially in length from one another, for example, one may be 40% or more shorter, 50% or more shorter, or sometimes 60% or more shorter, 70% or more shorter, 80% or more shorter, 90% or more shorter, or 95% or more shorter than the other homology arm.

[0092] In certain embodiments, the cells containing the modified genome generated in the first HDR step are identified in vitro or in vivo by including a selection marker expression cassette in the donor polynucleotide.The selection marker confers identifiable changes to cells, which allows for the positive selection of genetically modified cells with the donor polynucleotide integrated into the genome.For example, fluorescent or bioluminescent markers (e.g., green fluorescent protein (GFP), enhanced green fluorescent protein (EGFP), Dronpa, mCherry, mOrange, mPlum, Venus, YPet, phycoerythrin, or luciferase), cell surface markers (e.g., truncated CD8, NGFR, or CD19), reporter gene (e.g., GFP, dsRed, GUS, lacZ, CAT), or drug selection markers such as genes that confer resistance to neomycin, puromycin, hygromycin, DHFR, GPT, zeocin, or histidinol can be used to identify cells. Alternatively, enzymes such as herpes simplex virus thymidine kinase (tk) or chloramphenicol acetyltransferase (CAT) can be used.Any selectable marker can be used as long as it can be expressed after the integration of donor polynucleotide in the first HDR step, and can identify genetically modified cells.More examples of selectable markers are well known to those skilled in the art.

[0093] In certain embodiments, the selectable marker expression cassette encodes two or more selectable markers.Selectable markers may be used in combination, for example, a cell surface marker may be used with a fluorescent marker, or a drug resistance gene may be used with a suicide gene.In certain embodiments, the donor polynucleotide is provided by a multicistronic vector, which allows the expression of multiple selectable markers in combination.The multicistronic vector may also include an IRES or a viral 2A peptide, which allows the expression of more than one selectable marker from a single vector, as described further below.

[0094] Genome editing as described herein can result in either one allele or two alleles being modified in the genomic DNA of cell.In certain embodiments, at least one of the selection markers used for positive selection is a fluorescent marker, and can measure the fluorescence intensity to determine whether genetically modified cell contains single allele editing or both allele editing.

[0095] Exemplary expression cassettes encoding the fluorescent mCherry marker and the truncated CD19 (tCD19) cell surface marker are shown in Example 1. The selectable marker expression cassette comprises a UbC promoter operably linked to a polynucleotide encoding mCherry, a polynucleotide encoding a T2A peptide, a polynucleotide encoding truncated CD19 (tCD19), and a polyadenylation sequence.

[0096] In certain embodiments, a negative selection marker is used to identify cells that do not have the selection marker expression cassette (i.e., the sequence encoding the positive selection marker is deleted). For example, a suicide marker may be included as a negative selection marker in the selection marker expression cassette to facilitate the negative selection of cells after the second HDR step. A suicide gene can be used to selectively kill cells by inducing apoptosis in genetically modified cells or converting a non-toxic drug into a toxic compound. Examples include suicide genes encoding thymidine kinase, cytosine deaminase, intracellular antibody, telomerase, caspase, and DNase. In certain embodiments, a suicide gene is used in combination with one or more other selection markers, such as those listed above for use in positively selecting cells. The suicide gene may be removed in the second HDR step to achieve scarless editing. Alternatively, the suicide gene may be retained in the genetically modified cells after the second HDR step to improve safety by allowing its destruction at will. See, e.g., Jones et al. (2014) Front. Pharmacol. 5:254, Mitsui et al. (2017) Mol. Ther. Methods Clin. Dev. 5:51-58, Greco et al. (2015) Front. Pharmacol. 6:95, which are incorporated herein by reference.

[0097] Alternatively or additionally, cells can be tested for the absence of selectable marker sequences and other undesirable sequences in the genetically modified cells after the second HDR step using conventional methods such as polymerase chain reaction (PCR), fluorescence in situ hybridization (FISH), gene arrays, sequencing, or hybridization techniques (e.g., Southern blots).

[0098] Genome editing can be performed on a single cell or a population of cells of interest, and can be performed on any type of cell, including any cell from a prokaryotic, eukaryotic, or archaeal organism, including bacteria, archaea, fungi, protists, plants, and animals. Cells from tissues, organs, and biopsies, as well as recombinant cells, genetically modified cells, cells from in vitro cultured cell lines, and artificial cells (e.g., nanoparticles, liposomes, polymersomes, or microcapsules encapsulating nucleic acids) can all be used in the practice of the present invention. The methods of the present invention are also applicable to editing nucleic acids in cell fragments, cellular components, or organelles containing nucleic acids (e.g., mitochondria in animal and plant cells, plastids (e.g., chloroplasts) in plant cells and algae). Cells can be cultured or expanded before scarless genome editing as described herein, or at any time during the process, for example, between the first and second HDR steps, or after the second HDR step.

[0099] Any type II CRISPR system Cas9 endonuclease from any species, or its biologically active fragment, variant, analog or derivative that retains Cas9 endonuclease activity (i.e., catalyzes the site-specific cleavage of DNA to generate double-strand breaks) can be used to carry out HDR step.Cas9 does not need to be derived from organisms, and can be produced synthetically or recombinantly.Cas9 sequences from multiple bacterial species are well known in the art and are listed in the National Center for Biotechnology Information (NCBI) database.For example, Streptococcus pyogenes (WP_002989955, WP_038434062, WP_011528583); Campylobacter jejuni (WP_022552435, YP_002344900), Campylobacter coli (WP_060786116); Campylobacter fetus (WP_059434633); Corynebacterium ulcerans (NC_015683, NC_017317); Corynebacterium diphtheriae ria (NC_016782, NC_016786); Enterococcus faecalis (WP_033919308); Spiroplasma sylphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanense (NC_021846); Streptococcus iniae (NC_021314); Belleriella baltica (NC_018010); Cycloflexus torchis I (NC_01872 1); Streptococcus thermophilus (YP_820832), Streptococcus mutans (WP_061046374, WP_024786433); Listeria innocua (NP_472073); Listeria monocytogenes (WP_061665472); Legionella pneumophila (WP_062726656); Staphylococcus aureus (WP_001573634); Francisella tularensis (WP_0327 See NCBI entries for Cas9 from Enterococcus faecalis (WP_033919308); Lactobacillus rhamnosus (WP_048482595, WP_032965177); and Neisseria meningitidis (WP_061704949, YP_002342100); all of these sequences (deposited as of the filing date of this application) are incorporated herein by reference.Any of these sequences, or variants thereof, including sequences having at least about 70-100% sequence identity thereto, including any percent identity within this range, e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto, can be used for scarless genome editing as described herein. For sequence comparisons and discussions of genetic diversity and phylogenetic analyses of Cas9, see also Fonfara et al. (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al. (2015) J. Bacteriol. 198(5):797-807, Shmakov et al. (2015) Mol. Cell. 60(3):385-397, and Chylinski et al. (2014) Nucleic Acids Res. 42(10):6091-6105.

[0100] CRISPR-Cas systems naturally occur in bacteria and archaea, where they play a role in RNA-mediated adaptive immunity against foreign DNA. Bacterial type II CRISPR systems use the endonuclease Cas9, which forms a complex with a guide RNA (gRNA) that specifically hybridizes to a complementary genomic target sequence, where the Cas9 endonuclease catalyzes cleavage to generate a double-stranded break. Cas9 targeting further relies on the presence of a 5' protospacer adjacent motif (PAM) in the DNA at or near the gRNA binding site.

[0101] Cas9 can target a specific genomic sequence (i.e., the genomic target sequence to be modified) by modifying its guide RNA sequence. The target-specific guide RNA contains a nucleotide sequence complementary to the genomic target sequence, thereby mediating the binding of the Cas9-gRNA complex by hybridization at the target site. For example, the gRNA can be designed with a sequence complementary to the sequence of a minor allele to target the Cas9-gRNA complex to the site of the mutation. The mutation may include an insertion, deletion, or substitution. For example, the mutation may include a single nucleotide variation, gene fusion, translocation, inversion, duplication, frameshift, missense, nonsense, or other mutation associated with the phenotype or disease of interest. The targeted minor allele may be a common genetic variant or a rare genetic variant. In certain embodiments, the gRNA is designed to selectively bind to a minor allele with single-base-pair discrimination, for example, to enable the binding of the Cas9-gRNA complex to single nucleotide polymorphisms (SNPs). In particular, gRNA can be designed to target the disease-related mutation of interest, for the purpose of genome editing to remove mutation from gene.Alternatively, gRNA can be designed with the sequence complementary to the sequence of major allele or wild-type allele, so that Cas9-gRNA complex can target allele, for the purpose of genome editing to introduce mutation, for example, insertion, deletion or substitution into gene in genomic DNA of cell.Such genetically modified cell can be used, for example, to generate disease model for drug screening.

[0102] The genomic target site typically contains a nucleotide sequence complementary to the gRNA and may further contain a protospacer adjacent motif (PAM). In certain embodiments, the target site contains 20-30 base pairs in addition to the 3-base pair PAM. Typically, the first nucleotide of the PAM can be any nucleotide, while the other two nucleotides depend on the specific Cas9 protein selected. Exemplary PAM sequences are known to those of skill in the art and include, but are not limited to, NNG, NGN, NAG, and NGG, where N represents any nucleotide. In certain embodiments, the allele targeted by the gRNA contains a mutation that creates a PAM within the allele, which facilitates binding of the Cas9-gRNA complex to the allele.

[0103] In certain embodiments, the gRNA is 5-50 nucleotides in length, 10-30 nucleotides in length, 15-25 nucleotides in length, 18-22 nucleotides in length, or 19-21 nucleotides in length, or any length between the stated ranges, including, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides in length. The guide RNA may be a single guide RNA that includes the crRNA and tracrRNA sequences in a single RNA molecule, or the guide RNA may comprise two RNA molecules with the crRNA and tracrRNA sequences present in separate RNA molecules.

[0104] Donor polynucleotides and gRNAs can be prepared using standard techniques, e.g., U.S. Pat. Nos. 4,458,066 and 4,415,732; Beaucage et al., Tetrahedron (1992), which are incorporated herein by reference. 48: 2223-2311; and can be readily synthesized by solid-phase synthesis via phosphoramidite chemistry as disclosed in Applied Biosystems User Bulletin No. 13 (1 April 1987). Other chemical synthesis methods include, for example, the phosphotriester method described by Narang et al., Meth. Enzymol. (1979) 68 : the phosphotriester method described by Brown et al., Meth. Enzymol. (1979) 68 : and the phosphodiester method disclosed by Brown et al., Meth. Enzymol. (1979)

[0105] In some instances, a population of cells may be enriched for those containing a genetic modification by separating the genetically modified cells from the remaining population. Separation of genetically modified cells typically relies on the expression of a selectable marker that is incorporated simultaneously with the intended editing at the target locus. After the first HDR step, for example, positive selection is performed to isolate cells from the population such that an enriched population of cells containing the genetic modification is generated. After the second HDR step, negative selection is performed, for example, to remove cells that still contain unwanted sequences from the selectable marker expression cassette.

[0106] Cell separation can be achieved by any convenient separation technique appropriate for the selectable marker used, including but not limited to flow cytometry, fluorescence-activated cell sorting (FACS), magnetic-activated cell sorting (MACS), sedimentation, immunopurification, and affinity chromatography. For example, if a fluorescent marker is used, cells can be separated by fluorescence-activated cell sorting (FACS), whereas if a cell surface marker is used, cells can be separated from a heterogeneous population by an affinity separation technique, such as MACS, affinity chromatography, "panning" with an affinity reagent attached to a solid matrix, immunopurification with an antibody specific for the cell surface marker, or other convenient techniques.

[0107] In certain embodiments, positive or negative selection of genetically modified cells is performed using a binding substance that specifically binds to a selection marker on the cells (e.g., as generated from a selection marker expression cassette contained in a donor polynucleotide). Examples of binding substances include, but are not limited to, antibodies, antibody fragments, antibody mimics, and aptamers. In some embodiments, the binding substance binds to the selection marker with high affinity. The binding substance may be immobilized on a solid support to facilitate isolation of genetically modified cells from liquid culture. Exemplary solid supports include magnetic beads, non-magnetic beads, slides, gels, membranes, and microtiter plate wells.

[0108] In certain embodiments, the binding agent comprises an antibody that specifically binds to a selectable marker on a cell, including polyclonal and monoclonal antibodies, hybrid antibodies, altered antibodies, chimeric antibodies, and humanized antibodies, as well as hybrid (chimeric) antibody molecules (see, e.g., Winter et al. (1991) Nature 349:293-299; and U.S. Pat. No. 4,816,567); F(ab')2 and F(ab)2 fragments; F vmolecules (non-covalent heterodimers, see, e.g., Inbar et al. (1972) Proc Natl Acad Sci USA 69:2659-2662; and Ehrlich et al. (1980) Biochem 19:4091-4096); single-chain Fv molecules (sFv) (see, e.g., Huston et al. (1988) Proc Natl Acad Sci USA 85:5879-5883); nanobodies or single-domain antibodies (sdAb) (see, e.g., Wang et al. (2016) Int J Nanomedicine 11:3287-3303; Vincke et al. (2012) Methods Mol Biol 911:15-26); dimeric and trimeric antibody fragment constructs; minibodies (see, e.g., Pack et al. (1992) Biochem 31:1579-1584; Cumber et al. (1992) J Immunology 149B:120-126); humanized antibody molecules (see, e.g., Riechmann et al. (1988) Nature 332:323-327; Verhoeyan et al. (1988) Science 239:1534-1536; and UK Patent Publication No. GB ​​2,276,169, published September 21, 1994); and any functional fragments derived from such molecules, wherein such fragments retain the specific binding properties of the parent antibody molecule (i.e., specifically bind to a selectable marker on a cell).

[0109] In other embodiments, the binding substance comprises an aptamer that specifically binds to a selectable marker on a cell. Any type of aptamer may be used, including DNA, RNA, xenonucleic acid (XNA), or peptide aptamers that specifically bind to a target antibody isotype. Such aptamers can be identified, for example, by screening a combinatorial library. Nucleic acid aptamers (e.g., DNA or RNA aptamers) that selectively bind to a target antibody isotype can be generated by performing repeated rounds of in vitro selection or systematic evolution of ligands by exponential enrichment (SELEX). Peptide aptamers that bind to a selectable marker on a cell can be isolated from a combinatorial library and improved by repeated rounds of directed mutation or mutagenesis and selection.For descriptions of methods for generating aptamers, see, for example, Aptamers: Tools for Nanotherapy and Molecular Imaging (R.N. Veedu ed., Pan Stanford, 2016), Nucleic Acid and Peptide Aptamers: Methods and Protocols (Methods in Molecular Biology, G. Mayer ed., Humana Press, 2009), Nucleic Acid Aptamers: Selection, Characterization, and Application (Methods in Molecular Biology, G. Mayer ed., Humana Press, 2016), Aptamers Selected by Cell-SELEX for Theranostics (W. Tan, X. Fang eds., Springer, 2015), Cox et al. (2001) Bioorg. Med. Chem. 9(10):2525-2531; Cox et al. (2002) Nucleic Acids Res. 30(20): e108, Kenan et al. (1999) Methods Mol Biol. 118:217-231;Platella et al. (2016) Biochim. Biophys. Acta Nov 16 pii: S0304-4165(16)30447-0, and Lyu et al. (2016) Theranostics 6(9):1440-1452.

[0110] In yet other embodiments, the binding agent comprises an antibody mimetic. Affibody molecules (Nygren (2008) FEBS J. 275 (11):2668-2676), affilin (Ebersbach et al. (2007) J. Mol. Biol. 372 (1):172-185), affimer (Johnson et al. (2012) Anal. Chem. 84 (15):6553-6560), affitin (Krehenbrink et al. (2008) J. Mol. Biol. 383 (5):1058-1068), alphabody (Desmet et al. (2014) Nature Communications 5:5237), anticalin (Skerra (2008) FEBS J. 275 (11):2677-2683), avimer (Silverman et al. al. (2005) Nat. Biotechnol. 23 (12):1556-1561), darpins (Stumpp et al. (2008) Drug Discov. Today 13 (15-16):695-701), finomymers (Grabulovski et al. (2007) J. Biol. Chem. 282 (5):3196-3204), and monobodies (Koide et al. (2007) Methods Mol. Biol. 352:95-109).

[0111] In positive selection, the cells that carry selection markers are collected, whereas in negative selection, the cells that carry selection markers are removed from the cell population.For example, in positive selection, binding substances specific to surface markers can be immobilized on solid support (for example, column or magnetic beads) and used to collect cells of interest on the solid support.Cells that are not of interest do not bind to the solid support (for example, pass through the column or do not adhere to the magnetic beads).In negative selection, binding substances are used to deplete the cells that are not of interest, the cell population.Cells of interest are those that do not bind to binding substances (for example, pass through the column or remain after the magnetic beads are removed).

[0112] Dead cells may be selected by using a dye that preferentially stains dead cells (e.g., propidium iodide). Any technique may be used that is not unduly detrimental to the viability of the genetically modified cells.

[0113] In this manner, a composition that is highly enriched in cells with the desired genetic modification can be produced. By "highly enriched," it is meant that the genetically modified cells account for 70% or more, 75% or more, 80% or more, 85% or more, 90% or more, or 95% or more, or 98% or more of the cell composition. In other words, the composition can be a substantially pure composition of genetically modified cells.

[0114] The genetically modified cells produced by the method described herein can be used immediately. Alternatively, cells can be frozen at liquid nitrogen temperature and stored for a long period before thawing and use. In such a case, cells can be frozen in 10% DMSO, 50% serum, 40% buffered medium, or some other solution that is commonly used in the art to store cells at such freezing temperatures, and can be thawed in a manner commonly known in the art for thawing frozen cultured cells.

[0115] In certain embodiments, inhibitors of non-homologous end joining (NHEJ) pathway are used to increase the frequency of cells that are genetically modified by HDR.The example of inhibitors of NHEJ pathway includes any compound (substance) that inhibits or prevents any protein component in NHEJ pathway from expressing or activating.The protein components of NHEJ pathway include but are not limited to Ku70, Ku86, DNA protein kinase (DNA-PK), Rad50, MRE11, NBS1, DNA ligase IV and XRCC4.An exemplary inhibitor is wortmannin, which inhibits at least one protein component (such as DNA-PK) of NHEJ pathway. Another exemplary inhibitor is Scr7 (5,6-bis((E)-benzylideneamino)-2-mercaptopyrimidin-4-ol), which inhibits DSB junctions (Maruyama et al. (2015) Nat. Biotechnol. 33(5):538-542, Lin et al. (2016) Sci. Rep. 6:34531). RNA interference can also be used to prevent the expression of protein components of the NHEJ pathway (e.g., DNA-PK or DNA ligase IV). For example, small interfering RNA (siRNA), hairpin RNA, and other RNA or RNA:DNA species that can be cleaved or dissociated in vivo to form siRNA can be used. Alternatively, HDR enhancers such as RS-1 can be used to increase the frequency of HDR in cells (Song et al. (2016) Nat. Commun. 7:10548).

[0116] In certain embodiments, the first donor polynucleotide further comprises at least one expression cassette encoding a short hairpin RNA (shRNA) comprising a stem-loop structure that inhibits expression of a randomly integrated or episomal selectable marker. An exemplary construct is shown in Figure 4A, in which the first donor polynucleotide comprises a pair of expression cassettes encoding shRNAs, the first expression cassette of the pair being located 5' of the first left homology arm, and the second expression cassette of the pair being located 3' of the first right homology arm.

[0117] As described herein, the method steps of using Cas9 nuclease, donor polynucleotide and guide RNA can be repeated to provide any desired number of DNA modifications.Furthermore, the method can be adapted to provide the multiplex genome editing of cells.For example, by pooling multiple donor polynucleotides and guide RNAs that specifically target different genes, multiple genes can be edited simultaneously.

[0118] B. Donor Polynucleotide, Guide RNA, and Cas9-Encoding Nucleic Acid In certain embodiments, donor polynucleotide, guide RNA, and / or Cas9 are expressed in vivo from a vector. A "vector" is a material composition that can be used to deliver nucleic acid of interest into cells. Donor polynucleotide, guide RNA, and Cas9 can be introduced into cells in a single vector or in multiple separate vectors. The ability of constructs to produce donor polynucleotide, guide RNA, and Cas9 nuclease and genetically modify cells can be empirically determined (see, for example, Example 1, which describes the use of mCherry fluorescent marker to detect genetically modified cells and the immunomagnetic separation of cells that express tCD19 surface marker).

[0119] Numerous vectors are known in the art, including, but not limited to, linear polynucleotides, polynucleotides associated with ionic or amphipathic compounds, plasmids, and viruses. Thus, the term "vector" includes autonomously replicating plasmids or viruses. Examples of viral vectors include, but are not limited to, adenoviral vectors, adeno-associated viral vectors, retroviral vectors, lentiviral vectors, and the like. Expression constructs can be replicated in living cells or can be synthetically produced. For purposes of this application, the terms "expression construct," "expression vector," and "vector" are used interchangeably in a general, illustrative sense to demonstrate the application of the present invention and are not intended to limit the present invention.

[0120] In one embodiment, an expression vector for expressing a donor polynucleotide, gRNA, or Cas9 comprises a promoter "operably linked" to a polynucleotide encoding the donor polynucleotide, gRNA, or Cas9. As used herein, the phrases "operably linked" or "under transcriptional control" mean that the promoter is in the correct location and orientation relative to the polynucleotide to control initiation of transcription by RNA polymerase and expression of the polynucleotide.

[0121] In certain embodiments, the nucleic acid encoding the polynucleotide of interest is under the transcriptional control of a promoter. "Promoter" refers to the DNA sequence recognized by the cell's synthetic machinery or introduced synthetic machinery required to initiate the specific transcription of a gene. The term promoter is used herein to refer to a group of transcriptional control modules clustered around the initiation site for RNA polymerase I, II, or III. Typical promoters for mammalian cell expression include, among others, the SV40 early promoter, the CMV promoter such as the CMV immediate early promoter (see U.S. Patent Nos. 5,168,062 and 5,385,839, the entire contents of which are incorporated herein by reference), the mouse mammary tumor virus LTR promoter, the adenovirus major late promoter (Ad MLP), and the herpes simplex virus promoter. Other non-viral promoters, such as the promoter derived from the mouse metallothionein gene, can also be used for mammalian expression. These and other promoters can be obtained from commercially available plasmids using techniques well known in the art. See, for example, Sambrook et al., supra. Enhancer elements may be used in conjunction with the promoter to increase expression levels of the construct. Examples include those described in Dijkema et al., EMBO J. (1985). 4 SV40 early gene enhancer as described in Gorman et al., Proc. Natl. Acad. Sci. USA (1982b) 79 enhancer / promoter derived from the Rous sarcoma virus long terminal repeat (LTR) as described in Boshart et al., Cell (1985), such as elements contained in the CMV intron A sequence, as described in Boshart et al., Cell (1985), 41 :521.

[0122] Typically, a transcription terminator / polyadenylation signal is also present in the expression construct. Examples of such sequences include, but are not limited to, those derived from SV40, as described in Sambrook et al., supra, and the bovine growth hormone terminator sequence (see, for example, U.S. Patent No. 5,122,458). In addition, a 5'-UTR sequence can be placed adjacent to the coding sequence to enhance expression of the same. Such a sequence may include a UTR containing an internal ribosome entry site (IRES).

[0123] The inclusion of an IRES allows translation of one or more open reading frames from the vector. The IRES element attracts the eukaryotic ribosomal translation initiation complex and promotes translation initiation. See, e.g., Kaufman et al., Nuc. Acids Res. (1991). 19 :4485-4490;Gurtu et al., Biochem. Biophys. Res. Comm. (1996) 229 :295-298;Rees et al., BioTechniques (1996) 20 :102-110;Kobayashi et al., BioTechniques (1996) 21 :399-402; and Mosser et al., BioTechniques (1997 22 150-161. Numerous IRES sequences are known, including those derived from a wide variety of viruses, for example, picornavirus leader sequences such as the encephalomyocarditis virus (EMCV) UTR (Jang et al. J. Virol. (1989) 63 :1651-1660), polio leader sequence, hepatitis A virus leader, hepatitis C virus IRES, human rhinovirus type 2 IRES (Dobrikova et al., Proc. Natl. Acad. Sci. (2003) 100(25):15125-15130), the IRES element from foot-and-mouth disease virus (Ramesh et al., Nucl. Acid Res. (1996) 24 :2697-2700), Giardia virus IRES (Garlapati et al., J. Biol. Chem. (2004) 279 (5):3389-3397). IRES sequences from yeast were used, as well as the human angiotensin II receptor type 1 IRES (Martin et al., Mol. Cell Endocrinol. (2003) 212:51-61), fibroblast growth factor IRES (FGF-1 IRES and FGF-2 IRES, Martineau et al. (2004) Mol. Cell. Biol. 24(17):7622-7635), vascular endothelial growth factor IRES (Baranick et al. (2008) Proc. Natl. Acad. Sci. USA 105(12):4733-4738, Stein et al. (1998) Mol. Cell. Biol. 18(6):3112-3119, Bert et al. (2006) RNA 12(6):1074-1083), and insulin-like growth factor 2 IRES (Pedersen et al. (2002) Biochem. J. 363(Pt 1):37-44) are also used herein. These elements are readily commercially available, for example, in plasmids sold by Clontech (Mountain View, CA), Invivogen (San Diego, CA), Addgene (Cambridge, MA), and GeneCopoeia (Rockville, MD). See also IRESite (iresite.org), a database of experimentally validated IRES structures. IRES sequences may be included in vectors, for example, to express multiple selectable markers from an expression cassette or Cas9 in combination with one or more selectable markers.

[0124] Alternatively, polynucleotides encoding viral T2A peptides can be used to enable the production of multiple protein products (e.g., Cas9, one or more selectable markers) from a single vector. A 2A linker peptide is inserted between the coding sequences in a multicistronic construct. Self-cleaving 2A peptides allow co-expressed proteins from a multicistronic construct to be produced at equimolar levels. 2A peptides from various viruses may also be used, including, but not limited to, 2A peptides from foot-and-mouth disease virus, equine rhinitis A virus, Thosea asigna virus, and porcine teschovirus-1. See, e.g., Kim et al. (2011) PLoS One 6(4):e18556, Trichas et al. (2008) BMC Biol. 6:40, Provost et al. (2007) Genesis 45(10):625-629, Furler et al. (2001) Gene Ther. 8(11):864-873, the entire contents of which are incorporated herein by reference.

[0125] There are several ways that expression vectors can be introduced into cells. In certain embodiments, expression constructs comprise virus or engineered constructs derived from viral genomes. Several virus-based systems have been developed for gene transfer into mammalian cells. These include adenovirus, retrovirus (γ-retrovirus and lentivirus), poxvirus, adeno-associated virus, baculovirus, and herpes simplex virus (see, for example, Warnock et al. (2011) Methods Mol. Biol. 737:1-25; Walther et al. (2000) Drugs 60(2):249-271; and Lundstrom (2003) Trends Biotechnol. 21(3):117-122, which are incorporated herein by reference in their entirety). The ability of certain viruses to enter cells via receptor-mediated endocytosis, integrate into the host cell genome, and stably and efficiently express viral genes makes them attractive candidates for the transfer of foreign genes into mammalian cells.

[0126] For example, retroviruses provide a convenient platform for gene delivery systems. Using techniques known in the art, selected sequences can be inserted into vectors and packaged into retroviral particles. The recombinant virus can then be isolated and delivered to cells of a subject either in vivo or ex vivo. Several retroviral systems have been described (U.S. Pat. No. 5,219,740; Miller and Rosman (1989) BioTechniques 7:980-990; Miller, AD (1990) Human Gene Therapy 1:5-14; Scarpa et al. (1991) Virology 180:849-852; Burns et al. (1993) Proc. Natl. Acad. Sci. USA 90:8033-8037; Boris-Lawrie and Temin (1993) Cur. Opin. Genet. Develop. 3:102-109; and Ferry et al. (2011) Curr. Pharm. Des. 17(24):2516-2527). Lentiviruses are a class of retroviruses that are particularly useful for delivering polynucleotides to mammalian cells because they can infect both dividing and non-dividing cells (see, e.g., Lois et al (2002) Science 295:868-872; Durand et al. (2011) Viruses 3(2):132-159, incorporated herein by reference).

[0127] Several adenoviral vectors have also been described. Unlike retroviruses, which integrate into the host genome, adenoviruses exist extrachromosomally, thus minimizing the risks associated with insertional mutagenesis (Haj-Ahmad and Graham, J. Virol. (1986) 57:267-274; Bett et al., J. Virol. (1993) 67:5911-5921; Mittereder et al., Human Gene Therapy (1994) 5:717-729; Seth et al., J. Virol. (1994) 68:933-940; Barr et al., Gene Therapy (1994) 1:51-58; Berkner, KL BioTechniques (1988) 6:616-629; and Rich et al., Human Gene Therapy (1993) 4:461-476). In addition, various adeno-associated virus (AAV) vector systems have been developed for gene delivery. AAV vectors can be readily constructed using techniques well known in the art.See, e.g., U.S. Pat. Nos. 5,173,414 and 5,139,941; International Publication Nos. WO 92 / 01070 (published January 23, 1992) and WO 93 / 03769 (published March 4, 1993); Lebkowski et al., Molec. Cell. Biol. (1988) 8:3988-3996; Vincent et al., Vaccines 90 (1990) (Cold Spring Harbor Laboratory Press); Carter, B. J. Current Opinion in Biotechnology (1992) 3:533-539; Muzyczka, N. Current Topics in Microbiol. and Immunol. (1992) 158:97-129; Kotin, R. M. Human Gene Therapy (1994) 5:793-801; Shelling and Smith, Gene Therapy (1994) 1:165-169; and Zhou et al., J. Exp. Med. (1994) 179:1867-1875.

[0128] Another vector system useful for delivering the polynucleotides of the invention is a recombinant poxvirus vaccine administered to the intestinal tract, as described by Small, Jr., PA, et al. (U.S. Patent No. 5,676,950, issued October 14, 1997, incorporated herein by reference).

[0129] Additional viral vectors used to deliver nucleic acid molecules of interest include those derived from poxviruses, including vaccinia virus and avian poxvirus. For example, vaccinia virus recombinants expressing nucleic acid molecules of interest (e.g., donor polynucleotides, gRNA, or Cas9) can be constructed as follows: DNA encoding a specific nucleic acid sequence is first inserted into an appropriate vector so that it is adjacent to a vaccinia promoter and adjacent vaccinia DNA sequences, such as the sequence encoding thymidine kinase (TK). This vector is then used to transfect cells that are simultaneously infected with vaccinia. Homologous recombination occurs to insert the vaccinia promoter and the gene encoding the sequence of interest into the viral genome. The resulting TK recombinants can be selected by culturing cells in the presence of 5-bromodeoxyuridine and picking resistant viral plaques.

[0130] Alternatively, avipoxviruses such as fowlpox virus and canarypox virus can also be used to deliver nucleic acid molecules of interest.Because members of the genus Avipox can only productively replicate in susceptible bird species, and therefore are not infectious in mammalian cells, the use of avipox vectors is particularly desirable in humans and other mammalian species.The method for producing recombinant avipoxviruses is known in the art, and uses genetic recombination as described above for the production of vaccinia virus.For example, see WO 91 / 12882; WO 89 / 03429; and WO 92 / 03545.

[0131] Molecular conjugate vectors, such as the adenovirus chimeric vectors described in Michael et al., J. Biol. Chem. (1993) 268:6866-6869 and Wagner et al., Proc. Natl. Acad. Sci. USA (1992) 89:6099-6103, can also be used for gene delivery.

[0132] Members of the alphavirus genus, such as, but not limited to, vectors derived from Sindbis virus (SIN), Semliki Forest virus (SFV), and Venezuelan equine encephalitis virus (VEE), can also be used as viral vectors to deliver the polynucleotides of the present invention. For a description of Sindbis virus-derived vectors useful in practicing the methods of the present invention, see Dubensky et al. (1996) J. Virol. 70:508-519; and International Publication Nos. WO 95 / 07995 and WO 96 / 17072; and Dubensky, Jr., TW, et al., U.S. Patent No. 5,843,723, issued December 1, 1998, and Dubensky, Jr., TW, U.S. Patent No. 5,789,245, issued August 4, 1998, both of which are incorporated herein by reference. Particularly preferred are chimeric alphavirus vectors composed of sequences derived from Sindbis virus and Venezuelan equine encephalitis virus. See, e.g., Perri et al. (2003) J. Virol. 77: 10394-10403, and International Publication Nos. WO 02 / 099035, WO 02 / 080982, WO 01 / 81609, and WO 00 / 61772, which are incorporated by reference in their entireties.

[0133] A vaccinia-based infection / transfection system can be advantageously used to provide inducible, transient expression of a polynucleotide of interest (e.g., miR-181 or its mimic or inhibitor) in host cells. In this system, cells are first infected in vitro with a vaccinia virus recombinant encoding bacteriophage T7 RNA polymerase. This polymerase exhibits exquisite specificity in that it only transcribes templates with a T7 promoter. After infection, the cells are transfected with a polynucleotide of interest driven by a T7 promoter. The polymerase expressed in the cytoplasm from the vaccinia virus recombinant transcribes the transfected DNA into RNA. This method provides high-level, transient cytoplasmic production of large amounts of RNA. See, e.g., Elroy-Stein and Moss, Proc. Natl. Acad. Sci. USA (1990) 87:6743-6747; Fuerst et al., Proc. Natl. Acad. Sci. USA (1986) 83:8122-8126.

[0134] As an alternative approach to infection with vaccinia or avipox virus recombinants or delivery of nucleic acids using other viral vectors, an amplification system can be used that results in high-level expression after introduction into host cells. Specifically, a T7 RNA polymerase promoter can be engineered to precede the coding region of T7 RNA polymerase. Translation of RNA derived from this template produces T7 RNA polymerase, which then transcribes more templates. Concomitantly, there is a cDNA whose expression is under the control of the T7 promoter. Thus, some of the T7 RNA polymerase resulting from translation of the amplified template RNA will result in transcription of the desired gene. Because some T7 RNA polymerase is required to initiate amplification, T7 RNA polymerase can be introduced into cells together with the template to prime the transcription reaction. The polymerase can be introduced as a protein or on a plasmid encoding the RNA polymerase. For further discussion of the T7 system and its use to transform cells, see, e.g., International Publication No. WO 94 / 26911; Studier and Moffatt, J. Mol. Biol. (1986) 189:113-130; Deng and Wolff, Gene (1994) 143:245-249; Gao et al., Biochem. Biophys. Res. Commun. (1994) 200:1201-1206; Gao and Huang, Nuc. Acids Res. (1993) 21:2867-2872; Chen et al., Nuc. Acids Res. (1994) 22:2114-2120; and U.S. Patent No. 5,135,855.

[0135] In order to bring about the expression of sense or antisense gene construct, expression construct must be delivered into cell.This delivery can be achieved in vitro, such as in the experimental procedure for transforming cell lines, or in vivo or ex vivo, such as in the treatment of certain disease states.One mechanism for delivery is through viral infection, where expression construct is encapsulated in infectious viral particles.

[0136] Several non-viral methods for the transfer of expression constructs into cultured mammalian cells also are contemplated by the present invention. These include the use of calcium phosphate precipitation, DEAE-dextran, electroporation, direct microinjection, DNA-loaded liposomes, lipofectamine-DNA complexes, cell sonication, gene bombardment with high-velocity microprojectiles, and receptor-mediated transfection (see, e.g., Graham and Van Der Eb (1973) Virology 52:456-467; Chen and Okayama (1987) Mol. Cell Biol. 7:2745-2752; Rippe et al. (1990) Mol. Cell Biol. 10:689-695; Gopal (1985) Mol. Cell Biol. 5:1188-1190; Tur-Kaspa et al. (1986) Mol. Cell. Biol. 6:716-718; Potter et al. (1984) Proc. Natl. Acad. Sci. USA, incorporated herein by reference). 81:7161-7165); Harland and Weintraub (1985) J. Cell Biol. 101:1094-1099); Nicolau and Sene (1982) Biochim. Biophys. Acta 721:185-190; Fraley et al. (1979) Proc. Natl. Acad. Sci. USA 76:3348-3352;Fechheimer et al. (1987) Proc. Natl. Acad. Sci. USA 84:8463-8467;Yang et al. (1990) Proc. Natl. Acad. Sci. USA 87:9568-9572;Wu and Wu (1987) J. Biol. Chem. 262:4429-4432;Wu and Wu (1988) Biochemistry 27:887-892.) Some of these techniques can be successfully adapted for in vivo or ex vivo use.

[0137] Once the expression construct is delivered into the cell, the nucleic acid encoding the gene of interest can be located and expressed at various sites. In certain embodiments, the nucleic acid encoding the gene can be stably integrated into the genome of the cell. This integration can be in the same location and orientation via homologous recombination (gene replacement), or in a random, non-specific location (gene augmentation). In yet further embodiments, the nucleic acid can be stably maintained in the cell as a separate episomal segment of DNA. Such a nucleic acid segment, or "episome," encodes sufficient sequences to allow maintenance and replication independent of or synchronized with the host cell cycle. How the expression construct is delivered to the cell and where the nucleic acid remains in the cell depend on the type of expression construct used.

[0138] In another embodiment of the present invention, the expression construct may simply consist of naked recombinant DNA or a plasmid. Transfer of the construct may be carried out by any of the methods described above for physically or chemically permeabilizing the cell membrane. This is particularly applicable to in vitro transfer, but may also be applied to in vivo use. Dubensky et al. (Proc. Natl. Acad. Sci. USA (1984) 81:7529-7533) successfully injected polyomavirus DNA in the form of a calcium phosphate precipitate into the liver and spleen of adult and neonatal mice, demonstrating active viral replication and acute infection. Benvenisty and Neshif (Proc. Natl. Acad. Sci. USA (1986) 83:9551-9555) also demonstrated that direct intraperitoneal injection of calcium phosphate-precipitated plasmids results in expression of the transfected gene. It is envisioned that DNA encoding a gene of interest may also be transferred in a similar manner in vivo and the gene product expressed.

[0139] In yet another embodiment, naked DNA expression constructs can be transferred into cells by particle bombardment. This method relies on the ability to accelerate DNA-coated microprojectiles to high speeds, allowing them to penetrate cell membranes and enter cells without killing them (Klein et al. (1987) Nature 327:70-73). Several devices for accelerating small particles have been developed. One such device relies on a high-voltage discharge to generate an electric current, which then provides the motive force (Yang et al. (1990) Proc. Natl. Acad. Sci. USA 87:9568-9572). Microprojectiles can be made of biologically inert materials such as tungsten or gold beads.

[0140] In a further embodiment, the expression construct may be delivered using liposomes. Liposomes are vesicular structures characterized by a phospholipid bilayer membrane and an internal aqueous medium. Multilamellar liposomes have multiple lipid layers separated by aqueous medium. They spontaneously form when phospholipids are suspended in an excess of aqueous solution. The lipid components undergo self-rearrangement before the formation of a closed structure, trapping water and dissolved solutes between the lipid bilayers (Ghosh and Bachhawat (1991) Liver Diseases, Targeted Diagnosis and Therapy Using Specific Receptors and Ligands, Wu et al. (Eds.), Marcel Dekker, NY, pp. 87-104). The use of lipofectamine-DNA complexes is also contemplated.

[0141] In certain embodiments of the present invention, liposomes may be complexed with hemagglutinating virus (HVJ). This has been shown to facilitate fusion with the cell membrane and promote cell entry of liposome-encapsulated DNA (Kaneda et al. (1989) Science 243:375-378). In other embodiments, liposomes may be complexed with or used in conjunction with nuclear non-histone chromosomal proteins (HMG-I) (Kato et al. (1991) J. Biol. Chem. 266(6):3361-3364). In yet further embodiments, liposomes may be complexed with or used in conjunction with both HVJ and HMG-I. Insofar as such expression constructs have been successfully used in the transfer and expression of nucleic acids in vitro and in vivo, they are applicable to the present invention. When a bacterial promoter is used in the DNA construct, it is also desirable to include an appropriate bacterial polymerase within the liposome.

[0142] Other expression constructs that can be used to deliver nucleic acids into cells are receptor-mediated delivery vehicles. These utilize the selective uptake of macromolecules by receptor-mediated endocytosis in almost all eukaryotic organisms. Due to the cell-type-specific distribution of various receptors, delivery can be highly specific (Wu and Wu (1993) Adv. Drug Delivery Rev. 12:159-167).

[0143] Receptor-mediated gene targeting vehicle generally consists of two components: cell receptor specific ligand and DNA binding substance.Several ligands have been used for receptor-mediated gene transfer.The most extensively characterized ligands are asialoorosomucoid (ASOR) and transferrin (see, for example, Wu and Wu (1987) above; Wagner et al. (1990) Proc. Natl. Acad. Sci. USA 87(9):3410-3414). Recently, synthetic neoglycoproteins that recognize the same receptor as ASOR have been used as gene delivery vehicles (Ferkol et al. (1993) FASEB J. 7:1081-1091; Perales et al. (1994) Proc. Natl. Acad. Sci. USA 91(9):4086-4090), and epidermal growth factor (EGF) has also been used to deliver genes to squamous cell carcinoma cells (Myers, EPO 0273085).

[0144] In other embodiments, delivery vehicle may comprise ligand and liposome.For example, Nicolau et al. (Methods Enzymol. (1987) 149:157-176) used lactosylceramide, galactose-terminal asialganglioside, incorporated into liposome, and observed increased uptake of insulin gene by hepatocytes.Therefore, it is feasible that nucleic acid encoding specific gene can also be specifically delivered into cells by any number of receptor-ligand systems with or without liposome.In addition, antibodies against surface antigens on cells can also be used as targeting moieties.

[0145] In certain instances, the donor polynucleotide, gRNA, or recombinant polynucleotide encoding Cas9 may be administered in combination with a cationic lipid. Examples of cationic lipids include, but are not limited to, lipofectin, DOTMA, DOPE, and DOTAP. WO / 0071096, specifically incorporated by reference, describes various formulations, such as DOTAP:cholesterol or cholesterol derivative formulations, that can be effectively used for gene therapy. Other disclosures also discuss various lipid or liposome formulations, including nanoparticles, and methods of administration; these include, but are not limited to, U.S. Patent Application Publication Nos. 20030203865, 20020150626, 20030032615, and 20040048787, which are specifically incorporated by reference to the extent that they disclose formulations and other relevant aspects of nucleic acid administration and delivery. Methods used to form the particles are also disclosed in U.S. Pat. Nos. 5,844,107, 5,877,302, 6,008,336, 6,077,835, 5,972,901, 6,200,801, and 5,972,900, which are incorporated by reference for their aspects.

[0146] In certain embodiments, gene transfer can be more easily performed under ex vivo conditions.Ex vivo gene therapy refers to the isolation of cells from a subject, the delivery of nucleic acid into cells in vitro, and then returning the modified cells to the subject.This can include the collection of biological samples containing cells from a subject.For example, blood can be obtained by venipuncture, and solid tissue samples can be obtained by surgical techniques according to methods well known in the art.

[0147] Usually, but not always, the subject that receives cells (i.e., recipient) is also the subject from which the cells are collected or obtained, which provides the advantage that the donated cells are autologous.However, cells can be obtained from another subject (i.e., donor), from the culture of cells from a donor, or from an established cell culture line.Cells can be obtained from the same or different species as the subject to be treated, but preferably from the same species, and more preferably from the same immunological profile as the subject.Such cells can be obtained, for example, from a biological sample containing cells from a close relative or a compatible donor, and then transfected with a nucleic acid (e.g., encoding a donor polynucleotide, gRNA, or Cas9), and administered to a subject that needs genome modification, for example, for the treatment of a disease or condition.

[0148] C. Kit The above-mentioned reagents, including donor DNA and guide RNA for the first and second HDR steps, and Cas9, can be provided in a kit together with suitable instructions and other necessary reagents for scarless genome modification as described herein.The kit can also contain cells for genome modification, binding substances for positive and negative selection of cells, and transfection agents.The kit typically contains donor DNA, guide RNA, Cas9, and other necessary reagents in separate containers.The kit usually includes instructions (e.g., written instructions, CD-ROM, DVD, Blu-ray, flash drive, digital download, etc.) for performing genome editing as described herein.The kit can also contain other packaged reagents and materials (i.e., washing buffer, etc.) depending on the specific assay used.Genome editing of cells as described herein can be performed using these kits.

[0149] In another embodiment, the kit includes a first donor polynucleotide comprising a selectable marker expression cassette including a UbC promoter, a polynucleotide encoding mCherry, a polynucleotide encoding a T2A peptide, a polynucleotide encoding a truncated CD19 (tCD19), and a polyadenylation sequence. In one embodiment, the donor DNA comprises a selectable marker expression cassette comprising the sequence of SEQ ID NO: 1, or a sequence exhibiting at least about 80-100% sequence identity thereto, including any percent identity thereto within this range, for example, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% sequence identity thereto, and cells having the first donor polynucleotide integrated into the genomic DNA of the cell at the target site by Cas9-mediated HDR can be identified by positive selection of the selectable marker encoded by the expression cassette.

[0150] In another embodiment, the kit includes a second donor polynucleotide comprising a polynucleotide comprising the sequence of SEQ ID NO:2 or a sequence having at least 95% identity to the sequence of SEQ ID NO:2, wherein the second donor polynucleotide can be integrated into the modified genomic DNA of the cell (i.e., along with the integrated first donor polynucleotide) by Cas9-mediated HDR to remove the selectable marker expression cassette of the integrated first donor polynucleotide.

[0151] In certain embodiments, one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, and Cas9 in the kit are provided by a vector, for example, a plasmid or a viral vector.In another embodiment, the first guide RNA and Cas9 are provided by a single vector or multiple vectors.In another embodiment, the second donor polynucleotide, the second guide RNA, and Cas9 are provided by a single vector or multiple vectors.

[0152] D. Application The scarless genome editing method of the present invention finds many applications in basic research and development and regenerative medicine.This method can be used to introduce mutations (such as insertions, deletions, or substitutions) into any gene in the genomic DNA of a cell.For example, the method described herein can be used to inactivate genes in cells to determine the effect of gene knockout or to study the effect of known disease-causing mutations.Alternatively, the method described herein can be used to remove mutations, such as disease-causing mutations, from genes in the genomic DNA of a cell.

[0153] In particular, scarless genome editing as described herein can be used to develop cell lines with desired characteristics.This method can be used, for example, to develop genetic disease models in human ES / iPS cells, or to add differentiation marker genes to human ES / iPS cells.In addition, this method can be used to create transgenic animals, add reporter genes to cells at desired locations, or modify genomes to give cells desired properties, such as safety systems, enhanced efficacy, controllability, and / or improved graft survival.

[0154] In certain aspects, the method of the present invention can be used to improve, treat or prevent disease in individuals by genomic modification as described herein.For example, alleles may contribute to disease by increasing the individual's susceptibility to disease or by being a direct causative factor for disease.Therefore, by changing the sequence of alleles, disease can be improved, treated or prevented.The individual can be a mammal or other animal, preferably a human.

[0155] More than 3,000 diseases are caused by mutations, including sickle cell anemia, hemophilia, severe combined immunodeficiency (SCID), Tay-Sachs disease, Duchenne muscular dystrophy, Huntington's disease, alpha thalassaemia, and Lesch-Nyhan syndrome.Therefore, all such genetic diseases can benefit from the scarless genome editing described herein to correct the gene defect in cells.The method of the present invention is particularly suitable for diseases in which the cells corrected by genome modification have a significant selective advantage over mutant cells, but can also be useful for diseases in which the cells corrected by genome modification do not have a significant selective advantage over mutant cells.

[0156] In certain embodiments, this method can be used to modify the genome target sequence that makes the subject susceptible to infectious disease.For example, many viruses and bacterial pathogens enter cells by binding to and recruiting a set of cell surface protein and intracellular protein.Gene targeting can be used to eliminate or attenuate this binding site or entry mechanism.

[0157] Certain methods described herein may be applied to cells in vitro or ex vivo. Alternatively, the methods may be applied to a subject to effect genome modification in vivo. Donor polynucleotides, guide RNAs, and Cas9, or vectors encoding them, can be introduced into an individual using routes of administration generally known in the art (e.g., parenteral, mucosal, nasal, injection, systemic, implant, intraperitoneal, oral, intradermal, transdermal, intramuscular, intravenous including infusion and / or bolus injection, subcutaneous, topical, epidural, buccal, rectal, vaginal, etc.). In certain aspects, the donor polynucleotides, guide RNAs, Cas9, and vectors of the present invention can be formulated in combination with a suitable pharmaceutically acceptable carrier (excipient), such as saline, sterile water, dextrose, glycerol, ethanol, Ringer's solution, isotonic sodium chloride solution, and combinations thereof. The formulation must be appropriate for the mode of administration and is well within the skill of the art. The mode of administration is preferably the location of the target cells to be modified.

[0158] Donor polynucleotide, guide RNA, and Cas9, or the vectors that encode them, can be administered to individuals alone or together with other therapeutic agents.These various types of therapeutic agents can be administered in the same formulation or in separate formulations.The dosage of donor polynucleotide, guide RNA, and Cas9, or the vectors that encode them, that is administered to individuals, including the frequency of administration, will vary depending on various factors, including the mode and route of administration; the size, age, sex, health, weight, and diet of recipient; the nature and severity of the symptoms of the disease or disorder being treated; the type of concomitant treatment, frequency of treatment, and the desired effect; the nature of the formulation; and the judgment of the attending physician.These variations in dosage level can be adjusted using standard empirical routines for optimization, as is well understood in the art. [Example]

[0159] III. Experiment Below are examples of specific modes for carrying out the present invention. The examples are provided for illustrative purposes only and are not intended to limit the scope of the invention in any way.

[0160] Efforts have been made to ensure accuracy with respect to numbers used (eg, amounts, temperature, etc.), but some experimental error and deviation should, of course, be allowed for.

[0161] Example 1 Highly efficient scarless genome editing in human pluripotent stem cells In our method, we applied a two-step HR approach, which allows scarless editing on either one or both alleles without creating INDELs while generating either single nucleotide changes or insertion of larger DNA fragments (such as reporter genes), to four genes in both human embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs).

[0162] To overcome the limitations of previous scarless editing methods, we used a two-step HR strategy using the Cas9 / gRNA system in combination with positive-negative selection using magnetic beads (Figures 1A-1C). As proof of concept, we introduced a point mutation in TBX1 at nucleotide position 928 of the cDNA (c.928 G>A), which causes 22q11.2 deletion syndrome, also known as DiGeorge syndrome, into human ESCs (Yagi et al., Lancet 362, 1366-1373 (2003)). Because there is no TTAA site within 1 kb of TBX1 c.928 G, editing of this locus is not possible using the piggyBac system. For our method, we transfected hESCs with a plasmid co-expressing S. pyogenes Cas9 (Cas9) and a guide RNA (gRNA) targeting near TBX1 c.928 G (Figure 6A). In addition to the Cas9 / gRNA plasmid, we also transfected a plasmid carrying donor DNA to be used as a repair template after DSBs introduced by the Cas9-gRNA complex. The donor template also contains a bicistronic cassette expressing mCherry and truncated CD19 (tCD19) under the control of the human UbC promoter to serve as markers between the left and right homology arms (Figure 1B). Because mCherry and tCD19 are expressed episomally, transient expression of mCherry is observed, decreasing to 2.1% within 6 days after transfection (Figure 2A). Therefore, 6 days after transfection, we purified tCD19-positive cells by magnetic-assisted cell separation (MACS) selection using anti-CD19 conjugated to magnetic beads (Figure 2B). The cells were then cultured at low density (100 cells / cm) for single-cell cloning. 2) were plated. After an additional 8 days in culture, colonies with bright or dim mCherry fluorescence were observed (Figure 2C). Genotyping of TBX1 in these fluorescent cells revealed that 9 of 11 clones with bright mCherry expression possessed biallelic editing in the intended TBX1 site (Figure 2D). We also picked 3 clones with dim mCherry expression, all of which showed monoallelic insertion of our marker at the intended locus. For further confirmation, we verified the copy number of the targeting vector inserted into the genome by ddPCR quantification of the UbC promoter. All bright colonies contained four copies of the UbC promoter (derived from the endogenous two copies of the UbC promoter at the ubiquitin locus plus the exogenous two copies at our targeted locus, Figure 2E). This indicates that 9 of 11 clones were correctly targeted without random integration. In line with our hypothesis, all dim colonies contained three copies of the UbC promoter, derived from the two endogenous copies plus one from a single allele editing event within TBX1.

[0163] To demonstrate the ability of our method to generate cells with scarless biallelic and monoallelic editing after the first round of MACS selection, we performed a second round of HR on biallelic edited clone #2 and monoallelic edited clone #13. This was performed as before by transfecting a plasmid expressing Cas9 together with a custom gRNA that targets only the inserted marker (Guide 2, Figure 6A). Guide 2 was designed for high specificity based on bioinformatics analysis, targeting only the junction between the marker gene and the homology arm and therefore has no endogenous target site in the human genome. We co-transfected a plasmid carrying a donor template containing the TBX1 c.G928 G>A mutation along with the homology arm, which would result in a scarless, markerless editing event (Figure 1B). Cells carrying this second edit were purified by MACS negative selection 7 days after transfection. Using negative selection, we expanded the population of tCD19-negative cells to greater than 98% (Figure 2F). We then used single-cell cloning from the tCD19-negative population to identify marker-free cells that had undergone biallelic editing events. Analysis of the clones revealed seven TBX1 c.G928 G>A biallelic-edited lines from seven distinct clones derived from clone #2, and 12 monoallelic-edited lines from 16 clones derived from clone #13 (Figures 2G, 2H). All 16 clones derived from clone #13 underwent a second round of HR with monoallelic editing events, but afterward, four of the 16 clones possessed only the WT allele. We hypothesized that this loss of monoallelic editing events was due to an HR event using the WT homologous chromosome (not the sister chromatid) following the second Cas9-mediated DSB. Supporting this idea, a single clone was found with a homozygous single-base insertion after the second round of HR, which could not have resulted from contaminating non-edited cells (Figure 7).In addition, a second round of HR in clone 2 using a mixture of WT and TBX1 c.928 G>A donors yielded WT, heterozygous, and homozygous mutants (Figure 2G). Therefore, WT, heterozygous, and homozygous scarless clones in an isogenic background can be easily obtained for appropriate comparison when disease modeling is performed in genetically modified animals, such as litter comparisons. After the second round of HR and subsequent negative selection, we confirmed that the edited stem cells retained pluripotency markers (Figure 8).

[0164] We then applied the same approach to generate a RUNX1 reporter iPS cell line by integrating the mOrange gene followed by a 2A self-cleaving peptide sequence immediately before the stop codon of the RUNX1 gene. + To enable imaging analysis of directed differentiation into hematopoietic stem and progenitor cells, as well as RUNX1 in PSCs +The goal of this study was to generate an isogenic scarless RUNX1 reporter for screening culture conditions / small molecules that enhance differentiation into PSCs and for CRISPR- or shRNA-based genetic screening (Figure 3A). Because the RUNX1 gene is not expressed in PSCs, obtaining edited lines using conventional editing methods is not straightforward; unless marker selection is used, hundreds of colonies must usually be picked and analyzed. Through the first round of editing, we inserted an mCherry-tCD19 expression cassette 5' to the RUNX1 stop codon using the Cas9 / gRNA and donor plasmid shown in Figures 6B and 7. After MACS enrichment of tCD19-positive cells and single-cell culture, we collected eight bright clones. PCR-based genotyping showed that the marker was biallelically inserted at the targeted locus in clones 2 and 8 (Figure 3B), and ddPCR-based copy number analysis indicated that these clones had four copies of the UbC promoter. We then subjected clone 2 to a second round of editing using Cas9 / gRNA and a donor vector, as illustrated in Figure 3A. Subsequently, tCD19-negative cells were enriched by MACS separation, followed by single-cell cloning, resulting in the desired lineages carrying 2A-mOrange biallelically immediately before the RUNX1 stop codon (clones 1, 3, 5, 6, and 7, Figure 3C). Thus, we generated PSC lines into which a reporter for RUNX1 was introduced without disrupting the endogenous RUNX1 gene. To confirm whether the incorporated mOrange functioned as a reporter for RUNX1, we differentiated this iPS cell line into hematopoietic stem progenitor cells (HSPCs) (Nishimura et al. Cell Stem Cell 12, 114-126 (2013)). After 13 days of coculture with irradiated C3H10T1 / 2 cells, mOrange-positive HSPC-like cells were detected ( Fig. 3D ).FACS analysis shows that the CD34-positive and CD45-intermediately positive population, primarily containing HSPCs, expresses more bright mOrange than other populations containing undifferentiated or other cell types (Figure 3E). These results demonstrate that the mOrange reporter cells retain their differentiation potential and accurately report the expression of RUNX1, a non-cell surface transcription factor that regulates lineage differentiation. We also performed additional editing at the GFI1 locus using this RUNX1-mOrange line to generate a RUNX1-GFI1 double reporter line using the same editing strategy (Figure 9).

[0165] For both RUNX1 and GFI1 reporter integration, the frequency of biallelic targeting without random integration in the first editing step occurred in less than 30% of clones. Because the majority of clones with random integration had sequences derived from the plasmid backbone (Figure 10), we hypothesized that introducing a negative selection feature in the plasmid backbone would effectively reduce the enrichment of cells with random integration in the first round of editing. Instead of using the widely used thymidine kinase selection method, which is limited by the need for drug selection, we developed a novel negative selection method based on shRNA-based counterselection. In this system, an shRNA cassette is flanked by both homology arms and designed to suppress a positive marker selection gene. As long as the shRNA is expressed (either from episomal expression or random integration), the marker gene is suppressed and cells will not be scored as positive. Once episomal expression is lost, marker gene suppression is relieved, and positive marker clones in which the cassette has integrated only in a targeted manner can be easily identified. We have applied this system to the B2M-HLA-A subpopulation of the B2M locus in the ES H9 strain. * HLA-A was added with a linker 5' to the B2M stop codon to create the 24 fusion protein. *We applied this method to editing with a donor vector designed to integrate 24 cDNAs. As shown in Figure 4B, shGFP expression cassettes (0.3 kb each) were introduced into the outside of both the left and right homology arms of the first donor vector, and the selection markers were changed from mCherry and tCD19 to GFP and tCD8. As a result, four biallelic marker-integrated clones without random integration were obtained from five bright clones (Figure 4B). On the other hand, only one biallelic targeting clone was obtained from six bright clones using the donor vector without the shRNA selection system (Figure 4A). This indicated that the shRNA-based counterselection effectively reduced the contamination of cells with random integration during positive selection from the first round of editing. We were able to identify such clones without the need for drug selection, which made the system easier and less toxic to the remaining clones.

[0166] In addition, we found that the introduction of mutations (GG>AA) at positions 314 and 315 upstream from the ATG, which generate a new EcoRI site, does not affect human UbC promoter activity. This property allows copy number analysis of the new exogenous UbC promoter driving the selectable marker by regular PCR without the need for quantitative PCR strategies such as ddPCR (Figure 11). We then confirmed that a second round of editing combined with negative selection yielded eight correctly edited clones from the eight enriched clones (Figure 4C), demonstrating the desired HLA expression profile (knockout of endogenous HLA and HLA-A). *Expression of 24 genes was observed in the edited cells (Figure 4D). We also compared genomic copy number variation (CNV) between the original lines and the first- and / or second-round editing products in all editing experiments by high-density SNP array analysis. While there were no substantial chromosomal changes (deletions or gains) caused by the two-step editing process, we found copy-neutral loss of heterozygosity (LOH) in one of the TBX1 lines after the second round of editing and one line with integrated RUNX1 markers at each targeted locus, suggesting that these changes were caused by gRNA / Cas9-driven double-strand breaks. However, most lines after the two-step editing process did not have such changes. Assuming these changes were induced by gRNA / Cas9 cleavage, it is likely that such changes also occur using other editing strategies using engineered nucleases. G-banding analysis, which is often performed to examine genome integrity, cannot detect copy-neutral LOH and may therefore have been overlooked in previous studies that did not use high-density SNP array analysis. CNV analysis using high-density SNP arrays around the on-target site is particularly recommended for functional analysis or clinical cell therapy using genome-edited cells.

[0167] These results demonstrate that the method can be used to introduce either single nucleotide changes or large DNA fragments in a precise and scarless manner. In contrast to other methods, our method offers three key advantages in scarless genome editing. First, our method generates well-matched isogenic control (WT), homozygous, and heterozygous lines from the same clone (as shown in TBX1 editing). Because the editing process requires long-term culture, which may alter the phenotype of hPSCs, phenotypic comparisons between the original and edited lines can be misleading due to the acquired changes. Our strategy takes advantage of the extremely high targeting efficiency in the second editing step, making it easy to generate various genotypes (WT, heterozygous, homozygous) from the same clone at the same stage, thus reducing the variability induced by long-term culture. Second, our method allows for the scarless introduction of larger transgenes, including those into genes not expressed in hPSCs, with insert sizes that are not feasible for integration using the ssODN method. Other methods have achieved integration of reporter genes, for example, but only into genes expressed in hPSCs. Our editing of RUNX1 and GFI1 is an example of scarless marker integration in genes not expressed in hPSCs. Third, our method works efficiently at loci with long cut-to-edit distances, a feature not previously achieved using other methods. For example, for the B2M locus, we achieved 80% targeted integration of our marker with a cut-to-edit distance of 63 base pairs. This extension of the cut-to-edit distance significantly increases the range of scarless editing that can be achieved in hPSCs (Figure 12).The use of marker selection and subsequent marker removal increased the efficiency of both the first and second editing stages, thus creating a system that required analyzing only 5-10 clones instead of the hundreds required with other methods (Table 2) (Paquet et al., supra; Miyaoka et al. Nat. Methods 11, 291-3 (2014)). We additionally improved the efficiency of the first-stage editing process by including an shRNA cassette against a fluorescent marker protein, thereby reducing or even eliminating the collection of clones with random integration. For the second round of editing, the targeting frequency in the enriched marker cells can be, and consistently approaches, 100%, thus eliminating the laborious and tedious screening and analysis of hundreds of clones. This high frequency was previously achieved using a fluorescent marker, also used in Cre-loxP-based two-step editing (Xi et al. Genome Biol. 16, 1-17 (2015)), but it allows our novel two-step method to work without any failures to date. Additionally, negative selection, unlike positive selection, which identifies cells with both targeted and random integration, enriches for only targeted cells in the first step, contributing to 100% targeting frequency in the second editing step. This streamlined process resulted in the identification of heterozygous and homozygous clones in 6–8 weeks and was achieved at loci not expressed in human PSCs, demonstrating that it does not require expressed target sites for its speed and efficiency. This high efficiency eliminates the need for labor-intensive colony picking, genotyping, and sequencing required by alternative methods. In summary, our method generates human pluripotent cells with scarless genome editing, thus addressing multiple obstacles in advancing stem cell technology as disease and developmental models, drug screening tools, and therapeutics.

[0168] material and method Plasmid The sgRNA expression vector was constructed by inserting annealed oligonucleotides containing the target and adapter sequences into BbsI-digested px330 (Addgene plasmid no. 42230), which contains a human codon-optimized SpCas9 expression cassette and a human U6 promoter driving expression of the chimeric sgRNA. The target site is depicted in Figure 6.

[0169] The donor DNA plasmid vector was constructed using NEBuilder HiFi DNA Assembly (NEB). The left and right homology arms were amplified by nested PCR using genomic DNA extracts from K562 cells (ATCC) as templates. TBX1 mutations were generated by PCR and ligation.

[0170] cell culture Human ES H9 cells were used for TBX1 and B2M editing, and the TkDA3-4 iPSC line (Cell Applications Inc.), established from human skin fibroblasts as previously described (Takayama et al. J Exp Med. 207, 2817-2830 (2010)), was used for RUNX1 and GFI1 editing. The iAM9 iPSC line was transfected with HLA-A * This was used as a positive control for the detection of 24. hPSCs were maintained in mTeSR1 (STEMCELL technologies) on feeder-free Matrigel (Corning)-coated plates. Subculture was performed every 4–6 days using the EDTA method. After plating, 10 μM Y-27632 (Tocris) was added for 1 day.

[0171] K562 (ATCC) cells were maintained in RPMI 1640 (HyClone) supplemented with 10% bovine growth serum (HyClone), 100 mg / ml streptomycin, 100 units / ml penicillin, and 2 mM l -glutamine.

[0172] Transfection iPSCs (70-80% confluent) were harvested with Accutase (Life Technologies). 2 × 10 6 Cells were electroporated with 5 μg pX330 plasmid and 5 μg donor vector using the P3 Primary Cell 4D-Nucleofector L kit and 4D-Nucleofector system (Lonza) according to the manufacturer's protocol (program: CB-150 for TBX1 and RUNX1 editing, CA-137 for GFI1 and B2M editing). For transfections with a mixture of donor DNA (WT and TBX1 c.928G>A), 2.5 μg of each was used. After transfection, cells were plated in triplicate in Matrigel-coated 6-well plates and maintained in mTeSR1 supplemented with 10 μM Y-27632 for 3 days. Cells were then maintained in mTeSR without Y-27632.

[0173] K562 cells were nucleofected using Lonza Nucleofector 2b (program T-016) and nucleofection buffer (pH 7.4) containing 100 mM KH2PO4, 15 mM NaHCO3, 12 mM MgCl2 × 6H2O, 8 mM ATP, and 2 mM glucose.

[0174] HSPC differentiation using iPSCs iPSC differentiation into HSPCs was performed as previously reported with minor modifications. Briefly, small clumps of iPSCs (less than 100 cells) were transferred onto irradiated C3H10T1 / 2 cells and co-cultured in EB medium (Iscove's modified Dulbecco's medium supplemented with 15% fetal bovine serum and a cocktail of 10 μg / ml human insulin, 5.5 μg / ml human transferrin, 5 ng / ml sodium selenite, 2 mM L-glutamine, 0.45 μM α-monothioglycerol, and 50 μg / ml ascorbic acid) in the presence of VEGF, SCF, TPO, IL-3, and IL-6. The medium was changed every 3 days. After 14 days of culture, cells were harvested by treatment with TrypLE Select (Thermo Fisher Scientific) for 5 minutes at 37°C, and then subjected to FACS analysis.

[0175] Flow cytometry and fluorescence microscopy For measuring editing frequency, after harvesting with Accutase, cells were stained with anti-human CD19 antibody conjugated with APC (clone LT19, 1:20, Miltenyi Biotec) according to the manufacturer's protocol. Cells were then washed twice with washing buffer (PBS containing 1% human albumin and 0.5 mM EDTA). For HLA detection, cells were detached 3 days after 50 ng / mL INFγ treatment and stained with anti-HLA-A antibody conjugated with APC. * 02 (clone BB7.2, 1:20, eBioscience), anti-HLA-A conjugated with PE * 03 (clone GAP. A3, 1:20, eBioscience), or anti-HLA-A conjugated with FITC *Cells were stained with 24 (clone 22E1, 1:20, MBL) according to the manufacturer's protocol. Cells were then washed twice with wash buffer. Data were acquired using an Accuri C6 plus flow cytometer (BD Biosciences). For differentiation studies, propidium iodide was added after detachment to allow for the exclusion of dead cells. Cells were stained with anti-CD34 conjugated with APC-Cy7 (eBioscience) and anti-CD45 conjugated with APC (eBioscience) for 30 minutes at 4°C, and then the cells were washed twice with wash buffer. Data were acquired using a FACSAria II (BD Biosciences). All acquired data were analyzed using FlowJo (FlowJo, LLC).

[0176] To distinguish between clones with bright and dim mCherry fluorescence before colony picking, cells were observed using a fluorescence microscope IX-70 (Olympus). All fluorescence images were acquired using an EVOS FL cell imaging system (Thermo Fisher Scientific).

[0177] Off-target analysis The gene-edited ES / iPS cell lines were tested for predicted off-target editing events for each sgRNA using the COSMID46 tool (http: / / crispr.bme.gatech.edu), which considers mismatches, insertions, and deletions in the guide RNA target sequence. The results are shown in Table 2. All sequences at loci predicted as off-target candidates were examined in the final product, and no off-target editing was detected. No off-target candidates were predicted for the guide RNA sequences used for the first round of editing of TBX1 and B2M, or for all second round editing, using the default settings of the COSMID46 tool.

[0178] Magnetic Assisted Cell Separation (MACS) One hour before isolation, cells were treated with 10 μM Y-27632, and this treatment was maintained throughout isolation by adding Y-27632 to the wash buffer (PBS containing 1% human albumin and 0.5 μM EDTA). After detaching the cells with Accutase, positive and negative selections were performed on days 6 and 7 posttransfection, respectively, using MS and LD columns (Miltenyi Biotec) and magnetic bead-conjugated anti-human CD19 (Miltenyi Biotec) or anti-human CD8 (Miltenyi Biotec) according to the manufacturer's protocol. Positive selection was repeated twice using two MS columns in succession.

[0179] Genotyping and sequence analysis Marker insertion into the TBX1 locus was confirmed by PCR with Accuprime GC-rich DNA Polymerase (Thermo Fisher Scientific). Sequence analysis was performed at McLab (South San Francisco, CA, USA) using PCR amplicons generated by PrimeSTAR HS with GC-rich (Takara Bio). Genotyping of RUNX1, GFI1, and B2M was performed with PrimeSTAR GXL (Takara Bio). All primers are listed in Table 1. Genomic DNA samples were prepared with QuickExtract DNA Extraction Solution (Epicentre Madison) according to the manufacturer's instructions.

[0180] ddPCR-based copy number analysis Copy number analysis of the integrated exogenous UbC promoter derived from the selectable marker was performed by droplet digital PCR (ddPCR) to quantify the genomic UbC promoter and the reference sequence. The copy number of the UbC promoter reflects the sum of the two endogenous copies and the copy derived from the exogenous donor DNA, while only two copies of the reference sequence are present. Therefore, comparison of the amounts between the UbC promoter and the reference sequence provides the net copy number of the marker allele. ddPCR was performed using a QX200 (BioRad) and ddPCR Supermix for Probe (BioRad) according to the manufacturer's protocol. The primer and probe sequences were as follows: for the UbC promoter, primer 1: TIFF2026012842000002.tif5128, Primer 2 TIFF2026012842000003.tif4128, probe TIFF2026012842000004.tif4128, for the reference locus, primer 1 TIFF2026012842000005.tif4128, Primer 2 TIFF2026012842000006.tif4128, probe TIFF2026012842000007.tif4128. Genomic DNA samples were prepared using QuickExtract DNA Extraction Solution.

[0181] Alternative copy number analysis of the UbC promoter by standard PCR and restriction enzyme digestion PCR reactions were performed at a 15 μL scale using Phusion Green Hot Start II High-Fidelity PCR Master Mix (Thermo Fisher Scientific) with genomic DNA samples and primers. The PCR was performed using TIFF2026012842000008.tif13128. Then, 1 μL of EcoRI-HF (NEB) diluted to 10 U / μL with 1× CutSmart buffer was added directly to the PCR product. After 1.5 hours of incubation at 37°C, the sample was subjected to electrophoresis (30 minutes, 150 V) on a 2.0% agarose gel containing Midori Green Advance (Fast Gene). Images were captured using a ChemiDoc XRS+ (Bio-Rad) and then analyzed using Image J software.

[0182] Immunocytochemistry and microscopy For NANOG and OCT4 staining, cells were fixed in 4% paraformaldehyde, permeabilized in 0.2% Triton X-100 in PBS, blocked with blocking buffer (0.1% Triton-X and 2% FBS in PBS), and stained overnight at room temperature with antibodies diluted 1:200 in blocking buffer. Cells were then stained for 40 minutes at room temperature with an anti-rabbit IgG antibody conjugated with Alexa594 (R37117, Thermo Fisher Scientific) diluted 1:2000 in blocking buffer. For TRA-1-60 staining, cells were fixed in 4% paraformaldehyde, blocked with 2% FBS in PBS, and stained with antibodies diluted 1:200 in 2% FBS in PBS. The antibodies used in this study were obtained from STEMGENT and their catalog numbers are NANOG (catalog number 09-0020), OCT4 (catalog number 09-0023), and TRA-1-60 (catalog number 09-0068). Fluorescent images were acquired using an EVOS FL cell imaging system.

[0183] High-density SNP array analysis To examine genome integrity during the editing process, high-density SNP array analysis was performed using the CytoSNP-850K BeadChip (Illumina). 200 ng of genomic DNA samples prepared using the GeneJET Genomic DNA Purification Kit (Thermo Fisher Scientific) were processed and hybridized to microarray slides according to the manufacturer's protocol. They were then scanned using the iScan system (Illumina). The resulting data were analyzed using GenomeStudio software (Illumina) with the cnvPartition algorithm.

[0184] While the preferred embodiment of the invention has been illustrated and described, it will be recognized that various changes can be made therein without departing from the spirit and scope of the invention.

[0185] Table 1. Primer list for genotyping and sequencing TIFF2026012842000009.tif152139

[0186] Table 2. Off-target candidates predicted by the COSMID tool TIFF2026012842000010.tif23168 Mismatches and INDELs are underlined.

[0187] Sequence information SEQUENCE LISTING <110> The Board of Trustees of the Leland Stanford Junior University <120> SCARLESS GENOME EDITING THROUGH TWO-STEP HOMOLOGY DIRECTED REPAIR <150> US 62 / 533,780 <151> 2017-07-18 <160> 33 <170> PatentIn version 3.5 <210> 1 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> primer 1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 1 cgtcagtttc tttggtcggt 20 <210> 2 <211> 18 <212> DNA <213> Artificial Sequence <220> <223> primer 2 <220> <221> source <222> (1)..(18) <223> sequence is synthesized <400> 2 aaacacactc gccaaccc 18 <210> 3 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> probe <220> <221> source <222> (1)..(26) <223> sequence is synthesized <400> 3 tcttcttaag tagctgaagc tccggt 26 <210> 4 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> UbC promoter primer 1 <220> <221> source <222> (1)..(21) <223> sequence is synthesized <400> 4 ctctcctctt tgatacggcc c 21 <210> 5 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> UbC promoter primer 2 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 5 agtgttgtcc cagacagtgc 20 <210> 6 <211> 26 <212> DNA <213> Artificial Sequence <220> <223> UbC promoter probe <220> <221> source <222> (1)..(26) <223> sequence is synthesized <400> 6 ctgccaagtt gtggcctctg tcaaag 26 <210> 7 <211> 7151 <212> DNA <213> Artificial Sequence <220> <223> Tbx1 1st donor vector <400> 7 ccaatgctta atcagtgagg cacctatctc agcgatctgt ctatttcgtt catccatagt 60 tgcctgactc cccgtcgtgt agataactac gatacgggag ggcttaccat ctggccccag 120 tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag caataaacca 180 gccagccgga agggccgagc gcagaagtgg tctgcaact ttatccgcct ccatccagtc 240 tattaattgt tgccgggaag ctagagtaag tagttcgcca gttaatagtt tgcgcaacgt 300 tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg cttcattcag 360 ctccggttcc caacgatcaa ggcgagttac atgatccccc atgttgtgca aaaaagcggt 420 tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt tatcactcat 480 ggttatggca gcactgcata attctcttac tgtcatgcca tccgtaagat gcttttctgt 540 gactggtgag tactcaacca agtcattctg agaatagtgt atgcggcgac cgagttgctc 600 ttgcccggcg tcaatacggg ataataccgc gccacatagc agaactttaa aagtgctcat 660 cattggaaaa cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt tgagatccag 720 ttcgatgtaa cccactcgtg cacccaactg atcttcagca tcttttactt tcaccagcgt 780 ttctgggtga gcaaaaacag gaaggcaaaa tgccgcaaaa aagggaataa gggcgacacg 840 gaaatgttga atactcatac tcttcctttt tcaatattat tgaagcattt atcagggtta 900 ttgtctcatg agcggataca tatttgaatg tatttagaaa aataaacaaa taggggttcc 960 gcgcacattt ccccgaaaag tgccacctga cgttaactcg agtcgacggc ccagtgaccc 1020 agcctcatct tggaattaag ggttttgccc aactcatcca ggaaactcat tgccaactca 1080 gacctcagcc catttcctgg ctcccacccc agatcctcag cccagcccca ccgctggagc 1140 tgattcccca ccttgtcttc cagattattc tgaattccat gcacagatac cagccccgct 1200 tccacgtggt ctatgtggac ccacgcaaag atagcgagaa atatgccgag gagaacttca 1260 aaacctttgt gttcgaggag acacgattca ccgcggtcac tgcctaccag aaccatcggg 1320 tgagggcctg tggggaggac ctgagcggat tcaacgcctc tggaaaagcg ggtgtaattt 1380 tcagttgccg tttggggaca gtgggtccgc ttagacctgc aggctgtggt cccagtggag 1440 cccaacccaa ctggagcccc actcccaagg gcctcaggca gccccctccc tctcgaggct 1500 ggccggccca gcctcctatc agcttgacct ctccagcggc aactgtcact tcgtcctgaa 1560 agtttgtttt ccgaaccatt ccggaaactc cccatcaggg gcctgatctg aggtttaccc 1620 agattactag ggaacccgct ctgttcccca ccccccaccc cactgcacgt ggggggtggt 1680 gaccacattc ctgtcccagc gaggagcaca gggcctccat ccccacccac ctgggggaca 1740 ccagagaggg gttccctagt gagagaggag gttcctcaga cccccgcccc cctgcaggag 1800 ggagcaccag ctccgtagag gaggggcaga cgtggactgg ttcttgtcag ggcagcagaa 1860 aggcccttgg tgcgcttctc ctaacactcc cctatcctcc gccgaggtcg ggtggcccag 1920 gctgcagggc tccagcggct tgctcacacc cacctccctg cagatcacgc agctcaagat 1980 tgccagcaat cccttcgcga aaggcttccg ggactgtgac cctgaggact ggtgagtgtc 2040 ctcccccgag agagtgagcg ccgggcgcct ggcgcaggcg ccgccctgat ccgcctcccg 2100 cccgcaggcc ccggaaccac cggcccagcg cactgccgct catcgaaggt ctagactccg 2160 cgccgggttt tggcgcctcc cgcgggcgcc cccctcctca cggcgagcgc tgccacgtca 2220 gacgaagggc gcagcgagcg tcctgatcct tccgcccgga cgctcaggac agcggcccgc 2280 tgctcataag actcggcctt agaaccccag tatcagcaga aggacatttt aggacgggac 2340 ttgggtgact ctagggcact ggttttcttt ccagagagcg gaacaggcga ggaaaagtag 2400 tcccttctcg gcgattctgc ggagggatct ccgtggggcg gtgaacgccg atgattatat 2460 aaggacgcgc cgggtgtggc acagctagtt ccgtcgcagc cgggatttgg gtcgcggttc 2520 ttgtttgtgg atcgctgtga tcgtcacttg gtgagtagcg ggctgctggg ctggccgggg 2580 ctttcgtggc cgccgggccg ctcggtggga cggaagcgtg tggagagacc gccaagggct 2640 gtagtctggg tccgcgagca aggttgccct gaactggggg ttggggggag cgcagcaaaa 2700 tggcggctgt tcccgagtct tgaatggaag acgcttgtga ggcgggctgt gaggtcgttg 2760 aaacaaggtg gggggcatgg tgggcggcaa gaacccaagg tcttgaggcc ttcgctaatg 2820 cgggaaagct cttattcggg tgagatgggc tggggcacca tctggggacc ctgacgtgaa 2880 gtttgtcact gactggagaa ctcggtttgt cgtctgttgc gggggcggca gttatggcgg 2940 tgccgttggg cagtgcaccc gtacctttgg gagcgcgcgc cctcgtcgtg tcgtgacgtc 3000 acccgttctg ttggcttata atgcagggtg gggccacctg ccggtaggtg tgcggtaggc 3060 ttttctccgt cgcaggacgc agtgttcggg cctagggtag gctctcctga atcgacaggc 3120 gccggacctc tggtgagggg agggataagt gaggcgtcag tttctttggt cggttttatg 3180 tacctatctt cttaagtagc tgaagctccg gttttgaact atgcgctcgg ggttggcgag 3240 tgtgttttgt gaagtttttt aggcaccttt tgaaatgtaa tcatttgggt caatatgtaa 3300 ttttcagtgt tagactagta aattgtccgc taaattctgg ccgtttttgg cttttttgtt 3360 agacggtacc gagctcttcg aaggatccat cgccaccatg gtgagcaagg gcgaggagga 3420 taacatggcc atcatcaagg agttcatgcg cttcaaggtg cacatggagg gctccgtgaa 3480 cggccacgag ttcgagatcg agggcgaggg cgagggccgc ccctacgagg gcacccagac 3540 cgccaagctg aaggtgacca agggtggccc cctgcccttc gcctgggaca tcctgtcccc 3600 tcagttcatg tacggctcca aggcctacgt gaagcacccc gccgacatcc ccgactactt 3660 gaagctgtcc ttccccgagg gcttcaagtg ggagcgcgtg atgaacttcg aggacggcgg 3720 cgtggtgacc gtgacccagg actcctccct gcaggacggc gagttcatct acaaggtgaa 3780 gctgcgcggc accaacttcc cctccgacgg ccccgtaatg cagaagaaga ccatgggctg 3840 ggaggcctcc tccgagcgga tgtaccccga ggacggcgcc ctgaagggcg agatcaagca 3900 gaggctgaag ctgaaggacg gcggccacta cgacgctgag gtcaagacca cctacaaggc 3960 caagaagccc gtgcagctgc ccggcgccta caacgtcaac atcaagttgg acatcacctc 4020 ccacaacgag gactacacca tcgtggaaca gtacgaacgc gccgagggcc gccactccac 4080 cggcggcatg gacgagctgt acaaggaggg caggggcagc ctgctgacct gcggcgacgt 4140 ggaggagaac cccggcccca tgcccccccc caggctgctg ttcttcctgc tgttcctgac 4200 ccccatggag gtgaggcccg aggagcccct ggtggtgaag gtggaggagg gcgacaacgc 4260 cgtgctgcag tgcctgaagg gcaccagcga cggccccacc cagcagctga cctggagcag 4320 ggagagcccc ctgaagccct tcctgaagct gagcctgggc ctgcccggcc tgggcatcca 4380 catgaggccc ctggccatct ggctgttcat cttcaacgtg agccagcaga tgggcggctt 4440 ctacctgtgc cagcccggcc cccccagcga gaaggcctgg cagcccggct ggaccgtgaa 4500 cgtggagggc agcggcgagc tgttcaggtg gaacgtgagc gacctgggcg gcctgggctg 4560 cggcctgaag aacaggagca gcgagggccc cagcagcccc agcggcaagc tgatgagccc 4620 caagctgtac gtgtgggcca aggacaggcc cgagatctgg gagggcgagc ccccctgcct 4680 gccccccagg gacagcctga accagagcct gagccaggac ctgaccatgg cccccggcag 4740 caccctgtgg ctgagctgcg gcgtgccccc cgacagcgtg agcaggggcc ccctgagctg 4800 gacccacgtg caccccaagg gccccaagag cctgctgagc ctggagctga aggacgacag 4860 gcccgccagg gacatgtggg tgatggagac cggcctgctg ctgcccaggg ccaccgccca 4920 ggacgccggc aagtactact gccacagggg caacctgacc atgagcttcc acctggagat 4980 caccgccagg cccgtgctgt ggcactggct gctgaggacc ggcggctgga aggtgagcgc 5040 cgtgaccctg gcctacctga tcttctgcct gtgcagcctg gtgggcatcc tgcacctgca 5100 gagggccctg gtgctgagga ggaagaggaa gaggatgacc gaccccacca ggaggttctg 5160 aggcgcgccc cgctgatcag cctcgactgt gccttctagt tgccagccat ctgttgtttg 5220 cccctccccc gtgccttcct tgaccctgga aggtgccact cccactgtcc tttcctaata 5280 aaatgaggaa attgcatcgc attgtctgag taggtgtcat tctattctgg ggggtggggt 5340 ggggcaggac agcaaggggg aggattggga agacaatagc aggcatgctg gggatgcggt 5400 gggctctatg gcgtacgcct cacgagcgcc ttcgcgcgct cgcggaaccc cgtggcttcc 5460 ccgacgcagc ccagcggcac ggagaaaggt agggccgggg tcgtgggatc cgggttccgg 5520 ccctgtgcgc gctctacccc gggccggcgg cctcgcccga cctcgcctgc gcccccgggg 5580 cgctccaggc tttcgcgccg gttgcacaac ggccgcggcg gcgggcaagc gcgcactcgc 5640 ccgcccggcc cgacggctgc gccccgcccg ccgccgccgc cgcccgcaga ggggcgcggg 5700 ccccggggag ggctcggggc gccggcgact tggggtctcg ggcacgctgg caccgactgg 5760 tcggggaaca ccgagggcgg ccaagagcct tctctccgcc agggcctcgc atggggcgtc 5820 ggagctcctc ggcggccccg gccggccgcg ctcactcctc ggccctctcc gcagacgcgg 5880 ctgaggcccg gcgagaattc cagcgcgacg cgggcgggcc agcagtgctc ggggacccgg 5940 cgcatcctcc gcagctgctg gcccgggtgc taagcccctc gctgcccggg gccggcggcg 6000 ccggcggctt agtcccgctg cccggcgcgc ccggaggccg gcccagtccc ccgaaccccg 6060 agctgcgcct ggaggcgccc ggcgcatcgg agccgctgca ccaccacccc tacaaatatc 6120 cggccgccgc ctacgaccac tatctcgggg ccaagagccg gccggcgccc tacccgctgc 6180 ccggcctgcg tggccacggc taccacccgc acgcgcatcc gcaccaccac caccaccccg 6240 tgagtccagc cgccgcggcc gccgccgccg ctgccgcagc tgccgcggcc gccaacaatt 6300 gtccggaatt ctgcagagtg gcggccgcac atgtgagcaa aaggccagca aaaggccagg 6360 aaccgtaaaa aggccgcgtt gctggcgttt ttccataggc tccgcccccc tgacgagcat 6420 cacaaaaatc gacgctcaag tcagaggtgg cgaaacccga caggactata aagataccag 6480 gcgtttcccc ctggaagctc cctcgtgcgc tctcctgttc cgaccctgcc gcttaccgga 6540 tacctgtccg cctttctccc ttcgggaagc gtggcgcttt ctcatagctc acgctgtagg 6600 tatctcagtt cggtgtaggt cgttcgctcc aagctgggct gtgtgcacga accccccgtt 6660 cagcccgacc gctgcgcctt atccggtaac tatcgtcttg agtccaaccc ggtaagacac 6720 gacttatcgc cactggcagc agccactggt aacaggatta gcagagcgag gtatgtaggc 6780 ggtgctacag agttcttgaa gtggtggcct aactacggct acactagaag aacagtattt 6840 ggtatctgcg ctctgctgaa gccagttacc ttcggaaaaa gagttggtag ctcttgatcc 6900 ggcaaacaaa ccaccgctgg tagcggtggt ttttttgttt gcaagcagca gattacgcgc 6960 agaaaaaaag gatctcaaga agatcctttg atcttttcta cggggtctga cgctcagtgg 7020 aacgaaaact cacgttaagg gattttggtc atgagattat caaaaaggat cttcacctag 7080 atccttttaa attaaaaatg aagttttaaa tcaatctaaa gtatatatga gtaaacttgg 7140 tctgacagtt a 7151 <210> 8 <211> 3872 <212> DNA <213> Artificial Sequence <220> <223> Tbx1 2nd donor vector <400> 8 ccaatgctta atcagtgagg cacctatctc agcgatctgt ctatttcgtt catccatagt 60 tgcctgactc cccgtcgtgt agataactac gatacgggag ggcttaccat ctggccccag 120 tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag caataaacca 180 gccagccgga agggccgagc gcagaagtgg tctgcaact ttatccgcct ccatccagtc 240 tattaattgt tgccgggaag ctagagtaag tagttcgcca gttaatagtt tgcgcaacgt 300 tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg cttcattcag 360 ctccggttcc caacgatcaa ggcgagttac atgatccccc atgttgtgca aaaaagcggt 420 tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt tatcactcat 480 ggttatggca gcactgcata attctcttac tgtcatgcca tccgtaagat gcttttctgt 540 gactggtgag tactcaacca agtcattctg agaatagtgt atgcggcgac cgagttgctc 600 660 cattggaaaa cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt tgagatccag 720 ttcgatgtaa cccactcgtg cacccaactg atcttcagca tcttttactt tcaccagcgt 780 ttctgggtga gcaaaaacag gaaggcaaaa tgccgcaaaa aagggaataa gggcgacacg 840 gaaatgttga atactcatac tcttccttt tcaatattat tgaagcattt atcagggtta 900 ttgtctcatg agcggataca tatttgaatg tatttagaaa aataaacaaa taggggttcc 960 gcgcacattt ccccgaaaag tgccacctga cgttaactcg agtcgacggc ccagtgaccc 1020 agcctcatct tggaattaag ggttttgcc aactcatcca ggaaactcat tgccaactca 1080 gacctcagcc catttcctgg ctcccacccc agatcctcag cccagcccca ccgctggagc 1140 tgattcccca ccttgtcttc cagattattc tgaattccat gcacagatac cagccccgct 1200 tccacgtggt ctatgtggac ccacgcaaag atagcgagaa atatgccgag gagaacttca 1260 aaacctttgt gttcgaggag acacgattca ccgcggtcac tgcctaccag aaccatcggg 1320 tgagggcctg tggggaggac ctgagcggat tcaacgcctc tggaaaagcg ggtgtaattt 1380 tcagttgccg tttggggaca gtgggtccgc ttagacctgc aggctgtggt cccagtggag 1440 cccaacccaa ctggagcccc actcccaagg gcctcaggca gccccctccc tctcgaggct 1500 ggccggccca gcctcctatc agcttgacct ctccagcggc aactgtcact tcgtcctgaa 1560 agtttgtttt ccgaaccatt ccggaaactc cccatcaggg gcctgatctg aggtttaccc 1620 agattactag ggaacccgct ctgttcccca ccccccaccc cactgcacgt ggggggtggt 1680 gaccacattc ctgtcccagc gaggagcaca gggcctccat ccccacccac ctgggggaca 1740 ccagagaggg gttccctagt gagagaggag gttcctcaga cccccgcccc cctgcaggag 1800 ggagcaccag ctccgtagag gaggggcaga cgtggactgg ttcttgtcag ggcagcagaa 1860 aggcccttgg tgcgcttctc ctaacactcc cctatcctcc gccgaggtcg ggtggcccag 1920 gctgcagggc tccagcggct tgctcacacc cacctccctg cagatcacgc agctcaagat 1980 tgccagcaat cccttcgcga aaggcttccg ggactgtgac cctgaggact ggtgagtgtc 2040 ctcccccgag agagtgagcg ccgggcgcct ggcgcaggcg ccgccctgat ccgcctcccg 2100 cccgcaggcc ccggaaccac cggcccgagc gcactgccgc tcatgagcgc cttcgcgcgc 2160 tcgcggaacc ccgtggcttc cccgacgcag cccagcggca cggagaaagg tagggccggg 2220 gtcgtgggat ccgggttccg gccctgtgcg cgctctaccc cgggccggcg gcctcgcccg 2280 acctcgcctg cgcccccggg gcgctccagg ctttcgcgcc ggttgcacaa cggccgcggc 2340 ggcgggcaag cgcgcactcg cccgcccggc ccgacggctg cgccccgccc gccgccgccg 2400 ccgcccgcag aggggcgcgg gccccgggga gggctcgggg cgccggcgac ttggggtctc 2460 gggcacgctg gcaccgactg gtcggggaac accgagggcg gccaagagcc ttctctccgc 2520 cagggcctcg catggggcgt cggagctcct cggcggcccc ggccggccgc gctcactcct 2580 cggccctctc cgcagacgcg gctgaggccc ggcgagaatt ccagcgcgac gcgggcgggc 2640 cagcagtgct cggggacccg gcgcatcctc cgcagctgct ggcccgggtg ctaagcccct 2700 cgctgcccgg ggccggcggc gccggcggct tagtcccgct gcccggcgcg cccggaggcc 2760 ggcccagtcc cccgaacccc gagctgcgcc tggaggcgcc cggcgcatcg gagccgctgc 2820 accaccaccc ctacaaatat ccggccgccg cctacgacca ctatctcggg gccaagagcc 2880 ggccggcgcc ctacccgctg cccggcctgc gtggccacgg ctaccacccg cacgcgcatc 2940 cgcaccacca ccaccacccc gtgagtccag ccgccgcggc cgccgccgcc gctgccgcag 3000 ctgccgcggc cgccaacaat tgtccggaat tctgcagagt ggcggccgca catgtgagca 3060 aaaggccagc aaaaggccag gaaccgtaaa aaggccgcgt tgctggcgtt tttccatagg 3120 ctccgccccc ctgacgagca tcacaaaaat cgacgctcaa gtcagaggtg gcgaaacccg 3180 acaggactat aaagatacca ggcgtttccc cctggaagct ccctcgtgcg ctctcctgtt 3240 ccgaccctgc cgcttaccgg atacctgtcc gcctttctcc cttcgggaag cgtggcgctt 3300 tctcatagct cacgctgtag gtatctcagt tcggtgtagg tcgttcgctc caagctgggc 3360 tgtgtgcacg aaccccccgt tcagcccgac cgctgcgcct tatccggtaa ctatcgtctt 3420 gagtccaacc cggtaagaca cgacttatcg ccactggcag cagccactgg taacaggatt 3480 agcagagcga ggtatgtagg cggtgctaca gagttcttga agtggtggcc taactacggc 3540 tacactagaa gaacagtatt tggtatctgc gctctgctga agccagttac cttcggaaaa 3600 agagttggta gctcttgatc cggcaaacaa accaccgctg gtagcggtgg ttttttgtt 3660 tgcaagcagc agattacgcg cagaaaaaaa ggatctcaag aagatccttt gatctttct 3720 acggggtctg acgctcagtg gaacgaaaac tcacgttaag ggattttggt catgagatta 3780 tcaaaaagga tcttcaccta attaaaaat gaagttttaa atcaatctaa 3840 agtatatatg agtaaacttg gtctgacagt ta 3872 <210> 9 <211> 7132 <212> DNA <213> Artificial Sequence <220> <223> Runx1 1st donor vector <400> 9 ccaatgctta atcagtgagg cacctatctc agcgatctgt ctatttcgtt catccatagt 60 tgcctgactc cccgtcgtgt agataactac gatacgggag ggcttaccat ctggccccag 120 tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag caataaacca 180 gccagccgga agggccgagc gcagaagtgg tcctgcaact ttatccgcct ccatccagtc 240 tattaattgt tgccgggaag ctagagtaag tagttcgcca gttaatagtt tgcgcaacgt 300 tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg cttcattcag 360 ctccggttcc caacgatcaa ggcgagttac atgatcccccc atgttgtgca aaaaagcggt 420 tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt tatcactcat 480 ggttatggca gcactgcata attctcttac tgtcatgcca tccgtaagat gcttttctgt 540 gactggtgag tactcaacca agtcattctg agaatagtgt atgcggcgac cgagttgctc 600 660 cattggaaaa cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt tgagatccag 720 ttcgatgtaa cccactcgtg cacccaactg atcttcagca tcttttactt tcaccagcgt 780 ttctgggtga gcaaaaacag gaaggcaaaa tgccgcaaaaa aagggaata gggcgacacg gaatgttga atactcatac tcttcctttt tcaatattat tgaagcattt atcagggtta ttgtctcatg agcggataca tatttgaatg tatttagaaa aataaacaaa taggggttcc gcgcacattt ccccgaaaag tgccacctga cgttaactcg agtcgacaag cttccaatct tcctgggcgg taattctga tagaaaccca tctccctgga cagtagcatc ctgggtgtcc cccgtcctcc tagcggtag catcctgggt ggtctccatc ctcccaggtg gtgtcatcct 1140 gggtggtctc cgtcctctca ggaggtggca tcctgagtgg tccccgacct cctgggcata gcatcatggg tagtccccat cctcttggga ggtgacatgc tgggtgatcc tcgtcatctc aggaggtggc atcctgggtg gtccctgtcc ccctgggtat agcatcctgg gtaatcctcg 1320 tcctcttggg agtagcatcc cgggtggtcc ccgtcctccc cagcagtagc atcctgggtg 1380 gcttcccatc ctcctaggcg gtatcatcct gggtagcccc ctggggcaga gggaagagct 1440 gtggcctccg caacctccta ctcacttccg ctccgttctc ttgcccgccc tgcagcggca 1500 cccgacctga cagcgttcag cgacccgcgc cagttccccg cgctgccctc catctccgac 1560 ccccgcatgc actatccagg cgccttcacc tactccccga cgccggtcac ctcgggcatc 1620 ggcatcggca tgtcggccat gggctcggcc acgcgctacc acacctacct gccgccgccc 1680 taccccggct cgtcgcaagc gcagggaggc ccgttccaag ccagctcgcc ctcctaccac 1740 ctgtactacg gcgcctcggc cggctcctac cagttctcca tggtgggcgg cgagcgctcg 1800 ccgccgcgca tcctgccgcc ctgcaccaac gcctccaccg gctccgcgct gctcaacccc 1860 agcctcccga accagagcga cgtggtggag gccgagggca gccacagcaa ctcccccacc 1920 aacatggcgc cctccgcgcg cctggaggag gccgtgtgga ggccctaccc tgcgatgagc 1980 ggcagtgcgc tccgcgccgg gttttggcgc ctcccgcggg cgcccccctc ctcacggcga 2040 gcgctgccac gtcagacgaa gggcgcagcg agcgtcctga tccttccgcc cggacgctca 2100 ggacagcggc ccgctgctca taagactcgg ccttagaacc ccagtatcag cagaaggaca 2160 ttttaggacg ggacttgggt gactctaggg cactggtttt ctttccagag agcggaacag 2220 gcgaggaaaa gtagtccctt ctcggcgatt ctgcggaggg atctccgtgg ggcggtgaac 2280 gccgatgatt atataaggac gcgccgggtg tggcacagct agttccgtcg cagccgggat 2340 ttgggtcgcg gttcttgttt gtggatcgct gtgatcgtca cttggtgagt agcgggctgc 2400 tgggctggcc ggggctttcg tggccgccgg gccgctcggt gggacggaag cgtgtggaga 2460 gaccgccaag ggctgtagtc tgggtccgcg agcaaggttg ccctgaactg ggggttgggg 2520 ggagcgcagc aaaatggcgg ctgttcccga gtcttgaatg gaagacgctt gtgaggcggg 2580 ctgtgaggtc gttgaaacaa ggtggggggc atggtgggcg gcaagaaccc aaggtcttga 2640 ggccttcgct aatgcgggaa agctcttatt cgggtgagat gggctggggc accatctggg 2700 gaccctgacg tgaagtttgt cactgactgg agaactcggt ttgtcgtctg ttgcgggggc 2760 ggcagttatg gcggtgccgt tgggcagtgc acccgtacct ttgggagcgc gcgccctcgt 2820 cgtgtcgtga cgtcacccgt tctgttggct tataatgcag ggtggggcca cctgccggta 2880 ggtgtgcggt aggctttct ccgtcgcagg acgcagtgtt cgggcctagg gtaggctctc 2940 ctgaatcgac aggcgccgga cctctggtga ggggagggat aagtgaggcg tcagtttctt tggtcggttt tatgtaccta tcttcttaag tagctgaagc tccggttttg aactatgcgc tcggggttgg cgagtgtgtt ttgtgaagtt ttttaggcac cttttgaaat gtaatcattt 3120. gggtcaatat gtaattttca gtgttagact agtaaattgt ccgctaaatt ctggccgttt ttggcttttt tgttagacgg taccgagctc ttcgaaggat ccatcgccac catggtgagc 3240 aagggcgagg aggataacat ggccatcatc aaggagttca tgcgcttcaa ggtgcacatg gagggctccg tgaacggcca cgagttcgag atcgagggcg agggcgaggg ccgcccctac 3360. 3420. gagggcaccc agaccgccaa gctgaaggtg accaagggtg gccccctgcc cttcgcctgg 3480. gacatcctgt cccctcagtt catgtacggc tccaaggcct acgtgaagca ccccgccgac atccccgact acttgaagct gtccttcccc gagggcttca agtgggagcg cgtgagc 3540. ttcgaggacg gcggcgtggt gaccgtgacc caggactcct ccctgcagga cggcgagttc 3600 atctacaagg tgaagctgcg cggcaccaac ttcccctccg acggccccgt aatgcagaag 3660 aagaccatgg gctgggaggc ctcctccgag cggatgtacc ccgaggacgg cgccctgaag 3720 ggcgagatca agcagaggct gaagctgaag gacggcggcc actacgacgc tgaggtcaag 3780 accacctaca aggccaagaa gcccgtgcag ctgcccggcg cctacaacgt caacatcaag 3840 ttggacatca cctcccacaa cgaggactac accatcgtgg aacagtacga acgcgccgag 3900 ggccgccact ccaccggcgg catggacgag ctgtacaagg agggcagggg cagcctgctg 3960 acctgcggcg acgtggagga gaaccccggc cccatgcccc cccccaggct gctgttcttc 4020 ctgctgttcc tgacccccat ggaggtgagg cccgaggagc ccctggtggt gaaggtggag 4080 gagggcgaca acgccgtgct gcagtgcctg aagggcacca gcgacggccc cacccagcag 4140 ctgacctgga gcagggagag ccccctgaag cccttcctga agctgagcct gggcctgccc 4200 ggcctgggca tccacatgag gcccctggcc atctggctgt tcatcttcaa cgtgagccag 4260 cagatgggcg gcttctacct gtgccagccc ggccccccca gcgagaaggc ctggcagccc 4320 ggctggaccg tgaacgtgga gggcagcggc gagctgttca ggtggaacgt gagcgacctg 4380 ggcggcctgg gctgcggcct gaagaacagg agcagcgagg gccccagcag ccccagcggc 4440 aagctgatga gccccaagct gtacgtgtgg gccaaggaca ggcccgagat ctgggagggc 4500 gagcccccct gcctgccccc cagggacagc ctgaaccaga gcctgagcca ggacctgacc 4560 atggcccccg gcagcaccct gtggctgagc tgcggcgtgc cccccgacag cgtgagcagg 4620 ggccccctga gctggaccca cgtgcacccc aagggcccca agagcctgct gagcctggag 4680 ctgaaggacg acaggcccgc cagggacatg tgggtgatgg agaccggcct gctgctgccc 4740 agggccaccg cccaggacgc cggcaagtac tactgccaca ggggcaacct gaccatgagc 4800 ttccacctgg agatcaccgc caggcccgtg ctgtggcact ggctgctgag gaccggcggc 4860 tggaaggtga gcgccgtgac cctggcctac ctgatcttct gcctgtgcag cctggtgggc 4920 atcctgcacc tgcagagggc cctggtgctg aggaggaaga ggagaggat gaccgacccc 4980 accaggaggt tctgaggcgc gccccgctga tcagcctcga ctgtgccttc tagttgccag 5040 ccatctgttg tttgcccctc ccccgtgcct tccttgaccc tggaaggtgc cactcccact 5100. gtcctttcct aataaatga ggaaattgca tcgcattgtc tgagtaggtg tcattctatt ctggggggtg gggtggggca ggacagcaag ggggaggatt gggagacaa tagcaggcat 5220. gctggggatg cggtgggctc tatggcgtac gccttgaggc gccaggcctg gcccggctgg 5280 gccacgcggg ccgccgcctt cgcctccggg cgcgcgggcc tcctgttcgc gacaagcccg 5340 ccgggatccc gggccctggg cccggccacc gtcctggggc cgaggcgcc cgacggccag5400 gatctcgctg taggtcaggc ccgcgcagcc tcctgcgccc agaagcccac gccgccgccg 5460 tctgctggcg ccccggccct cgcggaggtg tccgaggcga cgcacctcga gggtgtccgc 5520 cggccccagc acccagggga cgcgctgga agcaaacagg aagattcccg gaggaact gtgaatgctt ctgatttagc aatgctgtga ataaaaaga agtttata cccttgactt aactttttaa ccaagttgtt tattccaaag agtgtggaat tttggttggg gtggggggag aggagggatg caactcgccc tgtttggcat ctaattctta tttttaatt ttccgcacct 5760 tatcaattgc aaaatgcgta tttgcatttg ggtggttttt atttttatat acgtttatat 5820 aaatatatat aaattgagct tgcttctttc ttgctttgac catggaaaga aatatgattc 5880 ccttttcttt aagttttatt taactttct tttggacttt tgggtagttg ttttttttg 5940 ttttgttttg ttttttgag aaacagctac agctttgggt catttttaac tactgtattc 6000 ccacaaggaa tccccagata tttatgtatc ttgatgttca gacatttatg tgttgataat 6060 ttttaatta tttaaatgta cttatattaa gaaaaatatc aagtactaca ttttcttttg 6120 ttcttgatag tagccaaagt taaatgtatc acattgaaga aggctagaaa aaaagaatga 6180 gtaatgtgat cgcttggtta tccagaagta ttgtttacat taaactccct ttcatgttaa 6240 tcaaacaagt gagtagctca cgcatgacgt acgcgtcaat tgtccggaat tctgcagagt 6300 ggcggccgca catgtgagca aaaggccagc aaaaggccag gaaccgtaaa aaggccgcgt 6360 tgctggcgtt tttccatagg ctccgccccc ctgacgagca tcacaaaaat cgacgctcaa 6420 gtcagaggtg gcgaaacccg acaggactat aaagatacca ggcgtttccc cctggaagct 6480 cccgtgcg ctctcctgtt ccgaccctgc cgcttaccgg atacctgtcc gcctttctcc 6540 cttcgggaag cgtggcgctt tctcatagct cacgctgtag gtatctcagt tcggtgtagg 6600 tcgttcgctc caagctgggc tgtgtgcacg aaccccccgt tcagcccgac cgctgcgcct 6660 tatccggtaa ctatcgtctt gagtccaacc cggtaagaca cgacttatcg ccactggcag 6720 cagccactgg taacaggatt agcagagcga ggtatgtagg cggtgctaca gagttcttga 6780 agtggtggcc taactacggc tacactagaa gaacagtatt tggtatctgc gctctgctga 6840 agccagttac cttcggaaaa agagttggta gctcttgatc cggcaaacaa accaccgctg 6900 gtagcggtgg ttttttgtt tgcaagcagc agattacgcg cagaaaaaaa ggatctcaag 6960 aagatccttt gatcttttct acggggtctg acgctcagtg gaacgaaaac tcacgttaag 7020 ggattttggt catgagatta tcaaaaagga tcttcaccta gatcctttta attaaaaat 7080 gaagttttaa atcaatctaa agtatatatg agtaaacttg gtctgacagt ta 7132 <210> 10 <211> 4617 <212> DNA <213> Artificial Sequence <220> <223> Runx1 2nd donor vector <400> 10 ccaatgctta atcagtgagg cacctatctc agcgatctgt ctatttcgtt catccatagt 60 tgcctgactc cccgtcgtgt agataactac gatacgggag ggcttaccat ctggccccag 120 tgctgcaatg ataccgcgag acccacgctc accggctcca gatttatcag caataaacca 180 gccagccgga agggccgagc gcagaagtgg tctgcaact ttatccgcct ccatccagtc 240 tattaattgt tgccgggaag ctagagtaag tagttcgcca gttaatagtt tgcgcaacgt 300 tgttgccatt gctacaggca tcgtggtgtc acgctcgtcg tttggtatgg cttcattcag 360 ctccggttcc caacgatcaa ggcgagttac atgatccccc atgttgtgca aaaaagcggt 420 tagctccttc ggtcctccga tcgttgtcag aagtaagttg gccgcagtgt tatcactcat 480 ggttatggca gcactgcata attctcttac tgtcatgcca tccgtaagat gcttttctgt 540 gactggtgag tactcaacca agtcattctg agaatagtgt atgcggcgac cgagttgctc 600 ttgcccggcg tcaatacggg fatherccgc gccacatagc agaactttaa aagtgctcat 660. cattggaaaa cgttcttcgg ggcgaaaact ctcaaggatc ttaccgctgt tgagatccag ttcgatgtaa cccactcgtg cacccaactg atcttcagca tcttttactt tcaccagcgt 780 ttctgggtga gcaaaaacag gaaggcaaaa tgccgcaaaaa aagggaata gggcgacacg gaatgttga atactcatac tcttcctttt tcaatattat tgaagcattt atcagggtta ttgtctcatg agcggataca tatttgaatg tatttagaaa aataaacaaa taggggttcc gcgcacattt ccccgaaaag tgccacctga cgttaactcg agtcgacaag cttccaatct tcctgggcgg taattctga tagaaaccca tctccctgga cagtagcatc ctgggtgtcc cccgtcctcc tagcggtag catcctgggt ggtctccatc ctcccaggtg gtgtcatcct 1140 gggtggtctc cgtcctctca ggaggtggca tcctgagtgg tccccgacct cctgggcata gcatcatggg tagtccccat cctcttggga ggtgacatgc tgggtgatcc tcgtcatctc aggaggtggc atcctgggtg gtccctgtcc ccctgggtat agcatcctgg gtaatcctcg 1320 tcctcttggg agtagcatcc cgggtggtcc ccgtcctccc cagcagtagc atcctgggtg 1380 gcttcccatc ctcctaggcg gtatcatcct gggtagcccc ctggggcaga gggaagagct 1440 gtggcctccg caacctccta ctcacttccg ctccgttctc ttgcccgccc tgcagcggca 1500 cccgacctga cagcgttcag cgacccgcgc cagttccccg cgctgccctc catctccgac 1560 ccccgcatgc actatccagg cgccttcacc tactccccga cgccggtcac ctcgggcatc 1620 ggcatcggca tgtcggccat gggctcggcc acgcgctacc acacctacct gccgccgccc 1680 taccccggct cgtcgcaagc gcagggaggc ccgttccaag ccagctcgcc ctcctaccac 1740 ctgtactacg gcgcctcggc cggctcctac cagttctcca tggtgggcgg cgagcgctcg 1800 ccgccgcgca tcctgccgcc ctgcaccaac gcctccaccg gctccgcgct gctcaacccc 1860 agcctcccga accagagcga cgtggtggag gccgagggca gccacagcaa ctcccccacc 1920 aacatggcgc cctccgcgcg cctggaggag gccgtgtgga ggccctacgg atcaggagag 1980 gggagaggat ccctgctgac ttgcggggat gtggaagaga accctggacc gatggtgagc 2040 aagggcgagg agataacat ggccatcatc aaggagttca tgcgcttcaa ggtgcgcatg 2100 gagggctccg tgaacggcca cgagttcgag atcgagggcg agggcgaggg ccgcccctac 2160 gagggctttc agaccgctaa gctgaaggtg accaagggtg gccccctgcc cttcgcctgg 2220 gacatcctgt cccctcagtt cacctacggc tccaaggcct acgtgaagca ccccgccgac 2280 atccccgact acttcaagct gtccttcccc gagggcttca agtgggagcg cgtgatgaac 2340 ttcgaggacg gcggcgtggt gaccgtgacc caggactcct ccctgcagga cggcgagttc 2400 atctacaagg tgaagctgcg cggcaccaac ttcccctccg acggccccgt aatgcagaag 2460 aagaccatgg gctgggaggc ctcctccgag cggatgtacc ccgaggacgg cgccctgaag 2520 ggcgagatca agatgaggct gaagctgaag gacggcggcc actacacctc cgaggtcaag 2580 accacctaca aggccaagaa gcccgtgcag ctgcccggcg cctacatcgt cggcatcaag 2640 2700 ggccgccact ccaccggcgg catggacgag ctgtacaagt gaggcgccag gcctggcccg 2760 gctgggccac gcgggccgcc gccttcgcct ccgggcgcgc gggcctcctg ttcgcgacaa 2820 gcccgccggg atcccgggcc ctgggccccgg ccaccgtcct ggggccgagg gcgcccgacg 2880 gccaggatct cgctgtaggt caggcccgcg cagcctcctg cgcccagaag cccacgccgc 2940 cgccgtctgc tggcgccccg gccctcgcgg aggtgtccga ggcgacgcac ctcgagggtg 3000 tccgccggcc ccagcaccca ggggacgcgc tggaaagcaa acaggaagat tcccggaggg 3060 aaactgtgaa tgcttctgat ttagcaatgc tgtgaataaa aagaaagatt ttataccctt 3120 gacttaactt tttaaccaag ttgtttattc caaagagtgt ggaattttgg ttggggtggg 3180 gggagggag ggatgcaact cgccctgttt ggcatctaat tcttattttt aatttttccg 3240 caccttatca attgcaaaat gcgtatttgc atttgggtgg tttttattt tatatacgtt 3300 tatataaata tatataaatt gagcttgctt cttcttgct ttgaccatgg aaagaaatat 3360 gattcccttt tctttaagtt ttatttaact tttcttttgg acttttgggt agttgtttt 3420 ttttgttttg ttttgttttt ttgagaaaca gctacagctt tgggtcattt ttaactactg 3480 tattcccaca aggaatcccc agatatttat gtatcttgat gttcagacat tttgtgttg 3540 ataattttt aattatttaa atgtacttat attaagaaaa atatcaagta ctacatttc 3600 ttttgttctt gatagtagcc aaagttaaat gtatcacatt gaagaaggct agaaaaaaag 3660 aatgagtaat gtgatcgctt ggttatccag aagtattgtt tacattaaac tccctttcat 3720 gttaatcaaa caagtgagta gctcacgcat gacgtacgcg tcaattgtcc ggaattctgc 3780 agagtggcgg ccgcacatgt gagcaaaagg ccagcaaaag gccaggaacc gtaaaaaggc 3840 cgcgttgctg gcgtttttcc ataggctccg cccccctgac gagcatcaca aaaatcgacg 3900 ctcaagtcag aggtggcgaa acccgacagg actataaaga taccaggcgt ttccccctgg 3960 aagctccctc gtgcgctctc ctgttccgac cctgccgctt accggatacc tgtccgcctt 4020 tctcccttcg ggaagcgtgg cgctttctca tagctcacgc tgtaggtatc tcagttcggt 4080 gtaggtcgtt cgctccaagc tgggctgtgt gcacgaaccc cccgttcagc ccgaccgctg 4140 cgccttatcc ggtaactatc gtcttgagtc caacccggta agacacgact tatcgccact 4200 ggcagcagcc actggtaaca ggattagcag agcgaggtat gtaggcggtg ctacagagtt 4260 cttgaagtgg tggcctaact acggctacac tagaagaaca gtatttggta tctgcgctct 4320 gctgaagcca gttaccttcg gaaaaagagt tggtagctct tgatccggca aaaaccac 4380 cgctggtagc ggtggtttt ttgtttgcaa gcagcagatt acgcgcagaa aaaaaggatc 4440 tcaagaagat cctttgatct tttctacggg gtctgacgct cagtggaacg aaaactcacg 4500 ttaagggatt ttggtcatga gattatcaaa aaggatcttc acctagatcc ttttaaatta 4560 aaaatgaagt tttaaatcaa tctaaagtat atatgagtaa acttggtctg acagtta 4617 <210> 11 <211> 24 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of marker-integrated Tbx1 <220> <221> source <222> (1)..(24) <223> sequence is synthesized <400> 11 gaggattggg aagacaatag cagg 24 <210> 12 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of marker-integrated Tbx1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 12 gcctccgacc gggcgctttg 20 <210> 13 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection and sequencing of marker-free Tbx1 <220> <221> source <222> (1)..(23) <223> sequence is synthesized <400> 13 ggcccagtga cccagcctca tct 23 <210> 14 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection and sequencing of marker-free Tbx1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 14 gcctccgacc gggcgctttg 20 <210> 15 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> marker-free Tbx1 read sequence <220> <221> source <222> (1)..(21) <223> sequence is synthesized <400> 15 gctgcagggc tccagcggct t 21 <210> 16 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of Runx1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 16 gggtggcaga ttctgggtag 20 <210> 17 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of Runx1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 17 cagcctggtg aaagcaacac 20 <210> 18 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> alternate UbC promoter forward primer <220> <221> source <222> (1)..(21) <223> sequence is synthesized <400> 18 tgcgggaaag ctcttattcg g 21 <210> 19 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> alternate UbC promoter reverse primer <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 19 caaaaacggc cagaatttag cg 22 <210> 20 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> RUNX1 gRNA <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 20 ccgtatggag tccctactga gg 22 <210> 21 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> GFI1 gRNA <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 21 agggctcaaa tgaggaccca gg 22 <210> 22 <211> 23 <212> DNA <213> Artificial Sequence <220> <223> GFI1 gRNA <220> <221> source <222> (1)..(23) <223> sequence is synthesized <400> 22 atgggttcaa atgagcaacc tgg 23 <210> 23 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> GFI1 gRNA <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 23 atgggctcaa atgagcctct gg 22 <210> 24 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of GFI1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 24 aatgccatgc tgggctattg 20 <210> 25 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of GFI1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 25 ccagctttcc ccctacagac 20 <210> 26 <211> 21 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of B2M <220> <221> source <222> (1)..(21) <223> sequence is synthesized <400> 26 atgcagcgca atctccagtg a 21 <210> 27 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of B2M <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 27 gtagctgcag acagttctcc aa 22 <210> 28 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of random integration 1 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 28 aaggatcagg acgctcgctg 20 <210> 29 <211> 28 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of random integration 1 <220> <221> source <222> (1)..(28) <223> sequence is synthesized <400> 29 gtctcatgag cggatacata tttgaatg 28 <210> 30 <211> 20 <212> DNA <213> Artificial Sequence <220> <223> forward primer for detection of random integration 2 <220> <221> source <222> (1)..(20) <223> sequence is synthesized <400> 30 cacctctgac ttgagcgtcg 20 <210> 31 <211> 27 <212> DNA <213> Artificial Sequence <220> <223> reverse primer for detection of random integration 2 <220> <221> source <222> (1)..(27) <223> sequence is synthesized <400> 31 ggaaattgca tcgcattgtc tgagtag 27 <210> 32 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> TBX1 reference cDNA <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 32 ccaccggccc ggcgcactgc cg 22 <210> 33 <211> 22 <212> DNA <213> Artificial Sequence <220> <223> TBX1 c. 928G>A <220> <221> source <222> (1)..(22) <223> sequence is synthesized <400> 33 ccaccggccc agcgcactgc cg 22

Claims

1. A method for scarless editing of genomic DNA of a cell, comprising the steps of: (a)(i) introducing into a cell a first donor polynucleotide, the first donor polynucleotide comprising a first left homology arm and a first right homology arm flanking a sequence comprising an intended edit to a target sequence to be modified in the genomic DNA of the cell, and an expression cassette encoding at least one selectable marker; and (ii) introducing into a cell a first guide RNA and Cas9, wherein the first guide RNA comprises a sequence complementary to a genomic target sequence to be modified in the genomic DNA of the cell, such that Cas9 forms a first complex with the first guide RNA, the guide RNA directs the first complex to the genomic target sequence, Cas9 creates a double-stranded break in the genomic DNA of the cell, and the first donor polynucleotide is integrated into the genomic DNA by Cas9-mediated homology-directed repair (HDR) to generate a genetically modified cell. performing a first Cas9-mediated homology-directed repair (HDR) step, comprising: (b) isolating the genetically modified cells based on positive selection of at least one selectable marker; (c)(i) introducing into the genetically modified cell a second donor polynucleotide, wherein the second donor polynucleotide comprises a second left homology arm and a second right homology arm flanking sequences complementary to the target genomic sequence as modified by integration of the first donor polynucleotide into the genomic DNA by Cas9-mediated HDR, except that the second donor polynucleotide comprises a deletion of an expression cassette encoding at least one selectable marker; (ii) introducing a second guide RNA and a Cas9 nuclease into the cell, wherein Cas9 forms a second complex with the second guide RNA, the second guide RNA directs the second complex to the modified target genomic sequence, Cas9 creates a double-stranded break in the modified genomic DNA, and the second donor polynucleotide is integrated into the genomic DNA by Cas9-mediated HDR, thereby removing an expression cassette encoding at least one selectable marker from the modified genomic DNA by Cas9-mediated HDR. performing a second Cas9-mediated homology-directed repair (HDR) step, comprising: (d) isolating genetically modified cells containing the intended edit based on negative selection of at least one selectable marker, wherein the expression cassette encoding the at least one selectable marker is deleted.

2. 2. The method of claim 1, wherein the sequence containing the intended edit is located within or near an expression cassette encoding at least one selectable marker.

3. 10. The method of claim 1, wherein the intended editing introduces a mutation into a gene in the genomic DNA of the cell.

4. 4. The method of claim 3, wherein the mutation is selected from the group consisting of an insertion, a deletion, and a substitution.

5. 2. The method of claim 1, wherein said intended editing removes a mutation from a gene in the genomic DNA of the cell.

6. 2. The method of claim 1, wherein said intended editing results in the inactivation of a gene in the genomic DNA of the cell.

7. The method of claim 1, wherein one allele is modified in the genomic DNA.

8. The method of claim 1, wherein both alleles are modified in the genomic DNA.

9. The method of claim 1, wherein the at least one selectable marker is a fluorescent marker.

10. 10. The method of claim 9, further comprising measuring fluorescence intensity to determine whether the genetically modified cells contain single-allelic edits or biallelic edits.

11. The method of claim 1, wherein the cell is derived from a eukaryotic, prokaryotic, or archaeal organism.

12. The method of claim 11 , wherein the cell is mammalian.

13. The method of claim 12, wherein the cell is human.

14. The method of claim 1, wherein the cells are derived from a cell line.

15. The method of claim 1, wherein the cell is an immortalized cell.

16. The method of claim 1 , wherein the cells are cancerous.

17. The method of claim 1, wherein the cell is in vitro or in vivo.

18. The method of claim 1, wherein the cells are selected from the group consisting of K562 cells, embryonic stem cells, and induced pluripotent stem cells.

19. The method of claim 1, wherein the at least one selectable marker is selected from the group consisting of a cell surface marker, a drug resistance gene, a reporter gene, and a suicide gene.

20. 20. The method of claim 19, wherein the cell surface marker is selected from the group consisting of cleaved CD8, NGFR, cleaved CD19 (tCD19), CCR5, and ABO antigens.

21. 20. The method of claim 19, wherein the cell surface marker is not essential for cell function.

22. 20. The method of claim 19, wherein the genetically modified cells generated from the first HDR step are isolated by positive selection using a binding agent that specifically binds to the selection marker.

23. 23. The method of claim 22, wherein the binding agent comprises an antibody, antibody mimetic, or aptamer that specifically binds to the selectable marker.

24. The antibody may be a monoclonal antibody, a polyclonal antibody, a chimeric antibody, a nanobody, a recombinant fragment of an antibody, a Fab fragment, a Fab' fragment, a F(ab') 2 Fragment, F v fragments, and scF v 24. The method of claim 23, wherein the fragment is selected from the group consisting of:

25. 23. The method of claim 22, wherein the antibody is selected from the group consisting of an anti-tCD19 antibody, an anti-CD8 antibody, and an anti-NGFR antibody.

26. 23. The method of claim 22, wherein the binding agent is immobilized on a solid support.

27. 27. The method of claim 26, wherein the solid support is a magnetic bead, a non-magnetic bead, a slide, a gel, a membrane, or a microtiter plate well.

28. 2. The method of claim 1, wherein the expression cassette encoding the at least one selectable marker comprises a UbC promoter, a polynucleotide encoding mCherry, a polynucleotide encoding a T2A peptide, a polynucleotide encoding a truncated CD19 (tCD19), and a polyadenylation sequence.

29. 10. The method of claim 1, wherein one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, or Cas9 are provided by a vector.

30. The method of claim 1, wherein the vector is a plasmid or a viral vector.

31. 30. The method of claim 29, wherein the first donor polynucleotide, the first guide RNA, and the Cas9 are provided by a single vector or multiple vectors.

32. 30. The method of claim 29, wherein the second donor polynucleotide, the second guide RNA, and the Cas9 are provided by a single vector or multiple vectors.

33. The method of claim 1, wherein the first donor polynucleotide further comprises at least one expression cassette encoding a short hairpin RNA (shRNA) that inhibits expression of a randomly integrated or episomal selectable marker.

34. The method of claim 33, wherein the first donor polynucleotide comprises a pair of expression cassettes encoding an shRNA, the first expression cassette of the pair being located 5' of the first left homology arm and the second expression cassette of the pair being located 3' of the first right homology arm.

35. 34. The method of claim 33, wherein the shRNA reduces selection of clones expressing a selectable marker that has been randomly integrated into the genome.

36. The method of claim 1, wherein the double-strand break is a non-gene-disrupting double-strand break.

37. 2. The method of claim 1, wherein the double-stranded break does not affect expression of the at least one selectable marker.

38. Scarless genome editing system, including: (a) a first donor polynucleotide comprising a first left homology arm and a first right homology arm flanking a sequence containing an intended edit to a target sequence to be modified in the genomic DNA of a cell, and an expression cassette encoding at least one selectable marker; (b) a second donor polynucleotide comprising a second left homology arm and a second right homology arm flanked by sequences complementary to the sequence of the first donor polynucleotide except for the deletion of an expression cassette encoding at least one selectable marker; (c) Cas9 nuclease; (d) a first guide RNA capable of forming a complex with Cas9 nuclease and directing the complex to a target sequence; and (e) a second guide RNA capable of forming a complex with a Cas9 nuclease and directing the complex to a target genomic sequence to be modified by integration of the first donor polynucleotide into genomic DNA by Cas9-mediated homology-directed repair (HDR).

39. 39. The scarless genome editing system of claim 38, wherein the expression cassette encoding the at least one selectable marker comprises the sequence of SEQ ID NO:1 or a sequence having at least 95% identity to the sequence of SEQ ID NO:1, wherein the first donor polynucleotide can be integrated into the genomic DNA of a cell at a target sequence by the Cas9-mediated HDR under conditions suitable for expression of the at least one selectable marker, and wherein cells having the first donor polynucleotide integrated into the genomic DNA of a cell at a target site by Cas9-mediated HDR can be identified by positive selection of the at least one selectable marker encoded by the expression cassette.

40. 39. The scarless genome editing system of claim 38, wherein the second donor polynucleotide comprises a polynucleotide comprising the sequence of SEQ ID NO:2 or a sequence having at least 95% identity to the sequence of SEQ ID NO:2, and wherein the second donor polynucleotide is capable of being integrated into the modified genomic DNA of the cell at the target site by Cas9-mediated HDR to remove an expression cassette encoding at least one selectable marker.

41. The scarless genome editing system of claim 38, wherein one or more of the first donor polynucleotide, the second donor polynucleotide, the first guide RNA, the second guide RNA, or Cas9 are provided by a vector.

42. The scarless genome editing system of claim 38, wherein the vector is a plasmid or a viral vector.

43. The scarless genome editing system of claim 38, wherein the first donor polynucleotide, the first guide RNA, and the Cas9 are provided by a single vector or multiple vectors.

44. The scarless genome editing system of claim 38, wherein the second donor polynucleotide, the second guide RNA, and the Cas9 are provided by a single vector or multiple vectors.

45. The scarless genome editing system of claim 38, wherein the first donor polynucleotide further comprises at least one expression cassette encoding a short hairpin RNA (shRNA) that inhibits expression of a randomly integrated or episomal selection marker.

46. The scarless genome editing system of claim 38, wherein the first donor polynucleotide comprises a pair of expression cassettes encoding shRNAs, the first expression cassette of the pair being located 5' of the first left homology arm and the second expression cassette of the pair being located 3' of the first right homology arm.

47. A host cell comprising the scarless genome editing system of claim 38.

48. A kit comprising the scarless genome editing system of claim 38 and instructions for performing scarless editing of genomic DNA.

49. 48. A kit comprising the host cell of claim 47.