Transgene genomic identification by nuclease-mediated long read sequencing

Long-read sequencing with nuclease-mediated cleavage of exogenous DNA segments allows for precise determination of genomic integration sites, addressing the challenge of transgene integration in undesired locations.

WO2025235388A1PCT designated stage Publication Date: 2025-11-13REGENERON PHARMACEUTICALS INC
View PDF 56 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/027765
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2025-05-05
Publication Date
2025-11-13

AI Technical Summary

Technical Problem

Identifying precise integration sites of genomically integrated exogenous DNA within the host genome remains challenging due to transgene integrations in undesired genomic locations during targeted genomic modifications.

Method used

A method involving long-read sequencing, using first and second nuclease agents to cleave specific sites within the genomically integrated exogenous DNA, followed by aligning sequences to a reference genome and exogenous DNA sequence to determine the genomic location, without amplification of the DNA segments.

Benefits of technology

Accurately determines the genomic location of integrated exogenous DNA, providing precise integration site identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025027765_13112025_PF_FP_ABST
    Figure US2025027765_13112025_PF_FP_ABST
Patent Text Reader

Abstract

Methods are provided for determining the genomic location of a genomically integrated exogenous DNA. The methods can use nuclease agents to create an enriched library for long- read sequencing to provide hybrid alignments to both a reference genome and the exogenous DNA sequence, thereby determining the location of the genomically integrated exogenous DNA.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.057766 / 629304 TRANSGENE GENOMIC IDENTIFICATION BY NUCLEASE-MEDIATED LONG READ SEQUENCING CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of US Application No.63 / 643,219, filed May 6, 2024, which is herein incorporated by reference in its entirety for all purposes. REFERENCE TO A SEQUENCE LISTING SUBMITTED AS AN XML FILE

[0002] The Sequence Listing written in file 629304SEQLIST.xml is 69,608 bytes, was created on May 2, 2025, and is hereby incorporated by reference in its entirety. BACKGROUND

[0003] Targeted genomic modifications are used in model organism development. The use of large targeting vectors allows for large targeted genomic modifications. These desired genomic modifications include gene knockouts, gene knock-ins, partial gene deletions, exon modifications, genomic alterations for alternate splice variants, and other modifications for generating partial or full foreign gene products. Targeted genomic modifications are traditionally mediated by homologous recombination (HR) in embryonic stem cells. This allows for all or majority of the cells that received the targeted genomic modification to have the desired genotype in the resulting organism. While homologous recombination is efficient, transgene integrations in genomic locations that were not desired can also occur. Identifying precise integration sites of these genetic segments within the host genome remains challenging. SUMMARY

[0004] Methods of determining the genomic location of a genomically integrated exogenous DNA are provided.

[0005] Some such methods comprise: (a) preparing a single long-read sequencing library, wherein the preparing comprises cleaving genomic DNA extracted from cells comprising a genomically integrated exogenous DNA with: (i) a first nuclease agent that cleaves a first nuclease cleavage site within the genomically integrated exogenous DNA to generate a firstAttorney Docket No.057766 / 629304 cleaved genomic DNA segment and (ii) a second nuclease agent that cleaves a second nuclease cleavage site within the genomically integrated exogenous DNA to generate second cleaved genomic DNA segment, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 50 bp; (b) performing long-read sequencing on the single long-read sequencing library to generate a plurality of sequences from the first cleaved genomic DNA segment and the second cleaved genomic DNA segment; and (c) aligning the plurality of sequences to both a reference genome and the exogenous DNA sequence to generate a plurality of hybrid alignments, thereby determining the genomic location of the genomically integrated exogenous DNA.

[0006] In some such methods, the long-read sequencing is nanopore sequencing or single molecule real-time sequencing. In some such methods, the long-read sequencing is nanopore sequencing.

[0007] In some such methods, the method does not comprise amplifying the genomic DNA, the first cleaved genomic DNA segment, or the second cleaved genomic DNA segment.

[0008] In some such methods, the exogenous DNA is at least 0.5 kb, at least 1 kb, at least 5 kb, at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb. Optionally, the exogenous DNA is at least 100 kb.

[0009] In some such methods, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp, at least 125 bp, at least 150 bp, at least 200 bp, or at least 250 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 150 bp or by about 150 bp to about 250 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by about 150 bp to about 500 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by about 100 bp to about 2000 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by about 100 bp to about 500 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 125 bp. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 250 bp. In some such methods, the first nuclease cleavage site and the second nuclease cleavage site are each at least 100 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some such methods, the first nuclease cleavage site and theAttorney Docket No.057766 / 629304 second nuclease cleavage site are each at least 500 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are each at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are each at least 1 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are each about 1 kb to about 250 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA.

[0010] In some such methods, the first nuclease agent is a first zinc finger nuclease (ZFN), a first transcription activator-like effector nuclease (TALEN), or a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) protein and a first guide RNA (gRNA), and the second nuclease agent is a second ZFN, a second TALEN, or a Cas protein and a second gRNA. Optionally, the first nuclease agent is a Cas9 protein and a first gRNA and the second nuclease agent is a Cas9 protein and a second gRNA. In some such methods, the first gRNA targets a first guide RNA target sequence and the second gRNA targets a second guide RNA target sequence, wherein the first guide RNA target sequence and the second guide RNA target sequence are on opposite strands, wherein the guide RNA target sequence on the sense strand is located 3′ of the guide RNA target sequence on the antisense strand.

[0011] In some such methods, the exogenous DNA sequence comprises a selection cassette or a reporter protein coding sequence, and the first nuclease cleavage site and the second nuclease cleavage site are within the selection cassette or the reporter protein coding sequence. Optionally, the selection cassette comprises a kanamycin resistance gene, a hygromycin resistance gene, or a puromycin resistance gene, and the first nuclease cleavage site and the second nuclease cleavage site are within the kanamycin resistance gene, the hygromycin resistance gene, or the puromycin resistance gene. Optionally, the first nuclease cleavage site and the second nuclease cleavage site are within a spectinomycin resistance gene, a kanamycin resistance gene, a hygromycin resistance gene, a puromycin resistance gene, a lacZ gene, a tamoxifen-inducible Cre gene, a luciferase gene, a human ubiquitin promoter, a mouse phosphoglycerate kinase promoter, or a mouse protamine promoter.

[0012] In some such methods, the amount of genomic DNA in step (a) is at least 400 ng, is less than 3 µg, or is from about 400 ng to about 3 µg.Attorney Docket No.057766 / 629304

[0013] In some such methods, the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a mapping quality (MAPQ) score of at least MAPQ40. Optionally, the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a MAPQ score of at least MAPQ50.

[0014] In some such methods, the cells are mammalian cells. Optionally, the cells are human cells. Optionally, the cells are human induced pluripotent stem cells. Optionally, the cells are rodent cells. In some such methods, the cells are mouse cells. Optionally, the cells are mouse embryonic stem (ES) cells, or optionally, the cells are from a living mouse. In some such methods, the cells are rat cells. Optionally the cells are rat ES cells, or optionally, the cells are from a living rat.

[0015] In some such methods, the cells are from a tissue sample from a subject.

[0016] Some such methods further comprise extracting genomic DNA from the cells comprising the genomically integrated exogenous DNA prior to step (a).

[0017] Some such methods further comprise genetically modifying a population of cells to generate the cells comprising the genomically integrated exogenous DNA prior to extracting the genomic DNA. Optionally, the genetically modifying comprises administering the exogenous DNA to the population of cells such that the exogenous DNA is integrated into the genome, optionally wherein the exogenous DNA that is administered is in the form of a viral vector, an adeno-associated virus (AAV) vector, a single-stranded oligodeoxynucleotide (ssODN), or a targeting vector comprising a 5′ homology arm and a 3′ homology arm. Optionally, the genetically modifying comprises administering a large targeting vector to the population of cells; wherein the large targeting vector comprises a 5′ homology arm, a 3′ homology arm, and optionally an insert nucleic acid flanked by the 5′ homology arm and the 3′ homology arm; wherein the large targeting vector is at least 10 kb in length; or wherein the sum total of the 5′ homology arm and the 3′ homology arm is at least 10 kb in length.

[0018] In some such methods, the large targeting vector is at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb. Optionally, the large targeting vector is at least 100 kb. In some such methods, the insert nucleic acid is at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb. Optionally, the insert nucleic acid is at least 100 kb. In some such methods, the 5′ homology arm, the 3′ homology arm, or the sum total of the 5′ homology arm and the 3′ homology arm is at least 10 kb, at least 50 kb, at least 100 kb, at leastAttorney Docket No.057766 / 629304 150 kb, or at least 200 kb. Optionally, the 5′ homology arm, the 3′ homology arm, or the sum total of the 5′ homology arm and the 3′ homology arm is at least 100 kb.

[0019] In some such methods, the long-read sequencing is nanopore sequencing; the method does not comprise amplifying the genomic DNA, the first cleaved genomic DNA segment, or the second cleaved genomic DNA segment; the exogenous DNA is at least 10 kb; and the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 250 bp.

[0020] In some such methods, the long-read sequencing is nanopore sequencing; the exogenous DNA is at least 10 kb; the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 250 bp; the first nuclease agent is Cas9 protein and a first gRNA and the second nuclease agent is the Cas9 protein and a second gRNA; the exogenous DNA sequence comprises a selection cassette or a reporter protein coding sequence; and the first nuclease cleavage site and the second nuclease cleavage site are within the selection cassette or the reporter protein coding sequence.

[0021] In some such methods, the long-read sequencing is nanopore sequencing; the method does not comprise amplifying the genomic DNA, the first cleaved genomic DNA segment, or the second cleaved genomic DNA segment; the exogenous DNA is at least 10 kb; and the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp.

[0022] In some such methods, the long-read sequencing is nanopore sequencing; the exogenous DNA is at least 10 kb; the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp; the first nuclease agent is Cas9 protein and a first gRNA and the second nuclease agent is the Cas9 protein and a second gRNA; the exogenous DNA sequence comprises a selection cassette or a reporter protein coding sequence; and the first nuclease cleavage site and the second nuclease cleavage site are within the selection cassette or the reporter protein coding sequence.

[0023] In some such methods, step (a) comprises cleaving a single genomic DNA sample with the first nuclease agent and the second nuclease agent simultaneously, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp, wherein the first nuclease cleavage site and the second nuclease cleavage site are each at least 500 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA, wherein the first nuclease agent is a Cas9 protein and a first gRNA and the second nuclease agent is a Cas9 protein and a second gRNA, wherein the first gRNA targets a first guide RNA target sequenceAttorney Docket No.057766 / 629304 and the second gRNA targets a second guide RNA target sequence, wherein the first guide RNA target sequence and the second guide RNA target sequence are on opposite strands, and wherein the guide RNA target sequence on the sense strand is located 3′ of the guide RNA target sequence on the antisense strand. BRIEF DESCRIPTION OF THE FIGURES

[0024] Figure 1 shows an overview of a non-limiting example of gRNA design for transgene integration identification. The DNA sequences and gRNA target sequences do not correspond to actual DNA sequences but were generated merely for the purposes of illustration. DEFINITIONS

[0025] The terms “protein,” “polypeptide,” and “peptide,” used interchangeably herein, include polymeric forms of amino acids of any length, including coded and non-coded amino acids and chemically or biochemically modified or derivatized amino acids. The terms also include polymers that have been modified, such as polypeptides having modified peptide backbones. The term “domain” refers to any part of a protein or polypeptide having a particular function or structure.

[0026] The terms “nucleic acid” and “polynucleotide,” used interchangeably herein, include polymeric forms of nucleotides of any length, including ribonucleotides, deoxyribonucleotides, or analogs or modified versions thereof. They include single-, double-, and multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers comprising purine bases, pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, or derivatized nucleotide bases. Nucleic acids are said to have “5' ends” and “3' ends” because mononucleotides are reacted to make oligonucleotides in a manner such that the 5' phosphate of one mononucleotide pentose ring is attached to the 3' oxygen of its neighbor in one direction via a phosphodiester linkage. An end of an oligonucleotide is referred to as the “5' end” if its 5' phosphate is not linked to the 3' oxygen of a mononucleotide pentose ring. An end of an oligonucleotide is referred to as the “3' end” if its 3' oxygen is not linked to a 5' phosphate of another mononucleotide pentose ring. A nucleic acid sequence, even if internal to a larger oligonucleotide, also may be said to have 5' and 3' ends. In either a linear or circular DNAAttorney Docket No.057766 / 629304 molecule, discrete elements are referred to as being “upstream” or 5' of the “downstream” or 3' elements.

[0027] Repair in response to double-strand breaks (DSBs) in DNA occurs principally through two conserved DNA repair pathways: homologous recombination (HR) and non- homologous end joining (NHEJ). See Kasparek & Humphrey (2011) Seminars in Cell & Dev. Biol.22:886-897, herein incorporated by reference in its entirety for all purposes. Likewise, repair of a target nucleic acid mediated by an exogenous donor nucleic acid can include any process of exchange of genetic information between the two polynucleotides.

[0028] The term “recombination” includes any process of exchange of genetic information between two polynucleotides and can occur by any mechanism. Recombination can occur via homology directed repair (HDR) or homologous recombination (HR). HDR or HR includes a form of nucleic acid repair that can require nucleotide sequence homology, uses a “donor” molecule as a template for repair of a “target” molecule (i.e., the one that experienced the double-strand break), and leads to transfer of genetic information from the donor to target. Without wishing to be bound by any particular theory, such transfer can involve mismatch correction of heteroduplex DNA that forms between the broken target and the donor, and / or synthesis-dependent strand annealing, in which the donor is used to resynthesize genetic information that will become part of the target, and / or related processes. In some cases, the donor polynucleotide, a portion of the donor polynucleotide, a copy of the donor polynucleotide, or a portion of a copy of the donor polynucleotide integrates into the target DNA. See Wang et al. (2013) Cell 153:910-918; Mandalos et al. (2012) PLoS ONE 7:e45768:1-9; and Wang et al. (2013) Nat Biotechnol.31:530-532, each of which is herein incorporated by reference in its entirety for all purposes.

[0029] The term “genomically integrated” refers to a nucleic acid that has been introduced into a cell such that the nucleotide sequence integrates into the genome of the cell. Any protocol may be used for the stable incorporation of a nucleic acid into the genome of a cell.

[0030] The term “targeting vector” refers to a recombinant nucleic acid that can be introduced by homologous recombination, non-homologous-end-joining-mediated ligation, or any other means of recombination to a target position in the genome of a cell.

[0031] The term “recombinant nucleic acid” or “nucleic acid construct” (and by analogy, a “recombinant polypeptide” produced by the expression of a recombinant nucleic acid) refers to aAttorney Docket No.057766 / 629304 polynucleotide that has a sequence that is not naturally occurring or has a sequence that is made by an artificial combination of two otherwise separated segments of sequence. This artificial combination can be accomplished by chemical synthesis or by the artificial manipulation of isolated segments of polynucleotides or nucleic acid molecules, such as by genetic engineering techniques that are well known to those of skill in the art.

[0032] The term “viral vector” refers to a recombinant nucleic acid that includes at least one element of viral origin and includes elements sufficient for or permissive of packaging into a viral vector particle. The vector and / or particle can be utilized for the purpose of transferring DNA, RNA, or other nucleic acids into cells either ex vivo or in vivo. Numerous forms of viral vectors are known.

[0033] The term “isolated” with respect to proteins, nucleic acids, and cells includes proteins, nucleic acids, and cells that are relatively purified with respect to other cellular or organism components that may normally be present in situ, up to and including a substantially pure preparation of the protein, nucleic acid, or cell. The term “isolated” may include proteins and nucleic acids that have no naturally occurring counterpart or proteins or nucleic acids that have been chemically synthesized and are thus substantially uncontaminated by other proteins or nucleic acids. The term “isolated” may include proteins, nucleic acids, or cells that have been separated or purified from most other cellular components or organism components with which they are naturally accompanied (e.g., but not limited to, other cellular proteins, nucleic acids, or cellular or extracellular components).

[0034] The term “wild type” includes entities having a structure and / or activity as found in a normal (as contrasted with mutant, diseased, altered, or so forth) state or context. Wild type genes and polypeptides often exist in multiple different forms (e.g., alleles).

[0035] The term “endogenous sequence” refers to a nucleic acid sequence that occurs naturally within a cell or animal. For example, an endogenous ALB sequence of an animal refers to a native ALB sequence that naturally occurs at the ALB locus in the animal.

[0036] “Exogenous” molecules or sequences include molecules or sequences that are not normally present in a cell in that form or that are introduced into a cell from an outside source. Normal presence includes presence with respect to the particular developmental stage and environmental conditions of the cell. An exogenous DNA, for example, can be a DNA sequence not normally present in a cell (e.g., a reporter gene, or a mouse DNA sequence in a human cell).Attorney Docket No.057766 / 629304 Alternatively, an exogenous DNA, for example, can include a mutated version of a corresponding endogenous DNA sequence within the cell (e.g., at the endogenous genomic locus), such as a humanized version of the endogenous DNA sequence, or can include a sequence corresponding to an endogenous DNA sequence within the cell but in a different form (i.e., not within a chromosome or inserted into a different genomic locus than the endogenous DNA sequence). In contrast, endogenous molecules or sequences include molecules or sequences that are normally present in that form in a particular cell at a particular developmental stage under particular environmental conditions.

[0037] The term “heterologous” when used in the context of a nucleic acid or a protein indicates that the nucleic acid or protein comprises at least two segments that do not naturally occur together in the same molecule. For example, the term “heterologous,” when used with reference to segments of a nucleic acid or segments of a protein, indicates that the nucleic acid or protein comprises two or more sub-sequences that are not found in the same relationship to each other (e.g., joined together) in nature. As one example, a “heterologous” region of a nucleic acid vector is a segment of nucleic acid within or attached to another nucleic acid molecule that is not found in association with the other molecule in nature. For example, a heterologous region of a nucleic acid vector could include a coding sequence flanked by a heterologous promoter not found in association with the coding sequence in nature. Likewise, a “heterologous” region of a protein is a segment of amino acids within or attached to another peptide molecule that is not found in association with the other peptide molecule in nature (e.g., a fusion protein, or a protein with a tag). Similarly, a nucleic acid or protein can comprise a heterologous label or a heterologous secretion or localization sequence.

[0038] “Codon optimization” takes advantage of the degeneracy of codons, as exhibited by the multiplicity of three-base pair codon combinations that specify an amino acid, and generally includes a process of modifying a nucleic acid sequence for enhanced expression in particular host cells by replacing at least one codon of the native sequence with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. For example, a nucleic acid encoding a protein can be modified to substitute codons having a higher frequency of usage in a given prokaryotic or eukaryotic cell, including a bacterial cell, a yeast cell, a human cell, a non-human cell, a mammalian cell, a rodent cell, a mouse cell, a rat cell, a hamster cell, or any other host cell, as compared to theAttorney Docket No.057766 / 629304 naturally occurring nucleic acid sequence. Codon usage tables are readily available, for example, at the “Codon Usage Database.” These tables can be adapted in a number of ways. See Nakamura et al. (2000) Nucleic Acids Research 28:292, herein incorporated by reference in its entirety for all purposes. Computer algorithms for codon optimization of a particular sequence for expression in a particular host are also available (see, e.g., Gene Forge).

[0039] The term “locus” refers to a specific location of a gene (or significant sequence), DNA sequence, polypeptide-encoding sequence, or position on a chromosome of the genome of an organism. For example, an “ALB locus” may refer to the specific location of an ALB gene, ALB DNA sequence, albumin-encoding sequence, or ALB position on a chromosome of the genome of an organism that has been identified as to where such a sequence resides. An “ALB locus” may comprise a regulatory element of an ALB gene, including, for example, an enhancer, a promoter, 5' and / or 3' untranslated region (UTR), or a combination thereof.

[0040] The term “gene” refers to a DNA sequence in a chromosome that codes for a product (e.g., an RNA product and / or a polypeptide product) and includes the coding region interrupted with non-coding introns and sequence located adjacent to the coding region on both the 5' and 3' ends such that the gene corresponds to the full-length mRNA (including the 5' and 3' untranslated sequences). The term “gene” also includes other non-coding sequences including regulatory sequences (e.g., promoters, enhancers, and transcription factor binding sites), polyadenylation signals, internal ribosome entry sites, silencers, insulating sequence, and matrix attachment regions. These sequences may be close to the coding region of the gene (e.g., within 10 kb) or at distant sites, and they influence the level or rate of transcription and translation of the gene.

[0041] The term “allele” refers to a variant form of a gene. Some genes have a variety of different forms, which are located at the same position, or genetic locus, on a chromosome. A diploid organism has two alleles at each genetic locus. Each pair of alleles represents the genotype of a specific genetic locus. Genotypes are described as homozygous if there are two identical alleles at a particular locus and as heterozygous if the two alleles differ.

[0042] A “promoter” is a regulatory region of DNA usually comprising a TATA box capable of directing RNA polymerase II to initiate RNA synthesis at the appropriate transcription initiation site for a particular polynucleotide sequence. A promoter may additionally comprise other regions which influence the transcription initiation rate. The promoter sequences disclosed herein modulate transcription of an operably linked polynucleotide. A promoter can be active inAttorney Docket No.057766 / 629304 one or more of the cell types disclosed herein (e.g., a eukaryotic cell, a non-human mammalian cell, a human cell, a rodent cell, a pluripotent cell, a one-cell stage embryo, a differentiated cell, or a combination thereof). A promoter can be, for example, a constitutively active promoter, a conditional promoter, an inducible promoter, a temporally restricted promoter (e.g., a developmentally regulated promoter), or a spatially restricted promoter (e.g., a cell-specific or tissue-specific promoter). Examples of promoters can be found, for example, in WO 2013 / 176772, herein incorporated by reference in its entirety for all purposes.

[0043] A constitutive promoter is one that is active in all tissues or particular tissues at all developing stages. Examples of constitutive promoters include the human cytomegalovirus immediate early (hCMV), mouse cytomegalovirus immediate early (mCMV), human elongation factor 1 alpha (hEF1α), mouse elongation factor 1 alpha (mEF1α), mouse phosphoglycerate kinase (PGK), chicken beta actin hybrid (CAG or CBh), SV40 early, and beta 2 tubulin promoters..

[0044] Examples of inducible promoters include, for example, chemically regulated promoters and physically-regulated promoters. Chemically regulated promoters include, for example, alcohol-regulated promoters (e.g., an alcohol dehydrogenase (alcA) gene promoter), tetracycline-regulated promoters (e.g., a tetracycline-responsive promoter, a tetracycline operator sequence (tetO), a tet-On promoter, or a tet-Off promoter), steroid regulated promoters (e.g., a rat glucocorticoid receptor, a promoter of an estrogen receptor, or a promoter of an ecdysone receptor), or metal-regulated promoters (e.g., a metalloprotein promoter). Physically regulated promoters include, for example temperature-regulated promoters (e.g., a heat shock promoter) and light-regulated promoters (e.g., a light-inducible promoter or a light-repressible promoter).

[0045] Tissue-specific promoters can be, for example, neuron-specific promoters or glial- specific promoters or muscle-specific promoters or liver-specific promoters.

[0046] Developmentally regulated promoters include, for example, promoters active only during an embryonic stage of development, or only in an adult cell.

[0047] “Operable linkage” or being “operably linked” includes juxtaposition of two or more components (e.g., a promoter and another sequence element) such that both components function normally and allow the possibility that at least one of the components can mediate a function that is exerted upon at least one of the other components. For example, a promoter can be operably linked to a coding sequence if the promoter controls the level of transcription of the codingAttorney Docket No.057766 / 629304 sequence in response to the presence or absence of one or more transcriptional regulatory factors. Operable linkage can include such sequences being contiguous with each other or acting in trans (e.g., a regulatory sequence can act at a distance to control transcription of the coding sequence).

[0048] The term “variant” refers to a nucleotide sequence differing from the sequence most prevalent in a population (e.g., by one nucleotide) or a protein sequence different from the sequence most prevalent in a population (e.g., by one amino acid).

[0049] The term “fragment,” when referring to a protein, means a protein that is shorter or has fewer amino acids than the full-length protein. The term “fragment,” when referring to a nucleic acid, means a nucleic acid that is shorter or has fewer nucleotides than the full-length nucleic acid. A fragment can be, for example, when referring to a protein fragment, an N- terminal fragment (i.e., removal of a portion of the C-terminal end of the protein), a C-terminal fragment (i.e., removal of a portion of the N-terminal end of the protein), or an internal fragment (i.e., removal of a portion of each of the N-terminal and C-terminal ends of the protein). A fragment can be, for example, when referring to a nucleic acid fragment, a 5' fragment (i.e., removal of a portion of the 3' end of the nucleic acid), a 3' fragment (i.e., removal of a portion of the 5' end of the nucleic acid), or an internal fragment (i.e., removal of a portion each of the 5' and 3' ends of the nucleic acid).

[0050] “Complementarity” of nucleic acids means that a nucleotide sequence in one strand of nucleic acid, due to orientation of its nucleobase groups, forms hydrogen bonds with another sequence on an opposing nucleic acid strand. The complementary bases in DNA are typically A with T and C with G. In RNA, they are typically C with G and U with A. Complementarity can be perfect or substantial / sufficient. Perfect complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which every base in the duplex is bonded to a complementary base by Watson-Crick pairing. “Substantial” or “sufficient” complementary means that a sequence in one strand is not completely and / or perfectly complementary to a sequence in an opposing strand, but that sufficient bonding occurs between bases on the two strands to form a stable hybrid complex in set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by using the sequences and standard mathematical calculations to predict the Tm (melting temperature) of hybridized strands, or by empirical determination of Tm by using routine methods. Tm includes the temperature at which a population of hybridization complexes formed between two nucleic acid strands are 50%Attorney Docket No.057766 / 629304 denatured (i.e., a population of double-stranded nucleic acid molecules becomes half dissociated into single strands). At a temperature below the Tm, formation of a hybridization complex is favored, whereas at a temperature above the Tm, melting or separation of the strands in the hybridization complex is favored. Tm may be estimated for a nucleic acid having a known G+C content in an aqueous 1 M NaCl solution by using, e.g., Tm=81.5+0.41(% G+C), although other known Tm computations consider nucleic acid structural characteristics.

[0051] “Hybridization condition” includes the cumulative environment in which one nucleic acid strand bonds to a second nucleic acid strand by complementary strand interactions and hydrogen bonding to produce a hybridization complex. Such conditions include the chemical components and their concentrations (e.g., salts, chelating agents, formamide) of an aqueous or organic solution containing the nucleic acids, and the temperature of the mixture. Other factors, such as the length of incubation time or reaction chamber dimensions may contribute to the environment. See, e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2.sup.nd ed., pp.1.90-1.91, 9.47-9.51, 11.47-11.57 (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989), herein incorporated by reference in its entirety for all purposes.

[0052] Hybridization requires that the two nucleic acids contain complementary sequences, although mismatches between bases are possible. The conditions appropriate for hybridization between two nucleic acids depend on the length of the nucleic acids and the degree of complementation, variables well known in the art. The greater the degree of complementation between two nucleotide sequences, the greater the value of the melting temperature (Tm) for hybrids of nucleic acids having those sequences. For hybridizations between nucleic acids with short stretches of complementarity (e.g., complementarity over 35 or fewer, 30 or fewer, 25 or fewer, 22 or fewer, 20 or fewer, or 18 or fewer nucleotides) the position of mismatches becomes important (see Sambrook et al., supra, 11.7-11.8). Typically, the length for a hybridizable nucleic acid is at least about 10 nucleotides. Illustrative minimum lengths for a hybridizable nucleic acid include at least about 15 nucleotides, at least about 20 nucleotides, at least about 22 nucleotides, at least about 25 nucleotides, and at least about 30 nucleotides. Furthermore, the temperature and wash solution salt concentration may be adjusted as necessary according to factors such as length of the region of complementation and the degree of complementation.

[0053] The sequence of polynucleotide need not be 100% complementary to that of its target nucleic acid to be specifically hybridizable. Moreover, a polynucleotide may hybridize over oneAttorney Docket No.057766 / 629304 or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a loop structure or hairpin structure). A polynucleotide (e.g., gRNA) can comprise at least 70%, at least 80%, at least 90%, at least 95%, at least 99%, or 100% sequence complementarity to a target within the target nucleic acid sequence to which they are targeted. For example, a gRNA in which 18 of 20 nucleotides are complementary to a target, and would therefore specifically hybridize, would represent 90% complementarity. In this example, the remaining noncomplementary nucleotides may be clustered or interspersed with complementary nucleotides and need not be contiguous to each other or to complementary nucleotides.

[0054] Percent complementarity between particular stretches of nucleic acid sequences within nucleic acids can be determined routinely using BLAST programs (basic local alignment search tools) and PowerBLAST programs known in the art (Altschul et al. (1990) J. Mol. Biol. 215:403-410; Zhang and Madden (1997) Genome Res.7:649-656, each of which is herein incorporated by reference in its entirety for all purposes) or by using the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), using default settings, which uses the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489), herein incorporated by reference in its entirety for all purposes.

[0055] The methods and compositions provided herein employ a variety of different components. It is recognized throughout the description that some components can have active variants and fragments. Such components include, for example, Cas9 proteins, CRISPR RNAs, tracrRNAs, and guide RNAs. Biological activity for each of these components is described elsewhere herein.

[0056] “Sequence identity” or “identity” in the context of two polynucleotides or polypeptide sequences refers to the residues in the two sequences that are the same when aligned for maximum correspondence over a specified comparison window. When percentage of sequence identity is used in reference to proteins, it is recognized that residue positions which are not identical often differ by conservative amino acid substitutions, where amino acid residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule. When sequences differ in conservative substitutions, the percent sequence identity may be adjustedAttorney Docket No.057766 / 629304 upwards to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have “sequence similarity” or “similarity.” Means for making this adjustment are well known to those of skill in the art. Typically, this involves scoring a conservative substitution as a partial rather than a full mismatch, thereby increasing the percentage sequence identity. Thus, for example, where an identical amino acid is given a score of 1 and a non-conservative substitution is given a score of zero, a conservative substitution is given a score between zero and 1. The scoring of conservative substitutions is calculated, e.g., as implemented in the program PC / GENE (Intelligenetics, Mountain View, California).

[0057] “Percentage of sequence identity” includes the value determined by comparing two optimally aligned sequences (greatest number of perfectly matched residues) over a comparison window, wherein the portion of the polynucleotide sequence in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence (which does not comprise additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleic acid base or amino acid residue occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison, and multiplying the result by 100 to yield the percentage of sequence identity. Unless otherwise specified (e.g., the shorter sequence includes a linked heterologous sequence), the comparison window is the full length of the shorter of the two sequences being compared.

[0058] Unless otherwise stated, sequence identity / similarity values include the value obtained using GAP Version 10 using the following parameters: % identity and % similarity for a nucleotide sequence using GAP Weight of 50 and Length Weight of 3, and the nwsgapdna.cmp scoring matrix; % identity and % similarity for an amino acid sequence using GAP Weight of 8 and Length Weight of 2, and the BLOSUM62 scoring matrix; or any equivalent program thereof. “Equivalent program” includes any sequence comparison program that, for any two sequences in question, generates an alignment having identical nucleotide or amino acid residue matches and an identical percent sequence identity when compared to the corresponding alignment generated by GAP Version 10.

[0059] The terms “substantial identity” and “substantially identical,” as used with reference to a nucleic acid or fragment thereof, indicates that, when optimally aligned with appropriate nucleotide insertions or deletions with another nucleic acid (or its complementary strand), thereAttorney Docket No.057766 / 629304 is nucleotide sequence identity in at least about 90%, e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%, of the nucleotide bases, as measured by any well-known algorithm of sequence identity, such as FASTA, BLAST or GAP, as discussed below. A nucleic acid molecule having substantial identity to a reference nucleic acid molecule may, in certain instances, encode a polypeptide having the same or substantially similar amino acid sequence as the polypeptide encoded by the reference nucleic acid molecule.

[0060] As applied to polypeptides, the terms “substantial identity” and “"substantially identical” mean that two peptide sequences, when optimally aligned, share at least about 90% sequence identity, e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity. In some embodiments, residue positions that are not identical differ by conservative amino acid substitutions. A “conservative amino acid substitution” is one in which an amino acid residue is substituted by another amino acid residue having a side chain (R group) with similar chemical properties (e.g., charge or hydrophobicity). In general, a conservative amino acid substitution will not substantially change the functional properties of a protein.

[0061] Sequence similarity for polypeptides is typically measured using sequence analysis software. Protein analysis software matches similar sequences using measures of similarity assigned to various substitutions, deletions, and other modifications, including conservative amino acid substitutions. For instance, GCG software contains programs such as GAP and BESTFIT which can be used with default parameters to determine sequence homology or sequence identity between closely related polypeptides, such as homologous polypeptides from different species of organisms or between a wild-type protein and a mutein thereof. See, e.g., GCG Version 6.1. Polypeptide sequences also can be compared using FASTA with default or recommended parameters; a program in GCG Version 6.1. FASTA (e.g., FASTA2 and FASTA3) provides alignments and percent sequence identity of the regions of the best overlap between the query and search sequences (Pearson, 2000 supra). Another preferred algorithm when comparing a sequence of the disclosure to a database containing a large number of sequences from different organisms is the computer program BLAST, especially BLASTP or TBLASTN, using default parameters. See, e.g., Altschul et al., 1990, J. Mol. Biol.215: 403-410 and 1997 Nucleic Acids Res.25:3389-3402, each of which is herein incorporated by reference in its entirety for all purposes.Attorney Docket No.057766 / 629304

[0062] The term “conservative amino acid substitution” refers to the substitution of an amino acid that is normally present in the sequence with a different amino acid of similar size, charge, or polarity. Examples of conservative substitutions include the substitution of a non-polar (hydrophobic) residue such as isoleucine, valine, or leucine for another non-polar residue. Likewise, examples of conservative substitutions include the substitution of one polar (hydrophilic) residue for another such as between arginine and lysine, between glutamine and asparagine, or between glycine and serine. Additionally, the substitution of a basic residue such as lysine, arginine, or histidine for another, or the substitution of one acidic residue such as aspartic acid or glutamic acid for another acidic residue are additional examples of conservative substitutions. Examples of non-conservative substitutions include the substitution of a non-polar (hydrophobic) amino acid residue such as isoleucine, valine, leucine, alanine, or methionine for a polar (hydrophilic) residue such as cysteine, glutamine, glutamic acid or lysine and / or a polar residue for a non-polar residue. Typical amino acid categorizations are summarized below.

[0063] Table 1. Amino Acid Categorizations. Alanine Ala A Nonpolar Neutral 1.8 Arginine Arg R Polar Positive -4.5 Asparagine Asn N Polar Neutral -3.5 Aspartic acid Asp D Polar Negative -3.5 Cysteine Cys C Nonpolar Neutral 2.5 Glutamic acid Glu E Polar Negative -3.5 Glutamine Gln Q Polar Neutral -3.5 Glycine Gly G Nonpolar Neutral -0.4 Histidine His H Polar Positive -3.2 Isoleucine Ile I Nonpolar Neutral 4.5 Leucine Leu L Nonpolar Neutral 3.8 Lysine Lys K Polar Positive -3.9 Methionine Met M Nonpolar Neutral 1.9 Phenylalanine Phe F Nonpolar Neutral 2.8 Proline Pro P Nonpolar Neutral -1.6 Serine Ser S Polar Neutral -0.8 Threonine Thr T Polar Neutral -0.7 Tryptophan Trp W Nonpolar Neutral -0.9 Tyrosine Tyr Y Polar Neutral -1.3 Valine Val V Nonpolar Neutral 4.2

[0064] A “homologous” sequence (e.g., nucleic acid sequence) includes a sequence that is either identical or substantially similar to a known reference sequence, such that it is, forAttorney Docket No.057766 / 629304 example, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the known reference sequence. Homologous sequences can include, for example, orthologous sequence and paralogous sequences. Homologous genes, for example, typically descend from a common ancestral DNA sequence, either through a speciation event (orthologous genes) or a genetic duplication event (paralogous genes). “Orthologous” genes include genes in different species that evolved from a common ancestral gene by speciation. Orthologs typically retain the same function in the course of evolution. “Paralogous” genes include genes related by duplication within a genome. Paralogs can evolve new functions in the course of evolution.

[0065] The term “in vitro” includes artificial environments and to processes or reactions that occur within an artificial environment (e.g., a test tube or an isolated cell or cell line). The term “in vivo” includes natural environments (e.g., a cell, organism, or body) and to processes or reactions that occur within a natural environment. The term “ex vivo” includes cells that have been removed from the body of an individual and processes or reactions that occur within such cells.

[0066] The term “reporter gene” refers to a nucleic acid having a sequence encoding a gene product (typically an enzyme) that is easily and quantifiably assayed when a construct comprising the reporter gene sequence operably linked to a heterologous promoter and / or enhancer element is introduced into cells containing (or which can be made to contain) the factors necessary for the activation of the promoter and / or enhancer elements. Examples of reporter genes include, but are not limited, to genes encoding beta-galactosidase (lacZ), the bacterial chloramphenicol acetyltransferase (cat) genes, firefly luciferase genes, genes encoding beta-glucuronidase (GUS), and genes encoding fluorescent proteins. A “reporter protein” refers to a protein encoded by a reporter gene.

[0067] The term “fluorescent reporter protein” as used herein means a reporter protein that is detectable based on fluorescence wherein the fluorescence may be either from the reporter protein directly, activity of the reporter protein on a fluorogenic substrate, or a protein with affinity for binding to a fluorescent tagged compound. Examples of fluorescent proteins include green fluorescent proteins (e.g., GFP, GFP-2, tagGFP, turboGFP, eGFP, Emerald, Azami Green, Monomeric Azami Green, CopGFP, AceGFP, and ZsGreenl), yellow fluorescent proteins (e.g.,Attorney Docket No.057766 / 629304 YFP, eYFP, Citrine, Venus, YPet, PhiYFP, and ZsYellowl), blue fluorescent proteins (e.g., BFP, eBFP, eBFP2, Azurite, mKalamal, GFPuv, Sapphire, and T-sapphire), cyan fluorescent proteins (e.g., CFP, eCFP, Cerulean, CyPet, AmCyanl, and Midoriishi-Cyan), red fluorescent proteins (e.g., RFP, mKate, mKate2, mPlum, DsRed monomer, mCherry, mRFP1, DsRed-Express, DsRed2, DsRed-Monomer, HcRed-Tandem, HcRedl, AsRed2, eqFP611, mRaspberry, mStrawberry, and Jred), orange fluorescent proteins (e.g., mOrange, mKO, Kusabira-Orange, Monomeric Kusabira-Orange, mTangerine, and tdTomato), and any other suitable fluorescent protein whose presence in cells can be detected by flow cytometry methods.

[0068] Compositions or methods “comprising” or “including” one or more recited elements may include other elements not specifically recited. For example, a composition that “comprises” or “includes” a protein may contain the protein alone or in combination with other ingredients. The transitional phrase “consisting essentially of” means that the scope of a claim is to be interpreted to encompass the specified elements recited in the claim and those that do not materially affect the basic and novel characteristic(s) of the claimed invention. Thus, the term “consisting essentially of” when used in a claim of this invention is not intended to be interpreted to be equivalent to “comprising.”

[0069] “Optional” or “optionally” means that the subsequently described event or circumstance may or may not occur and that the description includes instances in which the event or circumstance occurs and instances in which the event or circumstance does not.

[0070] Designation of a range of values includes all integers within or defining the range, and all subranges defined by integers within the range. For example, 5-10 nucleotides is understood as 5, 6, 7, 8, 9, or 10 nucleotides, whereas 5-10% is understood to contain 5% and all possible values through 10%.

[0071] At least 17 nucleotides of a 20 nucleotide sequence is understood to include 17, 18, 19, or 20 nucleotides of the sequence provided, thereby providing an upper limit even if one is not specifically provided as it would be clearly understood. Similarly, up to 3 nucleotides would be understood to encompass 0, 1, 2, or 3 nucleotides, providing a lower limit even if one is not specifically provided. When “at least,” “up to,” or other similar language modifies a number, it can be understood to modify each number in the series.

[0072] As used herein, “no more than” or “less than” is understood as the value adjacent to the phrase and logical lower values or integers, as logical from context, to zero. For example, aAttorney Docket No.057766 / 629304 duplex region of “no more than 2 nucleotide base pairs” has a 2, 1, or 0 nucleotide base pairs. When “no more than” or “less than” is present before a series of numbers or a range, it is understood that each of the numbers in the series or range is modified.

[0073] Unless otherwise apparent from the context, the term “about” encompasses values ± 5% of a stated value. In certain embodiments, the term “about” is understood to encompass tolerated variation or error within the art, e.g., 2 standard deviations from the mean, or the sensitivity of the method used to take a measurement, or a percent of a value as tolerated in the art, e.g., with age. When “about” is present before the first value of a series, it can be understood to modify each value in the series.

[0074] The term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).

[0075] The term “or” refers to any one member of a particular list and also includes any combination of members of that list.

[0076] The singular forms of the articles “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “a protein” or “at least one protein” can include a plurality of proteins, including mixtures thereof.

[0077] Statistically significant means p ≤0.05.

[0078] In the event of a conflict between a sequence in the application and an indicated accession number or position in an accession number, the sequence in the application predominates. DETAILED DESCRIPTION I. Overview

[0079] Provided herein are methods of determining the genomic location of a genomically integrated exogenous DNA. The methods can use nuclease agents to cleave genomic DNA (e.g., simultaneously cleave a single genomic DNA sample) at a first site and second site within the genomically integrated exogenous DNA to generated cleaved genomic DNA for the preparation of a single long-read sequencing library (i.e., to be used in a single long-read sequencing run). The long-read sequencing proceeds outward, in opposite directions, from the first and second cleavage sites within the genomically integrated exogenous DNA to generate a plurality ofAttorney Docket No.057766 / 629304 sequences. These long-read sequences are then used to generate hybrid alignments, meaning alignments of the sequences with both a reference genome and also the exogenous DNA sequence, thereby determining the genomic location of the genomically integrated exogenous DNA.

[0080] Transgenic modifications are frequently used in model organism development. The use of large targeting vectors allows for large genomic modifications (e.g., 100 base pairs (bp)- 200 kilobases (kb)), including but not limited to gene knockouts, gene knock-ins, partial gene deletions, exon modifications, genomic alterations for alternate splice variants, and transgenic constructs for producing foreign gene products or fragments thereof. Transgenic modifications are traditionally generated by homologous recombination (HR) in embryonic stem cells. This allows for all or a majority of the cells to have the desired genotype in the resulting organism. Although homologous recombination is efficient, transgene integration at off-target sites in the genome can occur. Identifying the precise location of genomically integrated exogenous DNA, particularly in the case of large genomic modifications, has remained a challenge.

[0081] Methods of identifying targeted genomic integration of exogenous DNA have traditionally relied upon searching in the expected transgenic integration locus, for example, using primers to known locations in a safe harbor locus being targeted, amplifying those regions by PCR, and sequencing to determine whether exogenous DNA has been inserted. However, even if a targeted insertion is detected, this biased method fails to identify other integration events that may have occurred outside the targeting region.

[0082] Other methods have relied upon developing PCR assays that would amplify specific regions of a transgene insert such that the amplicons produced would extend through the homology arms of the construct and into the endogenous locus. However, such assays only have limited applicability to instances in which the homology arms are relatively short. In the case of targeting vectors with large homology arms, the amplicons fail to extend into surrounding endogenous genomic sequence to identify the insertion site. Moreover, the necessity of positioning the primers or guides in close proximity to the homology arms results in very little sequence information provided on the transgene itself, which could contain any number of mutations, deletions, or rearrangements that might have occurred prior to or during recombination. The greatest drawback of this method, however, is the intensive labor and time required to develop these PCR assays for every transgenic project.Attorney Docket No.057766 / 629304

[0083] Next generation sequencing methods, including targeted locus amplification and whole genome sequencing, have also been employed to identify sites of genomic integration of exogenous DNA. However, these methods require amplification and / or significant amounts of starting material to support the method. In the case of whole genome sequencing, it is an entirely untargeted method that also requires significant sequencing output, extensive data processing, and intensive analysis. These methods also fail to provide genomic location with high fidelity because they rely upon the sequencing and subsequent reassembly of many short reads, which again, presents a significant issue in the case of transgene inserts containing large homology arms that overlap with endogenous sequence.

[0084] The advent of long-read sequencing has presented new possibilities for accurate identification of the site of genomic integration of exogenous DNA, particularly for large inserts. These sequencing technologies produce sequence data by generating individual reads that are each derived from a single polynucleotide that is >1,000 nucleotides in length. However, efforts to apply such technology to identifying the sites of genomic integration of exogenous DNA have largely been unsuccessful. In addition, long-read assays designed around the target integration site suffer the same bias as PCR-based assays and fail to identify off-target integrations. Moreover, long-read assays designed to originate within sequences specific to the exogenous DNA insert using nuclease target sequences or guide RNA target sequences that overlap are similarly labor and time intensive to develop for every transgenic project as PCR-based assays and / or require significant amounts of starting material (e.g., >10 μg genomic DNA) to support the production of multiple sequencing libraries for multiple sequencing runs required by such methods.

[0085] The methods disclosed herein provide nuclease-targeted enrichment for long-read sequencing to determine the genomic location of a genomically integrated exogenous DNA. The enrichment refers to the larger proportion of reads that originate from the genomic locus with the integrated exogenous DNA in the methods described herein compared to its copy number in the genome. In non-enriched sequencing, like whole genome sequencing, all loci of the genome have an equal availability to enter the pore and be sequenced. Therefore, if ten reads are required at any genomic locus, ten whole genomes need to be sequenced. For this reason, whole genome sequencing requires a large amount of data to ensure that the target or region of interest in the genome gets sequenced with enough coverage to understand the alleles. With the methodAttorney Docket No.057766 / 629304 described herein, we routinely get over 50x to over 100x coverage at the genomic locus with the integrated exogenous DNA with less than one whole genome worth of data. This increases the efficiency of the sequencing process, reduces the input DNA required, increases the coverage of the region of interest, and reduces the computational resources needed to make the transgenic identifications.

[0086] In some embodiments, these methods provide the ability to target common elements (e.g., selection cassettes, reporter genes, AAV vector sequences, etc.) found across transgenic projects, and through the use of nuclease-mediated enrichment using two cleavage sites separated by a sufficient distance to enable to separate cleavage events on the same genomic DNA, provide target-enriched, long-read sequences from a single sequencing library preparation. Using a hybrid alignment of these long read sequences with both a reference genome and also a known portion of the exogenous DNA, the methods provide the targeted and efficient, yet unbiased, identification of any and all genomic insertions of exogenous DNA. These methods require less starting material, allowing for analysis of animals without the need to sacrifice to obtain sufficient tissue for analysis, and can be performed using less time, fewer resources, and more sequence coverage than previous methods. II. Methods of Determining Genomic Location of Genomically Integrated Exogenous DNA

[0087] Some methods disclosed herein for determining the genomic location of genomically integrated exogenous DNA comprise cleaving genomic DNA extracted from cells (or tissue comprising the cells) comprising a genomically integrated exogenous DNA with: (i) a first nuclease agent that cleaves a first nuclease cleavage site within the genomically integrated exogenous DNA to generate a first cleaved genomic DNA segment (e.g., a first 5′ nuclease cleavage site to generate a first cleaved genomic DNA segment 5′ of the first 5′ nuclease cleavage site) and (ii) a second nuclease agent that cleaves a second nuclease cleavage site within the genomically integrated exogenous DNA to generate second cleaved genomic DNA segment (e.g., a second 3′ nuclease cleavage site to generate a second cleaved genomic DNA segment 3′ of the second 3′ nuclease cleavage site), wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 50 bp. In such methods, a single genomic DNA sample can be cleaved with both the first nuclease agent and the second nuclease agent (as opposed to cleaving a first genomic DNA sample with a first nuclease agent, andAttorney Docket No.057766 / 629304 cleaving a separate second genomic DNA sample with the second nuclease agent). For example, in such methods, a single genomic DNA sample can be cleaved simultaneously with both the first nuclease agent and the second nuclease agent (as opposed to cleaving a genomic DNA sample with a first nuclease agent, and then cleaving the genomic DNA sample with the second nuclease agent at a later time). By cleaved simultaneously, it is meant that both the first and second nuclease agents are mixed together in the same mixture with the genomic DNA sample. We have recognized that there is directionality to the sequencing reads that appear from the nuclease (e.g., Cas9) cleavage sites, where the forward guides (e.g., the guide with PAM sequence downstream) and the reverse guides (e.g., the guide with the PAM sequence upstream) produce the majority of its resulting sequencing data in the direction in which it is designed; i.e., the sequencing reads proceed outward in both directions from the nuclease cleavage sites within the exogenous DNA into the surrounding endogenous genomic DNA.

[0088] In some methods, the 5′ ends of the genomic DNA (e.g., after DNA extraction but before cleavage with the nuclease agents) are dephosphorylated to reduce ligation of sequencing adaptors to non-target strands. Cleavage by the nuclease agents then generates blunt ends or overhangs with ligatable 5′ phosphates. Ligation of two DNA fragments generally needs at least one of the DNA ends to have 5′ phosphate for the nucleophilic attack of the 3'OH end to the phosphate group. In some methods, the cleaved genomic DNA segments are dA-tailed, which prepares the cleaved ends for ligation to the sequencing adaptors. In such methods, sequencing adaptors are ligated primarily to nuclease agent cut sites, which are both 3′ dA-tailed and 5′ phosphorylated. For example, in the case of CRISPR / Cas (e.g., Cas9), after cleavage with the CRISPR / Cas (e.g., Cas9), the Cas enzyme (e.g., Cas9) can remain bound to the DNA on the 5′ side of the cleavage (e.g., upstream of the PAM), resulting in preferential ligation of adaptors onto the 3′ side of the cleavage. Therefore, in some embodiments in which CRISPR / Cas is used as the first and second nuclease agents, guide RNAs are designed to target guide RNA target sequences on opposite strands (i.e., the PAMs are on opposite strands), with the guide RNA target sequence on the sense strand being located 3′ of the guide RNA target sequence on the antisense strand.

[0089] The exogenous DNA can be of any size. In one example, the exogenous DNA is at least 0.5 kb, at least 1 kb, or at least 5 kb in length. In one example, the exogenous DNA is at least 10 kb in length. In another example, the exogenous DNA is at least 100 kb in length.Attorney Docket No.057766 / 629304 Alternatively, the exogenous DNA can be at least 15 kb, at least 20 kb, at least 25 kb, at least 50 kb, at least 75 kb, at least 100 kb, at least 125 kb, at least 150 kb, at least 175 kb, or at least 200 kb in length. The exogenous DNA can also be from about 10 kb to about 100 kb, about 50 kb to about 100 kb, about 50 kb to about 150 kb, about 10 kb to about 200 kb, about 15 kb to about 200 kb, about 20 kb to about 200 kb, about 25 kb to about 200 kb, about 50 kb to about 200 kb, about 75 kb to about 200 kb, about 100 kb to about 200 kb, about 125 kb to about 200 kb, or about 150 kb to about 200 kb in length. For example, the exogenous DNA could be about 10 kb or 10 kb in length. Alternatively, it could be about 100 kb or 100 kb in length. The exogenous DNA could also be about 25 kb, about 50 kb, about 75 kb, about 125 kb, about 150 kb, about 175 kb, or about 250 kb in length.

[0090] Some methods do not comprise amplifying the genomic DNA. Some methods do not comprise amplifying the first cleaved genomic DNA. Some methods do not comprise amplifying the second cleaved genomic DNA. In one example, the method does not comprise amplifying the genomic DNA, the first cleaved genomic DNA, or the second cleaved genomic DNA.

[0091] The methods can be performed on less or significantly less DNA than traditional whole genome sequencing methods, which would typically require at least 10 μg of genomic DNA. For example, the amount of genomic DNA in the method can be as little as 400 ng. In some methods, the amount of genomic DNA used is at least 400 ng. In some methods, the amount of genomic DNA used is at least 500 ng, at least 600 ng, at least 700 ng, at least 800 ng, at least 900 ng, at least 1 μg, at least 1.5 μg, at least 2 μg, at least 2.5 μg, or at least 3 μg. In some methods, the amount of genomic DNA used is less than 3 μg. In some methods, the amount of genomic DNA used is less than 2.5 μg, less than 5 μg, less than 1.5 μg, or less than 1 μg. In some methods, the amount of genomic DNA is from about 400 ng to about 3 μg, about 400 ng to about 500 ng, about 500 ng to about 600 ng, about 600 ng to about 700 ng, about 700 ng to about 800 ng, about 800 ng to about 900 ng, about 900 ng to about 1 μg, about 1 μg to about 1.5 μg, about 1.5 μg to about 2 μg, about 2 μg to about 2.5 μg, about 2.5 μg to about 3 μg, about 400 ng to about 2.5 μg, about 400 ng to about 2 μg, about 400 ng to about 1.5 μg, about 400 ng to about 1 μg, about 400 ng to about 900 ng, about 400 ng to about 800 ng, about 400 ng to about 700 ng, about 400 ng to about 600 ng, or about 400 ng to about 500 ng.

[0092] The methods can be performed on genomic DNA from any cell type. For example, the cells can be mammalian cells. The mammalian cells can be human cells. For example, theAttorney Docket No.057766 / 629304 human cells can be human induced pluripotent stem cells. Alternatively, the mammalian cells can be rodent cells. The rodent cells can be mouse cells, such as mouse embryonic stem (ES) cells. The rodent cells can also be rat cells, such as rat ES cells. The rodent cells can also be cells taken from a living rodent (e.g., a toe, ear, or tail clipping from a mouse or a rat). The amount of genomic DNA found in such a sample taken from a living rodent (e.g., a toe, ear, or tail clipping) is sufficient to perform the disclosed methods. That is, tissues can be used as well, such as biopsies, liver dissections (or other internal organ dissections), toe clips, ear punches, tail clips, and the like. The small amount of genomic DNA required for the methods (e.g., ~400 ng in some cases) may be advantageous, compared to the tissue requirements for traditional whole genome sequencing methods, because it allows the identification of the genomically integrated exogenous DNA without the need to sacrifice the organism (e.g., rodent), ultimately saving time, cost, and labor. The ability to perform the methods described herein with biopsies allow for identification of targeting events early and without complicated breeding needed, while also keeping the animal alive, thereby reducing cost and increasing efficiency. The methods can, optionally, further comprise extracting the exogenous DNA from the cells or tissues described above. Methods of extracting DNA from cells and / or tissues will be familiar to one of skill in the art and any such appropriate method of DNA extraction can be used.

[0093] The methods can also further comprise genetically modifying a population of cells to generate the cells comprising the genomically integrated exogenous DNA prior to extracting the genomic DNA.

[0094] In some embodiments, the genetically modifying can comprise introducing into the population of cells an exogenous donor nucleic acid comprising the exogenous DNA. The exogenous donor nucleic acid can be inserted into a genomic locus in the genome or can recombine with the genomic locus to generate the cells that comprise the genomically integrated exogenous DNA. In some embodiments, the genetically modifying can comprise introducing into the population of cells (1) a nuclease agent or one or more nucleic acids encoding the nuclease agent, wherein the nuclease agent targets a nuclease target sequence in a target genomic locus and (2) an exogenous donor nucleic acid comprising the exogenous DNA. The nuclease can cleave the target genomic locus, and the exogenous donor nucleic acid can be inserted into the target genomic locus or can recombine with the target genomic locus to generate the cells that comprise the genomically integrated exogenous DNA. However, the skilled person is awareAttorney Docket No.057766 / 629304 that alternative methods can also be used.

[0095] In methods using a nuclease agent, any suitable nuclease agent can be used. In some embodiments, for example, the methods can utilize nuclease agents such as Clustered Regularly Interspersed Short Palindromic Repeats (CRISPR) / CRISPR-associated (Cas) systems, zinc finger nuclease (ZFN) systems, or Transcription Activator-Like Effector Nuclease (TALEN) systems or components of such systems to modify a target genomic locus. Generally, the nuclease agents involve the use of engineered cleavage systems to induce a double strand break or a nick (i.e., a single strand break) in a nuclease target site. Cleavage or nicking can occur through the use of specific nucleases such as engineered ZFNs, TALENs, or CRISPR / Cas systems with an engineered guide RNA to guide specific cleavage or nicking of the nuclease target site. Any nuclease agent that induces a nick or double-strand break at a desired target sequence can be used in the methods and compositions disclosed herein.

[0096] In some embodiments, the nuclease agent is a CRISPR / Cas system. In some embodiments, the nuclease agent comprises one or more ZFNs. In some embodiments, the nuclease agent comprises one or more TALENs. These nuclease agents are described in more detail elsewhere herein.

[0097] Exogenous Donor Nucleic Acids. Any suitable exogenous donor nucleic acid can be used in the methods disclosed herein. In some embodiments, the exogenous donor nucleic acid recombines with a genomic locus via non-homologous end joining (NHEJ)-mediated ligation or through a homology-directed repair event.

[0098] Some exogenous donor nucleic acids comprise homology arms. Other exogenous donor nucleic acids do not comprise homology arms. The exogenous donor nucleic acids can be capable of insertion into a genomic locus by homology-directed repair, and / or they can be capable of insertion into a genomic locus by non-homologous end joining.

[0099] Exogenous donor nucleic acids can comprise deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), they can be single-stranded or double-stranded, and they can be in linear or circular form. For example, an exogenous donor nucleic acid can be a single-stranded oligodeoxynucleotide (ssODN). See, e.g., Yoshimi et al. (2016) Nat. Commun.7:10431, herein incorporated by reference in its entirety for all purposes. Exogenous donor nucleic acids can be naked nucleic acids or can be delivered by viruses, such as AAV. In a specific example, the exogenous donor nucleic acid can be delivered via AAV and can be capable of insertion into aAttorney Docket No.057766 / 629304 genomic locus by non-homologous end joining (e.g., the exogenous donor nucleic acid can be one that does not comprise homology arms).

[0100] In one example, an exogenous donor nucleic acid is between about 50 nucleotides to about 5 kb in length, is between about 50 nucleotides to about 3 kb in length, or is between about 50 to about 1,000 nucleotides in length. In some embodiments, other exogenous donor nucleic acids are between about 40 to about 200 nucleotides in length. For example, an exogenous donor nucleic acid can be between about 50-60, 60-70, 70-80, 80-90, 90-100, 100-110, 110-120, 120- 130, 130-140, 140-150, 150-160, 160-170, 170-180, 180-190, or 190-200 nucleotides in length. Alternatively, an exogenous donor nucleic acid can be between about 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 nucleotides in length. Alternatively, an exogenous donor nucleic acid can be between about 1-1.5, 1.5-2, 2-2.5, 2.5-3, 3-3.5, 3.5-4, 4-4.5, or 4.5-5 kb in length. Alternatively, an exogenous donor nucleic acid can be, for example, no more than 5 kb, 4.5 kb, 4 kb, 3.5 kb, 3 kb, 2.5 kb, 2 kb, 1.5 kb, 1 kb, 900 nucleotides, 800 nucleotides, 700 nucleotides, 600 nucleotides, 500 nucleotides, 400 nucleotides, 300 nucleotides, 200 nucleotides, 100 nucleotides, or 50 nucleotides in length. Exogenous donor nucleic acids (e.g., targeting vectors or large targeting vectors) can also be longer.

[0101] In one example, an exogenous donor nucleic acid is a single-stranded oligodeoxynucleotide (ssODN) that is between about 80 nucleotides and about 200 nucleotides in length. In another example, an exogenous donor nucleic acid is an ssODN that is between about 80 nucleotides and about 3 kb in length. Such an ssODN can have homology arms, for example, that are each between about 40 nucleotides and about 60 nucleotides in length. Such an ssODN can also have homology arms, for example, that are each between about 30 nucleotides and 100 nucleotides in length. The homology arms can be symmetrical (e.g., each 40 nucleotides or each 60 nucleotides in length), or they can be asymmetrical (e.g., one homology arm that is 36 nucleotides in length, and one homology arm that is 91 nucleotides in length).

[0102] Exogenous donor nucleic acids can include modifications or sequences that provide for additional desirable features (e.g., modified or regulated stability; tracking or detecting with a fluorescent label; a binding site for a protein or protein complex; and so forth). Exogenous donor nucleic acids can comprise one or more fluorescent labels, purification tags, epitope tags, or a combination thereof. For example, an exogenous donor nucleic acid can comprise one or more fluorescent labels (e.g., fluorescent proteins or other fluorophores or dyes), such as at least 1, atAttorney Docket No.057766 / 629304 least 2, at least 3, at least 4, or at least 5 fluorescent labels. Exemplary fluorescent labels include fluorophores such as fluorescein (e.g., 6-carboxyfluorescein (6-FAM)), Texas Red, HEX, Cy3, Cy5, Cy5.5, Pacific Blue, 5-(and-6)-carboxytetramethylrhodamine (TAMRA), and Cy7. A wide range of fluorescent dyes are available commercially for labeling oligonucleotides (e.g., from Integrated DNA Technologies). Such fluorescent labels (e.g., internal fluorescent labels) can be used, for example, to detect an exogenous donor nucleic acid that has been directly integrated into a cleaved target nucleic acid having protruding ends compatible with the ends of the exogenous donor nucleic acid. The label or tag can be at the 5′ end, the 3′ end, or internally within the exogenous donor nucleic acid. For example, an exogenous donor nucleic acid can be conjugated at 5′ end with the IR700 fluorophore from Integrated DNA Technologies (5′IRDYE®700).

[0103] Exogenous donor nucleic acids can also comprise nucleic acid inserts including segments of DNA to be integrated in a genomic locus. Integration of a nucleic acid insert in the genomic locus can result in addition of a nucleic acid sequence of interest to the genomic locus, deletion of a nucleic acid sequence of interest in the genomic locus, or replacement of a nucleic acid sequence of interest in the genomic locus (i.e., deletion and insertion; or substitution). Some exogenous donor nucleic acids are designed for insertion of a nucleic acid insert in the genomic locus without any corresponding deletion in the genomic locus. Other exogenous donor nucleic acids are designed to delete a nucleic acid sequence of interest in the genomic locus without any corresponding insertion of a nucleic acid insert. Yet other exogenous donor nucleic acids are designed to delete a nucleic acid sequence of interest in the genomic locus and replace it with a nucleic acid insert (e.g., a substitution).

[0104] The nucleic acid insert or the corresponding nucleic acid at the genomic locus being deleted and / or replaced can be various lengths. An exemplary nucleic acid insert or corresponding nucleic acid at the genomic locus being deleted and / or replaced is between about 1 nucleotide to about 5 kb in length or is between about 1 nucleotide to about 1,000 nucleotides in length. For example, a nucleic acid insert or a corresponding nucleic acid at the genomic locus being deleted and / or replaced can be between about 1 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, about 120 to about 130, about 130 to about 140, about 140 to about 150,Attorney Docket No.057766 / 629304 about 150 to about 160, about 160 to about 170, about 170 to about 180, about 180 to about 190, or about 190 to about 200 nucleotides in length. Likewise, a nucleic acid insert or a corresponding nucleic acid at the genomic locus being deleted and / or replaced can be between about 1 to about 100, about 100 to about 200, about 200 to about 300, about 300 to about 400, about 400 to about 500, about 500 to about 600, about 600 to about 700, about 700 to about 800, about 800 to about 900, or about 900 to about 1,000 nucleotides in length. Likewise, a nucleic acid insert or a corresponding nucleic acid at the genomic locus being deleted and / or replaced can be between about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, about 2 kb to about 2.5 kb, about 2.5 kb to about 3 kb, about 3 kb to about 3.5 kb, about 3.5 kb to about 4 kb, about 4 kb to about 4.5 kb, or about 4.5 kb to about 5 kb in length. A nucleic acid being deleted from a genomic locus can also be between about 1 kb to about 5 kb, about 5 kb to about 10 kb, about 10 kb to about 20 kb, about 20 kb to about 30 kb, about 30 kb to about 40 kb, about 40 kb to about 50 kb, about 50 kb to about 60 kb, about 60 kb to about 70 kb, about 70 kb to about 80 kb, about 80 kb to about 90 kb, about 90 kb to about 100 kb, about 100 kb to about 200 kb, about 200 kb to about 300 kb, about 300 kb to about 400 kb, about 400 kb to about 500 kb, about 500 kb to about 600 kb, about 600 kb to about 700 kb, about 700 kb to about 800 kb, about 800 kb to about 900 kb, about 900 kb to about 1 Mb or longer. Alternatively, a nucleic acid being deleted from a genomic locus can be between about 1 Mb to about 1.5 Mb, about 1.5 Mb to about 2 Mb, about 2 Mb to about 2.5 Mb, about 2.5 Mb to about 3 Mb, about 3 Mb to about 4 Mb, about 4 Mb to about 5 Mb, about 5 Mb to about 10 Mb, about 10 Mb to about 20 Mb, about 20 Mb to about 30 Mb, about 30 Mb to about 40 Mb, about 40 Mb to about 50 Mb, about 50 Mb to about 60 Mb, about 60 Mb to about 70 Mb, about 70 Mb to about 80 Mb, about 80 Mb to about 90 Mb, or about 90 Mb to about 100 Mb.

[0105] The nucleic acid insert or the corresponding nucleic acid at the genomic locus being deleted and / or replaced can be a coding region such as an exon; a non-coding region such as an intron, an untranslated region, or a regulatory region (e.g., a promoter, an enhancer, or a transcriptional repressor-binding element); or any combination thereof.

[0106] The nucleic acid insert can also comprise a conditional allele. The conditional allele can be a multifunctional allele, as described in US 2011 / 0104799, herein incorporated by reference in its entirety for all purposes.Attorney Docket No.057766 / 629304

[0107] Nucleic acid inserts can also comprise a polynucleotide encoding a selection marker. Alternatively, the nucleic acid inserts can lack a polynucleotide encoding a selection marker. The selection marker can be contained in a selection cassette. Optionally, the selection cassette can be a self-deleting cassette. See, e.g., US 8,697,851 and US 2013 / 0312129, each of which is herein incorporated by reference in its entirety for all purposes.

[0108] The nucleic acid insert can also comprise a reporter gene. Exemplary reporter genes include those encoding luciferase, β-galactosidase, green fluorescent protein (GFP), enhanced green fluorescent protein (eGFP), cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), blue fluorescent protein (BFP), enhanced blue fluorescent protein (eBFP), DsRed, ZsGreen, MmGFP, mPlum, mCherry, tdTomato, mStrawberry, J-Red, mOrange, mKO, mCitrine, Venus, YPet, Emerald, CyPet, Cerulean, T- Sapphire, and alkaline phosphatase. Such reporter genes can be operably linked to a promoter active in a cell being targeted.

[0109] The nucleic acid insert can also comprise one or more expression cassettes or deletion cassettes. A given cassette can comprise one or more of a nucleotide sequence of interest, a polynucleotide encoding a selection marker, and a reporter gene, along with various regulatory components that influence expression. Examples of selectable markers and reporter genes that can be included are discussed in detail elsewhere herein.

[0110] The nucleic acid insert can comprise a nucleic acid flanked with site-specific recombination target sequences. Alternatively, the nucleic acid insert can comprise one or more site-specific recombination target sequences. Although the entire nucleic acid insert can be flanked by such site-specific recombination target sequences, any region or individual polynucleotide of interest within the nucleic acid insert can also be flanked by such sites. Site- specific recombination target sequences, which can flank the nucleic acid insert or any polynucleotide of interest in the nucleic acid insert can include, for example, loxP, lox511, lox2272, lox66, lox71, loxM2, lox5171, FRT, FRT11, FRT71, attp, att, FRT, rox, or a combination thereof. In one example, the site-specific recombination sites flank a polynucleotide encoding a selection marker and / or a reporter gene contained within the nucleic acid insert. Following integration of the nucleic acid insert at a targeted locus, the sequences between the site-specific recombination sites can be removed.Attorney Docket No.057766 / 629304

[0111] Nucleic acid inserts can also comprise one or more restriction sites for restriction endonucleases (i.e., restriction enzymes), which include Type I, Type II, Type III, and Type IV endonucleases. Type I and Type III restriction endonucleases recognize specific recognition sites, but typically cleave at a variable position from the nuclease binding site, which can be hundreds of base pairs away from the cleavage site (recognition site). In Type II systems the restriction activity is independent of any methylase activity, and cleavage typically occurs at specific sites within or near to the binding site. Most Type II enzymes cut palindromic sequences, however Type IIa enzymes recognize non-palindromic recognition sites and cleave outside of the recognition site, Type IIb enzymes cut sequences twice with both sites outside of the recognition site, and Type IIs enzymes recognize an asymmetric recognition site and cleave on one side and at a defined distance of about 1-20 nucleotides from the recognition site. Type IV restriction enzymes target methylated DNA. Restriction enzymes are further described and classified, for example in the REBASE database (webpage at rebase.neb.com; Roberts et al., (2003) Nucleic Acids Res.31:418-420; Roberts et al., (2003) Nucleic Acids Res.31:1805-1812; and Belfort et al. (2002) in Mobile DNA II, pp.761-783, Eds. Craigie et al., (ASM Press, Washington, DC), each of which is herein incorporated by reference in its entirety for all purposes).

[0112] Donor Nucleic Acids for Non-Homologous-End-Joining-Mediated Insertion. Some exogenous donor nucleic acids are capable of insertion into a genomic locus by non-homologous end joining. In some cases, such exogenous donor nucleic acids do not comprise homology arms. For example, such exogenous donor nucleic acids can be inserted into a blunt end double-strand break following cleavage with a nuclease agent. In a specific example, the exogenous donor nucleic acid can be delivered via AAV and can be capable of insertion into a genomic locus by non-homologous end joining (e.g., the exogenous donor nucleic acid can be one that does not comprise homology arms).

[0113] In one example, the exogenous donor nucleic acid can be inserted via homology- independent targeted integration. For example, the insert sequence in the exogenous donor nucleic acid to be inserted into a target genomic locus can be flanked on each side by a target site for a nuclease agent (e.g., the same target site as in the target genomic locus, and the same nuclease agent being used to cleave the target site in the target genomic locus). The nuclease agent can then cleave the target sites flanking the insert sequence. In a specific example, the exogenous donor nucleic acid is delivered AAV-mediated delivery, and cleavage of the targetAttorney Docket No.057766 / 629304 sites flanking the insert sequence can remove the inverted terminal repeats (ITRs) of the AAV. In some methods, the target site in the target genomic locus (e.g., a gRNA target sequence including the flanking protospacer adjacent motif) is no longer present if the insert sequence is inserted into the target genomic locus in the correct orientation but it is reformed if the insert sequence is inserted into the target genomic locus in the opposite orientation. This can help ensure that the insert sequence is inserted in the correct orientation for expression.

[0114] Other exogenous donor nucleic acids have short single-stranded regions at the 5′ end and / or the 3′ end that are complementary to one or more overhangs created by nuclease-mediated cleavage in the target genomic locus ne. These overhangs can also be referred to as 5′ and 3′ homology arms. For example, some exogenous donor nucleic acids have short single-stranded regions at the 5′ end and / or the 3′ end that are complementary to one or more overhangs created by nuclease-mediated cleavage at 5′ and / or 3′ target sequences in the target genomic locus. Some such exogenous donor nucleic acids have a complementary region only at the 5′ end or only at the 3′ end. For example, some such exogenous donor nucleic acids have a complementary region only at the 5′ end complementary to an overhang created at a 5′ target sequence in the target genomic locus or only at the 3′ end complementary to an overhang created at a 3′ target sequence in the target genomic locus. Other such exogenous donor nucleic acids have complementary regions at both the 5′ and 3′ ends. For example, other such exogenous donor nucleic acids have complementary regions at both the 5′ and 3′ ends e.g., complementary to first and second overhangs, respectively, generated by nuclease-mediated cleavage in the target genomic locus. For example, if the exogenous donor nucleic acid is double-stranded, the single-stranded complementary regions can extend from the 5′ end of the top strand of the donor nucleic acid and the 5′ end of the bottom strand of the donor nucleic acid, creating 5′ overhangs on each end. Alternatively, the single-stranded complementary region can extend from the 3′ end of the top strand of the donor nucleic acid and from the 3′ end of the bottom strand of the template, creating 3′ overhangs.

[0115] The complementary regions can be of any length sufficient to promote ligation between the exogenous donor nucleic acid and the target nucleic acid. Exemplary complementary regions are between about 1 to about 5 nucleotides in length, between about 1 to about 25 nucleotides in length, or between about 5 to about 150 nucleotides in length. For example, a complementary region can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14,Attorney Docket No.057766 / 629304 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. Alternatively, the complementary region can be about 5-10, 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80- 90, 90-100, 100-110, 110-120, 120-130, 130-140, or 140-150 nucleotides in length, or longer.

[0116] Such complementary regions can be complementary to overhangs created by two pairs of nickases. Two double-strand breaks with staggered ends can be created by using first and second nickases that cleave opposite strands of DNA to create a first double-strand break, and third and fourth nickases that cleave opposite strands of DNA to create a second double-strand break. For example, a Cas protein can be used to nick first, second, third, and fourth guide RNA target sequences corresponding with first, second, third, and fourth guide RNAs. The first and second guide RNA target sequences can be positioned to create a first cleavage site such that the nicks created by the first and second nickases on the first and second strands of DNA create a double-strand break (i.e., the first cleavage site comprises the nicks within the first and second guide RNA target sequences). Likewise, the third and fourth guide RNA target sequences can be positioned to create a second cleavage site such that the nicks created by the third and fourth nickases on the first and second strands of DNA create a double-strand break (i.e., the second cleavage site comprises the nicks within the third and fourth guide RNA target sequences). Optionally, the nicks within the first and second guide RNA target sequences and / or the third and fourth guide RNA target sequences can be off-set nicks that create overhangs. The offset window can be, for example, at least about 5 bp, 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp or more. See Ran et al. (2013) Cell 154:1380-1389; Mali et al. (2013) Nat. Biotech.31:833-838; and Shen et al. (2014) Nat. Methods 11:399-404, each of which is herein incorporated by reference in its entirety for all purposes. In such cases, a double-stranded exogenous donor nucleic acid can be designed with single-stranded complementary regions that are complementary to the overhangs created by the nicks within the first and second guide RNA target sequences and by the nicks within the third and fourth guide RNA target sequences. Such an exogenous donor nucleic acid can then be inserted by non-homologous-end-joining-mediated ligation.

[0117] Donor Nucleic Acids for Insertion by Homology-Directed Repair. Some exogenous donor nucleic acids comprise homology arms. If the exogenous donor nucleic acid also comprises a nucleic acid insert, the homology arms can flank the nucleic acid insert. For ease of reference, the homology arms are referred to herein as 5′ and 3′ (i.e., upstream and downstream)Attorney Docket No.057766 / 629304 homology arms. This terminology relates to the relative position of the homology arms to the nucleic acid insert within the exogenous donor nucleic acid. The 5′ and 3′ homology arms correspond to regions within the target genomic locus, which are referred to herein as “5′ target sequence” and “3′ target sequence,” respectively.

[0118] A homology arm and a target sequence “correspond” or are “corresponding” to one another when the two regions share a sufficient level of sequence identity to one another to act as substrates for a homologous recombination reaction. The term “homology” includes DNA sequences that are either identical or share sequence identity to a corresponding sequence. The sequence identity between a given target sequence and the corresponding homology arm found in the exogenous donor nucleic acid can be any degree of sequence identity that allows for homologous recombination to occur. For example, the amount of sequence identity shared by the homology arm of the exogenous donor nucleic acid (or a fragment thereof) and the target sequence (or a fragment thereof) can be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity, such that the sequences undergo homologous recombination. Moreover, a corresponding region of homology between the homology arm and the corresponding target sequence can be of any length that is sufficient to promote homologous recombination. Exemplary homology arms are between about 25 nucleotides to about 2.5 kb in length, are between about 25 nucleotides to about 1.5 kb in length, or are between about 25 to about 500 nucleotides in length. For example, a given homology arm (or each of the homology arms) and / or corresponding target sequence can comprise corresponding regions of homology that are between about 25-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, 150- 200, 200-250, 250-300, 300-350, 350-400, 400-450, or 450-500 nucleotides in length, such that the homology arms have sufficient homology to undergo homologous recombination with the corresponding target sequences within the target nucleic acid. Alternatively, a given homology arm (or each homology arm) and / or corresponding target sequence can comprise corresponding regions of homology that are between about 0.5 kb to about 1 kb, about 1 kb to about 1.5 kb, about 1.5 kb to about 2 kb, or about 2 kb to about 2.5 kb in length. For example, the homology arms can each be about 750 nucleotides in length. The homology arms can be symmetrical (each about the same size in length), or they can be asymmetrical (one longer than the other).Attorney Docket No.057766 / 629304

[0119] When a nuclease agent is used in combination with an exogenous donor nucleic acid, the 5′ and 3′ target sequences are optionally located in sufficient proximity to the nuclease cleavage site (e.g., within sufficient proximity to the nuclease target sequence) so as to promote the occurrence of a homologous recombination event between the target sequences and the homology arms upon a single-strand break (nick) or double-strand break at the nuclease cleavage site. The term “nuclease cleavage site” includes a DNA sequence at which a nick or double- strand break is created by a nuclease agent (e.g., a Cas9 protein complexed with a guide RNA). The target sequences within the targeted locus that correspond to the 5′ and 3′ homology arms of the exogenous donor nucleic acid are “located in sufficient proximity” to a nuclease cleavage site if the distance is such as to promote the occurrence of a homologous recombination event between the 5′ and 3′ target sequences and the homology arms upon a single-strand break or double-strand break at the nuclease cleavage site. Thus, the target sequences corresponding to the 5′ and / or 3′ homology arms of the exogenous donor nucleic acid can be, for example, within at least 1 nucleotide of a given nuclease cleavage site or within at least 10 nucleotides to about 1,000 nucleotides of a given nuclease cleavage site. As an example, the nuclease cleavage site can be immediately adjacent to at least one or both of the target sequences.

[0120] The spatial relationship of the target sequences that correspond to the homology arms of the exogenous donor nucleic acid and the nuclease cleavage site can vary. For example, target sequences can be located 5′ to the nuclease cleavage site, target sequences can be located 3′ to the nuclease cleavage site, or the target sequences can flank the nuclease cleavage site.

[0121] In some such methods, genetically modifying the population of cells comprises administering a large targeting vector (LTVEC) to the population of cells, optionally in combination with a nuclease agent, wherein the large targeting vector comprises a 5′ homology arm, a 3′ homology arm, and optionally an insert nucleic acid flanked by the 5′ homology arm and the 3′ homology arm, wherein the large targeting vector is at least 10 kb in length, or wherein the sum total of the 5′ homology arm and the 3′ homology arm is at least 10 kb in length.

[0122] LTVECs include targeting vectors that comprise homology arms that correspond to and are derived from nucleic acid sequences larger than those typically used by other approaches intended to perform homologous recombination in cells. LTVECs also include targeting vectors comprising nucleic acid inserts having nucleic acid sequences larger than those typically used byAttorney Docket No.057766 / 629304 other approaches intended to perform homologous recombination in cells. For example, LTVECs make possible the modification of large loci that cannot be accommodated by traditional plasmid-based targeting vectors because of their size limitations. For example, the targeted locus can be (i.e., the 5′ and 3′ homology arms can correspond to) a locus of the cell that is not targetable using a conventional method or that can be targeted only incorrectly or only with significantly low efficiency in the absence of a nick or double-strand break induced by a nuclease agent (e.g., a Cas protein).

[0123] Examples of LTVECs include vectors derived from a bacterial artificial chromosome (BAC), a human artificial chromosome, or a yeast artificial chromosome (YAC). Non-limiting examples of LTVECs and methods for making them are described, e.g., in US Patent Nos. 6,586,251; 6,596,541; and 7,105,348; and in WO 2002 / 036789, each of which is herein incorporated by reference in its entirety for all purposes. LTVECs can be in linear form or in circular form.

[0124] LTVECs can be of any length and are typically at least 10 kb in length. For example, an LTVEC can be from about 50 kb to about 300 kb, from about 50 kb to about 75 kb, from about 75 kb to about 100 kb, from about 100 kb to 125 kb, from about 125 kb to about 150 kb, from about 150 kb to about 175 kb, from about 175 kb to about 200 kb, from about 200 kb to about 225 kb, from about 225 kb to about 250 kb, from about 250 kb to about 275 kb or from about 275 kb to about 300 kb. An LTVEC can also be from about 50 kb to about 500 kb, from about 100 kb to about 125 kb, from about 300 kb to about 325 kb, from about 325 kb to about 350 kb, from about 350 kb to about 375 kb, from about 375 kb to about 400 kb, from about 400 kb to about 425 kb, from about 425 kb to about 450 kb, from about 450 kb to about 475 kb, or from about 475 kb to about 500 kb. Alternatively, an LTVEC can be at least 10 kb, at least 15 kb, at least 20 kb, at least 30 kb, at least 40 kb, at least 50 kb, at least 60 kb, at least 70 kb, at least 80 kb, at least 90 kb, at least 100 kb, at least 150 kb, at least 200 kb, at least 250 kb, at least 300 kb, at least 350 kb, at least 400 kb, at least 450 kb, or at least 500 kb or greater.

[0125] The sum total of the 5′ homology arm and the 3′ homology arm in an LTVEC is typically at least 10 kb. As an example, the 5′ homology arm can range from about 5 kb to about 100 kb and / or the 3′ homology arm can range from about 5 kb to about 100 kb. As another example, the 5′ homology arm can range from about 5 kb to about 150 kb and / or the 3′ homology arm can range from about 5 kb to about 150 kb. Each homology arm can be, forAttorney Docket No.057766 / 629304 example, from about 5 kb to about 10 kb, from about 10 kb to about 20 kb, from about 20 kb to about 30 kb, from about 30 kb to about 40 kb, from about 40 kb to about 50 kb, from about 50 kb to about 60 kb, from about 60 kb to about 70 kb, from about 70 kb to about 80 kb, from about 80 kb to about 90 kb, from about 90 kb to about 100 kb, from about 100 kb to about 110 kb, from about 110 kb to about 120 kb, from about 120 kb to about 130 kb, from about 130 kb to about 140 kb, from about 140 kb to about 150 kb, from about 150 kb to about 160 kb, from about 160 kb to about 170 kb, from about 170 kb to about 180 kb, from about 180 kb to about 190 kb, or from about 190 kb to about 200 kb. The sum total of the 5′ and 3′ homology arms can be, for example, from about 10 kb to about 20 kb, from about 20 kb to about 30 kb, from about 30 kb to about 40 kb, from about 40 kb to about 50 kb, from about 50 kb to about 60 kb, from about 60 kb to about 70 kb, from about 70 kb to about 80 kb, from about 80 kb to about 90 kb, from about 90 kb to about 100 kb, from about 100 kb to about 110 kb, from about 110 kb to about 120 kb, from about 120 kb to about 130 kb, from about 130 kb to about 140 kb, from about 140 kb to about 150 kb, from about 150 kb to about 160 kb, from about 160 kb to about 170 kb, from about 170 kb to about 180 kb, from about 180 kb to about 190 kb, or from about 190 kb to about 200 kb. The sum total of the 5′ and 3′ homology arms can also be, for example, from about 200 kb to about 250 kb, from about 250 kb to about 300 kb, from about 300 kb to about 350 kb, or from about 350 kb to about 400 kb. Alternatively, each homology arm can be at least 5 kb, at least 10 kb, at least 15 kb, at least 20 kb, at least 30 kb, at least 40 kb, at least 50 kb, at least 60 kb, at least 70 kb, at least 80 kb, at least 90 kb, at least 100 kb, at least 110 kb, at least 120 kb, at least 130 kb, at least 140 kb, at least 150 kb, at least 160 kb, at least 170 kb, at least 180 kb, at least 190 kb, or at least 200 kb. Likewise, the sum total of the 5′ and 3′ homology arms can be at least 10 kb, at least 15 kb, at least 20 kb, at least 30 kb, at least 40 kb, at least 50 kb, at least 60 kb, at least 70 kb, at least 80 kb, at least 90 kb, at least 100 kb, at least 110 kb, at least 120 kb, at least 130 kb, at least 140 kb, at least 150 kb, at least 160 kb, at least 170 kb, at least 180 kb, at least 190 kb, or at least 200 kb. Each homology arm can also be at least 250 kb, at least 300 kb, at least 350 kb, or at least 400 kb.

[0126] LTVECs can comprise nucleic acid inserts having nucleic acid sequences larger than those typically used by other approaches intended to perform homologous recombination in cells. For example, an LTVEC can comprise a nucleic acid insert ranging from about 5 kb to about 10 kb, from about 10 kb to about 20 kb, from about 20 kb to about 40 kb, from about 40 kbAttorney Docket No.057766 / 629304 to about 60 kb, from about 60 kb to about 80 kb, from about 80 kb to about 100 kb, from about 100 kb to about 150 kb, from about 150 kb to about 200 kb, from about 200 kb to about 250 kb, from about 250 kb to about 300 kb, from about 300 kb to about 350 kb, from about 350 kb to about 400 kb, or greater. The LTVEC can also comprise a nucleic acid insert ranging, for example, from about 1 kb to about 5 kb, from about 400 kb to about 450 kb, from about 450 kb to about 500 kb, or greater. Alternatively, the nucleic acid insert can be at least 1 kb, at least 5 kb, at least 10 kb, at least 20 kb, at least 30 kb, at least 40 kb, at least 60 kb, at least 80 kb, at least 100 kb, at least 150 kb, at least 200 kb, at least 250 kb, at least 300 kb, at least 350 kb, at least 400 kb, at least 450 kb, or at least 500 kb.

[0127] The size of an LTVEC can be too large to enable screening of targeting events by conventional assays, e.g., southern blotting and long-range (e.g., 1 kb to 5 kb) PCR. In contrast, the present methods allow for identification of such genomically integrated LTVECs because sequencing reads begin at the nuclease cleavage sites in the nucleic acid insert. Thus, long-read sequencing can traverse the entire length of the surrounding portion of the LTVEC (e.g., not only the insert nucleic acid but also the homology arms) and provide sequence information for the surrounding genomic DNA to accurately determine the site of genomic integration.

[0128] Introduction of nucleic acids such as exogenous donor nucleic acids can, in some embodiments, be accomplished by virus-mediated delivery, such as AAV-mediated delivery or lentivirus-mediated delivery. The vectors can be, for example, viral vectors such as adeno- associated virus (AAV) vectors. The AAV may be any suitable serotype and may be a single- stranded AAV (ssAAV) or a self-complementary AAV (scAAV). Other exemplary viruses / viral vectors include retroviruses, lentiviruses, adenoviruses, vaccinia viruses, poxviruses, and herpes simplex viruses.

[0129] In some methods, the exogenous DNA is integrated into the genome via random genomic integration. In some methods, there is only partial genomic integration. A. Cleavage of Genomic DNA

[0130] The methods disclosed herein comprise cleaving genomic DNA extracted from cells comprising a genomically integrated exogenous DNA with a first nuclease agent and a second nuclease agent, at a first and second nuclease cleavage site, respectively, within the genomically integrated exogenous DNA to generate a first and a second cleaved genomic DNA segment. InAttorney Docket No.057766 / 629304 such methods, a single genomic DNA sample can be cleaved with both the first nuclease agent and the second nuclease agent (as opposed to cleaving a first genomic DNA sample with a first nuclease agent, and cleaving a separate second genomic DNA sample with the second nuclease agent). For example, in such methods, a single genomic DNA sample can be cleaved simultaneously with both the first nuclease agent and the second nuclease agent (as opposed to cleaving a genomic DNA sample with a first nuclease agent, and then cleaving the genomic DNA sample with the second nuclease agent at a later time).

[0131] In some methods, these nuclease cleavage sites are separated by at least 50 bp. In some methods, these nuclease cleavage sites are separated by at least 75 bp. In some methods, these nuclease cleavage sites are separated by at least 100 bp. In some methods, these nuclease cleavage sites are separated by at least 125 bp. In some methods, these nuclease cleavage sites are separated by at least 150 bp. In some methods, these nuclease cleavage sites are separated by at least 200 bp. In some methods, these nuclease cleavage sites are separated by at least 250 bp, or greater. Alternatively, the first nuclease cleavage site and the second nuclease cleavage site can be separated by about 50 bp to about 500 bp, about 75 bp to about 500 bp, about 100 bp to about 500 bp, about 125 bp to about 500 bp, about 150 bp to about 500 bp, about 200 bp to about 500 bp, about 250 bp to about 500 bp, about 300 bp to about 500 bp, about 350 bp to about 500 bp, about 400 bp to about 500 bp, or about 450 bp to about 500 bp. Alternatively, the first nuclease cleavage site and the second nuclease cleavage site can be separated by about 50 bp to about 1000 bp, about 75 bp to about 1000 bp, about 100 bp to about 1000 bp, about 125 bp to about 1000 bp, about 150 bp to about 1000 bp, about 200 bp to about 1000 bp, about 250 bp to about 1000 bp, about 300 bp to about 1000 bp, about 350 bp to about 1000 bp, about 400 bp to about 1000 bp, about 450 bp to about 1000 bp, or about 500 bp to about 1000 bp. Alternatively, the first nuclease cleavage site and the second nuclease cleavage site can be separated by about 50 bp to about 2000 bp, about 75 bp to about 2000 bp, about 100 bp to about 2000 bp, about 125 bp to about 2000 bp, about 150 bp to about 2000 bp, about 200 bp to about 2000 bp, about 250 bp to about 2000 bp, about 300 bp to about 2000 bp, about 350 bp to about 2000 bp, about 400 bp to about 2000 bp, about 450 bp to about 2000 bp, or about 500 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 50 bp to about 1000 bp or about 50 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 75 bp to about 1000 bp or about 75 bp to about 2000 bp. In some methods, these nuclease cleavageAttorney Docket No.057766 / 629304 sites are separated by about 100 bp to about 1000 bp or about 100 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 125 bp to about 1000 bp or about 125 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 150 bp to about 1000 bp or about 150 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 200 bp to about 1000 bp or about 200 bp to about 2000 bp. In some methods, these nuclease cleavage sites are separated by about 250 bp to about 1000 bp or about 250 bp to about 2000 bp. In one example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 150 bp or about 150 bp to about 250 bp. In another example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 125 bp or about 125 bp to about 500 bp. In another example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp or about 100 bp to about 500 bp. In another example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp. In another example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 125 bp. In another example, the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 250 bp.

[0132] In some methods, these nuclease cleavage sites are each at least 100 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the genomically integrated exogenous DNA and the genomic DNA). In some methods, these nuclease cleavage sites are each at least 200 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the genomically integrated exogenous DNA and the genomic DNA). In some methods, these nuclease cleavage sites are each at least 300 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the genomically integrated exogenous DNA and the genomic DNA). In some methods, these nuclease cleavage sites are each at least 400 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the genomically integrated exogenous DNA and the genomic DNA). In some methods, these nuclease cleavage sites are each at least 500 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the genomically integrated exogenous DNA and the genomic DNA). In some methods, these nuclease cleavage sites are each at least 600 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each at least 700 bp from the 5′ and 3′ ends ofAttorney Docket No.057766 / 629304 the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each at least 800 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each at least 900 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each at least 1 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each at least 1.5 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each about 100 bp to about 250 kb, about 200 bp to about 250 kb, about 300 bp to about 250 kb, about 400 bp to about 250 kb, about 500 bp to about 250 kb, about 600 bp to about 250 kb, about 700 bp to about 250 kb, about 800 bp to about 250 kb, about 900 bp to about 250 kb, about 1 kb to about 250 kb, or about 1.5 kb to about 250 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA. In some methods, these nuclease cleavage sites are each about 100 bp to about 200 kb, about 200 bp to about 200 kb, about 300 bp to about 200 kb, about 400 bp to about 200 kb, about 500 bp to about 200 kb, about 600 bp to about 200 kb, about 700 bp to about 200 kb, about 800 bp to about 200 kb, about 900 bp to about 200 kb, about 1 kb to about 200 kb, or about 1.5 kb to about 200 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA.

[0133] In some embodiments, the exogenous DNA sequence between the first nuclease cleavage site and the second nuclease cleavage site does not get sequenced or is sequenced with less frequency than the exogenous DNA in the outward direction from the cleavage sites. This is in contrast to the traditional design of nuclease cleavage sites to sequence an exogenous DNA insert, which would typically overlap to allow sequencing of the entirety of the exogenous DNA sequence. This novel approach to designing nuclease target sites separated by at least 50 bp, at least 100 bp, at least 150 bp, at least 200 bp, or at least 250 bp provides advantages for the methods disclosed herein. For example, providing a separation between the nuclease cleavage sites allows for deeper on-target sequencing in a single library preparation and sequencing run because the sequencing can proceed in both directions from the single library preparation. In contrast, the more traditional design of overlapping nuclease cleavage sites, a single library run would not be productive, as molecules in the library containing both double strand breaks would not produce long, genomic-spanning alignments for transgenic identification. The nucleaseAttorney Docket No.057766 / 629304 cleavage site design strategy of the methods also allows for efficient read alignment in the exogenous DNA for genomic location identification.

[0134] In some embodiments, in the case of CRISPR / Cas (e.g., Cas9), after cleavage with the CRISPR / Cas (e.g., Cas9), the Cas enzyme (e.g., Cas9) can remain bound to the DNA on the 5′ side of the cleavage (e.g., upstream of the PAM), resulting in preferential ligation of sequencing adaptors onto the 3′ side of the cleavage. Therefore, in some embodiments in which CRISPR / Cas is used as the first and second nuclease agents, guide RNAs are designed to target guide RNA target sequences on opposite strands (i.e., the PAMs are on opposite strands), with the guide RNA target sequence on the sense strand being located 3′ of the guide RNA target sequence on the antisense strand.

[0135] Separating the first nuclease cleavage site and the second nuclease cleavage site further allows for noncompetitive binding (i.e., no steric hinderance between binding of the first nuclease agent and the second nuclease agent) for nuclease-enriched targeted sequencing so that both the first nuclease agent and the second nuclease agent can cleave the same genomic DNA molecule. Although the nuclease cleavage sites are separated, they can remain within sufficient distance for the nuclease cleavage sites to be contained in common transgenic vector elements (i.e., both nuclease cleavage sites are contained in a single, common transgenic vector element), including but not limited to, reporter genes, polyA signal sequences, promoter sequences, antibiotic selection genes, AAV payload and backbone sequences, cloning vector sequences, or other vector sequences. In some embodiments, both nuclease cleavage sites are within a single antibiotic selection gene.

[0136] Providing sufficient distance between the nuclease cleavage sites and the 5′ and 3′ ends of the genomically integrated exogenous DNA allows for high quality alignments of the sequencing reads with the exogenous DNA sequence. Providing sufficient distance between the nuclease cleavage sites and the 5′ and 3′ ends of the genomically integrated exogenous DNA (i.e., the junctions of the exogenous DNA and the endogenous genomic DNA) is helpful because of the way sequence alignments build. Computationally, a sequence alignment is the result of a computational calculation of the likelihood of a query (read) aligning to a reference by chance. As the number of matched bases between the query and the reference increase, the likelihood of the alignment position occurring by chance decreases and the likelihood of correct alignment increases. The power of using long reads allows alignment scores to build by nature of moreAttorney Docket No.057766 / 629304 possibilities of matched bases. Although nanopore sequencing can have many single nucleotide base calling errors, the number of correct bases in an alignment builds a high score and can accommodate for the errors present. Preferably, the distance between the nuclease cleavage sites and the 5′ and 3′ ends of the genomically integrated exogenous DNA ensures that there is a large alignment score built from the nuclease cleavage sites through the exogenous DNA sequence and out to the genome. This creates a high alignment score, which allows of the identification of the exogenous DNA integration and gives confidence that the integration location is correctly identified, as a single read with high alignment scoring to the exogenous DNA and the genome in a specific locus is statistically improbable to occur by error. This improbability is amplified by the identification of the exogenous DNA integration at the same locus in different directions from the nuclease cleavage sites and with multiple reads.

[0137] Any rare-cutting nuclease agent can be used in the methods disclosed herein. A rare- cutting nuclease agent is a nuclease agent with a target sequence or recognition sequence that occurs rarely in a genome. Similarly, any nuclease agent with a target sequence or recognition sequence that does not occur outside of the intended cleavage site(s) in the targeting vectors described herein can be used. For example, any nuclease agent that does not have a target sequence or recognition sequence in the preexisting targeting vectors in the methods described herein can be used.

[0138] Any nuclease agent as described above that induces a double-strand break at a desired target sequence can be used in the methods and compositions disclosed herein. A naturally occurring or native nuclease agent can be employed so long as the nuclease agent induces a double-strand break in a desired target sequence. Alternatively, a modified or engineered nuclease agent can be employed. An “engineered nuclease agent” includes a nuclease that is engineered (modified or derived) from its native form to specifically recognize and induce a double-strand break in the desired target sequence. Thus, an engineered nuclease agent can be derived from a native, naturally occurring nuclease agent or it can be artificially created or synthesized. The engineered nuclease can induce a double-strand break in a target sequence, for example, wherein the target sequence is not a sequence that would have been recognized by a native (non-engineered or non-modified) nuclease agent. The modification of the nuclease agent can be as little as one amino acid in a protein cleavage agent or one nucleotide in a nucleic acidAttorney Docket No.057766 / 629304 cleavage agent. Producing a double-strand break in a target sequence or other DNA can be referred to herein as “cutting” or “cleaving” the target sequence or other DNA.

[0139] Active variants and fragments of the exemplified target sequences are also provided. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the given target sequence, wherein the active variants retain biological activity and hence are capable of being recognized and cleaved by a nuclease agent in a sequence-specific manner. Assays to measure the double- strand break of a target sequence by a nuclease agent are well-known. See, e.g., Frendewey et al. (2010) Methods in Enzymology 476:295-307, which is incorporated by reference herein in its entirety for all purposes.

[0140] Active variants and fragments of nuclease agents (i.e., an engineered nuclease agent) are also provided. Such active variants can comprise at least 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the native nuclease agent, wherein the active variants retain the ability to cut at a desired target sequence and hence retain double-strand-break-inducing activity. For example, any of the nuclease agents described herein can be modified from a native endonuclease sequence and designed to recognize and induce a double-strand break at a target sequence that was not recognized by the native nuclease agent. Thus, some engineered nucleases have a specificity to induce a double- strand break at a target sequence that is different from the corresponding native nuclease agent target sequence. Assays for double-strand-break-inducing activity are known and generally measure the overall activity and specificity of the endonuclease on DNA substrates containing the target sequence.

[0141] A nuclease target sequence includes a DNA sequence at which a double-strand break is induced by a nuclease agent. The length of the target sequence can vary, and includes, for example, target sequences that are about 30-36 bp for a zinc finger nuclease (ZFN) pair (i.e., about 15-18 bp for each ZFN), about 36 bp for a Transcription Activator-Like Effector Nuclease (TALEN), or about 20 bp for a CRISPR / Cas9 guide RNA.

[0142] CRISPR / Cas Systems. Clustered Regularly Interspersed Short Palindromic Repeats (CRISPR) / CRISPR-associated (Cas) systems can be used as the rare-cutting nuclease agents in the methods disclosed herein. CRISPR / Cas systems include transcripts and other elements involved in the expression of, or directing the activity of, Cas genes. A CRISPR / Cas system canAttorney Docket No.057766 / 629304 be, for example, a type I, a type II, a type III system, or a type V system (e.g., subtype V-A or subtype V-B). CRISPR / Cas systems used in the compositions and methods disclosed herein can be non-naturally occurring. A “non-naturally occurring” system includes anything indicating the involvement of the hand of man, such as one or more components of the system being altered or mutated from their naturally occurring state, being at least substantially free from at least one other component with which they are naturally associated in nature, or being associated with at least one other component with which they are not naturally associated. For example, some CRISPR / Cas systems employ non-naturally occurring CRISPR complexes comprising a gRNA and a Cas protein that do not naturally occur together, employ a Cas protein that does not occur naturally, or employ a gRNA that does not occur naturally.

[0143] Cas Proteins and Polynucleotides Encoding Cas Proteins. Cas proteins generally comprise at least one RNA recognition or binding domain that can interact with guide RNAs (gRNAs). Cas proteins can also comprise nuclease domains (e.g., DNase domains or RNase domains), DNA-binding domains, helicase domains, protein-protein interaction domains, dimerization domains, and other domains. Some such domains (e.g., DNase domains) can be from a native Cas protein. Other such domains can be added to make a modified Cas protein. A nuclease domain possesses catalytic activity for nucleic acid cleavage, which includes the breakage of the covalent bonds of a nucleic acid molecule. Cleavage can produce blunt ends or staggered ends, and it can be single-stranded or double-stranded. For example, a wild type Cas9 protein will typically create a blunt cleavage product. Alternatively, a wild type Cpf1 protein (e.g., FnCpf1) can result in a cleavage product with a 5-nucleotide 5′ overhang, with the cleavage occurring after the 18th base pair from the PAM sequence on the non-targeted strand and after the 23rd base on the targeted strand. A Cas protein can have full cleavage activity to create a double-strand break at a target genomic locus (e.g., a double-strand break with blunt ends), or it can be a nickase that creates a single-strand break at a target genomic locus (e.g., two nickases can be used to create a double-strand break).

[0144] Examples of Cas proteins include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Csy1, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4,Attorney Docket No.057766 / 629304 Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, and Cu1966, and homologs or modified versions thereof.

[0145] An exemplary Cas protein is a Cas9 protein or a protein derived from a Cas9 protein. Cas9 proteins are from a type II CRISPR / Cas system and typically share four key motifs with a conserved architecture. Motifs 1, 2, and 4 are RuvC-like motifs, and motif 3 is an HNH motif. Exemplary Cas9 proteins are from Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Neisseria meningitidis, or Campylobacter jejuni. Additional examples of the Cas9 family members are described in WO 2014 / 131833, herein incorporated by reference in its entirety for all purposes. Cas9 from S. pyogenes (SpCas9) (e.g., assigned UniProt accession number Q99ZW2) is an exemplary Cas9 protein. Smaller Cas9 proteins (e.g., Cas9 proteins whose coding sequences are compatible with the maximum AAV packaging capacity when combined with a guide RNA coding sequence and regulatory elements for the Cas9 and guide RNA, such as SaCas9 and CjCas9 and Nme2Cas9) are other exemplary Cas9 proteins. For example, Cas9 from S. aureus (SaCas9) (e.g., assigned UniProt accession number J7RUA5) is another exemplary Cas9 protein. Cas9 from Campylobacter jejuni (CjCas9) (e.g., assigned UniProt accession number Q0P897) is another exemplary Cas9 protein. See, e.g., Kim et al. (2017) Nat. Commun.8:14500, herein incorporated by reference in its entirety for all purposes.Attorney Docket No.057766 / 629304 SaCas9 is smaller than SpCas9, and CjCas9 is smaller than both SaCas9 and SpCas9. Cas9 from Neisseria meningitidis (Nme2Cas9) is another exemplary Cas9 protein. See, e.g., Edraki et al. (2019) Mol. Cell 73(4):714-726, herein incorporated by reference in its entirety for all purposes. Cas9 proteins from Streptococcus thermophilus (e.g., Streptococcus thermophilus LMD-9 Cas9 encoded by the CRISPR1 locus (St1Cas9) or Streptococcus thermophilus Cas9 from the CRISPR3 locus (St3Cas9)) are other exemplary Cas9 proteins. Cas9 from Francisella novicida (FnCas9) or the RHA Francisella novicida Cas9 variant that recognizes an alternative PAM (E1369R / E1449H / R1556A substitutions) are other exemplary Cas9 proteins. These and other exemplary Cas9 proteins are reviewed, e.g., in Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, herein incorporated by reference in its entirety for all purposes. Examples of Cas9 coding sequences, Cas9 mRNAs, and Cas9 protein sequences are provided in WO 2013 / 176772, WO 2014 / 065596, WO 2016 / 106121, WO 2019 / 067910, WO 2020 / 082042, US 2020 / 0270617, WO 2020 / 082041, US 2020 / 0268906, WO 2020 / 082046, and US 2020 / 0289628, each of which is herein incorporated by reference in its entirety for all purposes. Examples of ORFs and Cas9 amino acid sequences are provided in Table 30 at paragraph

[0449] WO 2019 / 067910, and examples of Cas9 mRNAs and ORFs are provided in paragraphs

[0214] -

[0234] of WO 2019 / 067910. See also WO 2020 / 082046 A2 (pp.84-85) and Table 24 in WO 2020 / 069296, each of which is herein incorporated by reference in its entirety for all purposes. An exemplary Cas9 protein sequence can comprise, consist essentially of, or consist of SEQ ID NO: 1. An exemplary DNA encoding the Cas9 protein can comprise, consist essentially of, or consist of SEQ ID NO: 2.

[0146] Another example of a Cas protein is a Cpf1 (CRISPR from Prevotella and Francisella 1) protein. Cpf1 is a large protein (about 1300 amino acids) that contains a RuvC- like nuclease domain homologous to the corresponding domain of Cas9 along with a counterpart to the characteristic arginine-rich cluster of Cas9. However, Cpf1 lacks the HNH nuclease domain that is present in Cas9 proteins, and the RuvC-like domain is contiguous in the Cpf1 sequence, in contrast to Cas9 where it contains long inserts including the HNH domain. See, e.g., Zetsche et al. (2015) Cell 163(3):759-771, herein incorporated by reference in its entirety for all purposes. Exemplary Cpf1 proteins are from Francisella tularensis 1, Francisella tularensis subsp. novicida, Prevotella albensis, Lachnospiraceae bacterium MC20171, Butyrivibrio proteoclasticus, Peregrinibacteria bacterium GW2011_GWA2_33_10, Parcubacteria bacteriumAttorney Docket No.057766 / 629304 GW2011_GWC2_44_17, Smithella sp. SCADC, Acidaminococcus sp. BV3L6, Lachnospiraceae bacterium MA2020, Candidatus Methanoplasma termitum, Eubacterium eligens, Moraxella bovoculi 237, Leptospira inadai, Lachnospiraceae bacterium ND2006, Porphyromonas crevioricanis 3, Prevotella disiens, and Porphyromonas macacae. Cpf1 from Francisella novicida U112 (FnCpf1; assigned UniProt accession number A0Q7Q2) is an exemplary Cpf1 protein.

[0147] Another example of a Cas protein is CasX (Cas12e). CasX is an RNA-guided DNA endonuclease that generates a staggered double-strand break in DNA. CasX is less than 1000 amino acids in size. Exemplary CasX proteins are from Deltaproteobacteria (DpbCasX or DpbCas12e) and Planctomycetes (PlmCasX or PlmCas12e). Like Cpf1, CasX uses a single RuvC active site for DNA cleavage. See, e.g., Liu et al. (2019) Nature 566(7743):218-223, herein incorporated by reference in its entirety for all purposes.

[0148] Another example of a Cas protein is CasΦ (CasPhi or Cas12j), which is uniquely found in bacteriophages. CasΦ is less than 1000 amino acids in size (e.g., 700-800 amino acids). CasΦ cleavage generates staggered 5′ overhangs. A single RuvC active site in CasΦ is capable of crRNA processing and DNA cutting. See, e.g., Pausch et al. (2020) Science 369(6501):333-337, herein incorporated by reference in its entirety for all purposes.

[0149] Cas proteins can be wild type proteins (i.e., those that occur in nature), modified Cas proteins (i.e., Cas protein variants), or fragments of wild type or modified Cas proteins. Cas proteins can also be active variants or fragments with respect to catalytic activity of wild type or modified Cas proteins. Active variants or fragments with respect to catalytic activity can comprise at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity to the wild type or modified Cas protein or a portion thereof, wherein the active variants retain the ability to cut at a desired cleavage site and hence retain nick-inducing or double-strand-break-inducing activity. Assays for nick-inducing or double-strand-break-inducing activity are known and generally measure the overall activity and specificity of the Cas protein on DNA substrates containing the cleavage site.

[0150] Cas proteins can be modified to increase or decrease one or more of nucleic acid binding affinity, nucleic acid binding specificity, and enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or aAttorney Docket No.057766 / 629304 Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity or a property of the Cas protein.

[0151] One example of a modified Cas protein is the modified SpCas9-HF1 protein, which is a high-fidelity variant of Streptococcus pyogenes Cas9 harboring alterations (N497A / R661A / Q695A / Q926A) designed to reduce non-specific DNA contacts. See, e.g., Kleinstiver et al. (2016) Nature 529(7587):490-495, herein incorporated by reference in its entirety for all purposes. Another example of a modified Cas protein is the modified eSpCas9 variant (K848A / K1003A / R1060A) designed to reduce off-target effects. See, e.g., Slaymaker et al. (2016) Science 351(6268):84-88, herein incorporated by reference in its entirety for all purposes. Other SpCas9 variants include K855A and K810A / K1003A / R1060A. These and other modified Cas proteins are reviewed, e.g., in Cebrian-Serrano and Davies (2017) Mamm. Genome 28(7):247-261, herein incorporated by reference in its entirety for all purposes. Another example of a modified Cas9 protein is xCas9, which is a SpCas9 variant that can recognize an expanded range of PAM sequences. See, e.g., Hu et al. (2018) Nature 556:57-63, herein incorporated by reference in its entirety for all purposes.

[0152] Cas proteins can comprise at least one nuclease domain, such as a DNase domain. For example, a wild type Cpf1 protein generally comprises a RuvC-like domain that cleaves both strands of target DNA, perhaps in a dimeric configuration. Likewise, CasX and CasΦ generally comprise a single RuvC-like domain that cleaves both strands of a target DNA. Cas proteins can also comprise at least two nuclease domains, such as DNase domains. For example, a wild type Cas9 protein generally comprises a RuvC-like nuclease domain and an HNH-like nuclease domain. The RuvC and HNH domains can each cut a different strand of double-stranded DNA to make a double-stranded break in the DNA. See, e.g., Jinek et al. (2012) Science 337:816-821, herein incorporated by reference in its entirety for all purposes.

[0153] One or more of the nuclease domains can be deleted or mutated so that they are no longer functional or have reduced nuclease activity. For example, if one of the nuclease domains is deleted or mutated in a Cas9 protein, the resulting Cas9 protein can be referred to as a nickase and can generate a single-strand break within a double-stranded target DNA but not a double- strand break (i.e., it can cleave the complementary strand or the non-complementary strand, but not both). If none of the nuclease domains is deleted or mutated in a Cas9 protein, the Cas9 protein will retain double-strand-break-inducing activity. An example of a mutation that convertsAttorney Docket No.057766 / 629304 Cas9 into a nickase is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. Likewise, H939A (histidine to alanine at amino acid position 839), H840A (histidine to alanine at amino acid position 840), or N863A (asparagine to alanine at amino acid position N863) in the HNH domain of Cas9 from S. pyogenes can convert the Cas9 into a nickase. Other examples of mutations that convert Cas9 into a nickase include the corresponding mutations to Cas9 from S. thermophilus. See, e.g., Sapranauskas et al. (2011) Nucleic Acids Res.39(21):9275-9282 and WO 2013 / 141680, each of which is herein incorporated by reference in its entirety for all purposes. Such mutations can be generated using methods such as site-directed mutagenesis, PCR-mediated mutagenesis, or total gene synthesis. Examples of other mutations creating nickases can be found, for example, in WO 2013 / 176772 and WO 2013 / 142578, each of which is herein incorporated by reference in its entirety for all purposes.

[0154] Examples of inactivating mutations in the catalytic domains of xCas9 are the same as those described above for SpCas9. Examples of inactivating mutations in the catalytic domains of Staphylococcus aureus Cas9 proteins are also known. For example, the Staphylococcus aureus Cas9 enzyme (SaCas9) may comprise a substitution at position N580 (e.g., N580A substitution) to create a nickase. Alternatively, the SaCas9 enzyme may comprise a substitution at position D10 (e.g., D10A substitution) to generate a nickase. See, e.g., WO 2016 / 106236, herein incorporated by reference in its entirety for all purposes. Examples of inactivating mutations in the catalytic domains of Nme2Cas9 are also known (e.g., D16A or H588A). Examples of inactivating mutations in the catalytic domains of St1Cas9 are also known (e.g., combination of D9A, D598A, H599A, and N622A). Examples of inactivating mutations in the catalytic domains of St3Cas9 are also known (e.g., D10A or N870A). Examples of inactivating mutations in the catalytic domains of CjCas9 are also known (e.g., D8A or H559A). Examples of inactivating mutations in the catalytic domains of FnCas9 and RHA FnCas9 are also known (e.g., N995A).

[0155] Examples of inactivating mutations in the catalytic domains of Cpf1 proteins are also known. With reference to Cpf1 proteins from Francisella novicida U112 (FnCpf1), Acidaminococcus sp. BV3L6 (AsCpf1), Lachnospiraceae bacterium ND2006 (LbCpf1), and Moraxella bovoculi 237 (MbCpf1 Cpf1), such mutations can include mutations at positions 908, 993, or 1263 of AsCpf1 or corresponding positions in Cpf1 orthologs, or positions 832, 925, 947, or 1180 of LbCpf1 or corresponding positions in Cpf1 orthologs. Such mutations can include, forAttorney Docket No.057766 / 629304 example one or more of mutations D908A, E993A, and D1263A of AsCpf1 or corresponding mutations in Cpf1 orthologs, or D832A, E925A, D947A, and D1180A of LbCpf1 or corresponding mutations in Cpf1 orthologs. See, e.g., US 2016 / 0208243, herein incorporated by reference in its entirety for all purposes.

[0156] Examples of inactivating mutations in the catalytic domains of CasX proteins are also known. With reference to CasX proteins from Deltaproteobacteria, D672A, E769A, and D935A (individually or in combination) or corresponding positions in other CasX orthologs are inactivating. See, e.g., Liu et al. (2019) Nature 566(7743):218-223, herein incorporated by reference in its entirety for all purposes.

[0157] Examples of inactivating mutations in the catalytic domains of CasΦ proteins are also known. For example, D371A and D394A, alone or in combination, are inactivating mutations. See, e.g., Pausch et al. (2020) Science 369(6501):333-337, herein incorporated by reference in its entirety for all purposes.

[0158] Cas proteins can also be operably linked to heterologous polypeptides as fusion proteins. For example, a Cas protein can be fused to a cleavage domain. See WO 2014 / 089290, herein incorporated by reference in its entirety for all purposes. Cas proteins can also be fused to a heterologous polypeptide providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.

[0159] As one example, a Cas protein can be fused to one or more heterologous polypeptides that provide for subcellular localization. Such heterologous polypeptides can include, for example, one or more nuclear localization signals (NLS) such as the monopartite SV40 NLS and / or a bipartite alpha-importin NLS for targeting to the nucleus, a mitochondrial localization signal for targeting to the mitochondria, an ER retention signal, and the like. See, e.g., Lange et al. (2007) J. Biol. Chem.282(8):5101-5105, herein incorporated by reference in its entirety for all purposes. Such subcellular localization signals can be located at the N-terminus, the C- terminus, or anywhere within the Cas protein. An NLS can comprise a stretch of basic amino acids, and can be a monopartite sequence or a bipartite sequence. Optionally, a Cas protein can comprise two or more NLSs, including an NLS (e.g., an alpha-importin NLS or a monopartite NLS) at the N-terminus and an NLS (e.g., an SV40 NLS or a bipartite NLS) at the C-terminus. AAttorney Docket No.057766 / 629304 Cas protein can also comprise two or more NLSs at the N-terminus and / or two or more NLSs at the C-terminus.

[0160] A Cas protein may, for example, be fused with 1-10 NLSs (e.g., fused with 1-5 NLSs or fused with one NLS. Where one NLS is used, the NLS may be linked at the N-terminus or the C-terminus of the Cas protein sequence. It may also be inserted within the Cas protein sequence. Alternatively, the Cas protein may be fused with more than one NLS. For example, the Cas protein may be fused with 2, 3, 4, or 5 NLSs. In one example, the Cas protein may be fused with two NLSs. In certain circumstances, the two NLSs may be the same (e.g., two SV40 NLSs) or different. For example, the Cas protein can be fused to two SV40 NLS sequences linked at the carboxy terminus. Alternatively, the Cas protein may be fused with two NLSs, one linked at the N-terminus and one at the C-terminus. In other examples, the Cas protein may be fused with 3 NLSs or with no NLS. The NLS may be a monopartite sequence, such as, e.g., the SV40 NLS. The NLS may be a bipartite sequence, such as the NLS of nucleoplasmin. In one example, a single monopartite NLS may be linked at the C-terminus of the Cas protein. One or more linkers are optionally included at the fusion site.

[0161] Cas proteins can also be operably linked to a cell-penetrating domain or protein transduction domain. For example, the cell-penetrating domain can be derived from the HIV-1 TAT protein, the TLM cell-penetrating motif from human hepatitis B virus, MPG, Pep-1, VP22, a cell penetrating peptide from Herpes simplex virus, or a polyarginine peptide sequence. See, e.g., WO 2014 / 089290 and WO 2013 / 176772, each of which is herein incorporated by reference in its entirety for all purposes. The cell-penetrating domain can be located at the N-terminus, the C- terminus, or anywhere within the Cas protein.

[0162] Cas proteins provided as mRNAs can be modified for improved stability and / or immunogenicity properties. The modifications may be made to one or more nucleosides within the mRNA. Examples of chemical modifications to mRNA nucleobases include pseudouridine, 1-methyl-pseudouridine, and 5-methyl-cytidine. For example, capped and polyadenylated Cas mRNA containing N1-methyl pseudouridine can be used. Likewise, Cas mRNAs can be modified by depletion of uridine using synonymous codons.

[0163] Guide RNAs. A “guide RNA” or “gRNA” is an RNA molecule that binds to a Cas protein (e.g., Cas9 protein) and targets the Cas protein to a specific location within a target DNA. Guide RNAs can comprise two segments: a “DNA-targeting segment” (also called a “guideAttorney Docket No.057766 / 629304 sequence”) and a “protein-binding segment.” “Segment” includes a section or region of a molecule, such as a contiguous stretch of nucleotides in an RNA. Some gRNAs, such as those for Cas9, can comprise two separate RNA molecules: an “activator-RNA” (e.g., tracrRNA) and a “targeter-RNA” (e.g., CRISPR RNA or crRNA). Other gRNAs are a single RNA molecule (single RNA polynucleotide), which can also be called a “single-molecule gRNA,” a “single- guide RNA,” or an “sgRNA.” See, e.g., WO 2013 / 176772, WO 2014 / 065596, WO 2014 / 089290, WO 2014 / 093622, WO 2014 / 099750, WO 2013 / 142578, and WO 2014 / 131833, each of which is herein incorporated by reference in its entirety for all purposes. A guide RNA can refer to either a CRISPR RNA (crRNA) or the combination of a crRNA and a trans-activating CRISPR RNA (tracrRNA). The crRNA and tracrRNA can be associated as a single RNA molecule (single guide RNA or sgRNA) or in two separate RNA molecules (dual guide RNA or dgRNA). For Cas9, for example, a single-guide RNA can comprise a crRNA fused to a tracrRNA (e.g., via a linker). For Cpf1 and CasΦ, for example, only a crRNA is needed to achieve binding to a target sequence. The terms “guide RNA” and “gRNA” include both double-molecule (i.e., modular) gRNAs and single-molecule gRNAs. In some of the methods and compositions disclosed herein, a gRNA is a S. pyogenes Cas9 gRNA or an equivalent thereof. In some of the methods and compositions disclosed herein, a gRNA is a S. aureus Cas9 gRNA or an equivalent thereof.

[0164] An exemplary two-molecule gRNA comprises a crRNA-like (“CRISPR RNA” or “targeter-RNA” or “crRNA” or “crRNA repeat”) molecule and a corresponding tracrRNA-like (“trans-activating CRISPR RNA” or “activator-RNA” or “tracrRNA”) molecule. A crRNA comprises both the DNA-targeting segment (single-stranded) of the gRNA and a stretch of nucleotides that forms one half of the dsRNA duplex of the protein-binding segment of the gRNA. An example of a crRNA tail (e.g., for use with S. pyogenes Cas9), located downstream (3′) of the DNA-targeting segment, comprises, consists essentially of, or consists of GUUUUAGAGCUAUGCU (SEQ ID NO: 3) or GUUUUAGAGCUAUGCUGUUUUG (SEQ ID NO: 4). Any of the DNA-targeting segments disclosed herein can be joined to the 5′ end of SEQ ID NO: 3 or 4 to form a crRNA.

[0165] A corresponding tracrRNA (activator-RNA) comprises a stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the gRNA. A stretch of nucleotides of a crRNA are complementary to and hybridize with a stretch of nucleotides of a tracrRNA to form the dsRNA duplex of the protein-binding domain of the gRNA. As such, eachAttorney Docket No.057766 / 629304 crRNA can be said to have a corresponding tracrRNA. Examples of tracrRNA sequences (e.g., for use with S. pyogenes Cas9) comprise, consist essentially of, or consist of any one of AGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACC GAGUCGGUGCUUU (SEQ ID NO: 5), AAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGG CACCGAGUCGGUGCUUUU (SEQ ID NO: 6), or GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA ACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 7).

[0166] In systems in which both a crRNA and a tracrRNA are needed, the crRNA and the corresponding tracrRNA hybridize to form a gRNA. In systems in which only a crRNA is needed, the crRNA can be the gRNA. The crRNA additionally provides the single-stranded DNA-targeting segment that hybridizes to the complementary strand of a target DNA. If used for modification within a cell, the exact sequence of a given crRNA or tracrRNA molecule can be designed to be specific to the species in which the RNA molecules will be used. See, e.g., Mali et al. (2013) Science 339(6121):823-826; Jinek et al. (2012) Science 337(6096):816-821; Hwang et al. (2013) Nat. Biotechnol.31(3):227-229; Jiang et al. (2013) Nat. Biotechnol.31(3):233-239; and Cong et al. (2013) Science 339(6121):819-823, each of which is herein incorporated by reference in its entirety for all purposes.

[0167] The DNA-targeting segment (crRNA) of a given gRNA comprises a nucleotide sequence that is complementary to a sequence on the complementary strand of the target DNA, as described in more detail below. The DNA-targeting segment of a gRNA interacts with the target DNA in a sequence-specific manner via hybridization (i.e., base pairing). As such, the nucleotide sequence of the DNA-targeting segment may vary and determines the location within the target DNA with which the gRNA and the target DNA will interact. The DNA-targeting segment of a subject gRNA can be modified to hybridize to any desired sequence within a target DNA. Naturally occurring crRNAs differ depending on the CRISPR / Cas system and organism but often contain a targeting segment of between 21 to 72 nucleotides length, flanked by two direct repeats (DR) of a length of between 21 to 46 nucleotides (see, e.g., WO 2014 / 131833, herein incorporated by reference in its entirety for all purposes). In the case of S. pyogenes, the DRs are 36 nucleotides long and the targeting segment is 30 nucleotides long. The 3′ located DRAttorney Docket No.057766 / 629304 is complementary to and hybridizes with the corresponding tracrRNA, which in turn binds to the Cas protein.

[0168] The DNA-targeting segment can have, for example, a length of at least about 12, 15, 17, 18, 19, 20, 25, 30, 35, or 40 nucleotides. Such DNA-targeting segments can have, for example, a length from about 12 to about 100, from about 12 to about 80, from about 12 to about 50, from about 12 to about 40, from about 12 to about 30, from about 12 to about 25, or from about 12 to about 20 nucleotides. For example, the DNA targeting segment can be from about 15 to about 25 nucleotides (e.g., from about 17 to about 20 nucleotides, or about 17, 18, 19, or 20 nucleotides). See, e.g., US 2016 / 0024523, herein incorporated by reference in its entirety for all purposes. For Cas9 from S. pyogenes, a typical DNA-targeting segment is between 16 and 20 nucleotides in length or between 17 and 20 nucleotides in length. For Cas9 from S. aureus, a typical DNA-targeting segment is between 21 and 23 nucleotides in length. For Cpf1, a typical DNA-targeting segment is at least 16 nucleotides in length or at least 18 nucleotides in length.

[0169] In one example, the DNA-targeting segment can be about 20 nucleotides in length. However, shorter and longer sequences can also be used for the targeting segment (e.g., 15-25 nucleotides in length, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length). The degree of identity between the DNA-targeting segment and the corresponding guide RNA target sequence (or degree of complementarity between the DNA-targeting segment and the other strand of the guide RNA target sequence) can be, for example, about 75%, about 80%, about 85%, about 90%, about 95%, or 100%. The DNA-targeting segment and the corresponding guide RNA target sequence can contain one or more mismatches. For example, the DNA- targeting segment of the guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches (e.g., where the total length of the guide RNA target sequence is at least 17, at least 18, at least 19, or at least 20 or more nucleotides). For example, the DNA-targeting segment of the guide RNA and the corresponding guide RNA target sequence can contain 1-4, 1-3, 1-2, 1, 2, 3, or 4 mismatches where the total length of the guide RNA target sequence 20 nucleotides.

[0170] In one example, the exogenous DNA comprises a selection cassette or a reporter protein coding sequence, and the nuclease target sequences or guide RNA target sequences (i.e., the first nuclease cleavage site and the second nuclease cleavage site) are within the selection cassette (e.g., within a polynucleotide encoding a selection marker or a disease resistance gene)Attorney Docket No.057766 / 629304 or the reporter protein coding sequence or a promoter used in the selection cassette or operably linked to the reporter protein coding sequence. Nucleic acid constructs can also comprise a polynucleotide encoding a selection marker. Alternatively, the nucleic acid constructs can lack a polynucleotide encoding a selection marker. The selection marker can be contained in a selection cassette. Optionally, the selection cassette can be a self-deleting cassette. See, e.g., US 8,697,851 and US 2013 / 0312129, each of which is herein incorporated by reference in its entirety for all purposes. As an example, the self-deleting cassette can comprise a Crei gene (comprises two exons encoding a Cre recombinase, which are separated by an intron) operably linked to a mouse Prm1 promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By employing the Prm1 promoter, the self-deleting cassette can be deleted specifically in male germ cells of F0 animals. Exemplary selection markers include neomycin phosphotransferase (neor), hygromycin B phosphotransferase (hygr), puromycin-N-acetyltransferase (puror), blasticidin S deaminase (bsrr), xanthine / guanine phosphoribosyl transferase (gpt), or herpes simplex virus thymidine kinase (HSV-k), or a combination thereof. The polynucleotide encoding the selection marker can be operably linked to a promoter active in a cell being targeted. Examples of promoters are described elsewhere herein. The exogenous DNA can comprise other elements that are common among targeting vectors to which guide RNAs can be targeted, including but not limited to, polyA signal sequences, promoter sequences, AAV payload and backbone sequences, or cloning vector sequences. For example, in some such methods, the first nuclease cleavage site and the second nuclease cleavage site are within a spectinomycin resistance gene or coding sequence, a kanamycin resistance gene or coding sequence, a hygromycin resistance gene or coding sequence, a puromycin resistance gene or coding sequence, a lacZ gene or coding sequence, a tamoxifen-inducible Cre gene, a luciferase gene or coding sequence, a human ubiquitin promoter, a mouse phosphoglycerate kinase promoter, or a mouse protamine promoter. In some embodiments, the first nuclease cleavage site and the second nuclease cleavage site are within a kanamycin resistance gene or coding sequence. In some embodiments, the first nuclease cleavage site and the second nuclease cleavage site are within a puromycin resistance gene or coding sequence. In some embodiments, the first nuclease cleavage site and the second nuclease cleavage site are within a hygromycin resistance gene or coding sequence.Attorney Docket No.057766 / 629304

[0171] An advantage of methods described herein in which the first nuclease cleavage site and second nuclease cleavage site are within a selection cassette or a reporter protein coding sequence is that the region between the cleavage sites (which may be sequenced less preferentially in some methods as described elsewhere herein) is minimal and can be assumed to be intact if the selection cassette or reporter protein was used to select the targeted cells whose genomic DNA is being sequenced. In other words, it is reasonable to assume that the only sequence in the exogenous DNA that does not have many sequencing reads would have to be correct, or the targeted cell could not proliferate on a selection plate (in the case of a selection cassette).

[0172] As one example, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34- 35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA- targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA- targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first andAttorney Docket No.057766 / 629304 second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34-35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 34 and 35. Alternatively, a guide RNA targeting a spectinomycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 34- 35, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting ofAttorney Docket No.057766 / 629304 sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 34 and 35.

[0173] As one example, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36- 38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment that is at leastAttorney Docket No.057766 / 629304 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36- 38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37). Alternatively, a guide RNA targeting a kanamycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 36-38, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at leastAttorney Docket No.057766 / 629304 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 36 and 37 or SEQ ID NOS: 36 and 38 (e.g., SEQ ID NOS: 36 and 37).

[0174] As one example, a guide RNA targeting a hygromycin resistance gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA- targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA- targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting aAttorney Docket No.057766 / 629304 hygromycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA- targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 39-40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 39 and 40. Alternatively, a guide RNA targeting a hygromycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 39- 40, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 39 and 40.

[0175] As one example, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can compriseAttorney Docket No.057766 / 629304 DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-Attorney Docket No.057766 / 629304 42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42. Alternatively, a guide RNA targeting a puromycin resistance gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 41-42, or a combination of first and second guide RNAs targeting a puromycin resistance gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 41 and 42.

[0176] As one example, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43-44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43-44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments comprising, consisting essentially of,Attorney Docket No.057766 / 629304 or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43-44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43-44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 43-44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43- 44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43-Attorney Docket No.057766 / 629304 44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 43 and 44. Alternatively, a guide RNA targeting a human ubiquitin promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 43- 44, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 43 and 44.

[0177] As one example, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinaseAttorney Docket No.057766 / 629304 promoter can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA- targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNA targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46. Alternatively, a guide RNAAttorney Docket No.057766 / 629304 targeting a mouse phosphoglycerate kinase promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 45-46, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 45 and 46.

[0178] As one example, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47- 48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA- targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to theAttorney Docket No.057766 / 629304 sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48. Alternatively, a guide RNA targeting a lacZ gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 47-48, or a combination of first and second guide RNAs targeting a lacZ gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 47 and 48.Attorney Docket No.057766 / 629304

[0179] As one example, a guide RNA targeting a mouse protamine promoter can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA- targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA- targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20Attorney Docket No.057766 / 629304 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA- targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50. Alternatively, a guide RNA targeting a mouse protamine promoter can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 49-50, or a combination of first and second guide RNAs targeting a mouse protamine promoter can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 49 and 50.

[0180] As one example, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth inAttorney Docket No.057766 / 629304 SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA- targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA- targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA- targeting segments that are at least 90% or at least 95% identical to the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51-Attorney Docket No.057766 / 629304 52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51-52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 51 and 52. Alternatively, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 51- 52, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 51 and 52.

[0181] As one example, a guide RNA targeting a luciferase gene can comprise a DNA- targeting segment (i.e., guide sequence) comprising, consisting essentially of, or consisting of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA- targeting segments (i.e., guide sequences) comprising, consisting essentially of, or consisting of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segmentAttorney Docket No.057766 / 629304 comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA- targeting segments that are at least 90% or at least 95% identical to the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment that is at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments that are at least 75%, at least 80%, at least 85%, atAttorney Docket No.057766 / 629304 least 90%, or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment that is at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments that are at least 90% or at least 95% identical to at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA- targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53-59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences that differ by no more than 3, no more than 2, or no more than 1 nucleotide from the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55). Alternatively, a guide RNA targeting a luciferase gene can comprise a DNA-targeting segment comprising, consisting essentially of, or consisting of a sequence that differs by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequence (DNA-targeting segment) set forth in any one of SEQ ID NOS: 53- 59, or a combination of first and second guide RNAs targeting a luciferase gene can comprise DNA-targeting segments comprising, consisting essentially of, or consisting of sequences thatAttorney Docket No.057766 / 629304 differ by no more than 3, no more than 2, or no more than 1 nucleotide from at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the sequences (DNA-targeting segments) set forth in SEQ ID NOS: 53 and 54, SEQ ID NOS: 53 and 55, SEQ ID NOS: 53 and 57, SEQ ID NOS: 53 and 59, SEQ ID NOS: 56 and 54, SEQ ID NOS: 56 and 55, SEQ ID NOS: 56 and 57, SEQ ID NOS: 56 and 59, SEQ ID NOS: 58 and 54, SEQ ID NOS: 58 and 57, or SEQ ID NOS: 58 and 59 (e.g., SEQ ID NOS: 53 and 55).

[0182] TracrRNAs can be in any form (e.g., full-length tracrRNAs or active partial tracrRNAs) and of varying lengths. They can include primary transcripts or processed forms. For example, tracrRNAs (as part of a single-guide RNA or as a separate molecule as part of a two- molecule gRNA) may comprise, consist essentially of, or consist of all or a portion of a wild type tracrRNA sequence (e.g., about or more than about 20, 26, 32, 45, 48, 54, 63, 67, 85, or more nucleotides of a wild type tracrRNA sequence). Examples of wild type tracrRNA sequences from S. pyogenes include 171-nucleotide, 89-nucleotide, 75-nucleotide, and 65-nucleotide versions. See, e.g., Deltcheva et al. (2011) Nature 471(7340):602-607; WO 2014 / 093661, each of which is herein incorporated by reference in its entirety for all purposes. Examples of tracrRNAs within single-guide RNAs (sgRNAs) include the tracrRNA segments found within +48, +54, +67, and +85 versions of sgRNAs, where “+n” indicates that up to the +n nucleotide of wild type tracrRNA is included in the sgRNA. See US 8,697,359, herein incorporated by reference in its entirety for all purposes.

[0183] The percent complementarity between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA can be at least 60% (e.g., at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%). The percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be at least 60% over about 20 contiguous nucleotides. As an example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over the 14 contiguous nucleotides at the 5′end of the complementary strand of the target DNA and as low as 0% over the remainder. In such a case, the DNA-targeting segment can be considered to be 14 nucleotides in length. As another example, the percent complementarity between the DNA-targeting segment and the complementary strand of the target DNA can be 100% over the seven contiguous nucleotides at the 5′ end of the complementary strand of the target DNA and as low as 0% over the remainder.Attorney Docket No.057766 / 629304 In such a case, the DNA-targeting segment can be considered to be 7 nucleotides in length. In some guide RNAs, at least 17 nucleotides within the DNA-targeting segment are complementary to the complementary strand of the target DNA. For example, the DNA-targeting segment can be 20 nucleotides in length and can comprise 1, 2, or 3 mismatches with the complementary strand of the target DNA. In one example, the mismatches are not adjacent to the region of the complementary strand corresponding to the protospacer adjacent motif (PAM) sequence (i.e., the reverse complement of the PAM sequence) (e.g., the mismatches are in the 5′ end of the DNA- targeting segment of the guide RNA, or the mismatches are at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 base pairs away from the region of the complementary strand corresponding to the PAM sequence).

[0184] The protein-binding segment of a gRNA can comprise two stretches of nucleotides that are complementary to one another. The complementary nucleotides of the protein-binding segment hybridize to form a double-stranded RNA duplex (dsRNA). The protein-binding segment of a subject gRNA interacts with a Cas protein, and the gRNA directs the bound Cas protein to a specific nucleotide sequence within target DNA via the DNA-targeting segment.

[0185] Single-guide RNAs can comprise a DNA-targeting segment and a scaffold sequence (i.e., the protein-binding or Cas-binding sequence of the guide RNA). For example, such guide RNAs can have a 5′ DNA-targeting segment joined to a 3′ scaffold sequence. Exemplary scaffold sequences (e.g., for use with S. pyogenes Cas9) comprise, consist essentially of, or consist of: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGA AAAAGUGGCACCGAGUCGGUGCU (version 1; SEQ ID NO: 60); GUUGGAACCAUUCAAAACAGCAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCA ACUUGAAAAAGUGGCACCGAGUCGGUGC (version 2; SEQ ID NO: 61); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGA AAAAGUGGCACCGAGUCGGUGC (version 3; SEQ ID NO: 62); and GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUU AUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (version 4; SEQ ID NO: 63). Other exemplary scaffold sequences comprise, consist essentially of, or consist of: GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGA AAAAGUGGCACCGAGUCGGUGCUUUUUUU (version 5; SEQ ID NO: 64); GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAttorney Docket No.057766 / 629304 AAAAGUGGCACCGAGUCGGUGCUUUU (version 6; SEQ ID NO: 65);GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCG UUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUUUUUU (version 7; SEQ ID NO: 66); or GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGG CACCGAGUCGGUGC (version 8; SEQ ID NO: 67). In some guide sgRNAs, the four terminal U residues of version 6 are not present. In some sgRNAs, only 1, 2, or 3 of the four terminal U residues of version 6 are present. Guide RNAs targeting any guide RNA target sequence (e.g., any of the guide RNA target sequences disclosed herein) can include, for example, a DNA- targeting segment on the 5′ end of the guide RNA fused to any of the exemplary guide RNA scaffold sequences on the 3′ end of the guide RNA. That is, a DNA-targeting segment (e.g., any of the DNA-targeting segments disclosed herein) can be joined to the 5′ end of any one of the above scaffold sequences to form a single guide RNA (chimeric guide RNA). Likewise, a DNA- targeting segment can be joined to the 5′ end of any one of the above scaffold sequences to form a single guide RNA (chimeric guide RNA).

[0186] Guide RNAs can include modifications or sequences that provide for additional desirable features (e.g., modified or regulated stability; subcellular targeting; tracking with a fluorescent label; a binding site for a protein or protein complex; and the like). That is, guide RNAs can include one or more modified nucleosides or nucleotides, or one or more non- naturally and / or naturally occurring components or configurations that are used instead of or in addition to the canonical A, G, C, and U residues. Examples of such modifications include, for example, a 5′ cap (e.g., a 7-methylguanylate cap (m7G)); a 3′ polyadenylated tail (i.e., a 3′ poly(A) tail); a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (i.e., a hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, and so forth); a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and the like); and combinations thereof. OtherAttorney Docket No.057766 / 629304 examples of modifications include engineered stem loop duplex structures, engineered bulge regions, engineered hairpins 3′ of the stem loop duplex structure, or any combination thereof. See, e.g., US 2015 / 0376586, herein incorporated by reference in its entirety for all purposes. A bulge can be an unpaired region of nucleotides within the duplex made up of the crRNA-like region and the minimum tracrRNA-like region. A bulge can comprise, on one side of the duplex, an unpaired 5′-XXXY-3′ where X is any purine and Y can be a nucleotide that can form a wobble pair with a nucleotide on the opposite strand, and an unpaired nucleotide region on the other side of the duplex.

[0187] Unmodified nucleic acids can be prone to degradation. Exogenous nucleic acids can also induce an innate immune response. Modifications can help introduce stability and reduce immunogenicity. Guide RNAs can comprise modified nucleosides and modified nucleotides including, for example, one or more of the following: (1) alteration or replacement of one or both of the non-linking phosphate oxygens and / or of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage; (2) alteration or replacement of a constituent of the ribose sugar such as alteration or replacement of the 2′ hydroxyl on the ribose sugar; (3) replacement of the phosphate moiety with dephospho linkers; (4) modification or replacement of a naturally occurring nucleobase; (5) replacement or modification of the ribose-phosphate backbone; (6) modification of the 3′ end or 5′ end of the oligonucleotide (e.g., removal, modification or replacement of a terminal phosphate group or conjugation of a moiety); and (7) modification of the sugar. Other possible guide RNA modifications include modifications of or replacement of uracils or poly-uracil tracts. See, e.g., WO 2015 / 048577 and US 2016 / 0237455, each of which is herein incorporated by reference in its entirety for all purposes. Similar modifications can be made to Cas-encoding nucleic acids, such as Cas mRNAs.

[0188] As one example, nucleotides at the 5′ or 3′ end of a guide RNA can include phosphorothioate linkages (e.g., the bases can have a modified phosphate group that is a phosphorothioate group). For example, a guide RNA can include phosphorothioate linkages between the 2, 3, or 4 terminal nucleotides at the 5′ or 3′ end of the guide RNA. As another example, nucleotides at the 5′ and / or 3′ end of a guide RNA can have 2′-O-methyl modifications. For example, a guide RNA can include 2′-O-methyl modifications at the 2, 3, or 4 terminal nucleotides at the 5′ and / or 3′ end of the guide RNA (e.g., the 5′ end). See, e.g., WO 2017 / 173054 A1 and Finn et al. (2018) Cell Rep.22(9):2227-2235, each of which is hereinAttorney Docket No.057766 / 629304 incorporated by reference in its entirety for all purposes. In one example, the guide RNA comprises 2′-O-methyl analogs and 3′ phosphorothioate internucleotide linkages at the first three 5′ and 3′ terminal RNA residues. In another example, the guide RNA is modified such that all 2′OH groups that do not interact with the Cas9 protein are replaced with 2′-O-methyl analogs, and the tail region of the guide RNA, which has minimal interaction with Cas9, is modified with 5′ and 3′ phosphorothioate internucleotide linkages. See, e.g., Yin et al. (2017) Nat. Biotech. 35(12):1179-1187, herein incorporated by reference in its entirety for all purposes. Other examples of modified guide RNAs are provided, e.g., in WO 2018 / 107028 A1, herein incorporated by reference in its entirety for all purposes.

[0189] gRNAs can be prepared by various other methods. For example, gRNAs can be prepared by in vitro transcription using, for example, T7 RNA polymerase (see, e.g., WO 2014 / 089290 and WO 2014 / 065596, each of which is herein incorporated by reference in its entirety for all purposes). Guide RNAs can also be a synthetically produced molecule prepared by chemical synthesis. For example, a guide RNA can be chemically synthesized to include 2′-O- methyl analogs and 3′ phosphorothioate internucleotide linkages at the first three 5′ and 3′ terminal RNA residues.

[0190] Guide RNA Target Sequences. Target DNAs for guide RNAs include nucleic acid sequences present in a DNA to which a DNA-targeting segment of a gRNA will bind, provided sufficient conditions for binding exist. Suitable DNA / RNA binding conditions include physiological conditions normally present in a cell. Other suitable DNA / RNA binding conditions (e.g., conditions in a cell-free system) are known in the art (see, e.g., Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), herein incorporated by reference in its entirety for all purposes). The strand of the target DNA that is complementary to and hybridizes with the gRNA can be called the “complementary strand,” and the strand of the target DNA that is complementary to the “complementary strand” (and is therefore not complementary to the Cas protein or gRNA) can be called “noncomplementary strand” or “template strand.”

[0191] The target DNA includes both the sequence on the complementary strand to which the guide RNA hybridizes and the corresponding sequence on the non-complementary strand (e.g., adjacent to the protospacer adjacent motif (PAM)). The term “guide RNA target sequence” as used herein refers specifically to the sequence on the non-complementary strand correspondingAttorney Docket No.057766 / 629304 to (i.e., the reverse complement of) the sequence to which the guide RNA hybridizes on the complementary strand. That is, the guide RNA target sequence refers to the sequence on the non- complementary strand adjacent to the PAM (e.g., upstream or 5′ of the PAM in the case of Cas9). A guide RNA target sequence is equivalent to the DNA-targeting segment of a guide RNA, but with thymines instead of uracils. As one example, a guide RNA target sequence for an SpCas9 enzyme can refer to the sequence upstream of the 5′-NGG-3′ PAM on the non-complementary strand. A guide RNA is designed to have complementarity to the complementary strand of a target DNA, where hybridization between the DNA-targeting segment of the guide RNA and the complementary strand of the target DNA promotes the formation of a CRISPR complex. Full complementarity is not necessarily required, provided that there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex. If a guide RNA is referred to herein as targeting a guide RNA target sequence, what is meant is that the guide RNA hybridizes to the complementary strand sequence of the target DNA that is the reverse complement of the guide RNA target sequence on the non-complementary strand.

[0192] A target DNA or guide RNA target sequence can comprise any polynucleotide, and can be located, for example, in the nucleus or cytoplasm of a cell or within an organelle of a cell, such as a mitochondrion or chloroplast. A target DNA or guide RNA target sequence can be any nucleic acid sequence endogenous or exogenous to a cell. The guide RNA target sequence can be a sequence coding a gene product (e.g., a protein) or a non-coding sequence (e.g., a regulatory sequence) or can include both.

[0193] Site-specific binding and cleavage of a target DNA by a Cas protein can occur at locations determined by both (i) base-pairing complementarity between the guide RNA and the complementary strand of the target DNA and (ii) a short motif, called the protospacer adjacent motif (PAM), in the non-complementary strand of the target DNA. The PAM can flank the guide RNA target sequence. Optionally, the guide RNA target sequence can be flanked on the 3′ end by the PAM (e.g., for Cas9). Alternatively, the guide RNA target sequence can be flanked on the 5′ end by the PAM (e.g., for Cpf1). For example, the cleavage site of Cas proteins can be about 1 to about 10 or about 2 to about 5 base pairs (e.g., 3 base pairs) upstream or downstream of the PAM sequence (e.g., within the guide RNA target sequence). In the case of SpCas9, the PAM sequence (i.e., on the non-complementary strand) can be 5′-N1GG-3′, where N1is any DNA nucleotide, and where the PAM is immediately 3′ of the guide RNA target sequence on the non-Attorney Docket No.057766 / 629304 complementary strand of the target DNA. As such, the sequence corresponding to the PAM on the complementary strand (i.e., the reverse complement) would be 5′-CCN2-3′, where N2 is any DNA nucleotide and is immediately 5′ of the sequence to which the DNA-targeting segment of the guide RNA hybridizes on the complementary strand of the target DNA. In some such cases, N1 and N2 can be complementary and the N1- N2 base pair can be any base pair (e.g., N1=C and N2=G; N1=G and N2=C; N1=A and N2=T; or N1=T, and N2=A). In the case of Cas9 from S. aureus, the PAM can be NNGRRT or NNGRR, where N can A, G, C, or T, and R can be G or A. In the case of Cas9 from C. jejuni, the PAM can be, for example, NNNNACAC or NNNNRYAC, where N can be A, G, C, or T, and R can be G or A. In some cases (e.g., for FnCpf1), the PAM sequence can be upstream of the 5′ end and have the sequence 5′-TTN-3′. In the case of DpbCasX, the PAM can have the sequence 5′-TTCN-3′. In the case of CasΦ, the PAM can have the sequence 5′-TBN-3′, wherein B is G, T, or C.

[0194] An example of a guide RNA target sequence is a 20-nucleotide DNA sequence immediately preceding an NGG motif recognized by an SpCas9 protein. For example, two examples of guide RNA target sequences plus PAMs are GN19NGG or N20NGG. See, e.g., WO 2014 / 165825, herein incorporated by reference in its entirety for all purposes. The guanine at the 5′ end can facilitate transcription by RNA polymerase in cells. Other examples of guide RNA target sequences plus PAMs can include two guanine nucleotides at the 5′ end (e.g., GGN20NGG) to facilitate efficient transcription by T7 polymerase in vitro. See, e.g., WO 2014 / 065596, herein incorporated by reference in its entirety for all purposes. Other guide RNA target sequences plus PAMs can have between 4-22 nucleotides in length of any of the above guide RNA target sequences plus PAMs, including the 5′ G or GG and the 3′ GG or NGG. Yet other guide RNA target sequences plus PAMs can have between 14 and 20 nucleotides in length of any of the above guide RNA target sequences plus PAMs.

[0195] In one example, the exogenous DNA comprises a selection cassette or a reporter protein coding sequence, and the guide RNA target sequences (i.e., the first nuclease cleavage site and the second nuclease cleavage site) are within the selection cassette (e.g., within a polynucleotide encoding a selection marker or a disease resistance gene) or the reporter protein coding sequence or a promoter used in the selection cassette or operably linked to the reporter protein coding sequence. Nucleic acid constructs can also comprise a polynucleotide encoding a selection marker. Alternatively, the nucleic acid constructs can lack a polynucleotide encoding aAttorney Docket No.057766 / 629304 selection marker. The selection marker can be contained in a selection cassette. Optionally, the selection cassette can be a self-deleting cassette. See, e.g., US 8,697,851 and US 2013 / 0312129, each of which is herein incorporated by reference in its entirety for all purposes. As an example, the self-deleting cassette can comprise a Crei gene (comprises two exons encoding a Cre recombinase, which are separated by an intron) operably linked to a mouse Prm1 promoter and a neomycin resistance gene operably linked to a human ubiquitin promoter. By employing the Prm1 promoter, the self-deleting cassette can be deleted specifically in male germ cells of F0 animals. Exemplary selection markers include neomycin phosphotransferase (neor), hygromycin B phosphotransferase (hygr), puromycin-N-acetyltransferase (puror), blasticidin S deaminase (bsrr), xanthine / guanine phosphoribosyl transferase (gpt), or herpes simplex virus thymidine kinase (HSV-k), or a combination thereof. The polynucleotide encoding the selection marker can be operably linked to a promoter active in a cell being targeted. Examples of promoters are described elsewhere herein. The exogenous DNA can comprise other elements that are common among targeting vectors to which guide RNAs can be targeted, including but not limited to, polyA signal sequences, promoter sequences, AAV payload and backbone sequences, or cloning vector sequences. For example, in some such methods, the first nuclease cleavage site and the second nuclease cleavage site are within a spectinomycin resistance gene or coding sequence, a kanamycin resistance gene or coding sequence, a hygromycin resistance gene or coding sequence, a puromycin resistance gene or coding sequence, a lacZ gene or coding sequence, a tamoxifen-inducible Cre gene, a luciferase gene or coding sequence, a human ubiquitin promoter, a mouse phosphoglycerate kinase promoter, or a mouse protamine promoter.

[0196] As one example, a guide RNA targeting a spectinomycin resistance gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 8-9, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can target the guide RNA target sequences set forth in SEQ ID NOS: 8 and 9. As another example, a guide RNA targeting a spectinomycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 8-9, or a combination of first and second guide RNAs targeting a spectinomycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 8 and 9.Attorney Docket No.057766 / 629304

[0197] As one example, a guide RNA targeting a kanamycin resistance gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 10-12, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can target the guide RNA target sequences set forth in SEQ ID NOS: 10 and 11 or SEQ ID NOS: 10 and 12 (e.g., SEQ ID NOS: 10 and 11). As another example, a guide RNA targeting a kanamycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 10-12, or a combination of first and second guide RNAs targeting a kanamycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 10 and 11 or SEQ ID NOS: 10 and 12 (e.g., SEQ ID NOS: 10 and 11).

[0198] As one example, a guide RNA targeting a hygromycin resistance gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 13-14, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can target the guide RNA target sequences set forth in SEQ ID NOS: 13 and 14. As another example, a guide RNA targeting a hygromycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 13-14, or a combination of first and second guide RNAs targeting a hygromycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 13 and 14.

[0199] As one example, a guide RNA targeting a puromycin resistance gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 15-16, or a combination of first and second guide RNAs targeting a puromycin resistance gene can target the guide RNA target sequences set forth in SEQ ID NOS: 15 and 16. As another example, a guide RNA targeting a puromycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 15-16, or a combination of first and second guide RNAs targeting a puromycin resistance gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 15 and 16.

[0200] As one example, a guide RNA targeting a human ubiquitin promoter can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 17-18, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can target the guide RNA targetAttorney Docket No.057766 / 629304 sequences set forth in SEQ ID NOS: 17 and 18. As another example, a guide RNA targeting a human ubiquitin promoter can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 17-18, or a combination of first and second guide RNAs targeting a human ubiquitin promoter can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 17 and 18.

[0201] As one example, a guide RNA targeting a mouse phosphoglycerate kinase promoter can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 19-20, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can target the guide RNA target sequences set forth in SEQ ID NOS: 19 and 20. As another example, a guide RNA targeting a mouse phosphoglycerate kinase promoter can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 19-20, or a combination of first and second guide RNAs targeting a mouse phosphoglycerate kinase promoter can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 19 and 20.

[0202] As one example, a guide RNA targeting a lacZ gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 21-22, or a combination of first and second guide RNAs targeting a lacZ gene can target the guide RNA target sequences set forth in SEQ ID NOS: 21 and 22. As another example, a guide RNA targeting a lacZ gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 21-22, or a combination of first and second guide RNAs targeting a lacZ gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 21 and 22.

[0203] As one example, a guide RNA targeting a mouse protamine promoter can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 23-24, or a combination of first and second guide RNAs targeting a mouse protamine promoter can target the guide RNA target sequences set forth in SEQ ID NOS: 23 and 24. As another example, a guide RNA targeting a mouse protamine promoter can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 23-24, or a combination of first and second guide RNAs targeting a mouse protamine promoter can target atAttorney Docket No.057766 / 629304 least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 23 and 24.

[0204] As one example, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 25-26, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can target the guide RNA target sequences set forth in SEQ ID NOS: 25 and 26. As another example, a guide RNA targeting a tamoxifen-inducible Cre (CreERT) gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 25-26, or a combination of first and second guide RNAs targeting a tamoxifen-inducible Cre (CreERT) gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 25 and 26.

[0205] As one example, a guide RNA targeting a luciferase gene can target the guide RNA target sequence set forth in any one of SEQ ID NOS: 27-33, or a combination of first and second guide RNAs targeting a luciferase gene can target the guide RNA target sequences set forth in SEQ ID NOS: 27 and 28, SEQ ID NOS: 27 and 29, SEQ ID NOS: 27 and 31, SEQ ID NOS: 27 and 33, SEQ ID NOS: 30 and 28, SEQ ID NOS: 30 and 29, SEQ ID NOS: 30 and 31, SEQ ID NOS: 30 and 33, SEQ ID NOS: 32 and 28, SEQ ID NOS: 32 and 31, or SEQ ID NOS: 32 and 33 (e.g., SEQ ID NOS: 27 and 29). As another example, a guide RNA targeting a luciferase gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequence set forth in any one of SEQ ID NOS: 27-33, or a combination of first and second guide RNAs targeting a luciferase gene can target at least 17, at least 18, at least 19, or at least 20 contiguous nucleotides of the guide RNA target sequences set forth in SEQ ID NOS: 27 and 28, SEQ ID NOS: 27 and 29, SEQ ID NOS: 27 and 31, SEQ ID NOS: 27 and 33, SEQ ID NOS: 30 and 28, SEQ ID NOS: 30 and 29, SEQ ID NOS: 30 and 31, SEQ ID NOS: 30 and 33, SEQ ID NOS: 32 and 28, SEQ ID NOS: 32 and 31, or SEQ ID NOS: 32 and 33 (e.g., SEQ ID NOS: 27 and 29).

[0206] Formation of a CRISPR complex hybridized to a target DNA can result in cleavage of one or both strands of the target DNA within or near the region corresponding to the guide RNA target sequence (i.e., the guide RNA target sequence on the non-complementary strand of the target DNA and the reverse complement on the complementary strand to which the guide RNAAttorney Docket No.057766 / 629304 hybridizes). For example, the cleavage site can be within the guide RNA target sequence (e.g., at a defined location relative to the PAM sequence). The “cleavage site” includes the position of a target DNA at which a Cas protein produces a single-strand break or a double-strand break. Cleavage sites can be at the same position on both strands (producing blunt ends; e.g., Cas9)) or can be at different sites on each strand (producing staggered ends (i.e., overhangs); e.g., Cpf1). Staggered ends can be produced, for example, by using two Cas proteins, each of which produces a single-strand break at a different cleavage site on a different strand, thereby producing a double-strand break. For example, a first nickase can create a single-strand break on the first strand of double-stranded DNA (dsDNA), and a second nickase can create a single- strand break on the second strand of dsDNA such that overhanging sequences are created. In some cases, the guide RNA target sequence or cleavage site of the nickase on the first strand is separated from the guide RNA target sequence or cleavage site of the nickase on the second strand by at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 40, 50, 75, 100, 250, 500, or 1,000 base pairs.

[0207] Other Nuclease Agents. Any other type of known rare-cutting nuclease agent can also be used in the methods described herein. One example of such a nuclease agent is a Transcription Activator-Like Effector Nuclease (TALEN). TAL effector nucleases are a class of sequence-specific nucleases that can be used to make double-strand breaks at specific target sequences in DNA. TAL effector nucleases are created by fusing a native or engineered transcription activator-like (TAL) effector, or functional part thereof, to the catalytic domain of an endonuclease, such as, for example, FokI. The unique, modular TAL effector DNA binding domain allows for the design of proteins with potentially any given DNA recognition specificity. Thus, the DNA binding domains of the TAL effector nucleases can be engineered to recognize specific DNA target sites and thus, used to make double-strand breaks at desired target sequences. See WO 2010 / 079430; Morbitzer et al. (2010) Proc. Natl. Acad. Sci. U.S.A. 107(50):21617-21622; Scholze & Boch (2010) Virulence 1:428-432; Christian et al. Genetics (2010) 186:757-761; Li et al. (2010) Nucleic Acids Res. (2010) 39(1):359-372; and Miller et al. (2011) Nat. Biotechnol.29:143–148, each of which is herein incorporated by reference in its entirety for all purposes.

[0208] Examples of suitable TAL nucleases, and methods for preparing suitable TAL nucleases, are disclosed, e.g., in US 2011 / 0239315, US 2011 / 0269234, US 2011 / 0145940, USAttorney Docket No.057766 / 629304 2003 / 0232410, US 2005 / 0208489, US 2005 / 0026157, US 2005 / 0064474, US 2006 / 0188987, and US 2006 / 0063231, each of which is herein incorporated by reference in its entirety for all purposes.

[0209] In some TALENs, each monomer of the TALEN comprises 33-35 TAL repeats that recognize a single base pair via two hypervariable residues. The TALEN can be a chimeric protein comprising a TAL-repeat-based DNA binding domain operably linked to an independent nuclease such as a FokI endonuclease. For example, the nuclease agent can comprise a first TAL-repeat-based DNA binding domain and a second TAL-repeat-based DNA binding domain, wherein each of the first and the second TAL-repeat-based DNA binding domains is operably linked to a FokI nuclease, wherein the first and the second TAL-repeat-based DNA binding domain recognize two contiguous target DNA sequences in each strand of the target DNA sequence separated by a spacer sequence of varying length (12-20 bp), and wherein the FokI nuclease subunits dimerize to create an active nuclease that makes a double strand break at a target sequence.

[0210] Another example of a suitable nuclease agent is a zinc-finger nuclease (ZFN). In some ZFNs, each monomer of the ZFN comprises 3 or more zinc finger-based DNA binding domains, wherein each zinc finger-based DNA binding domain binds to a 3 bp subsite. In other ZFNs, the ZFN is a chimeric protein comprising a zinc finger-based DNA binding domain operably linked to an independent nuclease such as a FokI endonuclease. For example, the nuclease agent can comprise a first ZFN and a second ZFN, wherein each of the first ZFN and the second ZFN is operably linked to a FokI nuclease subunit, wherein the first and the second ZFN recognize two contiguous target DNA sequences in each strand of the target DNA sequence separated by about 5-7 bp spacer, and wherein the FokI nuclease subunits dimerize to create an active nuclease that makes a double strand break. See, e.g., US 2006 / 0246567; US 2008 / 0182332; US 2002 / 0081614; US 2003 / 0021776; WO 2002 / 057308; US 2013 / 0123484; US 2010 / 0291048; WO 2011 / 017293; and Gaj et al. (2013) Trends Biotechnol., 31(7):397-405, each of which is herein incorporated by reference in its entirety for all purposes. B. Creating a Sequencing Library

[0211] Following cleavage of the genomic DNA at the first nuclease cleavage site and the second nuclease cleavage site in the exogenous DNA, by any of the above nuclease agents,Attorney Docket No.057766 / 629304 comprises a step to generate a sequencing library comprising a first cleaved genomic DNA and a second cleaved genomic DNA. In some cases, the methods for generating a sequencing library do not comprise amplification of the genomic DNA, the first cleaved genomic DNA, or the second cleaved genomic DNA. In some embodiments, the methods comprise preparing a single long-read sequencing library for use in a single long-read sequencing run (i.e., only one long- read sequencing library is prepared). In some embodiments, the method does not comprise preparing a second long-read sequencing library (and does not comprise performing multiple long-read sequencing runs). The methods can comprise preparing the single long-read sequencing library (i.e., for use in a single long-read sequencing run) by ligating a first sequencing adaptor to the first cleaved genomic DNA segment and a second sequencing adaptor to the second cleaved genomic DNA segment. This single library can then be used to perform long-read sequencing. In some methods, only one long-read sequencing library is used (i.e., for a single long-read sequencing run).

[0212] The ability to use a single long-read sequencing library is made possible by the separated design of the first nuclease cleavage site and the second nuclease cleavage site in the disclosed methods. The sequencing target coverage can be achieved in a single, unamplified, long-read library preparation because the noncompetitive cleavages allow sequencing in both directions from a single library. In contrast, in the case that overlapping nuclease or guide RNA target sequences are used, two sequencing libraries would be required (e.g., for two long-read sequencing runs)—one library to read in each direction. Further, because there is no amplification of the genomic DNA, the first cleaved genomic DNA segment, or the second cleaved genomic DNA segment, the maximum number of sequencing reads at a particular locus is limited by the number of copies of that locus in the genome. In some embodiments, amplification is not used because it would not allow for long reads from the exogenous DNA target that extend into the surrounding reference genome sequence. For example, amplification by PCR may not extend past the homology arms or the inserted exogenous DNA sequence to allow for identification of integration sites. In some embodiments, use of two nuclease agents also allows for sequencing in both directions. For example, after cleavage with a Cas protein (e.g., Cas9), the Cas protein (e.g., Cas9) may remain bound to the DNA on the 5′ side of the cleavage site, resulting in preferential ligation of sequencing adaptors onto the 3′ side of the cleavage site. Particularly in the case of limited input DNA, methods of nuclease target siteAttorney Docket No.057766 / 629304 design with overlapping nuclease or guide RNA target sequences would fail to generate sufficient sequencing target coverage because they would require that limited input DNA be split into two libraries in order to accommodate reads in either direction. This creates more data that can negatively affect analysis (e.g., background / non-targeted) and consumes more reagents (e.g., two flow cells instead of one, two sequencing kits instead of one, etc.). Thus, the methods disclosed herein provide advantages over any previous method for identifying the site of genomic integration of exogenous DNA and allow for a more cost-effective, rapid, broadly applicable (i.e., targeting common components of targeting vectors), and accurate means of identifying the genomic location of genomically integrated exogenous DNA. C. Long-Read Sequencing

[0213] Following preparation of the single long-read sequencing library, the methods further comprise performing long-read sequencing on the library (i.e., a single long-read sequencing run) to generate a plurality of sequences from the first cleaved genomic DNA segment and the second cleaved genomic DNA segment. Long-read sequencing can be performed by any such appropriate long-read sequencing method. In one example, the long-read sequencing is nanopore sequencing. In another example, the long-read sequencing is single molecule real-time sequencing (SMRT sequencing). In some embodiments, a single long-read sequencing run is performed on a single long-read sequencing library. In some embodiments, the method does not comprise performing multiple long-read sequencing runs (i.e., in some methods, only one long- read sequencing run is performed).

[0214] Long-read sequencing is a form of next-generation sequencing (NGS) in which individual reads are each derived from a single molecule that is >1 kb in length, as opposed to traditional short-read sequencing methods in which <400 bp lengths of DNA are sequenced and reassembled. Long-read sequencing has technical advantages over short-read sequencing for the detection of specific types of genetic variation because it can provide much higher fidelity sequences. There are currently two major producers of long-read sequencing technology: Oxford Nanopore Technologies (ONT) sequencing (nanopore sequencing) and Pacific Biosciences (PacBio) single molecule real-time sequencing (SMRT sequencing).

[0215] Nanopore Sequencing. In nanopore sequencing (e.g., ONT nanopore sequencing), DNA bases are detected as they pass through a very small hole, e.g., modified protein nanoporeAttorney Docket No.057766 / 629304 (known as a nanopore) in a membrane. Linear DNA is linked to an enzyme that ‘unzips’ the double helix so that a single strand can be fed into the nanopore. Nanopores are present in a thin membrane submerged in a salt solution. An electrical potential is applied. This causes salt ions to pass through the pore, establishing a current. As the single-stranded DNA passes through the nanopore, each base (A, C, G or T) disrupts the voltage with a different electrical profile. The order of these types of current flow disruptions is read and translated into the order of bases in the DNA strand. Oxford Nanopore Technologies (ONT) nanopore sequencing is one example of nanopore sequencing. Other examples of nanopore sequencing include Genia Technologies’s nanotag-based real-time sequencing by synthesis (Nano-SBS) technology, NobleGen Biosciences’ optipore system, and Quantum Biosystems’s sequencing by electronic tunneling (SBET) technology.

[0216] Single Molecule Real-Time (SMRT) Sequencing. In SMRT sequencing (e.g., PacBio SMRT), a long chain of DNA is synthesized using the DNA to be sequenced as a template. Fluorescence is detected when a nucleotide is incorporated into the growing DNA strand. The DNA to be sequenced is first made circular. The circular DNA is applied to a surface (e.g., a ‘SMRT Cell’) patterned with thousands of tiny wells called ‘zero mode waveguides.’ Each well contains a single DNA polymerase enzyme working on a single circular DNA molecule. Fluorescently labelled nucleotides are used to generate a new strand of DNA from the circular DNA. As each nucleotide is incorporated by the polymerase into the new strand, its fluorescence is measured. This happens in each of the thousands of wells. Each circular DNA molecule is sequenced multiple times because there is no end to stop sequencing.

[0217] Synthetic Long-Read Sequencing. In addition to the long-read sequencing methods described above, it is also possible to perform short-read sequencing with modifications that allow the assembly of larger fragments after sequencing (synthetic long reads). Large fragments of DNA are separated from each other, for example, by attaching each fragment to a different bead. While separated from each other, the large fragments are sheared into shorter fragments and labeled with synthetic DNA barcodes. Short-read sequencing is performed for all short fragments together. During data analysis, short reads with the same barcode are identified and assembled back into the larger fragment from which they are derived.Attorney Docket No.057766 / 629304 D. Sequence Alignment and Analysis to Determine Genomic Location of Genomically Integrated Exogenous DNA

[0218] Following long-read sequencing, the methods further comprise aligning the plurality of sequences to both a reference genome and the exogenous DNA sequence to generate a plurality of hybrid alignments, thereby determining the genomic location of the genomically integrated exogenous DNA. Aligning to both the exogenous DNA sequence and the host reference genome simultaneously is helpful for the identification of reads that have this hybrid alignment mentioned below. Because the methods described herein are using long reads which begin in the genomically integrated exogenous DNA, the methods allow for high alignment scores to both references in a single sequencing run and produce efficient analysis. If only the background reference genome was used for alignments, identifying reads with hybrid alignments would be more difficult. This would require more bioinformatic sorting and statistical analysis to find reads with matched alignment scores to the exogenous DNA target. Specifically, the breakpoints in the genome alignment would have to be identified and subsequently this sequence would have to be investigated for the presence of the exogenous DNA sequence. The methods described herein identify the exogenous DNA simultaneously with the hybrid alignment.

[0219] Hybrid Alignments. Analysis of the long-read sequencing reads comprises generating hybrid alignments by aligning the reads that result from the library prep to not only the exogenous DNA reference sequence (target reference), but also the reference genome. This combination produces hybrid alignments for identification of the integration site. Exemplary reference genomes include the mouse mm9 and mm10 genomes, the human hg19 and hg38 genomes, and Chinese Hamster Ovary (CHO) genomic contigs. Target reference sequences can comprise common transgenic vector elements, including but not limited to, reporter genes, polyA signal sequences, promoter sequences, antibiotic selection genes, AAV payload and backbone sequences, cloning vector sequences, or any synthetic or exogenous nucleic acid that may be integrated in the genome, as well as nucleic acid insert sequence and / or homology arm sequence. Any such suitable reference genome corresponding to the organism in which the exogenous DNA has been inserted can be used, and likewise, any such suitable target reference corresponding to the exogenous DNA can be used.

[0220] After alignment of the sequencing read files, sequence alignment maps (SAM files) are produced containing read IDs, alignment positions, and alignment scores for each alignment position on a genome. Reads containing high map quality (good MAPQ alignment scores,Attorney Docket No.057766 / 629304 described below) are identified. Reads that contain high alignment scores to the target reference and the reference genome are expected to come from the designed nuclease agent cleavage site if the read alignment originates from the nuclease agent cleavage site. These long reads with multiple, strong alignments allow the methods to accurately identify the genomic break points between the exogenous DNA and the genome. Further, these methods are capable of identifying structural variations in the alignments for further characterization of the integration pattern. As described in the Examples, such methods have been demonstrated to successfully accommodate the complexity of large homology arms in LTVECs and accurately identify the genomic location of genomically integrated exogenous DNA in targeting constructs with homology arms >100 kb in length.

[0221] MAPping Quality Scores. The Phred score or MAPping Quality (MAPQ) score is the standard sequence alignment score in bioinformatics for the sequence alignment map (SAM) format. The SAM format is the standard sequence mapping format for aligned reads from a data set. The MAPQ score is calculated when each read alignment is made. When aligning a read to a genome reference or target reference, a local alignment is done where every read from a data set is queried for alignment to every coordinate in the reference genome. For example, a 3 kb nucleotide read is aligned to the mouse mm10 reference genome, which contains approximately 2.5 billion bases. There may be multiple locations on the mouse genome where parts of the single, 3 kb read may align. There may be alignment for the whole read at a specific place in the genome. The MAPQ score is designed to provide a probabilistic determination of the likelihood that a mapping position identified in the SAM file does not occur by chance, i.e., the alignment is correctly mapped. A 100% probability of correct alignment technically does not exist due to inherent error rates in current sequencing technologies.

[0222] MAPQ scoring is on a log scale, where the MAPQ Score = int(-10 log P) and P is the probability that the mapping position is correct. This calculation is done to the integer and is provided for each alignment. For example, a MAPQ10 = 0.9 probability of correct genomic alignment, or a 1 in 10 likelihood of the alignment incorrectly positioned. Likewise, a MAPQ20 = 0.99 probability of correct genomic alignment, a MAPQ30 = 0.999 probability of correct genomic alignment, a MAPQ40 = 0.9999 probability of correct genomic alignment, a MAPQ50 = 0.99999 probability of correct genomic alignment, and a MAPQ60 = 0.999999 probability ofAttorney Docket No.057766 / 629304 correct genomic alignment, which computes to a 1 in 1,000,000 probability of the alignment in question to be incorrectly aligned.

[0223] These MAPQ scores are calculated, by default, to any whole integer from 0 to 60. Because alignment scores are a product of computational calculation for each alignment, all alignment probabilities for each alignment above 1:1000000 have their calculations terminated and are represented as 1:1000000 or a MAPQ of 60. This is done both for speed of alignment and to limit computational resources for alignment runs. In other words, a MAPQ60 refers to a map quality score of 60 or greater.

[0224] The MAPQ score is an important data point assessing the identity and quality of read calls. In one example, the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a MAPQ score of at least MAPQ40. In another example, the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a MAPQ score of at least MAPQ50. A single, continuous long- read sequence that has genomic alignment in both the transgene target and the genome background at a specific position with such a MAPQ score should not exist in a wildtype organism that was not genetically engineered at that locus.

[0225] All patent filings, websites, other publications, accession numbers and the like cited above or below are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be so incorporated by reference. If different versions of a sequence are associated with an accession number at different times, the version associated with the accession number at the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or filing date of a priority application referring to the accession number if applicable. Likewise, if different versions of a publication, website or the like are published at different times, the version most recently published at the effective filing date of the application is meant unless otherwise indicated. Any feature, step, element, embodiment, or aspect of the invention can be used in combination with any other unless specifically indicated otherwise. Although the present invention has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims.Attorney Docket No.057766 / 629304 BRIEF DESCRIPTION OF THE SEQUENCES

[0226] The nucleotide and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and three-letter code for amino acids. The nucleotide sequences follow the standard convention of beginning at the 5′ end of the sequence and proceeding forward (i.e., from left to right in each line) to the 3′ end. Only one strand of each nucleotide sequence is shown, but the complementary strand is understood to be included by any reference to the displayed strand. When a nucleotide sequence encoding an amino acid sequence is provided, it is understood that codon degenerate variants thereof that encode the same amino acid sequence are also provided. The amino acid sequences follow the standard convention of beginning at the amino terminus of the sequence and proceeding forward (i.e., from left to right in each line) to the carboxy terminus.

[0227] Table 2. Description of Sequences. SEQ ID NO Type Description 1 Protein Cas9 proteinAttorney Docket No.057766 / 629304 SEQ ID NO Type Description 33 DNA Luc_guide_1rev_(target) 34 RNA Spec guide fw (guide)EXAMPLES Example 1. A Process for Identifying Transgene Integrations in Mouse and Human Genomes Using Long-Read Oxford Nanopore Sequencing

[0228] Regeneron’s VELOCIGENE® has been used to make genetically modified mouse models for early-stage target discovery and other purposes. To accomplish this, targeting vector DNA is sourced from Bacterial Artificial Chromosomes (BACs) or is synthetically derived. These large DNA constructs are introduced into the genome of the mice or other model organismAttorney Docket No.057766 / 629304 via targeted homologous recombination. The targeting vector is designed with homology for the target genomic locus of interest with the intent of a site-specific integration into the target genomic locus. CRISPR / Cas9 can also be used to mediate this process by introducing a double- strand break to promote targeted transgene integration at the desired locus. This process can also be repeated to generate several targeted modifications to generate a precise large genomic modification of interest.

[0229] When genomic recombination occurs with targeting vectors, the targeting vector insert sequence can sometimes integrate into other genomic locations in addition to or instead of the desired locus. The general targeting efficiency is ~10%, where ~90% of clones produced are non-targeted transgenic integrations characterized by TAQMAN® qPCR assays located in targeting vector selection cassette and / or human / mutant sequence to be introduced. In some special cases this random transgenic integration is acceptable, specifically when integrated in safe harbor sites where there is no gene or regulatory sequence around the transgenic integration site. Integrations can also happen in sites that are active due to proximity of promoters, enhancers, or other regulatory regions. TAQMAN® qPCR is a strong quantitative method for determining copy number but must be designed specifically for the target and cannot provide information about location. Short read exome and whole genome sequencing is not able to determine the transgenic integration site as precisely and requires a large amount of sequence coverage and processing to do it successfully.

[0230] To address these deficiencies, we developed a process for transgene insertion identification using long-read sequencing and CRISPR / Cas digests and sequence alignments. With the ability to use long-read sequencing, specifically using reads with a known starting location (CRISPR cut and adaptor ligation within the transgene), long reads from the expected transgene sequence can be aligned to the reference genome and transgene sequence, allowing for efficient and precise genomic integration site identification. In the process that we developed, guide RNAs are designed to target a guide RNA target sequence in the transgene (e.g., inside of selection cassette) in an orientation designed to add sequencing adaptors and originate long reads in the transgene. Because of our guide RNA target sequence design and distance, we can control the strands that the sequencing adaptors can efficiently bind to and facilitate optimal sequencing output. This method has been used to identify multiple genomic integration sites from a single targeting event. This method is also able to identify multiple concatenated copies of genomicAttorney Docket No.057766 / 629304 integrations. Targeted integrations from 4 kb to160 kb have been successfully identified with this method. AAV integrations and safe harbor sites have also been identified using this process. This is done without the need for whole genome sequencing. This process has been done with genomic DNA sourced from mouse ES cells, tail snips, and ear clippings. This allows the mice to be correctly characterized earlier and without the need to sacrifice.

[0231] The method described herein uses long-read sequencing to generate sequencing libraries enriched for the regions of interest. After alignment to the reference genome, the long- read capability of long-read nanopore sequencing spans entire transgene and vector sequences, providing a comprehensive view of the integration sites. Subsequent read alignment and breakpoint identification are performed using a custom bioinformatics pipeline designed to handle the unique challenges posed by long-read sequencing data, such as sequencing error rates and how to interpret complex transgene rearrangements.

[0232] In the process, CRISPR guide RNAs are first designed for a target locus (e.g., within the transgene) for the purposes of originating long reads at this locus. After DNA extraction, 5′ ends are dephosphorylated to reduce ligation of sequencing adaptors to non-target strands. Cas9 ribonucleoprotein particles (RNPs), with bound crRNA and tracrRNA (or sgRNA), are added to the genomic DNA, then bind and cleave the region of interest in the genomically integrated exogenous DNA. dsDNA cleavage by Cas9 reveals blunt ends with ligatable 5′ phosphates. All the DNA in the samples is dA-tailed, which prepares the blunt ends for ligation to the sequencing adaptors. Sequencing adaptors are ligated primarily to Cas9 cut sites, which are both 3′ dA-tailed and 5′ phosphorylated. After cleavage with Cas9, the enzyme remains bound to the DNA on the 5′ side of the guide RNA, resulting in preferential ligation of adaptors onto the 3′ side of the cut. Long reads are then aligned to both the target and the genomic reference for identification of integration. Reads that align to both target and genome are identified as transgenic sequence. An informatics process then finds reads that are contained in both and identifies genomic location and breakpoints of alignments. An example of a specific process is as follows (see Figure 1): • Design CRISPR guides specifically for target sequence (suspected sequence integrated). Guides are designed to target guide RNA target sequences on opposite strands (i.e., the PAMs are on opposite strands), with the guide RNA target sequence on the sense strand being located 3′ of the guide RNA target sequence on the antisense strand. • Extract high molecular weight genomic DNA from embryonic stem cells or tissues.Attorney Docket No.057766 / 629304 • Align to a concatenated file containing full reference genome (e.g., full mouse reference genome) and target sequence. • Identify individual reads that align to both references with high Phred quality scores (Q- score). • >Q40 alignments for both references at a particular location are accepted, where Qscore = -10 log10P where P is probability of alignment error. Q40 = 1 in 10000 probability of incorrect alignment or 99.99% likelihood of correct alignment. • When Qscores are not high enough quality, lower alignment parameters are used to generate alignments. • The confidence in the transgenic call is associated with this alignment score and probability, where higher scored alignments (for both references) are more confident than lower. • This alignment probability is calculated for both the target sequence and the genomic reference sequence. • Long reads in this process have a systematic advantage over shorter reads because of the length of the alignment query allowing for higher scored alignments, regardless of nucleotide level base call inaccuracies. • Identify large variants associated with these highly scored alignments. • Identify and report alignments associated with modified alleles and their coordinates in the reference genome. Materials and Methods.

[0233] DNA Extractions. DNA extractions were done to maintain high molecular weight genomic DNA. The NEB Monarch kit was used for the high molecular weight DNA extractions from tissues and cells.

[0234] Sample Preparation: Cells were resuspended in phosphate-buffered saline (PBS). Tissue samples were dissected into small pieces (<5mm) using a sterile scalpel.

[0235] Lysis: The appropriate volume of Monarch HMW DNA Tissue Lysis Buffer and Proteinase K (as per manufacturer’s instructions) was added to the samples. The samples were mixed by inversion and incubated at 56°C for 35 minutes (cells) or 60 minutes (tissues).Attorney Docket No.057766 / 629304

[0236] RNase Treatment: RNase A was added to the lysate followed by incubation at 37°C for 30 minutes.

[0237] Protein Removal: Protein Precipitation Buffer was added to the lysate. The samples were centrifuged at maximum speed for 10 minutes to pellet the proteins.

[0238] DNA Precipitation: The supernatant was transferred to a new tube, avoiding the protein pellet. Isopropanol was added to the supernatant and mixed by inversion. The samples were centrifuged at maximum speed for 10 minutes to pellet the DNA.

[0239] DNA Wash: The supernatant was removed, and DNA Wash Buffer was added to the pellet. After inversion, the samples were centrifuged at maximum speed for 5 minutes. The wash buffer was carefully removed.

[0240] DNA Elution: Elution Buffer or nuclease-free water was added to the DNA pellet. After incubation at room temperature for 5 minutes, the samples were centrifuged at maximum speed for 1 minute to elute the DNA.

[0241] Storage: The extracted HMW DNA was stored at 4°C for short term storage and - 20°C for long-term storage.

[0242] This protocol was performed under aseptic conditions to prevent contamination. The exact volumes and incubation times were adjusted according to the specific sample type and the manufacturer’s instructions.

[0243] Sequencing: The ONT Cas9 Sequencing Kit was used.

[0244] Analysis: all read data are Fastq basecalled files. These files are first concatenated into a single file and aligned to the reference genome.

[0245] Guide RNA Design and Synthesis: Guide RNAs (gRNAs) were designed to target the genomic regions of interest. The gRNAs were synthesized according to the manufacturer's instructions.

[0246] Guide RNAs were designed and validated in common selection genes used in selection cassettes. The idea is that these guide RNAs can direct the targeted sequencing in the genome wherever these selection cassettes are integrated. These genes would be selected for in the samples as they are grown on resistance plates. Only clones containing the resistance will be able to grow on these plates. The following guide RNA sequences were designed and validated.Attorney Docket No.057766 / 629304

[0247] Table 3. Designed and Validated Guide RNAs. Transgen SEQ SEQ Strand e Element Guide Name Guide RNA Target Sequence ID NO ID NO (target) (guide)Attorney Docket No.057766 / 629304

[0248] Table 4. Distances Between Guide RNA Target Sequences. Distance Between Targets (bp) Strand Transgene Element Guide Name + Spectinomycin resistance gene Spec guide fw

[0249] Cas9 Digestion: The genomic DNA was mixed with the two guide RNAs designed to target guide RNA target sequences on opposite strands and Cas9 nuclease in the provided reaction buffer. The mixture was incubated at 37°C for 1 hour in a thermocycler to allow for Cas9 cleavage at both guide RNA target sequences.

[0250] Digestion Cleanup: The Cas9 digestion mixture was cleaned up using the provided cleanup reagents and spin columns, according to the manufacturer’s instructions (Oxford Nanopore SQK-CS9109). The cleaned-up DNA was eluted in nuclease-free water.

[0251] Library Preparation: The cleaned-up DNA was used for library preparation. The DNA was end-repaired, A-tailed, and ligated to sequencing adaptors using the provided reagents and following the manufacturer’s instructions (Oxford Nanopore SQK-CS9109).

[0252] Library Cleanup: The amplified library was cleaned up using the provided cleanup reagents and spin columns. The cleaned-up library was eluted in nuclease-free water or provided elution buffer.Attorney Docket No.057766 / 629304

[0253] Library Quantification and Validation: The library concentration was quantified using a Qubit fluorometer or similar instrument. The library size distribution was validated using the Femto Pulse DNA analyzer. Results.

[0254] In a first proof-of-concept experiment, a gene humanization was done into B6 mice following VELOCIGENE’s normal humanization targeting vector design and construction. For this gene humanization, the targeting vector size was 7185 bp, and the guide RNAs used for nanopore sequencing were designed to target the humanized region with 338 bp between the two sequencing guide cuts. After pronuclear injection, the mice were born and 2 toe clips were taken from 10 mice samples from this cohort. The samples were eluted in 25 µL of elution buffer. After gDNA extraction was done for these samples, two of the samples gave usable concentration for sequencing, the yields from the extractions were 1.4 µg and 0.98 µg with concentrations of 55 ng / µL and 39 ng / µL (samples 5 and 6, respectively). The lower concentration inputs were not used for sequencing in this experiment. The full 1.4 µg and 0.9 8µg inputs were used for the genomic sequencing with guides designed in our design strategy for maximization of results with minimal DNA input. Although the protocol from Oxford Nanopore calls for 5 µg of input DNA, we have had success at sequencing in a single run with roughly 1 / 5 of the recommended input. No integration was observed in sample #5 integration, and this was confirmed with TAQMAN data suggesting that vector did not integrate in sample #5. Using the process described above, we were able to identify the genomic integration at the desired targeting locus on chromosome 6 and an off target, transgenic integration on chromosome 17 in sample #6. There were 33x reads aligned in the genome for the integration in chromosome 6. With whole genome sequencing, 33x coverage at a single locus in the mouse genome (2.5 gigabases of genomic sequence) would require 2.5Gb x 33 = ~83 gigabases of sequencing data, assuming equal coverage across the genome. The total sequencing output for this run was 2.32 gigabases of data. This dramatically reduces the need for more input DNA which would require sacrificing the mouse. The mouse remained alive throughout this experiment. Further, we were able to determine that there were reads that were continuous across a concatenated reference file that aligns between a forward guide cut and a reverse guide cut, indicating that the targeting event was a tandem integration. With our CRISPR design and sequencing process, we were able toAttorney Docket No.057766 / 629304 positively identify the desired targeting event and an off-target transgenic integration from the data acquired. We were also able to characterize the pattern of genomic integration for the samples in the experiment as well. We were able to attain this data without the normal DNA amounts needed for whole genome sequencing or the recommended CRISPR Targeting enrichment sequencing and acquire the data in a single library prep, furthering the efficiency of the process. Although this first proof-of-concept experiment is one example using a pair of guide RNAs targeting one specific humanized sequence, guide RNA designed against other humanized sequences in other humanized mouse models have been validated in the methods described above to identify genomic integration sites of transgenes (data not shown).

[0255] In a second proof-of-concept experiment (AAV integration), the objective was to determine if AAVs containing luciferase were successfully integrated at the designated safe harbor loci in the human genome with our process. The AAV was 3.7 kb (including internal tandem repeats) and contained an enhancer, a luciferase gene, and a polyA signal. To test this, cells that were first targeted with CRISPR cuts at the desired safe harbor sites 1, 2, and 3, together with the AAV-luciferase. Because the cells did fluoresce in the luciferase assay, the AAV luciferase reporter should be either genomically integrated or episomal in the cytoplasm of the human cells. We hypothesized that the data that we obtain from the luciferase guide RNAs cuts according to our design should be able to span the AAV and map the genomic integrations as well as provide information about the patterns of the AAV integrations. The guide RNAs used for nanopore sequencing were designed to target the luciferase coding sequence with 364 bp between the two sequencing guide cuts (Fwd guide RNA target sequence = TATTATCATGGATTCTAAAA (SEQ ID NO: 68); Rev guide RNA target sequence = CTTATGCAGTTGCTCTCCAG (SEQ ID NO: 69)). The distances between the cut sites and the junctions of the transgene with the endogenous genome were 1764 bp and 1583 bp, respectively. The reads from the project identified the genome integrations from the luciferase guide RNA cuts and showed alignment on the hg38 human reference genome. The specific genomic integrations were determined. We were also able to determine the patterns of the chromosomal genomic integrations and the episomal integrations from the reads in the sequencing run. As in the first POC experiment we were able to identify genomic integration patterns and sites for AAVs in human samples. We were able to do this using long reads which are able to have sequence alignments while genomic rearrangements persist. Although this second proof-of-Attorney Docket No.057766 / 629304 concept experiment is one example using a pair of luciferase guide RNAs from Table 3, all of the guide RNAs in Table 3 have been validated in the methods described above to identify genomic integration sites of transgenes (data not shown).

[0256] Our results demonstrate the effectiveness of this process in accurately identifying transgene integration sites. We have also characterized tandem / inverted integrations and off target integrations of AAVs in the host genome. This method allows for the precise and more comprehensive understanding of vector and transgene integration, an important consideration given the recent advances in AAV based gene therapies. Other methods like targeted locus amplification use amplification and whole genome sequencing. Other long-read methods use guides to sequence at known genomic locations or rely on whole genome sequencing to identify breakpoints in foreign sequence in the genome. In contrast, the approach described herein is targeted and efficient and reduces the amount of time, recourses, and sequence coverage needed to identify transgenic sequences. This process needs only a partial known sequence expected to be integrated for guide RNA design for targeting. Example 2. Identifying Transgene Integrations in Mouse Genomes Using Long-Read Oxford Nanopore Sequencing and Comparison to Other Methods

[0257] The method described in Example 1 (hereinafter referred to as VelociTRansgenic Allele Characterization [VelociTRAC]) was used to identify transgene integration sites in mouse cells. In this experiment, a targeting vector was designed for targeting at the genomic locus chr1:40,797,406 of the mouse genome. This vector contained 2 kb targeting arms for homologous recombination into the genome. Four clones were tested: A1, C5, D3, and G3. After performing the VelociTRAC method outlined in Example 1 using a pair of guide RNAs designed to target the neomycin selection cassette (Kan_guide_fw and Kan_guide_rev1, with 367 bp between the cuts), three different genomic integrations were found for the four tested clones. The genome viewer alignments of reads (containing targeting vector) aligned to four different positions of the mouse (mm10) reference genome. The positions in the mouse genome were the loci at chr1:40,797,406; chr13:81,211,876; chr2:25,429,563; chr4:147,963,387; and chr9: 119,645,098. We were able to identify these genomic integrations at these positions with relatively deep coverage from 15x to 900x at locus. The alignment coverage and directionality of the long-read sequencing reads is shown in Table 4. As explained in Example 1, after cleavage with Cas9, the CRISPR / Cas ribonucleoprotein complex remains bound to the DNA on the 5′ sideAttorney Docket No.057766 / 629304 of the guide RNA complementary target DNA, resulting in preferential ligation of sequencing adaptors onto the 3′ side (of the sense strand) or downstream of the double strand cut. Thus, the guide RNAs used in this experiment were designed to target guide RNA target DNA sequences on opposite strands (i.e., the PAMs are on opposite strands), with the guide RNA target sequence on the sense strand being located downstream of the guide RNA target sequence on the antisense strand. This should result in a larger number of sequencing reads from the Cas9 cuts oriented outward into the genome, which is preferable for determining the location of the genomic integration, as compared to sequencing reads inward between the two cuts. This preferred directionality is confirmed in the aligned sequencing reads in Table 4, demonstrating that, as designed, there is more alignment coverage for sequencing reads outward from the cuts than inward between the cuts, allowing for better and more efficient identification of genomic integration sites.

[0258] Table 4. Alignment Coverage for Sequencing Reads for Four Clones Alignment Coverage Clone # Sequencing Reads Aligned # Se uencin Reads Ali ned # Sequencing Reads Aligned Cut

[0259] What was achieved through the VelociTRAC method is beneficial over what could be achieved with whole genome sequencing. The genomic sequencing for this project with this method yielded at locus peak alignment coverage of 900x, 27x, and 14x for the clones A1, C5 and D3, respectively. This coverage was achieved with the VelociTRAC method from 2.44 gigabases (gb), 1.79 gb, and 1.41 gb of genomic sequencing data, respectively. The mouse haploid mm10 reference genome (19 autosomes, XY allosomes, and mitochondrial genome) is approximately 2.7 gigabases total. The amount of sequencing data used for genomic integration identification in these runs equates to 0.9x, 0.7x, and 0.5x full genome coverage, respectively. This low level of sequencing coverage from exhaustive whole genome sequencing would likely not yield the ability to identify these transgenic integrations. To achieve 900x on-target coverage with whole genome sequencing would require approximately 2.4 terabases of genomic data. This is expensive in animal sample material, extracted (and quality checked) high quality highAttorney Docket No.057766 / 629304 molecular weight DNA, sequencing resources (hardware, reagents), data acquired, and computational resources. The on-target enrichment of sequencing with the VelociTRAC method makes this method more cost efficient and effective to identify transgenic genomic integrations. Example 3. Identifying Transgene Integrations in Human Genomes Using Long-Read Oxford Nanopore Sequencing

[0260] The VelociTRAC method described in Example 1 was used for identification of targeting by a plasmid in human induced pluripotent stem cells. In this experiment, the goal was to validate the targeting of the plasmid at a target genomic locus (hAAVS1 genomic locus) of a human induced pluripotent stem cell line. For this run, a pair of guide RNAs targeting the puromycin selection cassette was used (Puro_guide_fw and Puro_guide_rev, with 241 bp between the cuts), and sequencing out into the genome was performed to identify the location of the genomic targeting. Six clones were sequenced, including two previously validated control samples. The alignment coverage and directionality of the long-read sequencing reads is shown in Table 5. As explained in Example 1, after cleavage with Cas9, the CRISPR / Cas ribonucleoprotein complex remains bound to the DNA on the 5′ side of the guide RNA complementary target DNA, resulting in preferential ligation of sequencing adaptors onto the 3′ side (of the sense strand) or downstream of the double strand cut. Thus, the guide RNAs used in this experiment were designed to target guide RNA target DNA sequences on opposite strands (i.e., the PAMs are on opposite strands), with the guide RNA target sequence on the sense strand being located downstream of the guide RNA target sequence on the antisense strand. This should result in a larger number of sequencing reads from the Cas9 cuts oriented outward into the genome, which is preferable for determining the location of the genomic integration, as compared to sequencing reads inward between the two cuts. This preferred directionality is confirmed in the aligned sequencing reads in Table 5, demonstrating that, as designed, there is more alignment coverage for sequencing reads outward from the cuts than inward between the cuts, allowing for better and more efficient identification of genomic integration sites.Attorney Docket No.057766 / 629304

[0261] Table 5. Alignment Coverage for Sequencing Reads for Four Clones Alignment Coverage Clone # Sequencing Reads Aligned everse Cut # Sequenci # Sequencing Reads Aligned Upstream of R ng Reads Aligned Downstream of Forward Cutlocus at Chr19: 55,115,766 in the GRCh38 human reference genome. All clones had a single genomic integration called at the target genomic locus. From these data, we were also able to identify that some of the integrations contained the backbone and Amp sequence from the plasmid used for targeting. Another significant finding was that one of the clones (C2) contained a two-copy tandem insertion in the target genomic locus. This finding was further validated by obtaining these reads and aligning them to a tandem reference file of the vector to visualize the alignments spanning two vector references. This result was significant as it can have an impact on expression levels from the targeted sequence at the locus. Example 4. Sensitivity of Assay

[0263] An experiment was done to determine the sensitivity of the VelociTRAC method used in Example 1. In this experiment, four mixtures of transgenic mouse DNA combined with wild type mouse genomic DNA were tested, with mixtures of 100% transgenic and 0% wild type, 75% transgenic and 25% wild type, 50% transgenic and 50% wild type and 25% transgenic and 75% wild type genomic DNA. While the reads produced did not show a linear correlation of transgene targeted reads as the concentration of the transgenic DNA increased, the transgene identifications were detected in the 25% transgene DNA sample. For this experiment, the total input across all dilutions was fixed to 3 µg genomic DNA input. This demonstrates that the VelociTRAC method may be able to detect transgenic integrations in a mosaic or heterogenous DNA sample.

Claims

Attorney Docket No.057766 / 629304 We claim:

1. A method of determining the genomic location of a genomically integrated exogenous DNA, comprising: (a) preparing a single long-read sequencing library, wherein the preparing comprises cleaving genomic DNA extracted from cells comprising a genomically integrated exogenous DNA with: (i) a first nuclease agent that cleaves a first nuclease cleavage site within the genomically integrated exogenous DNA to generate a first cleaved genomic DNA segment and (ii) a second nuclease agent that cleaves a second nuclease cleavage site within the genomically integrated exogenous DNA to generate second cleaved genomic DNA segment, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 50 bp; (b) performing long-read sequencing on the single long-read sequencing library to generate a plurality of sequences from the first cleaved genomic DNA segment and the second cleaved genomic DNA segment; and (c) aligning the plurality of sequences to both a reference genome and the exogenous DNA sequence to generate a plurality of hybrid alignments, thereby determining the genomic location of the genomically integrated exogenous DNA.

2. The method of claim 1, wherein the long-read sequencing is nanopore sequencing or single molecule real-time sequencing.

3. The method of claim 1 or 2, wherein the long-read sequencing is nanopore sequencing.

4. The method of any preceding claim, wherein the method does not comprise amplifying the genomic DNA, the first cleaved genomic DNA segment, or the second cleaved genomic DNA segment.

5. The method of any preceding claim, wherein the exogenous DNA is at least 0.5 kb, at least 1 kb, at least 5 kb, at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb.Attorney Docket No.057766 / 629304 6. The method of any preceding claim, wherein the exogenous DNA is at least 100 kb.

7. The method of any preceding claim, wherein step (a) comprises cleaving a single genomic DNA sample with the first nuclease agent and the second nuclease agent simultaneously.

8. The method of any preceding claim, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp, at least 125 bp, at least 150 bp, at least 200 bp, or at least 250 bp.

9. The method of any preceding claim, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by about 100 bp to about 2000 bp or about 100 bp to about 500 bp.

10. The method of any preceding claim, wherein the first nuclease cleavage site and the second nuclease cleavage site are each at least 100 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA.

11. The method of any preceding claim, wherein the first nuclease cleavage site and the second nuclease cleavage site are each at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA.

12. The method of any preceding claim, wherein the first nuclease cleavage site and the second nuclease cleavage site are each about 1 kb to about 250 kb from the 5′ and 3′ ends of the genomically integrated exogenous DNA.

13. The method of any preceding claim, wherein the first nuclease agent is a first zinc finger nuclease (ZFN), a first transcription activator-like effector nuclease (TALEN), orAttorney Docket No.057766 / 629304 a Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-associated (Cas) protein and a first guide RNA (gRNA), and the second nuclease agent is a second ZFN, a second TALEN, or a Cas protein and a second gRNA.

14. The method of any preceding claim, wherein the first nuclease agent is a Cas9 protein and a first gRNA and the second nuclease agent is a Cas9 protein and a second gRNA.

15. The method of claim 14, wherein the first gRNA targets a first guide RNA target sequence and the second gRNA targets a second guide RNA target sequence, wherein the first guide RNA target sequence and the second guide RNA target sequence are on opposite strands, wherein the guide RNA target sequence on the sense strand is located 3′ of the guide RNA target sequence on the antisense strand.

16. The method of any preceding claim, wherein the exogenous DNA sequence comprises a selection cassette or a reporter protein coding sequence, and the first nuclease cleavage site and the second nuclease cleavage site are within the selection cassette or the reporter protein coding sequence, optionally wherein the selection cassette comprises a kanamycin resistance gene, a hygromycin resistance gene, or a puromycin resistance gene, and the first nuclease cleavage site and the second nuclease cleavage site are within the kanamycin resistance gene, the hygromycin resistance gene, or the puromycin resistance gene.

17. The method of any one of claims 1-15, wherein the first nuclease cleavage site and the second nuclease cleavage site are within a spectinomycin resistance gene, a kanamycin resistance gene, a hygromycin resistance gene, a puromycin resistance gene, a lacZ gene, a tamoxifen-inducible Cre gene, a luciferase gene, a human ubiquitin promoter, a mouse phosphoglycerate kinase promoter, or a mouse protamine promoter.

18. The method of any preceding claim, wherein the amount of genomic DNA in step (a) is at least 400 ng, is less than 3 µg, or is from about 400 ng to about 3 µg.Attorney Docket No.057766 / 629304 19. The method of any preceding claim, wherein the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a mapping quality (MAPQ) score of at least MAPQ40.

20. The method of any preceding claim, wherein the hybrid alignments used to determine the genomic location of the genomically integrated exogenous DNA have a MAPQ score of at least MAPQ50.

21. The method of any preceding claim, wherein the cells are mammalian cells.

22. The method of any preceding claim, wherein the cells are human cells, optionally wherein the cells are human induced pluripotent stem cells.

23. The method of any one of claims 1-21, wherein the cells are rodent cells.

24. The method of claim 23, wherein the cells are mouse cells, optionally wherein the cells are mouse embryonic stem (ES) cells or optionally wherein the cells are from a living mouse.

25. The method of claim 23, wherein the cells are rat cells, optionally wherein the cells are rat ES cells or optionally wherein the cells are from a living rat.

26. The method of any preceding claim, wherein the cells are from a tissue sample from a subject.

27. The method of any preceding claim, further comprising extracting genomic DNA from the cells comprising the genomically integrated exogenous DNA prior to step (a).Attorney Docket No.057766 / 629304 28. The method of claim 27, further comprising genetically modifying a population of cells to generate the cells comprising the genomically integrated exogenous DNA prior to extracting the genomic DNA.

29. The method of claim 28, wherein the genetically modifying comprises administering the exogenous DNA to the population of cells such that the exogenous DNA is integrated into the genome, optionally wherein the exogenous DNA that is administered is in the form of a viral vector, an adeno-associated virus (AAV) vector, a single-stranded oligodeoxynucleotide (ssODN), or a targeting vector comprising a 5′ homology arm and a 3′ homology arm.

30. The method of claim 28, wherein the genetically modifying comprises administering a large targeting vector to the population of cells, wherein the large targeting vector comprises a 5′ homology arm, a 3′ homology arm, and optionally an insert nucleic acid flanked by the 5′ homology arm and the 3′ homology arm, wherein the large targeting vector is at least 10 kb in length or wherein the sum total of the 5′ homology arm and the 3′ homology arm is at least 10 kb in length.

31. The method of claim 30, wherein the large targeting vector is at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb.

32. The method of claim 30 or 31, wherein the large targeting vector is at least 100 kb.

33. The method of any one of claims 30-32, wherein the insert nucleic acid is at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb.

34. The method of any one of claims 30-33, wherein the insert nucleic acid is at least 100 kb.Attorney Docket No.057766 / 629304 35. The method of any one of claims 30-34, wherein the 5′ homology arm, the 3′ homology arm, or the sum total of the 5′ homology arm and the 3′ homology arm is at least 10 kb, at least 50 kb, at least 100 kb, at least 150 kb, or at least 200 kb.

36. The method of any one of claims 30-35, wherein the 5′ homology arm, the 3′ homology arm, or the sum total of the 5′ homology arm and the 3′ homology arm is at least 100 kb.

37. The method of any preceding claim, wherein the long-read sequencing is nanopore sequencing, wherein the method does not comprise amplifying the genomic DNA segment, the first cleaved genomic DNA, or the second cleaved genomic DNA segment, wherein the exogenous DNA is at least 10 kb, and wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp.

38. The method of any preceding claim, wherein the long-read sequencing is nanopore sequencing, wherein the exogenous DNA is at least 10 kb, wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp, wherein the first nuclease agent is Cas9 protein and a first gRNA and the second nuclease agent is the Cas9 protein and a second gRNA, and wherein the exogenous DNA sequence comprises a selection cassette or a reporter protein coding sequence, and the first nuclease cleavage site and the second nuclease cleavage site are within the selection cassette or the reporter protein coding sequence.

39. The method of any preceding claim, wherein step (a) comprises cleaving a single genomic DNA sample with the first nuclease agent and the second nuclease agent simultaneously,Attorney Docket No.057766 / 629304 wherein the first nuclease cleavage site and the second nuclease cleavage site are separated by at least 100 bp, wherein the first nuclease cleavage site and the second nuclease cleavage site are each at least 500 bp from the 5′ and 3′ ends of the genomically integrated exogenous DNA, wherein the first nuclease agent is a Cas9 protein and a first gRNA and the second nuclease agent is a Cas9 protein and a second gRNA, wherein the first gRNA targets a first guide RNA target sequence and the second gRNA targets a second guide RNA target sequence, wherein the first guide RNA target sequence and the second guide RNA target sequence are on opposite strands, and wherein the guide RNA target sequence on the sense strand is located 3′ of the guide RNA target sequence on the antisense strand.

Citation Information

Patent Citations

  • Functional genomics using zinc finger proteins

    US20020081614A1

  • Regulation of angiogenesis with zinc finger proteins

    US20030021776A1

  • Methods and compositions for using zinc finger endonucleases to enhance homologous recombination

    US20030232410A1

  • Use of chimeric nucleases to stimulate gene targeting

    US20050026157A1

  • Methods and compositions for targeted cleavage and recombination

    US20050064474A1