Nuclease profiling system

A 'one-cut' screening strategy for nucleases improves specificity by identifying target sites with minimal off-target activity, addressing the cytotoxicity issues of current endonucleases and enhancing genome manipulation safety and efficacy.

JP7850445B2Active Publication Date: 2026-04-23PRESIDENT & FELLOWS OF HARVARD COLLEGE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PRESIDENT & FELLOWS OF HARVARD COLLEGE
Filing Date
2023-02-03
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Current site-specific endonucleases exhibit considerable off-target activity, leading to cytotoxicity and undesirable genomic modifications, making them unsuitable for clinical use and efficient genome manipulation.

Method used

A 'one-cut' screening strategy for evaluating nuclease specificity, compatible with single-end high-throughput sequencing, is employed to identify and design nucleases with improved specificity by selecting target sites that are distinct from other genomic sequences and using guide RNAs to minimize off-target binding and cleavage.

Benefits of technology

The method enables the identification of highly specific endonucleases for clinical use by reducing off-target cleavage, enhancing the safety and efficacy of genome manipulation techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850445000026
    Figure 0007850445000026
  • Figure 0007850445000027
    Figure 0007850445000027
  • Figure 0007850445000028
    Figure 0007850445000028
Patent Text Reader

Abstract

Strategies, methods, and reagents are provided for determining the nuclease target site preference and specificity of site-specific endonucleases. [0003] Some methods provided herein utilize a novel "single cleavage" strategy to screen a library of concatemers containing repeat units of candidate nuclease target sites and constant insertion regions to identify library members cleaved by a nuclease of interest by sequencing intact target sites adjacent to and identical to the cleaved target site. Some aspects of the present disclosure provide strategies, methods, and reagents for selecting site-specific endonucleases based on determining their target site preference and specificity. Methods and reagents for determining target site preference and specificity are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims priority under 365(c) of U.S. Patent Act § 365(c) to U.S. Patent Application USSN 14 / 320,370 filed June 30, 2014 and U.S. Patent Application USSN 14 / 320,413 filed June 30, 2014, and also claims priority under 119(e) of U.S. Patent Act § 119(e) to U.S. Provisional Patent Application USSN 61 / 864,289 filed August 9, 2013, each of which is incorporated herein by reference.

[0002] Government aid This invention was made with the assistance of the U.S. Government under grants HR0011-11-2-0003 and N66001-12-C-4207, awarded by the Defense Advanced Research Projects Agency. The U.S. Government reserves the rights to this invention. [Background technology]

[0003] Site-specific endonucleases theoretically enable targeted manipulation of single sites within the genome, making them useful in the context of gene targeting and even for therapeutic applications. In a variety of organisms, including mammals, site-specific endonucleases have been used for genomic manipulation by stimulating either non-homologous end joining or homologous recombination. In addition to providing a powerful research tool, site-specific nucleases also have potential as gene therapies, and two site-specific endonucleases have recently entered clinical trials. One is CCR5-2246, which targets the human CCR-5 allele as part of an anti-HIV therapy (NCT00842634, NCT01044654, NCT01252641), and the other is VF24684, which targets the human VEGF-A promoter as part of an anti-cancer therapy (NCT01082926).

[0004] Specific cleavage of the target nuclease site with no or minimal off-target activity is a necessary condition for the clinical use of site-specific endonucleases, as well as for highly efficient genome manipulation in basic research. This is because imperfect specificity of the manipulated site-specific binding domain is linked to cytotoxicity and undesirable modification of genomic loci other than the target. However, most nucleases available today exhibit considerable off-target activity and may therefore be unsuitable for clinical use. Therefore, techniques are needed to evaluate nuclease specificity and to manipulate nucleases with improved specificity. [Overview of the project]

[0005] Some aspects of this disclosure are based on the recognition that the reported toxicity of some manipulated site-specific endonucleases is based not only on off-target binding but also on off-target DNA cleavage. Some aspects of this disclosure provide strategies, compositions, systems, and methods for evaluating and characterizing the sequence specificity of site-specific nucleases (e.g., RNA-programmable endonucleases, e.g., Cas9 endonucleases, zinc finger nucleases (ZNFs), homing endonucleases, or transcription activator-like element nucleases (TALENs)).

[0006] The strategies, methods, and reagents of this disclosure represent, in several respects, improvements over previous methods for assaying nuclease specificity. For example, several previously reported methods for determining nuclease target site specificity profiles by screening libraries of nucleic acid molecules containing candidate target sites rely on a "two-cut" in vitro selection method, which requires the indirect reconstruction of the target site from the sequences of two half-sites resulting from two adjacent cuts of the nuclease in the library member nucleic acid (see, e.g., PCT application WO2013 / 066438 and Pattanayak, V., Ramirez, CL, Joung, JK & Liu, DR Revealing off-target cleavage specificities of zinc-finger nucleases by in vitro selection. Nature Methods 8, 765-770 (2011). The overall contents of each are incorporated herein by reference). In contrast to such "two-cut" strategies, the method of the present disclosure utilizes an optimized "one-cut" screening strategy, which enables the identification of library members that have been cut at least once by a nuclease. As described in more detail elsewhere herein, the "one-cut" selection strategy provided herein is compatible with single-end high-throughput sequencing methods and does not require computer reconstruction of the cut target site from the cut half-site, thus streamlining the nuclease profiling process.

[0007] Several aspects of this disclosure provide in vitro selection methods for evaluating the cleavage specificity of endonucleases and for selecting nucleases having a desired level of specificity. Such methods are useful, for example, for characterizing an endonuclease of interest and for identifying nucleases that exhibit a desired level of specificity (e.g., for identifying highly specific endonucleases for clinical use).

[0008] Several aspects of this disclosure provide a method for identifying a suitable nuclease target site that is sufficiently distinct from any other site in the genome to achieve specific cleavage by a given nuclease without any or at least minimal off-target cleavage. Such a method is useful, for example, for identifying candidate nuclease target sites that can be cleaved with high specificity on a genomic background when selecting a target site for in vitro or in vivo genomic manipulation.

[0009] Several aspects of this disclosure provide methods for evaluating, selecting, and / or designing site-specific nucleases having improved specificity compared to current nucleases. For example, methods useful in selecting and / or designing site-specific nucleases having minimal off-target cleavage activity are provided herein, for example, by designing variant nucleases having binding domains with reduced binding affinity, by lowering the final concentration of the nuclease, by selecting target sites that are at least three base pairs apart from their nearest related sequences in the genome, and, in the case of RNA-programmable nucleases, by selecting guide RNAs that result in the minimum number of off-target sites binding and / or cleaving.

[0010] Compositions and kits useful for carrying out the methods described herein are also provided.

[0011] Several aspects of this disclosure provide methods for identifying target sites of nucleases. In some embodiments, the method includes (a) providing a nuclease that cleaves a double-stranded nucleic acid target site, wherein the cleavage of the target site results in a cleaved nucleic acid strand containing a 5' phosphate moiety; (b) contacting a library of candidate nucleic acid molecules containing the nuclease target site with the nuclease of (a) under conditions suitable for the nuclease to cleave the nuclease, wherein each nucleic acid molecule contains a concatemer of a sequence containing the candidate nuclease target site and a constant insertion sequence; and (c) identifying the nuclease target site cleaved by the nuclease in step (b) by determining the sequence of the uncleaved nuclease target site on the nucleic acid strand cleaved by the nuclease in step (b). In some embodiments, the nuclease produces a blunt end. In some embodiments, the nuclease produces a 5' overhang. In some embodiments, determining step (c) involves ligating the first nucleic acid adapter to the 5' end of the nucleic acid strand cleaved by the nuclease in step (b) by 5' phosphate-dependent ligation. In some embodiments, the nucleic acid adapter is provided in a double-stranded form. In some embodiments, the 5' phosphate-dependent ligation is blunt-end ligation. In some embodiments, the method involves filling the 5' overhang before ligating the first nucleic acid adapter to the nucleic acid strand cleaved by the nuclease. In some embodiments, determining step (c) further involves amplifying a fragment of the concatemer cleaved by the nuclease containing an uncleaved target site by a PCR reaction using PCR primers that hybridize to a constant insertion sequence and PCR primers that hybridize to the adapter. In some embodiments, the method further involves enriching the amplified nucleic acid molecule for a molecule containing a single uncleaved target sequence. In some embodiments, the enrichment step includes size fractionation.In some embodiments, the determination in step (c) includes sequencing the nucleic acid strand cleaved by the nuclease in step (b), or a copy thereof obtained by PCR. In some embodiments, the library of candidate nucleic acid molecules is at least 10. 8 , at least 10 9 , at least 10 10 , at least 10 11 , or at least 10 12The method includes several different candidate nuclease cleavage sites. In some embodiments, the nuclease is a therapeutic nuclease that cleaves a specific nuclease target site in a disease-associated gene. In some embodiments, the method further includes determining a maximum concentration of the therapeutic nuclease in which the therapeutic nuclease cleaves the specific nuclease target site and does not cleave more than 10 additional nuclease target sites, more than 5 additional nuclease target sites, more than 4 additional nuclease target sites, more than 3 additional nuclease target sites, more than 2 additional nuclease target sites, more than 1 additional nuclease target sites, or no additional nuclease target sites at all. In some embodiments, the method further includes administering the therapeutic nuclease to a subject in an amount effective to produce a final concentration that is lower than or equal to the maximum concentration. In some embodiments, the nuclease is an RNA-programmable nuclease that forms a complex with an RNA molecule, and the nuclease:RNA complex specifically binds to a nucleic acid sequence complementary to the sequence of the RNA molecule. In some embodiments, the RNA molecule is a single guide RNA (sgRNA). In some embodiments, the sgRNA contains 5-50 nucleotides, 10-30 nucleotides, 15-25 nucleotides, 18-22 nucleotides, 19-21 nucleotides, e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the nuclease is a Cas9 nuclease. In some embodiments, the nuclease target site contains a [sgRNA complementary sequence]-[protospacer adjacent motif (PAM)] structure, and the nuclease cleaves the target site within the sgRNA complementary sequence. In some embodiments, the sgRNA complementary sequence contains 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the nuclease contains a nonspecific nucleic acid cleavage domain.In some embodiments, the nuclease includes a FokI cleavage domain. In some embodiments, the nuclease includes a nucleic acid cleavage domain that cleaves a target sequence upon cleavage domain dimerization. In some embodiments, the nuclease includes a binding domain that specifically binds to a nucleic acid sequence. In some embodiments, the binding domain includes zinc fingers. In some embodiments, the binding domain includes at least two, at least three, at least four, or at least five zinc fingers. In some embodiments, the nuclease is a zinc finger nuclease. In some embodiments, the binding domain includes a transcription activator-like element. In some embodiments, the nuclease is a transcription activator-like element nuclease (TALEN). In some embodiments, the nuclease is an organic compound. In some embodiments, the nuclease includes an enediyne functional group. In some embodiments, the nuclease is an antibiotic. In some embodiments, the compound is dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin, or a derivative thereof. In some embodiments, the nuclease is a homing endonuclease.

[0012] Some aspects of the present disclosure provide a library of nucleic acid molecules, each nucleic acid molecule comprising a concatemer of sequences comprising a constant insertion sequence of 10 to 100 nucleotides and a candidate nuclease target site. In some embodiments, the constant insertion sequence is at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, or at least 95 nucleotides in length. In some embodiments, the constant insertion sequence is 15 or less, 20 or less, 25 or less, 30 or less, 35 or less, 40 or less, 45 or less, 50 or less, 55 or less, 60 or less, 65 or less, 70 or less, 75 or less, 80 or less, or 95 or less nucleotides in length. In some embodiments, the candidate nuclease target site is a site that can be cleaved by an RNA-programmable nuclease, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a homing endonuclease, an organic compound nuclease, or an enediyne antibiotic (e.g., dynemicin, neocarzinostatin, calicheamicin, esperamicin, bleomycin). In some embodiments, the candidate nuclease target site can be cleaved by Cas9 nuclease. In some embodiments, the library is at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , or at least 10 12The library contains several different candidate nuclease target sites. In some embodiments, the library contains nucleic acid molecules with molecular weights of at least 0.5 kDa, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 4 kDa, at least 5 kDa, at least 6 kDa, at least 7 kDa, at least 8 kDa, at least 9 kDa, at least 10 kDa, at least 12 kDa, or at least 15 kDa. In some embodiments, the library contains candidate nuclease target sites which are variations of known target sites of the nuclease of interest. In some embodiments, the variations of known nuclease target sites contain 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, or 2 or fewer mutations compared to the known nuclease target site. In some embodiments, the deformation differs from the known target site of the nuclease of interest by an average of more than 5%, more than 10%, more than 15%, more than 20%, more than 25%, or more than 30% (binomial distribution). In some embodiments, the deformation differs from the known target site by an average of less than 10%, less than or equal to 15%, less than or equal to 20%, less than or equal to 25%, less than or equal to 30%, less than or equal to 40%, or less than or equal to 50% (binomial distribution). In some embodiments, the nuclease of interest is a Cas9 nuclease, a zinc finger nuclease, a TALEN, a homing endonuclease, an organic compound nuclease, or an enediyne antibiotic (e.g., dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin). In some embodiments, the candidate nuclease target site is a Cas9 nuclease target site containing a [sgRNA complementary sequence]-[PAM] structure, where the sgRNA complementary sequence contains 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.

[0013] Several aspects of this disclosure provide a method for selecting a nuclease that specifically cleaves a consensus target site from a plurality of nucleases. In some embodiments, the method includes (a) providing a plurality of candidate nucleases that cleave the same consensus sequence, (b) for each of the candidate nucleases in step (a), identifying nuclease target sites that are cleaved by a candidate nuclease different from the consensus target site using the method provided herein, and (c) selecting a nuclease based on the nuclease target site(s) identified in step (b). In some embodiments, the nuclease selected in step (c) is the nuclease that cleaves the consensus target site with the highest specificity. In some embodiments, the nuclease that cleaves the consensus target site with the highest specificity is the candidate nuclease that cleaves the fewest number of target sites different from the consensus site. In some embodiments, a candidate nuclease that cleaves the consensus target site with the highest specificity is a candidate nuclease that cleaves the minimum number of target sites different from the consensus site in the context of the target genome. In some embodiments, the candidate nuclease selected in step (c) is a nuclease that does not cleave any target sites other than the consensus target site. In some embodiments, the candidate nuclease selected in step (c) is a nuclease that does not cleave any target sites other than the consensus target site in the genome of interest at a therapeutically effective concentration of the nuclease. In some embodiments, the method further comprises contacting the genome with the nuclease selected in step (c). In some embodiments, the genome is that of a vertebrate, mammal, human, non-human primate, rodent, mouse, rat, hamster, goat, sheep, cattle, dog, cat, reptile, amphibian, fish, nematode, insect, or fly. In some embodiments, the genome is within a living cell. In some embodiments, the genome is within the subject. In some embodiments, the consensus target site lies within an allele associated with the disease or disorder.In some embodiments, cleavage of the consensus target site results in the treatment or prevention of a disease or disorder (e.g., recovery from or prevention of at least one sign and / or symptom of the disease or disorder). In some embodiments, cleavage of the consensus target site results in the alleviation of signs and / or symptoms of the disease or disorder. In some embodiments, cleavage of the consensus target site results in the prevention of a disease or disorder. In some embodiments, the disease is HIV / AIDS. In some embodiments, the allele is the CCR5 allele. In some embodiments, the disease is a proliferative disorder. In some embodiments, the disease is cancer. In some embodiments, the allele is the VEGFA allele.

[0014] Several aspects of this disclosure provide isolated nucleases selected according to methods provided herein. In some embodiments, the nucleases are engineered to cleave target sites in the genome. In some embodiments, the nuclease is a Cas9 nuclease containing an sgRNA complementary to the target site in the genome. In some embodiments, the nuclease is a zinc finger nuclease (ZFN), or a transcription activator-like effector nuclease (TALEN), a homing endonuclease, or an organic compound nuclease (e.g., endiyne, antibiotic nuclease, dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin, or a derivative thereof). In some embodiments, the nuclease is selected based on the fact that it does not cleave other candidate target sites in addition to its known nuclease target sites, or that it cleaves one or fewer candidate target sites, two or fewer candidate target sites, three or fewer candidate target sites, four or fewer candidate target sites, five or fewer candidate target sites, six or fewer candidate target sites, seven or fewer candidate target sites, eight or fewer candidate target sites, eight or fewer candidate target sites, nine or fewer candidate target sites, or ten or fewer candidate target sites.

[0015] Some aspects of this disclosure provide kits comprising a library of nucleic acid molecules containing candidate nuclease target sites provided herein. Some aspects of this disclosure provide kits comprising isolated nucleases provided herein. In some embodiments, the nuclease is a Cas9 nuclease. In some embodiments, the kit further comprises a nucleic acid molecule containing the target site of the isolated nuclease. In some embodiments, the kit comprises an excipient and instructions for contacting the nuclease with the excipient to produce a composition suitable for contacting the nuclease with the nucleic acid. In some embodiments, the composition is suitable for contacting nucleic acids in a genome. In some embodiments, the composition is suitable for contacting nucleic acids in a cell. In some embodiments, the composition is suitable for contacting nucleic acids in a subject. In some embodiments, the excipient is a pharmaceutically acceptable excipient.

[0016] Several aspects of this disclosure provide pharmaceutical compositions suitable for administration to a subject. In some embodiments, the composition comprises an isolated nuclease provided herein. In some embodiments, the composition comprises a nucleic acid encoding such nuclease. In some embodiments, the composition comprises a pharmaceutically acceptable excipient.

[0017] Other advantages, features, and uses of the present invention will become apparent from the detailed description of certain non-limiting embodiments of the invention, the drawings (which are schematic and not intended to be drawn to any particular scale), and the claims. [Brief explanation of the drawing]

[0018] [Figure 1]Outline of in vitro selection. (A) Cas9, complexed with short guide RNA (sgRNA), recognizes approximately 20 bases of the target DNA substrate complementary to the sgRNA sequence and cleaves both DNA strands. White triangles indicate cleavage sites. (B) A modified version of the inventor's previously described in vitro selection was used to comprehensively profile Cas9 specificity. Concatemer pre-selection DNA libraries, each containing one of 10¹² distinct variants of the target DNA sequence (white rectangles), were generated from synthetic DNA oligonucleotides by ligation and rolling circle amplification. These libraries were incubated together with the Cas9:sgRNA complex of interest. The cleaved library members contain a 5' phosphate group (circle with "P") and are therefore substrates for adapter ligation and PCR. The resulting amplicons were subjected to high-throughput DNA sequencing and computer analysis.

[0019] [Figure 2]In vitro selection results for Cas9:CLTA1 sgRNA. Heatmap 21 shows the specificity profiles of Cas9:CLTA1 sgRNA v2.1 (A, B) under enzyme restriction conditions, Cas9:CLTA1 sgRNA v1.0 (C, D) under enzyme saturation conditions, and Cas9:CLTA1 sgRNA v2.1 (E, F) under enzyme saturation conditions. The heatmap shows all selected sequences (A, C, E) or sequences containing only single mutations in the target site and 2-base pair PAM defined by the 20-base pair sgRNA (B, D, F). Specificity scores of 1.0 and -1.0 correspond to positive or negative 100% enrichment of a specific base pair at a specific position, respectively. Black cells indicate the target nucleotide. (G) Effect of Cas9:sgRNA concentration on specificity. Changes in site specificity (normalized to the maximum possible change in site specificity) between enzyme restriction (200 nM DNA, 100 nM Cas9:sgRNA v2.1) and enzyme saturation (200 nM DNA, 1000 nM Cas9:sgRNA v2.1) conditions are shown for CLTA1. (H) Effect of sgRNA composition on specificity. Changes in site specificity (normalized to the maximum possible change in site specificity) between sgRNA v1.0 and sgRNA v2.1 under enzyme saturation conditions are shown for CLTA1. See Figures 6-8, 25, and 26 for corresponding data for CLTA2, CLTA3, and CLTA4. Sequence identifiers: The sgRNA sequences shown in (A-F) correspond to Sequence ID No. 1.

[0020] [Figure 3]Target sites profiled in this study. (A) The 5' end of the sgRNA has 20 nucleotides complementary to the target site. The target site contains an NGG motif (PAM) adjacent to the RNA:DNA complementarity region. (B) Four human clathrin gene (CLTA) target sites are shown. (C, D) Four human clathrin gene (CLTA) target sites are shown together with sgRNA. sgRNA v1.0 is shorter than sgRNA v2.1. PAMs are shown for each site. The non-PAM end of the target site corresponds to the 5' end of the sgRNA. Sequence identifiers: The sequences shown in (B), from top to bottom, are SEQ ID NOs. 2, 3, 4, 5, 6, and 7. The sequences shown in (C), from top to bottom, are SEQ ID NOs. 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, and 19. The sequences shown in (D), from top to bottom, are sequence numbers 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, and 31.

[0021] [Figure 4] In vitro Cas9:guide RNA cleavage of on-target DNA sequences. Individual DNA cleavage assays were performed on approximately 1kb linear substrates for each of the four CLTA target sites using 200nM on-target sites and 100nM Cas9:v1.0 sgRNA, 100nM Cas9:v2.1 sgRNA, 1000nM Cas9:v1.0 sgRNA, and 1000nM Cas9:v2.1 sgRNA. For CLTA1, CLTA2, and CLTA4, Cas9:v2.1 sgRNA showed higher activity than Cas9:v1.0 sgRNA. For CLTA3, the activity of Cas9:v1.0 sgRNA and Cas9:v2.1 sgRNA was equivalent.

[0022] [Figure 5] In vitro selection results for four target sites. In vitro selection was performed using 100 nM Cas9:sgRNA v2.1, 1000 nM Cas9:sgRNA v1.0, or 1000 nM Cas9:sgRNA v2.1 in a 200 nM pre-selection library. (A) Post-selection PCR products are shown for the 12 selections performed. DNA containing 1.5 repeats was quantified for each selection and pooled in equimolar amounts before gel purification and sequencing. (B-E) Variant distributions are shown for the pre-selection (black) and post-selection (colored) libraries. The post-selection library is enriched with sequences containing fewer mutations than the pre-selection library. Mutations are counted from 20-base pair and 2-base pair PAMs defined by sgRNA. P-values ​​are <0.01 for all pairwise comparisons between the distributions of each panel. P-values ​​were calculated using t-tests, assuming unequal size and unequal variances.

[0023] [Figure 6] In vitro selection results for Cas9:CLTA2 sgRNA. Heatmap 24 shows the specificity profiles of Cas9:CLTA2 sgRNA v2.1 under enzyme restriction conditions (A, B), Cas9:CLTA2 sgRNA v1.0 under enzyme overload conditions (C, D), and Cas9:CLTA2 sgRNA v2.1 under enzyme overload conditions (E, F). The heatmap shows all selected sequences (A, C, E) or sequences containing only single mutations in the target site and 2-base pair PAM defined by the 20-base pair sgRNA (B, D, F). Specificity scores of 1.0 and -1.0 correspond to positive or negative 100% enrichment of a specific base pair at a specific position, respectively. Black cells indicate the target nucleotide. Sequence identifiers: The sgRNA sequences shown in (A-F) correspond to Sequence ID No. 32.

[0024] [Figure 7]In vitro selection results for Cas9:CLTA3 sgRNA. Heatmap 24 shows the specificity profiles of Cas9:CLTA3 sgRNA v2.1 under enzyme restriction conditions (A, B), Cas9:CLTA3 sgRNA v1.0 under enzyme excess conditions (C, D), and Cas9:CLTA3 sgRNA v2.1 under enzyme saturation conditions (E, F). The heatmap shows all selected sequences (A, C, E) or sequences containing only single mutations in the target site and 2-base pair PAM defined by the 20-base pair sgRNA (B, D, F). Specificity scores of 1.0 and -1.0 correspond to positive or negative 100% enrichment of a specific base pair at a specific position, respectively. Black cells indicate the target nucleotide. Sequence identifiers: The sgRNA sequences shown in (A-F) correspond to Sequence ID No. 32.

[0025] [Figure 8] In vitro selection results for Cas9:CLTA4 sgRNA. Heatmap 24 shows the specificity profiles of Cas9:CLTA4 sgRNA v2.1 under enzyme restriction conditions (A, B), Cas9:CLTA4 sgRNA v1.0 under enzyme excess conditions (C, D), and Cas9:CLTA4 sgRNA v2.1 under enzyme saturation conditions (E, F). The heatmap shows all selected sequences (A, C, E) or sequences containing only single mutations in the target site and 2-base pair PAM defined by the 20-base pair sgRNA (B, D, F). Specificity scores of 1.0 and -1.0 correspond to positive or negative 100% enrichment of a specific base pair at a specific location, respectively. Black cells indicate the target nucleotide. Sequence identifiers: The sgRNA sequences shown in (A-F) correspond to Sequence ID No. 33.

[0026] [Figure 9]In vitro selection results as sequence logos. The amount of information is plotted for each target site position (1-20) defined by CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) sgRNA v2.1 under enzyme restriction conditions.25 Positions in PAM are labeled "P1", "P2", and "P3". The amount of information is plotted in bits. 2.0 bits indicates absolute specificity, and 0 bits indicates no specificity.

[0027] [Figure 10] The tolerance of distal mutations in the PAM of CLTA1. The highest specificity score at each position is shown for Cas9:CLTA1 v2.1 sgRNA selection. Here, only sequences with on-target base pairs (gray) are considered, while mutations in the first 1-12 base pairs are allowed (A-L). Positions not bound to on-target base pairs are shown by dark bars. Higher specificity scores indicate higher specificity at a given position. Positions that could not contain any mutations (gray) are plotted with a specificity score of +1. For all panels, specificity scores were calculated from pre-selection and post-selection library sequences with n≧5,130 and n≧74,538, respectively.

[0028] [Figure 11] The tolerance of distal mutations in the PAM of CLTA2. The highest specificity score at each position is shown for Cas9:CLTA2 v2.1 sgRNA selection. Here, only sequences with on-target base pairs (gray) are considered, while mutations in the first 1-12 base pairs are allowed (A-L). Positions not bound to on-target base pairs are shown by dark bars. Higher specificity scores indicate higher specificity at a given position. Positions that could not contain any mutations (gray) are plotted with a specificity score of +1. For all panels, specificity scores were calculated from pre-selection and post-selection library sequences with n≧3,190 and n≧25,365, respectively.

[0029] [Figure 12] The tolerance of distal mutations in the PAM of CLTA3. The highest specificity score at each position is shown for Cas9:CLTA3 v2.1 sgRNA selection. Here, only sequences with on-target base pairs (gray) are considered, while mutations in the first 1-12 base pairs are allowed (A-L). Positions not bound to on-target base pairs are shown by dark bars. Higher specificity scores indicate higher specificity at a given position. Positions that could not contain any mutations (gray) are plotted with a specificity score of +1. For all panels, specificity scores were calculated from pre-selection and post-selection library sequences with n≧5,604 and n≧158,424, respectively.

[0030] [Figure 13] The tolerance of distal mutations in the PAM of CLTA4. The highest specificity score at each position is shown for Cas9:CLTA4 v2.1 sgRNA selection. Here, only sequences with on-target base pairs (gray) are considered, while mutations in the first 1-12 base pairs are allowed (A-L). Positions not bound to on-target base pairs are shown by dark bars. Higher specificity scores indicate higher specificity at a given position. Positions that could not contain any mutations (gray) are plotted with a specificity score of +1. For all panels, specificity scores were calculated from pre-selection and post-selection library sequences with n≧2,323 and n≧21,819, respectively.

[0031] [Figure 14]Tolerance of distal mutations in the PAM of the CLTA1 target site. The mutation distribution is shown for in vitro selection of a 200 nM pre-selection library with 1000 nM Cas9:CLTA1 sgRNA v2.1. The number of mutations shown is in the target site subsequences furthest from the PAM (A-L), where the remaining target site containing the PAM contains only on-target base pairs. The pre- and post-selection distributions are similar up to 3 base pairs, demonstrating tolerance for target sites with mutations in the 3 base pairs furthest from the PAM. In this case, the remaining target site has optimal interaction with Cas9:sgRNA. For all panels, graphs were generated from pre- and post-selection library sequences with n≧5,130 and n≧74,538, respectively.

[0032] [Figure 15] Tolerance of distal mutations in the PAM of the CLTA2 target site. The mutation distribution is shown for in vitro selection of a 200 nM pre-selection library with 1000 nM Cas9:CLTA2 sgRNA v2.1. The number of mutations shown is in the target site subsequences furthest from the PAM (A-L), where the remaining target site containing the PAM contains only on-target base pairs. The pre- and post-selection distributions are similar up to 3 base pairs, demonstrating tolerance for target sites with mutations in the 3 base pairs furthest from the PAM. In this case, the remaining target site has optimal interaction with Cas9:sgRNA. For all panels, graphs were generated from pre- and post-selection library sequences with n≧3,190 and n≧21,265, respectively.

[0033] [Figure 16]Tolerance of distal mutations in the PAM of the CLTA3 target site. The mutation distribution is shown for in vitro selection of a 200 nM pre-selection library with 1000 nM Cas9:CLTA3 sgRNA v2.1. The number of mutations shown is in the target site subsequence of 1-12 base pairs furthest from the PAM (A-L), in which case the remainder of the target site encompassing the PAM contains only on-target base pairs. The pre- and post-selection distributions are similar up to 3 base pairs, demonstrating tolerance for target sites with mutations in the 3 base pairs furthest from the PAM. In this case, the remainder of the target site has optimal interaction with Cas9:sgRNA. For all panels, graphs were generated from pre- and post-selection library sequences with n≧5,604 and n≧158,424, respectively.

[0034] [Figure 17] Tolerance of distal mutations in the PAM of the CLTA4 target site. The mutation distribution is shown for in vitro selection of a 200 nM pre-selection library with 1000 nM Cas9:CLTA4 sgRNA v2.1. The number of mutations shown is in the target site subsequences furthest from the PAM (A-L), where the remaining target site containing the PAM contains only on-target base pairs. The pre- and post-selection distributions are similar up to 3 base pairs, demonstrating tolerance for target sites with mutations in the 3 base pairs furthest from the PAM. In this case, the remaining target site has optimal interaction with Cas9:sgRNA. For all panels, graphs were generated from pre- and post-selection library sequences with n≧2,323 and n≧21,819, respectively.

[0035] [Figure 18]The site-specificity pattern of 100 nM Cas9:sgRNA v2.1. Site-specificity (defined as the sum of the specificity scores for each of the four possible base pairs recognized at a given site within the target site) is plotted for each target site under enzyme restriction conditions for sgRNA v2.1. Site-specificity is shown as a value normalized to the maximum site-specificity value of the target site. Site-specificity is highest at the terminal of the target site proximal to the PAM and lowest in the middle of the target site and at the few nucleotides most distal to the PAM.

[0036] [Figure 19] The site-specificity pattern of 1000 nM Cas9:sgRNA v1.0. Site-specificity (defined as the sum of the specificity scores for each of the four possible base pairs recognized at a given site within the target site) is plotted for each target site under enzyme-over-enzyme conditions of sgRNA v1.0. Site-specificity is shown as a value normalized to the maximum site-specificity value of the target site. Site-specificity is relatively constant throughout the target site, but is lowest in the middle of the target site and at the few nucleotides most distal to the PAM.

[0037] [Figure 20] Site-specificity patterns of 1000 nM Cas9:sgRNA v2.1. Site-specificity (defined as the sum of the specificity scores for each of the four possible base pairs recognized at a given position within the target site) is plotted for each target site under enzyme-over-enzyme conditions of sgRNA v2.1. Site-specificity is shown as a value normalized to the maximum site-specificity value of the target site. Site-specificity is relatively constant throughout the target site, but is lowest in the middle of the target site and at the few nucleotides most distal to the PAM.

[0038] [Figure 21]PAM nucleotide preference. Abundance ratios in pre- and post-selection libraries under enzyme restriction or excess conditions are shown for all 16 possible PAM dinucleotides for selection by CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) sgRNA v2.1. GG dinucleotides showed increased abundance in the post-selection library, while other possible PAM dinucleotides showed decreased abundance after selection.

[0039] [Figure 22] PAM nucleotide preference at the on-target site. Only post-selection library members that did not contain mutations within the 20 base pairs defined by the guide RNA were included in this analysis. Abundance ratios in the pre-selection and post-selection libraries under enzyme restriction and enzyme overload conditions are shown for all 16 possible PAM dinucleotides for selection by CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) sgRNA v2.1. GG dinucleotides showed increased abundance in the post-selection library, while other possible PAM dinucleotides generally decreased abundance after selection; however, for enzyme overload concentrations of Cas9:sgRNA, this effect was negligible or absent for many dinucleotides.

[0040] [Figure 23]PAM dinucleotide specificity scores. Specificity scores under enzyme restriction and enzyme excess conditions are shown for all 16 possible PAM dinucleotides (positions 2 and 3 of the 3-nucleotide NGG PAM) selected by CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) sgRNA v2.1. The specificity scores indicate the enrichment of PAM dinucleotides in the pre-selection library and the post-selection library relative to each other, and are normalized to the maximum possible enrichment of the dinucleotide. A specificity score of +1.0 indicates that the dinucleotide is 100% enriched in the post-selection library, and a specificity score of -1.0 indicates that the dinucleotide is 100% de-enriched. GG dinucleotides are the most enriched in the post-selection library, while AG, GA, GC, GT, and TG show less relative de-enrichment compared to other possible PAM dinucleotides.

[0041] [Figure 24]PAM dinucleotide specificity scores at the on-target site. Only post-selection library members that did not contain mutations within the 20 base pairs defined by the guide RNA were included in this analysis. Specificity scores under enzyme restriction and enzyme over-excess conditions are shown for all 16 possible PAM dinucleotides (positions 2 and 3 of the 3-nucleotide NGG PAM) for selection by CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) sgRNA v2.1. The specificity scores indicate the enrichment of PAM dinucleotides in the post-selection library relative to the pre-selection library and are normalized to the maximum possible enrichment of the dinucleotide. A specificity score of +1.0 indicates 100% enrichment of the dinucleotide in the post-selection library, and a specificity score of -1.0 indicates 100% de-enrichment of the dinucleotide. GG dinucleotides were the most enriched in the post-selection library, AG and GA nucleotides were neither enriched nor deenriched under at least one selection criterion, and GC, GT, and TG showed less relative deenrichment compared to other possible PAM dinucleotides.

[0042] [Figure 25] The effect of Cas9:sgRNA concentration on specificity. Changes in site specificity between enzyme restriction (200 nM DNA, 100 nM Cas9:sgRNA v2.1) and enzyme overload (200 nM DNA, 1000 nM Cas9:sgRNA v2.1) conditions are shown for sgRNA selection targeting CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) target sites. Lines indicate the maximum possible change in site specificity for a given site. The greatest change in specificity occurs proximal to the PAM as enzyme concentration increases.

[0043] [Figure 26]The effect of sgRNA composition on specificity. Changes in site specificity between Cas9:sgRNA v1.0 and Cas9:sgRNA v2.1 under enzyme overload conditions (200 nM DNA, 1000 nM Cas9:sgRNA v2.1) are shown for sgRNA selection targeting CLTA1(A), CLTA2(B), CLTA3(C), and CLTA4(D) target sites. Lines indicate the maximum possible change in site specificity for a given site.

[0044] [Figure 27] Cas9:guide RNA cleavage of off-target DNA sequences in vitro. Individual DNA cleavage assays on a 96bp linear substrate were performed with 200 nM DNA and 1000 nM Cas9:CLTA4 v2.1 sgRNA for the on-target CLTA4 site (CLTA4-0) and five CLTA4 off-target sites identified by in vitro selection. The enrichment values ​​shown are from in vitro selection with 1000 nM Cas9:CLTA4 v2.1 sgRNA. CLTA4-1 and CLTA4-3 were the most highly enriched sequences under those conditions. CLTA4-2a, CLTA4-2b, and CLTA4-2c are two variant sequences representing enrichment values ​​ranging from highly enriched to no enrichment to highly deenriched. Lowercase letters indicate mutations relative to the on-target CLTA4 site. The enrichment values ​​qualitatively agree with the observed amount of cleavage in vitro. Sequence identifiers: The sequences shown, from top to bottom, are sequence number 34, sequence number 35, sequence number 36, sequence number 37, sequence number 38, and sequence number 39.

[0045] [Figure 28]The effect of guide RNA composition and Cas9:sgRNA concentration on in vitro cleavage of off-target sites. Individual DNA cleavage assays were performed on a 96bp linear substrate for the CLTA4-3 off-target site (5'GggGATGTAGTGTTTCCACtGGG 3' (SEQ ID NO: 39); mutations are shown in lowercase) with 200nM DNA and 100nM Cas9:v1.0 sgRNA, 100nM Cas9:v2.1 sgRNA, 1000nM Cas9:v1.0 sgRNA, or 1000nM Cas9:v2.1 sgRNA. DNA cleavage was observed under all four conditions tested, and the cleavage rate was higher under enzyme-over-enzyme conditions or with v2.1 sgRNA compared to v1.0 sgRNA.

[0046] definition As used herein and in the claims, the singular forms "a," "an," and "the" encompass both singular and plural references unless the context otherwise clearly indicates. Thus, for example, the reference "drug" encompasses both a single drug and multiple such drugs.

[0047] The term "Cas9" or "Cas9 nuclease" refers to an RNA guide nuclease containing the Cas9 protein or a fragment thereof. Cas9 nucleases are sometimes also called cason1 nucleases or CRISPR (clustered regularly interspaced short palindromic repeat)-associated nucleases. CRISPR is an adaptive immune system that provides defense against mobile genetic elements (e.g., viruses, transposable elements, and conjugative plasmids). CRISPR clusters contain spacers, sequences complementary to the preceding mobile element, and target the invading nucleic acid. CRISPR clusters are transcribed and processed into CRISPR-RNA (crRNA). In the type II CRISPR system, the correct processing of pre-crRNA requires tracrRNA (trans-encoded small RNA), endogenous ribonuclease 3 (rnc), and the Cas9 protein. tracrRNA acts as a guide for the processing of pre-crRNA, which is facilitated by ribonuclease 3. Subsequently, Cas9 / crRNA / tracrRNA endonucleolytically cleaves a linear or circular dsDNA target complementary to the spacer. Target strands not complementary to crRNA are first endonucleolytically cleaved and then exonucleolytically trimmed from 3' to 5'. In nature, DNA binding and cleavage typically require proteins and both RNA species. However, single guide RNAs ("sgRNA" or simply "gNRA") can be manipulated to incorporate aspects of both crRNA and tracrRNA into a single RNA molecule. See, for example, Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). Its complete content is incorporated here by reference. Cas9 recognizes short motifs (PAM or protospacer adjacency motifs) within CRISPR repeat sequences, helping to distinguish between self and non-self.The sequence and structure of Cas9 nuclease are well known to those skilled in the art (e.g., "Complete genome sequence of an M1 strain of Streptococcus pyogenes." Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G., Lyon K., Primeaux C., Sezate S., Suvorov AN, Kenton S., Lai HS, Lin SP, Qian Y., Jia HG, Najar FZ, Ren Q., Zhu H., Song L. expand / collapse author list McLaughlin RE, Proc. Natl. Acad. Sci. USA 98:4658-4663(2001), "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III." Deltcheva E., Chylinski K., Sharma CM, Gonzales K., Chao Y., Pirzada ZA, Eckert MR, See Vogel J., Charpentier E., Nature 471:602-607 (2011), and “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821 (2012). (The overall content is incorporated herein by reference). Cas9 orthologs have been described in various species (including, but not limited to, S. pyogenes and S. thermophilus). Additional preferred Cas9 nucleases and sequences will be apparent to those skilled in the art based on this disclosure.Such Cas9 nucleases and sequences include Cas9 sequences from organisms and loci disclosed in Chylinski, Rhun, and Charpentier, “The tracrRNA and Cas9 families of type II CRISPR-Cas immune systems” (2013) RNA Biology 10:5, 726-737 (the entire content of which is incorporated herein by reference). In some embodiments, proteins containing Cas9 or a fragment of Cas9 protein are referred to as “Cas9 variants.” Cas9 variants share homology with Cas9 or its fragments. For example, a Cas9 variant is at least about 70% identical, at least about 80% identical, at least about 90% identical, at least about 95% identical, at least about 98% identical, at least about 99% identical, at least about 99.5% identical, or at least about 99.9% identical to wild-type Cas9. In some embodiments, Cas9 variants contain a fragment of Cas9 (e.g., a gRNA-binding domain or a DNA-cleaving domain) as a result, the fragment is at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, at least about 99.5%, or at least about 99.9% identical to the corresponding fragment of wild-type Cas9. In some embodiments, wild-type Cas9 corresponds to Cas9 from Streptococcus pyogenes (NCBI reference sequence: NC_017053.1, SEQ ID NO: 40 (nucleotide), SEQ ID NO: 41 (amino acid)).

[0048] The term "concatemer," as used herein in the context of nucleic acid molecules, refers to a nucleic acid molecule containing multiple copies of the same DNA sequence linked together in a series. For example, 10 copies of a specific sequence of nucleotides (e.g., [XYZ]). 10 A concatemer containing a nuclease target site and a constant insertion sequence will contain 10 copies of the same specific sequence linked together in a series, e.g., 5'-XYZXYZXYZXYZXYZXYZXYZXYZXYZXYZ-3'. A concatemer may contain a copy number of either repeat units or sequences (e.g., at least 2 copies, at least 3 copies, at least 4 copies, at least 5 copies, at least 10 copies, at least 100 copies, at least 1000 copies, etc.). An example of a concatemer of a nucleic acid sequence containing a nuclease target site and a constant insertion sequence is [(target site)-(constant insertion sequence)]. 300 Concatemers can be linear nucleic acid molecules or cyclic molecules.

[0049] The terms “conjugation,” “conjugated,” and “conjugation” refer to the association of two entities (e.g., two molecules, e.g., two proteins, two domains (e.g., a binding domain and a cleavage domain)), or a protein and a drug, e.g., a protein-binding domain and a small molecule. In some aspects, the association is between a protein (e.g., an RNA-programmable nuclease) and a nucleic acid (e.g., guide RNA). The association can be achieved, for example, by covalent binding (e.g., by a linker) or by non-covalent interactions, direct or indirect. In some embodiments, the association is covalent. In some embodiments, two molecules are conjugated by a linker that connects both molecules. For example, in some embodiments, where two proteins (e.g., the binding domain and cleavage domain of an engineered nuclease) are conjugated to each other to form a protein fusion, the two proteins may be conjugated by a polypeptide linker (e.g., an amino acid sequence that connects the C-terminus of one protein to the N-terminus of another protein).

[0050] When used herein in the context of nucleic acid sequences, the term “consensus sequence” refers to a calculated sequence that represents the most frequently found nucleotide residue at each position among several similar sequences. Typically, the consensus sequence is determined by sequence alignment, where similar sequences are compared to one another to calculate similar sequence motifs. In the context of nuclease target site sequences, the consensus sequence of a nuclease target site may, in some embodiments, be the sequence that binds most frequently or with the highest affinity by a given nuclease. With respect to target site sequences of RNA-programmable nucleases (e.g., Cas9), the consensus sequence may, in some embodiments, be a sequence or region that is designed or expected to bind to a gRNA or multiple gRNAs (e.g., based on complementary base pairing).

[0051] The term “effective dose,” as used herein, may mean the amount of a bioactive agent sufficient to elicit a desired biological response. For example, in some embodiments, the effective dose of a nuclease means the amount of nuclease sufficient to induce cleavage of a target site to which the nuclease specifically binds and cleaves. As will be recognized by those skilled in the art, the effective dose of an agent (e.g., a nuclease, hybrid protein, or polynucleotide) can vary depending on various factors (e.g., the desired biological response, the specific allele being targeted, the genome, the target site, cell, or tissue, and the agent to be used).

[0052] The term "enediyne," as used herein, refers to a class of bacterial natural products characterized by either a 9- or 10-membered ring containing two triple bonds separated by a double bond (see, for example, KC Nicolaou; AL Smith; EW Yue (1993). "Chemistry and biology of natural and designed enediynes." PNAS 90 (13): 5881–5888, the entire content of which is incorporated herein by reference). Some engineers can undergo Bergmann cyclization, and the resulting diradical (1,4-dehydrobenzene derivative) can extract a hydrogen atom from the sugar backbone of DNA, resulting in DNA strand breaks (see, e.g., S. Walker; R. Landovitz; WD Ding; GA Ellestad; D. Kahne (1992). "Cleavage behavior of calicheamicin gamma 1 and calicheamicin T". Proc Natl Acad Sci USA 89 (10): 4608-12. Its overall content is incorporated herein by reference). Their reactivity with DNA confers antibiotic character to many engineers, and some engineers have been clinically studied as anticancer antibiotics. Examples of enediyne antibiotics that are not limited to enediyne include dynemycin, neocardinostatin, calichemycin, and esperamicin (see, for example, Adrian L. Smith and KC Bicolaou, "The Enediyne Antibiotics," J. Med. Chem., 1996, 39 (11), pp 2103-2117 and Donald Borders, "Enediyne antibiotics as antitumor agents," Informa Healthcare; 1 st (See edition November 23, 1994, ISBN-10: 0824789385. Its overall content is incorporated herein by reference.)

[0053] The term "homing endonuclease," as used herein, refers to a type of restriction enzyme typically encoded by an intron or intein. Edgell DR (February 2009). "Selfish DNA: homing endonucleases find a home." Curr Biol 19 (3): R115-R117; Jasin M (Jun 1996). "Genetic manipulation of genomes with rare-cutting endonucleases." Trends Genet 12 (6): 224-8; Burt A, Koufopanou V (December 2004). "Homing endonuclease genes: the rise and fall and rise again of a selfish element." Curr Opin Genet Dev 14 (6): 609-15 (its entire content is incorporated herein by reference). Homing endonuclease recognition sequences are found with a very low probability (approximately 7 × 10⁻⁶). 10 They are long enough to occur randomly only once per bp, and usually only one is found per genome.

[0054] The term “library,” as used herein in the context of nucleic acids or proteins, means a collection of two or more different nucleic acids or proteins, respectively. For example, a library of nuclease target sites comprises at least two nucleic acid molecules containing different nuclease target sites. In some embodiments, the library comprises at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11, at least 10 12 , at least 10 13 , at least 10 14 , or at least 10 15 The library contains individual different nucleic acids or proteins. In some embodiments, members of the library may include randomized sequences, e.g., completely or partially randomized sequences. In some embodiments, the library contains nucleic acid molecules that are unrelated to each other (e.g., nucleic acids containing completely randomized sequences). In other embodiments, at least some members of the library may be related, e.g., they may be variants or derivatives of a particular sequence, such as a consensus target site sequence.

[0055] The term "linker," as used herein, refers to a chemical group or molecule that links two adjacent molecules or parts (e.g., the binding and cleaving domains of a nuclease). Typically, a linker is located between or adjacent to two groups, molecules, or other parts, and is linked to one another by a covalent bond, thus linking the two. In some embodiments, the linker is an amino acid or a group of amino acids (e.g., a peptide or protein). In some embodiments, the linker is an organic molecule, group, polymer, or chemical part.

[0056] The term “nuclease,” as used herein, refers to an agent (e.g., a protein or small molecule) capable of cleaving phosphodiester bonds that connect nucleotide residues in a nucleic acid molecule. In some embodiments, a nuclease is a protein (e.g., an enzyme that can bind to a nucleic acid molecule and cleave phosphodiester bonds that connect nucleotide residues in the nucleic acid molecule). A nuclease may be an endonuclease that cleaves phosphodiester bonds within a polynucleotide chain, or an exonuclease that cleaves phosphodiester bonds at the ends of a polynucleotide chain. In some embodiments, a nuclease is a site-specific nuclease that binds and / or cleaves specific phosphodiester bonds in a specific nucleotide sequence, which is also referred herein to as the “recognition sequence,” “nuclease target site,” or “target site.” In some embodiments, a nuclease is an RNA-guided (i.e., RNA-programmable) nuclease that complexes (e.g., binds) to an RNA having a sequence complementary to the target site, thereby providing the sequence specificity of the nuclease. In some embodiments, nucleases recognize single-stranded target sites, while in other embodiments, nucleases recognize double-stranded target sites (e.g., double-stranded DNA target sites). The target sites of many naturally occurring nucleases (e.g., many naturally occurring DNA restriction nucleases) are well known to those skilled in the art. In many cases, DNA nucleases (e.g., EcoRI, HindIII, or BamHI) recognize palindromic double-stranded DNA target sites of 4–10 base pairs in length and cleave each of the two DNA strands at specific locations within the target site. Some endonucleases symmetrically cleave double-stranded nucleic acid target sites, i.e., cleave both strands at the same location, resulting in ends containing base-paired nucleotides (also referred to herein as blunt ends). Other endonucleases asymmetrically cleave double-stranded nucleic acid target sites, i.e., cleave each strand at different locations, resulting in ends containing unpaired nucleotides.Unpaired nucleotides at the ends of a double-stranded DNA molecule are also called "overhangs" (e.g., "5' overhangs" or "3' overhangs," depending on whether the unpaired nucleotide(s) or other nucleotides form the 5' or 3' end of each DNA strand). Ends of a double-stranded DNA molecule ending with unpaired nucleotide(s) or other nucleotide(s) are also called sticky ends because they can "stick" to the ends of other double-stranded DNA molecules containing complementary unpaired nucleotide(s) or other nucleotides. Nuclease proteins typically contain a "binding domain" (which mediates the interaction of the protein with the nucleic acid substrate and, in some cases, also specifically binds to a target site) and a "cleavage domain" (which catalyzes the cleavage of phosphodiester bonds in the nucleic acid backbone). In some embodiments, nuclease proteins can bind to and cleave nucleic acid molecules in monomeric form, while in other embodiments, nuclease proteins must dimerize or polymerize to cleave target nucleic acid molecules. Naturally occurring nuclease binding and cleavage domains, as well as modular binding and cleavage domains that can be fused to produce nucleases that bind to specific target sites, are well known to those skilled in the art. For example, a zinc finger or transcription activator-like element may be used as a binding domain to specifically bind to a desired target site and fused or conjugated to a cleavage domain (e.g., the cleavage domain of FokI) to produce an engineered nuclease that cleaves the target site.

[0057] The terms “nucleic acid” and “nucleic acid molecule,” as used herein, refer to compounds comprising nucleic acid bases and acidic moieties, such as nucleosides, nucleotides, or polymers of nucleotides. Typically, polymeric nucleic acids, such as nucleic acid molecules containing three or four or more nucleotides, are linear molecules in which adjacent nucleotides are linked to one another by phosphodiester bonds. In some embodiments, “nucleic acid” refers to nucleic acid residues of an individual (e.g., nucleotides and / or nucleosides). In some embodiments, “nucleic acid” refers to oligonucleotide chains containing three or four or more nucleotide residues of an individual. As used herein, the terms “oligonucleotide” and “polynucleotide” may be used interchangeably to refer to polymers of nucleotides (e.g., strings of at least three nucleotides). In some embodiments, “nucleic acid” encompasses RNA, and even single- and / or double-stranded DNA. Nucleic acids may exist naturally in the context of genomes, transcripts, mRNA, tRNA, rRNA, siRNA, snRNA, plasmids, cosmids, chromosomes, chromatids, or other naturally occurring nucleic acid molecules. On the other hand, nucleic acid molecules may be molecules that do not exist in nature, such as recombinant DNA or RNA, artificial chromosomes, engineered genomes or fragments thereof, or synthetic DNA, RNA, DNA / RNA hybrids, or may include nucleotides or nucleosides that do not exist in nature. Furthermore, the terms “nucleic acid,” “DNA,” “RNA,” and / or similar terms include nucleic acid analogs, i.e., analogs having something other than a phosphodiester backbone. Nucleic acids may be purified from natural sources, produced using recombinant expression systems and optionally purified, or chemically synthesized. Where appropriate, for example in the case of chemically synthesized molecules, nucleic acids may include nucleoside analogs (e.g., analogs having chemically modified bases or sugars) and backbone modifications. Nucleic acid sequences are shown in the 5' to 3' direction unless otherwise indicated.In some aspects, nucleic acids include natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyladenosine, 5-methylcytidine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2 -aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), chemically modified bases, biologically modified bases (e.g., methylated bases), intercalation bases, modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose), and / or modified phosphate groups (e.g., phosphorothioates and 5'-N-phosphoamidite bonds), or comprising them.

[0058] The term “pharmaceutical composition,” as used herein, means a composition that can be administered to a subject in the context of treating a disease or disorder. In some embodiments, a pharmaceutical composition comprises an active ingredient (e.g., a nuclease or a nucleic acid encoding a nuclease) and pharmaceutically acceptable excipients.

[0059] The term "proliferative disorder," as used herein, refers to any disorder in which the homeostasis of a cell or tissue is disrupted by an abnormally high rate of proliferation of cells or cell populations. Proliferative disorders include hyperproliferative disorders such as preneoplastic hyperplasia and neoplastic disorders. Neoplastic disorders are characterized by abnormal cell proliferation and include both benign and malignant neoplasms. Malignant neoplasms are also known as cancer.

[0060] The terms “protein,” “peptide,” and “polypeptide” are used interchangeably herein and refer to polymers of amino acid residues linked together by peptide (amide) bonds. The terms refer to proteins, peptides, or polypeptides of any size, structure, or function. Typically, a protein, peptide, or polypeptide will be at least three amino acid long. A protein, peptide, or polypeptide can refer to an individual protein or a collection of proteins. One or more amino acids in a protein, peptide, or polypeptide can be modified, for example, by the addition of chemical entities (e.g., carbohydrate groups, hydroxyl groups, phosphate groups, farnesyl groups, isofarnesyl groups, fatty acid groups, linkers, etc., for conjugation, functionalization, or other modification). A protein, peptide, or polypeptide can be a single molecule or a multimolecular complex. A protein, peptide, or polypeptide can be simply a fragment of a naturally occurring protein or peptide. A protein, peptide, or polypeptide can be naturally occurring, recombinant, synthetic, or a combination of any of these. A protein may contain different domains (e.g., nucleic acid binding domains and nucleic acid cleavage domains). In some embodiments, a protein comprises a proteinaceous part (e.g., an amino acid sequence that forms a nucleic acid binding domain) and an organic compound (e.g., a compound that can act as a nucleic acid cleavage agent). In some embodiments, the protein is complexed with or associated with a nucleic acid (e.g., RNA).

[0061] The term "randomized," as used herein in the context of nucleic acid sequences, refers to a sequence or residue in a sequence that has been synthesized to incorporate a mixture of free nucleotides (e.g., a mixture of all four nucleotides A, T, G, and C). Randomized residues are typically represented by the letter N in nucleotide sequences. In some embodiments, the randomized sequence or residue is fully randomized, in which case the randomized residue is synthesized by adding equal amounts of the nucleotides to be incorporated (e.g., 25% T, 25% A, 25% G, and 25% C) during the synthesis step of each sequence residue. In some embodiments, the randomized sequence or residue is partially randomized, in which case the randomized residue is synthesized by adding unequal amounts of the nucleotides to be incorporated (e.g., 79% T, 7% A, 7% G, and 7% C) during the synthesis step of each sequence residue. Partial randomization allows for the generation of sequences that use a given sequence as a template but have incorporated mutations at a desired frequency. For example, if a known nuclease target site is used as a synthesis template, partial randomization in which the nucleotide represented by each residue is added to the synthesis with 79% probability at each step, and the other three nucleotides are added with 7% probability each, will result in the synthesis of a mixture of partially randomized target sites. This still represents the consensus sequence of the original target site, but each residue thus synthesized differs from the original target site at a statistical frequency of 21% (binomial distribution). In some embodiments, the partially randomized sequence differs from the consensus sequence by an average of more than 5%, more than 10%, more than 15%, more than 20%, more than 25%, or more than 30% (binomial distribution). In some embodiments, the partially randomized sequence differs from the consensus site by an average of less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, or less than 50% (binomial distribution).

[0062] The terms “RNA-programmable nuclease” and “RNA-guided nuclease” are used interchangeably herein and refer to a nuclease that forms a complex with (e.g., binds to or associates with) one or more RNAs that are not the target of cleavage. In some embodiments, an RNA-programmable nuclease may be referred to as a nuclease:RNA complex when it is a complex with RNA. Typically, the bound RNA(s) are referred to as guide RNA (gRNA) or single guide RNA (sgRNA). The gRNA / sgRNA contains a nucleotide sequence complementary to the target site, which mediates the binding of the nuclease / RNA complex to that target site and provides sequence specificity for the nuclease:RNA complex.Additionally, RNA-induced RNA-induced polymerization (CRIS PR vaccine) Cas9 strain strain Streptococcus Ferretti JJ, McShan WM, Ajdic DJ, Savic DJ, Savic G, Lyon K, Primeaux C, Sezate S, Complete genome sequence of an M1 strain of Streptococcus pyogenes. Suvorov AN, Kenton S, Lai HS, Qian Y, Jia HG, Najar FZ, Ren Q, Zhu L. expand / collapse author list McLaughlin RE, Proc Natl small RNA and host factor RNase III.Deltcheva E, Chylinski K, Sharma CM, Gonzales K, Chao Y, Pirzada ZA, Eckert MR, Vogel J, Charpentier E, Nature 471 :602–607(2011). endonuclease in adaptive bacterial immunity.” Jinek M., Chylinski K., Fonfara I., Hauer M., Doudna JA, Charpentier E. Science 337:816-821(2012) and some of these articles have been published in a suitable manner.

[0063] RNA-programmable nucleases (e.g., Cas9) use RNA:DNA hybridization to determine the target DNA cleavage site; therefore, in principle, these proteins have the ability to cleave any sequence defined by the guide RNA. Methods using RNA-programmable nucleases such as Cas9 for site-specific cleavage (for example, to modify the genome) are well known in this field (e.g., Cong, L. et al. Multiplex genome engineering using CRISPR / Cas systems. Science 339, 819-823 (2013); Mali, P. et al. RNA-guided human genome engineering via Cas9. Science 339, 823-826 (2013); Hwang, WY et al. Efficient genome editing in zebrafish using a CRISPR-Cas system. Nature biotechnology 31, 227-229 (2013); Jinek, M. et al. RNA-programmed genome editing in human cells. eLife 2, e00471 (2013); Dicarlo, JE et al. Genome engineering in Saccharomyces cerevisiae using CRISPR-Cas systems. Nucleic acids research (2013), See Jiang, W. et al. RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Nature biotechnology 31, 233-239 (2013). Its entire content is incorporated herein by reference.

[0064] The terms “small molecule” and “organic compound” are used interchangeably herein and refer to molecules having a relatively low molecular weight, whether naturally occurring or artificially produced (e.g., by chemical synthesis). Typically, organic compounds contain carbon. Organic compounds may contain multiple carbon-carbon bonds, stereocenters, and other functional groups (e.g., amines, hydroxyls, carbonyls, or heterocyclic rings). In some embodiments, organic compounds are monomers with a molecular weight less than about 1500 g / mol. In some embodiments, the molecular weight of a small molecule is less than about 1000 g / mol or less than about 500 g / mol. In some embodiments, a small molecule is a drug (e.g., a drug already deemed safe and effective for use in humans or animals by the appropriate government agency or regulatory body). In some embodiments, organic molecules are known to bind to and / or cleave nucleic acids. In some embodiments, organic compounds are enediynes. In some embodiments, the organic compound is an antibiotic drug, such as an anticancer antibiotic, such as dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin, or a derivative thereof.

[0065] The term "subject," as used herein, means an individual organism, such as an individual mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal. In some embodiments, the subject is a non-human primate. In some embodiments, the subject is a rodent. In some embodiments, the subject is a sheep, goat, cow, cat, or dog. In some embodiments, the subject is a vertebrate, amphibian, reptile, fish, insect, fly, or nematode.

[0066] When used herein in the context of nucleases, the terms “target nucleic acid” and “target genome” refer, respectively, to a nucleic acid molecule or genome containing at least one target site of a given nuclease.

[0067] The term “target site” is used herein interchangeably with the term “nuclease target site” and refers to a sequence within a nucleic acid molecule that is bound and cleaved by a nuclease. The target site may be single-stranded or double-stranded. In the context of dimerizing nucleases (e.g., nucleases containing a FokI DNA cleavage domain), the target site typically comprises a left half-site (bound by one monomer of the nuclease), a right half-site (bound by a second monomer of the nuclease), and a spacer sequence between the halves where cleavage occurs. This structure ([left half-site]-[spacer sequence]-[right half-site]) is referred herein to as the LSR structure. In some embodiments, the left half-site and / or the right half-site are 10–18 nucleotides long. In some embodiments, either or both halves are shorter or longer. In some embodiments, the left and right halves contain different nucleic acid sequences. In the context of zinc finger nucleases, the target site may, in some embodiments, consist of two half-sites, each 6–18 bp long, adjacent to an undefined spacer region 4–8 bp long. In the context of TALENs, the target site may, in some embodiments, consist of two half-sites, each 10–23 bp long, adjacent to an undefined spacer region 10–30 bp long. In the context of RNA-guided (e.g., RNA-programmable) nucleases, the target site typically consists of a nucleotide sequence complementary to the sgRNA of the RNA-programmable nuclease and a 3'-terminal protospacer-adjacent motif (PAM) adjacent to the sgRNA complementary sequence. For the RNA-guided nuclease Cas9, the target site may, in some embodiments, be a 20-base pair + 3-base pair PAM (e.g., NNN, where N represents any nucleotide). Typically, the first nucleotide of a PAM can be any nucleotide, while the two downstream nucleotides are determined depending on a specific RNA guide nuclease.Exemplary target sites for RNA guide nucleases such as Cas9 are known to those skilled in the art and include, without limitation, NNG, NGN, NAG, and NGG, where N represents any of the nucleotides. In addition, Cas9 nucleases from different species (e.g., S. thermophilus instead of S. pyogenes) recognize PAMs containing the sequence NGGNG. Additional PAM sequences are known and include, but are not limited to, NNAGAAW and NAAR (see, for example, Esvelt and Wang, Molecular Systems Biology, 9:641 (2013), the entire content of which is incorporated herein by reference). For example, target sites for RNA guide nucleases (e.g., Cas9) are structured [N. Z It may contain ]-[PAM], where each N is independently any of the nucleotides. Z is an integer between 1 and 50. In some embodiments, Z is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50. In some embodiments, Z Z is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50. In some embodiments, Z is 20.

[0068] The term "transmission activator-like effector" (TALE), as used herein, refers to a bacterial protein containing a DNA-binding domain, which contains a highly conserved 33-34 amino acid sequence including a highly variable two-amino acid motif (repeat variable duo, RVD).The RVD motif determines its binding specificity to nucleic acid sequences and can be manipulated according to methods well known to those skilled in the art to specifically bind to desired DNA sequences (e.g., Miller, Jeffrey; et.al. (February 2011). "A TALE nuclease architecture for efficient genome editing". Nature Biotechnology 29 (2): 143-8; Zhang, Feng; et.al. (February 2011). "Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription". Nature Biotechnology 29 (2): 149-53; GeiBler, R.; Scholze, H.; Hahn, S.; Streubel, J.; Bonas, U.; Behrens, SE; Boch, J. (2011), Shiu, Shin-Han. ed. "Transcriptional Activators of Human Genes with Programmable DNA-Specificity". PLoS ONE 6 (5): See el9509, Boch, Jens (February 2011). "TALEs of genome targeting." Nature Biotechnology 29 (2): 135-6, Boch, Jens; et.al. (December 2009). "Breaking the Code of DNA Binding Specificity of TAL-Type III Effectors." Science 326 (5959): 1509-12, and Moscou, Matthew J.; Adam J. Bogdanove (December 2009). "A Simple Cipher Governs DNA Recognition by TAL Effectors." Science 326 (5959): 1501. (Their respective overall contents are incorporated herein by reference).The simple relationship between amino acid sequences and DNA recognition allowed for the manipulation of specific DNA-binding domains by selecting combinations of repeat segments containing appropriate RVDs.

[0069] The term "transmission activator-like element nuclease" (TALEN), as used herein, refers to an artificial nuclease containing a transcription activator-like effector DNA binding domain to a DNA cleavage domain (e.g., a FokI domain). Several modular assembly schemes have been reported for generating manipulated TAL constructs (e.g., Zhang, Feng; et.al. (February 2011). "Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription." Nature Biotechnology 29 (2): 149-53; GeiBler, R.; Scholze, H.; Hahn, S.; Streubel, J.; Bonas, U.; Behrens, SE; Boch, J. (2011), Shiu, Shin-Han. ed. "Transcriptional Activators of Human Genes with Programmable DNA-Specificity." PLoS ONE 6 (5): e19509; Cermak, T.; Doyle, EL; Christian, M.; Wang, L.; Zhang, Y.; Schmidt, C.; Baller, JA; Somia, NV et al. (2011)). “Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting”. Nucleic Acids Research, Morbitzer, R.; Elsaesser, J.; Hausner, J.; Lahaye, T. (2011). “Assembly of custom TALE-type DNA binding domains by modular cloning”. Nucleic Acids Research, Li, T.; Huang, S.; Zhao, X.; Wright, D.A.; Carpenter, S.; Spalding, MH; Weeks, DP; Yang, B. (2011). "Modularly assembled designer TAL effector nucleases for targeted gene knockout and gene replacement in eukaryotes." Nucleic Acids Research. See Weber, E.; Gruetzner, R.; Werner, S.; Engler, C; Marillonnet, S. (2011), Bendahmane, Mohammed, ed. "Assembly of Designer TAL Effectors by Golden Gate Cloning." PLoS ONE 6 (5): e19722. The complete contents of each are incorporated herein by reference).

[0070] The terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms described herein. As used herein, the terms “treatment,” “to treat,” and “to treat” refer to a clinical intervention aimed at reversing, alleviating, delaying the onset of, or inhibiting the progression of a disease or disorder or one or more symptoms described herein. In some embodiments, a treatment may be administered after the onset of one or more symptoms and / or after the disease has been diagnosed. In other embodiments, a treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms or to inhibit the onset or progression of the disease. For example, a treatment may be administered to a susceptible individual prior to the onset of symptoms (for example, in light of a history of symptoms and / or in light of genetic or other susceptibility factors). A treatment may also be continued after the symptoms have subsided, for example, to prevent or delay their recurrence.

[0071] The term "zinc finger," as used herein, refers to a small nucleic acid-binding protein structural motif characterized by a fold and the coordination of one or more zinc ions that stabilize the fold. Zinc fingers encompass a wide variety of different protein structures (see, e.g., Klug A, Rhodes D (1987). "Zinc fingers: a novel protein fold for nucleic acid recognition." Cold Spring Harb. Symp. Quant. Biol. 52: 473-82; its overall content is incorporated herein by reference). Zinc fingers may be designed to bind to specific sequences of nucleotides, and zinc finger arrays containing fusions of a series of zinc fingers may be designed to bind substantially to any desired target sequence. Such zinc finger arrays may form binding domains on proteins (e.g., nucleases when conjugated to nucleic acid cleavage domains). Different types of zinc finger motifs are known to those skilled in the art and include, but are not limited to, Cys2His2, Gag knuckle, Treble clef, zinc ribbon, Zn2 / Cys6, and TAZ2 domain-like motifs (see, for example, Krishna SS, Majumdar I, Grishin NV (January 2003). "Structural classification of zinc fingers: survey and summary." Nucleic Acids Res. 31 (2): 532-50). Typically, a single zinc finger motif binds to 3 or 4 nucleotides of a nucleic acid molecule. Thus, a zinc finger domain containing two zinc finger motifs can bind to 6-8 nucleotides, a zinc finger domain containing three zinc finger motifs can bind to 9-12 nucleotides, a zinc finger domain containing four zinc finger motifs can bind to 12-16 nucleotides, and so on.Any suitable protein manipulation technique can be used to modify the DNA binding specificity of zinc fingers and / or design novel zinc finger fusions that can bind to substantially any desired target sequence of length 3–30 nucleotides (see, for example, Pabo CO, Peisach E, Grant RA (2001). "Design and selection of novel cys2His2 Zinc finger proteins." Annual Review of Biochemistry 70: 313–340; Jamieson AC, Miller JC, Pabo CO (2003). "Drug discovery with engineered zinc-finger proteins." Nature Reviews Drug Discovery 2 (5): 361–368; and Liu Q, Segal DJ, Ghiara JB, Barbas CF (May 1997). "Design of polydactyl zinc-finger proteins for unique addressing within complex genomes." Proc. Natl. Acad. Sci. USA 94 (11). The overall contents of each are incorporated herein by reference). A fusion between an engineered zinc finger array and a protein domain that cleaves nucleic acids can be used to generate a "zinc finger nuclease." A zinc finger nuclease typically contains a zinc finger domain that binds to a specific target site within a nucleic acid molecule, and a nucleic acid cleavage domain that cleaves the nucleic acid molecule within or near the target site bound by the binding domain. A typical engineered zinc finger nuclease contains a binding domain with 3–6 individual zinc finger motifs that binds to a target site spanning 9–18 base pairs in length. Longer target sites are particularly attractive in situations where binding to and cleaving a unique target site within a given genome is desired.

[0072] The term "zinc finger nuclease," as used herein, refers to a nuclease comprising a nucleic acid cleavage domain conjugated to a binding domain containing a zinc finger array. In some embodiments, the cleavage domain is the cleavage domain of the type II restriction endonuclease FokI. Zinc finger nucleases may be designed to target substantially any desired sequence in a given nucleic acid molecule for cleavage, and the possibility of designing the zinc finger binding domain to bind to a unique site in the context of a complex genome allows for targeted cleavage of a single genomic site in a living cell, for example, achieving targeted genomic modifications with therapeutic value. Targeting double-strand breaks to desired genomic loci may be used to introduce frameshift mutations into the coding sequence of a gene due to the erroneous nature of non-homologous DNA repair pathways. Zinc finger nucleases may be generated to target a desired site by methods well known to those skilled in the art. For example, a zinc finger binding domain with desired specificity may be designed by combining zinc finger motifs of individuals with known specificity. The structure of the DNA-bound zinc finger protein Zif268 has provided information for a significant portion of research in this field. The idea of ​​obtaining a zinc finger for each of the 64 possible base pair triplets, and then mixing and matching these modular zinc fingers to design a protein with any desired sequence specificity, has been described (Pavletich NP, Pabo CO (May 1991). "Zinc finger-DNA recognition: crystal structure of a Zif268-DNA complex at 2.1 A". Science 252 (5007): 809-17. The entire content is incorporated herein). In some embodiments, separate zinc fingers, each recognizing a 3-base pair DNA sequence, are combined to generate 3, 4, 5, or 6-finger arrays that recognize target sites spanning 9 to 18 base pairs in length. In some embodiments, longer arrays are conceivable.In other embodiments, two-finger modules that recognize 6-8 nucleotides are combined to generate a 4, 6, or 8 zinc finger array. In some embodiments, a bacterial or phage display is used to generate zinc finger domains that recognize a desired nucleic acid sequence (e.g., a desired nuclease target site of length 3-30 bp). In some embodiments, the zinc finger nuclease includes a zinc finger binding domain and a cleavage domain fused to or otherwise conjugated together by a linker (e.g., a polypeptide linker). The length of the linker determines the distance of cleavage from the nucleic acid sequence bound by the zinc finger domain. If a shorter linker is used, the cleavage domain will cleave nucleic acids closer to the bound nucleic acid sequence. On the other hand, a longer linker will result in a greater distance between the bound and cleaved nucleic acid sequences. In some embodiments, the cleavage domain of the zinc finger nuclease must dimerize to cleave the bound nucleic acid. In some such embodiments, the dimer is a heterodimer of two monomers, each containing a different zinc finger binding domain. For example, in some embodiments, the dimer may comprise one monomer containing a zinc finger domain A conjugated to a FokI cleavage domain, and another monomer containing a zinc finger domain B conjugated to a FokI cleavage domain. In this non-limiting example, zinc finger domain A binds to the nucleic acid sequence on one side of the target site, zinc finger domain B binds to the nucleic acid sequence on the other side of the target site, and the dimerized FokI domain cleaves the nucleic acid between the zinc finger domain binding sites.

[0073] Detailed description of a certain aspect of the present invention Introduction Site-specific nucleases are powerful tools for targeted genome modification in vitro or in vivo. Some site-specific nucleases can theoretically achieve levels of specificity to target cleavage sites, allowing for the targeting of a single unique site in the genome for cleavage without affecting any other genomic regions. It has been reported that nuclease cleavage in living cells triggers DNA repair mechanisms, which frequently lead to modification of the cleaved and repaired genomic sequence (e.g., by homologous recombination). Therefore, targeted cleavage of specific, unique sequences in the genome opens new avenues for gene targeting and modification in living cells (including cells that are difficult to manipulate by conventional gene targeting methods, such as many human somatic or embryonic stem cells). Nuclease-mediated modifications of disease-related sequences (e.g., the CCR-5 allele in HIV / AIDS patients, or genes necessary for tumor angiogenesis) can be used in a clinical context, and two site-specific nucleases are currently in clinical trials.

[0074] One important aspect in the field of site-directed nuclease-mediated modifications is off-target nuclease effects (e.g., cleavage of genomic sequences that differ by one or more nucleotides from the intended target sequence). Unwanted side effects of off-target cleavage range from insertions into unwanted loci during gene targeting events to severe complications in clinical scenarios. Off-target cleavage of sequences encoding essential gene functions or tumor suppressor genes by an endonuclease administered to a subject can lead to disease or even death in the subject. Therefore, it is desirable to characterize the cleavage preference of a nuclease before using it in the laboratory or hospital to determine its efficacy and safety. Furthermore, characterization of nuclease cleavage properties allows for the selection of the nuclease best suited to a specific task from a group of candidate nucleases, or the selection of an evolutionary product derived from multiple nucleases. Such characterization of nuclease cleavage properties can also inform the de novo design of nucleases with improved properties (e.g., improved specificity or efficiency).

[0075] In many scenarios where nucleases are used for targeted manipulation of nucleic acids, cleavage specificity is a critical feature. Imperfect specificity of some manipulated nuclease-binding domains can lead to off-target cleavage and undesirable effects both in vitro and in vivo. Current methods for evaluating site-specific nuclease specificity (including ELISA assays, microarrays, one-hybrid systems, SELEX and its variants, as well as computer predictions based on Rosetta) all assume that the binding specificity of nucleases is equivalent to or proportional to their cleavage specificity.

[0076] It has been previously discovered that predicting the off-target binding effects of nucleases is an imperfect approximation of the off-target cleavage effects of nucleases, which can result in undesirable biological effects (see PCT application WO2013 / 066438 and Pattanayak, V., Ramirez, CL, Joung, JK & Liu, DR Revealing off-target cleavage specificities of zinc-finger nucleases by in vitro selection. Nature Methods 8, 765-770 (2011). The full contents of each are incorporated herein by reference). This finding is consistent with the view that the reported toxicity of some site-specific DNA nucleases stems not from off-target binding alone, but from off-target DNA cleavage.

[0077] The methods and reagents of this disclosure represent improvements over previous methods in several respects, enabling precise evaluation of the target site specificity of a given nuclease and providing strategies for selecting suitable unique target sites and designing or selecting highly specific nucleases for targeted cleavage of single sites in the context of complex genomes. For example, some previously reported methods for determining the target site specificity profile of a nuclease by screening a library of nucleic acid molecules containing candidate target sites rely on a “double cleavage” in vitro selection method, which requires the indirect reconstruction of the target site from the sequences of two half-sites resulting from two adjacent cleavages of the nuclease in the library member nucleic acid (see, e.g., Pattanayak, V. et al., Nature Methods 8, 765-770 (2011)). In contrast to such a “double cleavage” strategy, the methods of this disclosure utilize a “single cleavage” screening strategy, which enables the identification of library members that have been cleaved at least once by the nuclease. The “single-cut” selection strategies provided herein are compatible with single-end high-throughput sequencing methods and do not require computer reconstruction of the cleaved target site from the cleaved half-site, because, in some embodiments, they feature direct sequencing of the intact target nuclease sequence in the cleaved library member nucleic acid.

[0078] In addition, the “single-cleavage” screening method of this disclosure utilizes a concatemer of a candidate nuclease target site and a constant insertion region, which is approximately 10 times shorter than previously reported constructs used in two-cleavage strategies (approximately 50 bp repeat sequence length vs. approximately 500 bp repeat sequence length in previously reported constructs). This difference in repeat sequence length in the concatemer of the library is a factor in highly complex libraries of candidate nuclease target sites (e.g., 10 12This enables the generation of libraries containing several different candidate nuclease target sequences. As described herein, an exemplary library of such complexity was generated using a known Cas9 nuclease target site as a template by altering the sequence of the known target site. The exemplary library demonstrated that coverage greater than 10 times greater than all sequences with eight or seven or fewer mutations of the known target site can be achieved using the strategies provided herein. The use of shorter repeat sequences also enables the use of single-end sequencing, because both the cleaved half-site and the adjacent uncleaved site of the same library member are contained within a 100-nucleotide sequencing read.

[0079] The strategies, methods, libraries, and reagents provided herein may be used to analyze the sequence preference and specificity of any site-specific nuclease (e.g., zinc finger nucleases (ZFNs), activator-like effector nucleases (TALENs), homing endonucleases, organic compound nucleases, and enediine antibiotics (e.g., dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin)). In addition to those described herein, suitable nucleases will be apparent to those skilled in the art based on this disclosure.

[0080] Furthermore, the methods, reagents, and strategies provided herein enable those skilled in the art to identify, design, and / or select nucleases with improved specificity and minimize off-target effects of any given nuclease (e.g., site-specific nucleases such as ZFNs, and TALENs that produce cleavage products with sticky ends, and even RNA-programmable nucleases such as Cas9, which produces cleavage products with blunt ends). While particularly relating to DNA and DNA cleavage nucleases, the concepts, methods, strategies, and reagents provided herein are not limited in this respect and may be applied to any nucleic acid:nuclease pair.

[0081] Identification of nuclease target sites cleaved by site-specific nucleases. Several aspects of this disclosure provide improved methods and reagents for determining nucleic acid target sites cleaved by any site-specific nuclease. The methods provided herein can be used to evaluate the target site preference and specificity of both blunt-end and sticky-end nucleases. Generally, such methods involve contacting a library of target sites with a given nuclease under conditions favorable for the nuclease to bind to and cleave the target sites, and determining which target sites the nuclease actually cleaves. Determining the target site profile of a nuclease based on actual cleavage has advantages over binding-dependent methods in that it measures parameters involved by mediating undesirable off-target effects of site-specific nucleases. Generally, the methods provided herein involve ligating an adapter of a known sequence to the nucleic acid molecule cleaved by the nuclease of interest by 5'-phosphate-dependent ligation. Therefore, the methods provided herein are particularly useful for identifying target sites that are cleaved by nucleases that leave a phosphate moiety at the 5' end of the cleaved nucleic acid strand when cleaving those target sites. After ligating the adapter to the 5' end of the cleaved nucleic acid strand, the cleaved strand can be directly sequenced using the adapter as a sequencing linker. Alternatively, a portion of the cleaved library member concatemer containing the same intact target site as the cleaved target site can be amplified by PCR, and the amplified product can then be sequenced.

[0082] In some embodiments, the method includes (a) providing a nuclease that cleaves a double-stranded nucleic acid target site, wherein cleavage of the target site results in a cleaved nucleic acid strand containing a 5' phosphate moiety; (b) contacting a library of candidate nucleic acid molecules containing the target site of the nuclease with the nuclease of (a) under conditions suitable for the nuclease to cleave the nuclease, wherein each nucleic acid molecule contains a concatemer of a sequence containing the candidate nuclease target site and a constant insertion sequence; and (c) identifying the nuclease target site to be cleaved by the nuclease in (b) by determining the sequence of the uncleaved nuclease target site on the nucleic acid strand cleaved by the nuclease in step (b).

[0083] In some embodiments, the method comprises providing a nuclease and contacting the nuclease with a library of candidate nucleic acid molecules containing a candidate target site. In some embodiments, the candidate nucleic acid molecules are double-stranded nucleic acid molecules. In some embodiments, the candidate nucleic acid molecules are DNA molecules. In some embodiments, each nucleic acid molecule in the library contains a concatemer of a sequence containing a candidate nuclease target site and a constant insertion sequence. For example, in some embodiments, the library has the structure R1-[(candidate nuclease target site)-(constant insertion sequence)] n -The nucleic acid molecule comprises R2, where R1 and R2 are nucleic acid sequences that independently may contain fragments of the structure [(candidate nuclease target site)-(constant insertion sequence)], and n is an integer from 2 to y. In some embodiments, y is at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 1013 , at least 10 14 , or at least 10 15 In some embodiments, y is 10 2 A little more than 10 3 A little more than 10 4 A little more than 10 5 A little more than 10 6 A little more than 10 7 A little more than 10 8 A little more than 10 9 A little more than 10 10 A little more than 10 11 A little more than 10 12 A little more than 10 13 A little more than 10 14 A little less than, or 10 15 It is slightly less than that.

[0084] For example, in some embodiments, candidate nucleic acid molecules in a library have a structure [(N Z The candidate nuclease target site is R1-(PAM), and therefore the nucleic acid molecules in the library have the structure R1-[(N Z )-(PAM)-(steady-state region)] X -Includes R2, and R1 and R2 are independently [(N Z A nucleic acid sequence that may contain fragments of a repeat unit ()-(PAM)-(steady region), where each N independently represents any nucleotide, Z is an integer from 1 to 50, and X is an integer from 2 to y. In some embodiments, y is at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 , or at least 1015 is. In some embodiments, y is 10 2 a little less than, 10 3 a little less than, 10 4 a little less than, 10 5 a little less than, 10 6 a little less than, 10 7 a little less than, 10 8 a little less than, 10 9 a little less than, 10 10 a little less than, 10 11 a little less than, 10 12 a little less than, 10 13 a little less than, 10 14 a little less than, or 10 15 a little less than. In some embodiments, Z is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50. In some embodiments, Z is 20. Each N represents any nucleotide. Thus, Z N having = 2 Z The array provided as will be NN, and each N will independently represent A, T, G, or C. Thus, Z N having = 2 Z can represent AA, AT, AG, AC, TA, TT, TG, TC, GA, GT, GG, GC, CA, CT, CG, and CC.

[0085] In other embodiments, the candidate nucleic acid molecules of the library include a candidate nuclease target site of the structure [left half site]-[spacer sequence]-[right half site] ("LSR"), and thus, the nucleic acid molecules of the library have the structure R1-[(LSR)-(constant region)] X-R2 is included, where R1 and R2 independently may contain fragments of the [(LSR)-(steady region)] repeat unit, and X is an integer between 2 and y. In some embodiments, y is at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 , or at least 10 15 In some embodiments, y is 10 2 A little more than 10 3 A little more than 10 4 A little more than 10 5 A little more than 10 6 A little more than 10 7 A little more than 10 8 A little more than 10 9 A little more than 10 10 A little more than 10 11 A little more than 10 12 A little more than 10 13 A little more than 10 14 A little less than, or 10 15It is slightly less than that. The constant region is, in some embodiments, a length that allows for efficient self-ligation of a single repeat unit. Preferred lengths will be obvious to those skilled in the art. For example, in some embodiments, the constant region is 5 to 100 base pairs long, e.g., about 5 base pairs, about 10 base pairs, about 15 base pairs, about 20 base pairs, about 25 base pairs, about 30 base pairs, about 35 base pairs, about 40 base pairs, about 50 base pairs, about 60 base pairs, about 70 base pairs, about 80 base pairs, about 90 base pairs, or about 100 base pairs long. In some embodiments, the constant region is 16 base pairs long. In some embodiments, the nuclease cleaves the double-stranded nucleic acid target site, producing a blunt end. In other embodiments, the nuclease produces a 5' overhang. In some such embodiments, the target site includes a [left half-site]-[spacer sequence]-[right half-site] (LSR) structure, and the nuclease cleaves the target site within the spacer sequence.

[0086] In some embodiments, the nuclease cleaves a double-stranded target site, producing a blunt end. In some embodiments, the nuclease cleaves a double-stranded target site, producing an overhang or sticky end (e.g., a 5' overhang). In some such embodiments, the method comprises filling the 5' overhang of a nucleic acid molecule produced from a nucleic acid molecule cleaved once by a nuclease, wherein the nucleic acid molecule comprises a constant insertion sequence flanked by a left or right half-site and a cleaved spacer sequence on one side and an uncleaved target site sequence on the other side, thereby producing a blunt end.

[0087] In some embodiments, determining step (c) involves ligating the first nucleic acid adapter to the 5' end of the nucleic acid strand cleaved by the nuclease in step (b) by 5' phosphate-dependent ligation. In some embodiments, the nuclease produces a blunt end. In such embodiments, the adapter can be directly ligated to the blunt end resulting from the nuclease cleavage at the target site by contacting the cleaved library member with a double-stranded blunt-end adapter lacking 5' phosphorylation. In some embodiments, the nuclease produces an overhang (sticky end). In some such embodiments, the adapter can be ligated to the cleaved site by contacting the cleaved library member with an excess of adapter having a suitable sticky end. When a nuclease that cleaves within a constant spacer sequence between variable half-sites is used, the sticky end can be designed to match the 5' overhang resulting from the spacer sequence. In embodiments in which the nuclease cleaves within a variable sequence, a population of adapters having a variable overhang sequence and a constant annealed sequence (for use as a sequencing linker or PCR primer) may be used, or the 5' overhang may be filled in to form a blunt end before adapter ligation.

[0088] In some embodiments, determining step (c) further includes amplifying the fragment of the concatemer cleaved by a nuclease containing the uncleaved target site by PCR using PCR primers that hybridize to an adapter and PCR primers that hybridize to a constant insert sequence. Typically, amplification of the concatemer by PCR will produce an amplicon containing at least one intact candidate target site identical to the cleaved target site, because the target sites in each concatemer are identical. For unidirectional sequencing, enrichment of amplicons containing one intact target site, two or fewer intact target sites, three or fewer intact target sites, four or fewer intact target sites, or five or fewer intact target sites may be desirable. In embodiments in which PCR is used to amplify nucleic acid molecules cleaved by PCR, the PCR parameters may be optimized to prefer amplification of shorter sequences and not longer sequences, for example by using short extension times in the PCR cycle. Another possibility for concentrating short amplicons is size fractionation, for example by gel electrophoresis or size exclusion chromatography. Size fractionation may be performed before and / or after amplification. Other suitable methods for concentrating short amplicons will be obvious to those skilled in the art. This disclosure is not limited in this respect.

[0089] In some embodiments, determining in step (c) involves sequencing the nucleic acid strand cleaved by the nuclease in step (b) or a copy thereof obtained by amplification (e.g., by PCR). Sequencing methods are well known to those skilled in the art. This disclosure is not limited in this respect.

[0090] In some embodiments, the nuclease to be profiled using the present invention system is an RNA-programmable nuclease, which forms a complex with an RNA molecule, and the nuclease:RNA complex specifically binds to a nucleic acid sequence complementary to the sequence of the RNA molecule. In some embodiments, the RNA molecule is a single guide RNA (sgRNA). In some embodiments, the sgRNA contains 5-50 nucleotides, 10-30 nucleotides, 15-25 nucleotides, 18-22 nucleotides, 19-21 nucleotides, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. In some embodiments, the sgRNA contains 5–50 nucleotides, 10–30 nucleotides, 15–25 nucleotides, 18–22 nucleotides, 19–21 nucleotides, e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides complementary to the nuclease target site. In some embodiments, the sgRNA contains 20 nucleotides complementary to the nuclease target site. In some embodiments, the nuclease is a Cas9 nuclease. In some embodiments, the nuclease target site contains a [sgRNA complementary sequence]-[protospacer adjacent motif (PAM)] structure, and the nuclease cleaves the target site within the sgRNA complementary sequence. In some embodiments, the sgRNA complementary sequence contains 5–50 nucleotides, 10–30 nucleotides, 15–25 nucleotides, 18–22 nucleotides, 19–21 nucleotides, for example, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides.

[0091] In some embodiments, the RNA-programmable nuclease is a Cas9 nuclease. The RNA-programmable Cas9 endonuclease cleaves double-stranded DNA (dsDNA) at a site adjacent to a two-base pair PAM motif and complementary to a guide RNA sequence (sgRNA). Typically, the sgRNA sequence complementary to the target site is about 20 nucleotides long, but shorter and longer complementary sgRNA sequences can also be used. For example, in some embodiments, the sgRNA may contain 5–50 nucleotides, 10–30 nucleotides, 15–25 nucleotides, 18–22 nucleotides, 19–21 nucleotides, e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides. The Cas9 system has been used to modify the genomes of multiple cell types, demonstrating its potential as an easy-to-use genome manipulation tool.

[0092] In some embodiments, the nuclease includes a nonspecific nucleic acid cleavage domain. In some embodiments, the nuclease includes a FokI cleavage domain. In some embodiments, the nuclease includes a nucleic acid cleavage domain that cleaves a target sequence upon cleavage domain dimerization. In some embodiments, the nuclease includes a binding domain that specifically binds to a nucleic acid sequence. In some embodiments, the binding domain includes zinc fingers. In some embodiments, the binding domain includes at least two, at least three, at least four, or at least five zinc fingers. In some embodiments, the nuclease is a zinc finger nuclease. In some embodiments, the binding domain includes a transcription activator-like element. In some embodiments, the nuclease is a transcription activator-like element nuclease (TALEN). In some embodiments, the nuclease is a homing endonuclease. In some embodiments, the nuclease is an organic compound. In some embodiments, the nuclease includes an enediyne functional group. In some embodiments, the nuclease is an antibiotic. In some aspects, the compound is dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin, or a derivative thereof.

[0093] Incubation of a nuclease together with a library nucleic acid will result in the cleavage of concatemers in the library that contain target sites that can be bound and cleaved by the nuclease. If a given nuclease cleaves a specific target site with high efficiency, the concatemer containing the target site will be cleaved, for example, once or multiple times, resulting in the generation of fragments containing cleaved target sites adjacent to one or more repeat units. Depending on the structure of the library member, an exemplary cleaved nucleic acid molecule released from a library member concatemer by a single nuclease cleavage is, for example, structure (cleaved target site)-(constant region)-[(target site)-(constant region)]. X-R2 is possible. For example, in the context of RNA guide nucleases, an exemplary cleaved nucleic acid molecule released from a library member concatemer by a single nuclease cleavage is, for example, structure (PAM)-(constant region)-[(N Z )-(PAM)-(steady-state region)] X -R2 is possible. In the context of a nuclease that cleaves the LSR structure within the spacer region, an exemplary cleaved nucleic acid molecule released from the library member concatemer by a single nuclease cleavage is, for example, structure (cleaved spacer region)-(right half region)-(constant region)-[(LSR)-(constant region)]. X -R2 is possible. Such cleavage fragments released from the library candidate molecule can be isolated therefrom, and / or the sequence of the target site cleaved by the nuclease can be identified by sequencing the intact target site (e.g., the intact (N) of the released repeat unit). Z )-(PAM) site. See Figure 1B for an example.

[0094] Preferred conditions for exposure of a library of nucleic acid molecules will be apparent to those skilled in the art. In some embodiments, preferred conditions do not result in denaturation of the library nucleic acids or nucleases, and allow the nucleases to exhibit at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98% of their nuclease activity.

[0095] In addition, if a given nuclease cleaves a specific target site, some cleavage products will include a cleaved half-site and an intact or uncleaved target site. As described herein, such products can be isolated by conventional methods, and since the insertion sequence is slightly less than 100 base pairs in some aspects, such isolated cleavage products can be sequenced in a single read-through, allowing for identification of the target site sequence without, for example, reconstructing the sequence from the cleaved half-site.

[0096] Any suitable method for isolating and sequencing repeat units can be used to elucidate the LSR sequences cleaved by the nuclease. For example, since the length of the constant region is known, free repeat units of an organism can be separated based on their size from larger, uncleaved library nucleic acid molecules, and even from fragments of library nucleic acid molecules containing multiple repeat units (indicating inefficiently targeted cleavage by the nuclease). Suitable methods for separating and / or isolating nucleic acid molecules based on their size are well known to those skilled in the art and include, for example, size fractionation methods (e.g., gel electrophoresis, density gradient centrifugation, and dialysis through a semipermeable membrane with a suitable molecular cutoff value). The separated / isolated nucleic acid molecules can then be further characterized, for example, by ligating PCR and / or sequencing adapters to the cleaved ends and amplifying and / or sequencing each nucleic acid. Furthermore, if the length of the steady-state region is selected to favor self-ligation of the free repeat units of the organism, such free repeat units can be enriched by contacting a library molecule treated with a ligase and a nuclease, and subsequent amplification and / or sequencing based on the cyclic nature of the self-ligated repeat units of the organism.

[0097] In some embodiments, a nuclease is used that generates a 5' overhang as a result of cleaving a target nucleic acid, and the 5' overhang of the cleaved nucleic acid molecule is then filled. Methods for filling the 5' overhang are well known to those skilled in the art and include, for example, methods using a DNA polymerase I Klenow fragment lacking exonuclease activity (Klenow(3'->5'exo-)). Filling the 5' overhang results in an extension using the overhang of the sink strand as a template, which in turn results in a blunt end. In the case of a single repeat unit released from a library concatemer, the resulting structure is a blunt-ended S2'R-(constant region)-LS1', where S1' and S2' contain blunt ends. A PCR and / or sequencing adapter can then be added to the ends by blunt-end ligation, and each repeat unit (containing the S2'R and LS1' regions) can be sequenced. From the sequence data, the original LSR region can be deduced. Smoothing of overhangs generated during the nuclease cleavage process also makes it possible to distinguish between target sites that have been properly cleaved by each nuclease and target sites that have been nonspecifically cleaved based on non-nuclease effects, such as physical shear. Correctly cleaved nuclease target sites may be recognized by the presence of complementary S2'R and LS1' regions, which include replication of overhang nucleotides as a result of filling the overhang, whereas target sites that were not cleaved by each nuclease are unlikely to include replication of overhang nucleotides. In some embodiments, the method includes identifying nuclease target sites cleaved by the nuclease by determining the sequences of the left half, right half, and / or spacer sequences of the repeat unit of the free organism. Any preferred method for amplification and / or sequencing may be used to identify the LSR sequences of target sites cleaved by each nuclease. Methods for amplifying and / or sequencing nucleic acids are well known to those skilled in the art, and this disclosure is not limited in this respect.In the case of nucleic acids released from a library concatemer containing cleaved half-regions and uncleaved target regions (e.g., containing at least approximately 1.5 repeat sequences), filling the 5' overhang also provides assurance that the nucleic acid has been cleaved by the nuclease. Since the nucleic acid also contains intact or uncleaved target regions, the sequence of those regions can be determined without having to reconstruct the sequence from the left half-region, the right half-region, and / or spacer sequences.

[0098] Some of the methods and strategies provided herein allow for the simultaneous evaluation of multiple candidate target sites as possible cleavage targets for any given nuclease. Thus, data obtained from such methods can be used to create a list of target sites cleaved by a given nuclease, also referred to herein as a target site profile. When sequencing methods are used that enable the generation of quantitative sequencing data, it is also possible to record the relative abundance of any of the nuclease target sites detected as cleaved by each nuclease. Target sites that are more efficiently cleaved by a nuclease will be detected more frequently in the sequencing step. On the other hand, target sites that are not efficiently cleaved will rarely release individual repeat units from candidate concatemers, and therefore will generate at most only a small number of sequencing reads. Such quantitative sequencing data can be integrated into the target site profile to generate a ranking list of highly preferred and less preferred nuclease target sites.

[0099] The methods and strategies for nuclease target site profiling provided herein may be applied to any site-specific nuclease, including, for example, ZFNs, TALENs, homing endonucleases, and RNA-programmable nucleases, such as Cas9 nucleases. As described in more detail herein, nuclease specificity typically decreases with increasing nuclease concentrations, and the methods described herein may be used to determine the concentration at which a given nuclease efficiently cleaves its intended target site but not any off-target sequences. In some embodiments, a maximum concentration of therapeutic nuclease is determined at which the therapeutic nuclease cleaves its intended nuclease target site but not more than 10, more than 5, more than 4, more than 3, more than 2, more than 1, or any additional sites. In some embodiments, the therapeutic nuclease is administered to a subject in an amount effective to produce a final concentration that is lower than or equal to the maximum concentration determined above.

[0100] In some embodiments, the library of candidate nucleic acid molecules used in the methods provided herein comprises at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , or at least 10 12 It contains several different candidate nuclease target sites.

[0101] In some embodiments, the nuclease is a therapeutic nuclease that cleaves a specific nuclease target site in a disease-associated gene. In some embodiments, the method further comprises determining a maximum concentration of the therapeutic nuclease in which the therapeutic nuclease cleaves a specific nuclease target site and does not cleave more than 10 additional nuclease target sites, more than 5 additional nuclease target sites, more than 4 additional nuclease target sites, more than 3 additional nuclease target sites, more than 2 additional nuclease target sites, more than 1 additional nuclease target sites, or does not cleave any additional sites. In some embodiments, the method further comprises administering the therapeutic nuclease to a subject in an amount effective to produce a final concentration that is lower than or equal to the maximum concentration.

[0102] Nuclease Target Site Library Some aspects of this disclosure provide a library of nucleic acid molecules for nuclease target site profiling. In some aspects, candidate nucleic acid molecules in the library have the structure R1-[(N Z )-(PAM)-(steady-state region)] X -Includes R2, and R1 and R2 are independently [(N Z A nucleic acid sequence that may contain fragments of a repeat unit ()-(PAM)-(steady region), where each N independently represents any nucleotide, Z is an integer from 1 to 50, and X is an integer from 2 to y. In some embodiments, y is at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12, at least 10 13 , at least 10 14 , or at least 10 15 In some embodiments, y is 10 2 A little more than 10 3 A little more than 10 4 A little more than 10 5 A little more than 10 6 A little more than 10 7 A little more than 10 8 A little more than 10 9 A little more than 10 10 A little more than 10 11 A little more than 10 12 A little more than 10 13 A little more than 10 14 A little less than, or 10 15 It is slightly less than that. In some embodiments, Z is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50. In some embodiments, Z is 20. Each N independently represents any of the nucleotides. Thus, Z N with =2 Z The sequence provided is NN, where each N independently represents A, T, G, or C. Therefore, Z N with =2 Z This can represent AA, AT, AG, AC, TA, TT, TG, TC, GA, GT, GG, GC, CA, CT, CG, and CC.

[0103] In some embodiments, a library is provided comprising candidate nucleic acid molecules comprising a target site having a partially randomized left half region, a partially randomized right half region, and / or a partially randomized spacer sequence. In some embodiments, a library is provided comprising candidate nucleic acid molecules comprising a target site having a partially randomized left half region, a fully randomized spacer sequence, and a partially randomized right half region. In some embodiments, a library is provided comprising candidate nucleic acid molecules comprising a target site having a partially or fully randomized sequence, wherein the target site has a structure [N] as described herein, for example. Z This includes (PAM). In some embodiments, the partially randomized region differs from the consensus region by an average of more than 5%, more than 10%, more than 15%, more than 20%, more than 25%, or more than 30% (binomial distribution).

[0104] In some embodiments, such a library comprises multiple nucleic acid molecules, each containing a concatemer of a candidate nuclease target site and a constant insertion sequence (also referred to herein as a constant region). For example, in some embodiments, the candidate nucleic acid molecule of the library has the structure R1-[(sgRNA complementary sequence)-(PAM)-(constant region)] X -R2, or structure R1-[(LSR)-(steady region)] X -R2 is included, and the structure in brackets ("[...]") is called a repeat unit or repeat sequence, where R1 and R2 are nucleic acid sequences that may independently contain fragments of a repeat unit, and X is an integer between 2 and y. In some embodiments, y is at least 10 1 , at least 10 2 , at least 10 3 , at least 10 4 , at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 1010 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 , or at least 10 15 In some embodiments, y is 10 2 A little more than 10 3 A little more than 10 4 A little more than 10 5 A little more than 10 6 A little more than 10 7 A little more than 10 8 A little more than 10 9 A little more than 10 10 A little more than 10 11 A little more than 10 12 A little more than 10 13 A little more than 10 14 A little less than, or 10 15It is slightly less than that. In some embodiments, the steady region is a length that allows for efficient self-ligation of a single repeat unit. In some embodiments, the steady region is a length that allows for efficient separation of a single repeat unit from a fragment containing two or more repeat units. In some embodiments, the steady region is a length that allows for efficient sequencing of a complete repeat unit in a single sequencing read. Preferred lengths will be obvious to those skilled in the art. For example, in some embodiments, the steady region is 5 to 100 base pairs long, e.g., about 5 base pairs, about 10 base pairs, about 15 base pairs, about 20 base pairs, about 25 base pairs, about 30 base pairs, about 35 base pairs, about 40 base pairs, about 50 base pairs, about 60 base pairs, about 70 base pairs, about 80 base pairs, about 90 base pairs, or about 100 base pairs long. In some embodiments, the constant region is 1, 2, 3, 4, 5, 6, 7, 8, 9, 0, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 base pairs long.

[0105] The LSR site typically contains a [left half-site]-[spacer sequence]-[right half-site] structure. The length of the half-size and spacer sequence will depend on the specific nuclease to be evaluated. Generally, the half-site will be 6–30 nucleotides long, preferably 10–18 nucleotides long. For example, each half-site may be individually 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. In some embodiments, the LSR site may be longer than 30 nucleotides. In some embodiments, the left and right half-sites of the LSR are the same length. In some embodiments, the left and right half-sites of the LSR are of different lengths. In some embodiments, the left and right half-sites of the LSR are of different sequences. In some embodiments, a library is provided comprising candidate nucleic acids containing LSRs that can be cleaved by a FokI cleavage domain, a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), a homing endonuclease, or an organic compound (e.g., enediine antibiotics, e.g., dynemycin, neocardinostatin, calichemycin, and esperamicin I, as well as bleomycin).

[0106] In some embodiments, at least 10 5 , at least 10 6 , at least 10 7 , at least 10 8 , at least 10 9 , at least 10 10 , at least 10 11 , at least 10 12 , at least 10 13 , at least 10 14 , or at least 10 15A library of candidate nucleic acid molecules containing several different candidate nuclease target sites is provided. In some embodiments, the candidate nucleic acid molecules in the library are concatemers produced from a template cyclically formed by rolling cycle amplification. In some embodiments, the library contains nucleic acid molecules (e.g., concatemers) with molecular weights of at least 5 kDa, at least 6 kDa, at least 7 kDa, at least 8 kDa, at least 9 kDa, at least 10 kDa, at least 12 kDa, or at least 15 kDa. In some embodiments, the molecular weight of the nucleic acid molecules in the library may be greater than 15 kDa. In some embodiments, the library contains nucleic acid molecules within a specific size range (e.g., in the range of 5–7 kDa, 5–10 kDa, 8–12 kDa, 10–15 kDa, or 12–15 kDa, or 5–10 kDa, or any possible partial range). Several methods suitable for generating nucleic acid concatemers according to some aspects of this disclosure result in the generation of nucleic acid molecules of vastly different molecular weights, while such mixtures of nucleic acid molecules can be size-fractionated to obtain a desired size distribution. Suitable methods for concentrating nucleic acid molecules of a desired size or excluding nucleic acid molecules of a desired size are well known to those skilled in the art, and the disclosure is not limited in this respect.

[0107] In some embodiments, partially randomized sites differ from the consensus site by an average of less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, or less than 50% (binomial distribution). For example, in some embodiments, partially randomized sites differ from the consensus site by more than 5% but less than 10%, more than 10% but less than 20%, more than 20% but less than 25%, more than 5% but less than 20%, etc. Using partially randomized nuclease target sites in a library is useful for increasing the concentration of library members containing target sites that are closely related to the consensus site (e.g., differing from the consensus site by only 1, 2, 3, 4, or 5 residues). The logic behind this is that a given nuclease (e.g., a given ZFN or RNA-programmable nuclease) is likely to cleave its intended target site and any closely related target sites, but unlikely to cleave target sites that are far removed from or completely unrelated to its intended target site. Therefore, using a library containing partially randomized target sites may be more efficient than using a library containing fully randomized target sites, without compromising sensitivity in detecting any off-target cleavage events for any given nuclease. Hence, the use of a partially randomized library significantly reduces the cost and effort required to produce a library with a high probability of covering substantially all off-target sites for a given nuclease. However, in some embodiments, for example, in embodiments where the specificity of a given nuclease is to be evaluated in the context of any possible site in a given genome, it may be preferable to use a library with fully randomized target sites.

[0108] Selection and design of site-specific nucleases Several aspects of this disclosure provide methods and strategies for selecting and designing site-specific nucleases that enable targeted cleavage of a single unique site in the context of a complex genome. In some embodiments, a method is provided, comprising: providing a plurality of candidate nucleases that are designed to cleave or are known to cleave the same consensus sequence; profiling the target sites actually cleaved by each candidate nuclease and therefore detecting any cleaved off-target sites (target sites different from the consensus target site); and selecting a candidate nuclease based on the off-target site(s) thus identified. In some embodiments, this method is used to select the most specific nuclease from the group of candidate nucleases (e.g., a nuclease that cleaves the consensus target site with the highest specificity, a nuclease that cleaves the fewest number of off-target sites, a nuclease that cleaves the fewest number of off-target sites in the context of the target genome, or a nuclease that cleaves no target sites other than the consensus target site). In some embodiments, this method is used to select a nuclease that does not cleave any off-target sites in the context of the target genome at concentrations higher than or equal to the therapeutically effective concentration of the nuclease.

[0109] The methods and reagents provided herein can be used, for example, to evaluate multiple different nucleases that target the same target site (e.g., multiple variations of a given site-specific nuclease, e.g., a given zinc finger nuclease). Thus, such methods can be used as a selection step in evolving or designing novel site-specific nucleases with improved specificity.

[0110] Identification of unique nuclease target sites in the genome Several aspects of this disclosure provide methods for selecting nuclease target sites in a genome. As described in more detail elsewhere herein, it has been surprisingly discovered that off-target sites cleaved by a given nuclease are typically highly similar to consensus target sites, differing, for example, by only 1, 2, 3, 4, or 5 nucleotide residues. Based on this discovery, nuclease target sites in the genome can be selected to increase the likelihood that a nuclease targeting this site will not cleave any off-target sites in the genome. For example, in several embodiments, methods are provided that include identifying candidate nuclease target sites and comparing the candidate nuclease target sites with other sequences in the genome. Methods for comparing candidate nuclease target sites with other sequences in the genome are well known to those skilled in the art and include, for example, sequence alignment methods, which use, for example, sequence alignment software or algorithms (e.g., BLAST) on a general-purpose computer. A suitable unique nuclease target site can then be selected based on the results of the sequence comparison. In some embodiments, a candidate nuclease target site is selected as a unique site in the genome if it differs from any other sequence in the genome by at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 nucleotides; however, if the site does not meet this criterion, the site may be discarded. In some embodiments, as briefly described above, once a site has been selected based on sequence comparison, a site-specific nuclease is designed to target the selected site. For example, a zinc finger nuclease may be designed to target any selected nuclease target site by constructing a zinc finger array that binds to the target site and conjugating the zinc finger array to a DNA cleavage domain.In embodiments where the DNA cleavage domain needs to dimerize to cleave DNA, a zinc finger array is designed, each binding to a half-site of the nuclease target site, and each conjugating to the cleavage domain. In some embodiments, nuclease design and / or production is done by recombinant techniques. Preferred recombinant techniques are well known to those skilled in the art, and the disclosure is not limited in this regard.

[0111] In some embodiments, site-specific nucleases designed or produced in accordance with aspects of this disclosure are isolated and / or purified. Methods and strategies for designing site-specific nucleases in accordance with aspects of this disclosure may be applied to design or produce any site-specific nuclease (including, but not limited to, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), homing endonucleases, organic compound nucleases, or enediine antibiotics (e.g., dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin)).

[0112] Isolated nuclease Several aspects of this disclosure provide isolated site-specific nucleases with enhanced specificity, designed using the methods and strategies described herein. Several aspects of this disclosure provide nucleic acids encoding such nucleases. Several aspects of this disclosure provide expression constructs comprising such encoding nucleic acids. For example, in some embodiments, an isolated nuclease is provided which has been engineered to cleave a desired target site in the genome and has been evaluated according to the methods provided herein to cleave a few more than one, a few more than two, a few more than three, a few more than four, a few more than five, a few more than six, a few more than seven, a few more than eight, a few more than nine, or a few more than ten off-target sites at concentrations effective for the nuclease to cleave its intended target site. In some embodiments, an isolated nuclease is provided which has been engineered to cleave a desired unique target site selected to differ from any other site in the genome by at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten nucleotide residues. In some embodiments, the isolated nuclease is an RNA-programmable nuclease, e.g., Cas9 nuclease, zinc finger nuclease (ZFN), or transcription activator-like effector nuclease (TALEN), homing endonuclease, organic compound nuclease, or enediine antibiotic (e.g., dynemycin, neocardinostatin, calichemycin, esperamicin, bleomycin). In some embodiments, the isolated nuclease cleaves a target site within an allele associated with a disease or disorder. In some embodiments, the isolated nuclease cleaves a target site, and this cleavage results in treatment or prevention of the disease or disorder. In some embodiments, the disease is HIV / AIDS or a proliferative disorder. In some embodiments, the allele is CCR5 (for treating HIV / AIDS) or VEGFA allele (for treating a proliferative disorder).

[0113] In some embodiments, isolated nucleases are provided as part of a pharmaceutical composition. For example, some embodiments provide a pharmaceutical composition comprising a nuclease provided herein or a nucleic acid encoding such nuclease and a pharmaceutically acceptable excipient. The pharmaceutical composition may optionally contain one or more additional therapeutic agents.

[0114] In some embodiments, the compositions provided herein are administered to a subject (e.g., a human subject) to carry out targeted genomic modifications within the subject. In some embodiments, cells are obtained from the subject, ex vivo contacted with a nuclease or nucleic acid encoding a nuclease, and then re-administered to the subject after the desired genomic modification has been carried out or detected in the cells. While the descriptions of pharmaceutical compositions provided herein are primarily directed toward pharmaceutical compositions suitable for administration to humans, it will be understood by those skilled in the art that such compositions are generally suitable for administration to all types of animals. Modifications of pharmaceutical compositions suitable for administration to humans to make them suitable for administration to various animals are well known, and veterinary pharmacologists can design and / or make such modifications by at best ordinary experimental methods. The subjects to whom the pharmaceutical composition may be administered include, but are not limited to, humans and / or other primates, mammals (including commercially relevant mammals such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats) and / or birds (including commercially relevant birds such as chickens, ducks, geese, and / or turkeys).

[0115] Formulations of the pharmaceutical compositions described herein may be prepared by any method known or subsequently developed in the field of pharmacology. Generally, such preparations involve the steps of associating the active ingredient with excipients and / or one or more other auxiliary components, and then, if necessary and / or desirable, shaping and / or packaging the product into desired single or multiple dose units.

[0116] In addition, pharmaceutical formulations may contain pharmaceutically acceptable excipients, which, as used herein, include any and all solvents, dispersions, diluents, or other liquid bases, dispersion or suspension aids, surfactants, isotonic agents, thickeners or emulsifiers, preservatives, solid binders, lubricants, and the like, in a manner suitable for the desired specific dosage form. (Remington's The Science and Practice of Pharmacy, 21) st Edition, AR Gennaro (Lippincott, Williams & Wilkins, Baltimore, MD, 2006; incorporated herein by reference in its entirety) discloses various excipients used in formulating pharmaceutical compositions and known techniques for their preparation. For additional preferred methods, reagents, excipients, and solvents for producing pharmaceutical compositions comprising nucleases, also refer to PCT application PCT / US2010 / 055131, incorporated herein by reference in its entirety. The use of any conventional excipient medium is considered to be within the scope of this disclosure, except insofar as any conventional excipient medium is incompatible with the substance or its derivatives (e.g., by producing any undesirable biological effect or otherwise by interacting in an adverse manner with any other component(s) of the pharmaceutical composition.

[0117] The functions and advantages of these and other embodiments of the present invention will be better understood from the following examples. The following examples are intended to illustrate the benefits of the present invention and to describe specific embodiments, but are not intended to illustrate the entire scope of the invention. It will be understood that the examples are not intended to limit the scope of the invention. [Examples]

[0118] material and method Oligonucleotides. All oligonucleotides used in this study were purchased from Integrated DNA Technologies. The oligonucleotide sequences are listed in Table 9.

[0119] Expression and purification of S. pyogenes Cas9. E. coli Rosetta (DE3) cells were fused to the N-terminal 6×His tag / maltose-binding protein in plasmid pMJ806 encoding the S. pyogenes Cas9 gene. 11 The cells were transformed by [method]. The resulting expression strains were inoculated overnight at 37°C in Luria Bertani (LB) broth containing 100 μg / mL ampicillin and 30 μg / mL chloramphenicol. The cells were diluted 1:100 in the same growth medium and [method]. 600The cells were grown at 37°C until the saturation was approximately 0.6. The culture was incubated at 18°C ​​for 30 minutes, and Cas9 expression was induced by adding 0.2 mM isopropyl β-D-1-thiogalactopyranoside (IPTG). After approximately 17 hours, the cells were harvested by centrifugation at 8,000 g and resuspended in lysis buffer (20 mM tris(hydroxymethyl)-aminomethane (Tris)-HCl (pH 8.0), 1 M KCl, 20% glycerol, 1 mM tris(2-carboxyethyl)phosphine (TCEP)). The cells were lysed by ultrasonic treatment (6 W output for a total of 10 minutes, with 10 sec pulse-on and 30 sec pulse-off), and the soluble lysate was obtained by centrifugation at 20,000 g for 30 minutes. Cell lysates were incubated with nickel-nitriloacetate (nickel-NTA) resin (Qiagen) at 4°C for 20 minutes to capture His-tagged Cas9. The resin was transferred to a 20 mL column and washed with 20 column volumes of lysis buffer. Cas9 was eluted in 20 mM Tris-HCl (pH 8), 0.1 M KCl, 20% glycerol, 1 mM TCEP, and 250 mM imidazole, and concentrated to approximately 50 mg / mL using an Amicon ultra centrifugation filter (Millipore, 30 kDa molecular weight cutoff). 6×His-tagged and maltose-binding proteins were removed by TEV protease treatment at 4°C for 20 hours and captured by a second Ni affinity purification step. The eluate containing Cas9 was injected into a HiTrap SP FF column (GE Healthcare) in a purification buffer (containing 20 mM Tris-HCl (pH 8), 0.1 M KCl, 20% glycerol, and 1 mM TCEP). Cas9 was eluted with a purification buffer containing a 0.1 M to 1 M linear KCl gradient in 5 column volumes. The eluted Cas9 was further purified in the purification buffer using a HiLoad Superdex 200 column, flash-frozen in liquid nitrogen, and stored in small portions at -80°C.

[0120] In vitro RNA transcription: 100 pmol of CLTA(#) v2.1 fwd and v2.1 template rev were incubated at 95°C and cooled to 37°C at 0.1°C / s in NEBuffer 2 (50 mM sodium chloride, 10 mM Tris-HCl, 10 mM magnesium chloride, 1 mM dithiothreitol (pH 7.9)) supplemented with 10 μM dNTP mix (Bio-Rad). 10 U of Klenow fragment (3'->5' exo-) (NEB) was added to the reaction mixture to obtain the double-stranded CLTA(#) v2.1 template by overlap extension at 37°C for 1 hour. A 200 nM CLTA(#) v2.1 template alone or a 100 nM CLTA(#) template with a 100 nM T7 promoter oligo was incubated overnight at 37°C with 0.16 U / μL of T7 RNA polymerase (NEB) in NEB RNAPol buffer (40 mM Tris-HCl (pH 7.9), 6 mM magnesium chloride, 10 mM dithiothreitol, 2 mM spermidine) supplemented with a 1 mM rNTP mix (1 mM rNTP, 1 mM rCTP, 1 mM rGTP, 1 mM rUTP). The in vitro transcribed RNA was precipitated with ethanol and purified by gel electrophoresis on a Criterion 10% polyacrylamide TBE-urea gel (Bio-Rad). The gel-purified sgRNA was precipitated with ethanol and redissolved in water.

[0121] In vitro library construction. 10 pmol of CLTA(#) lib oligonucleotides were separately circularized by incubation at 60°C for 16 hours in a total reaction volume of 20 μL with 100 units of CircLigase II ssDNA ligase (Epicentre) in 1×CircLigase II reaction buffer supplemented with 2.5 mM manganese chloride (33 mM Tris-acetic acid, 66 mM potassium acetate, 0.5 mM dithiothreitol (pH 7.5)). The reaction mixture was incubated at 85°C for 10 minutes to inactivate the enzyme. 5 μL (5 pmol) of crude circular single-stranded DNA was converted to a pre-concatemer library using the illustra TempliPhi amplification kit (GE Healthcare) according to the manufacturer's protocol. The pre-concatemer library was quantified using the Quant-it PicoGreen dsDNA assay kit (Invitrogen).

[0122] Plasmid templates for in vitro cleavage of on-target and off-target substrates were constructed by ligation of annealed oligonucleotides CLTA(#) site fwd / rev into HindIII / XbaI double-digested pUC19(NEB). On-target substrate DNA was generated by PCR with the plasmid template and test fwd and test rev primers, and then purified using the QIAquick PCR purification kit (Qiagen). Off-target substrate DNA was generated by primer extension. 100 pmol of off-target(#) fwd and off-target(#) rev primers were incubated at 95°C and cooled to 37°C at 0.1°C / s in NEBuffer2 (50 mM sodium chloride, 10 mM Tris-HCl, 10 mM magnesium chloride, 1 mM dithiothreitol (pH 7.9)) supplemented with 10 μM dNTP mix (Bio-Rad). 10U of Klenow fragment (3'->5' exo-)(NEB) was added to the reaction mixture to obtain a double-stranded off-target template by overlap extension at 37°C for 1 hour, followed by enzyme inactivation at 75°C for 20 minutes, and then purified using the QIAquick PCR purification kit (Qiagen). 200nM substrate DNA was incubated at 37°C for 10 minutes with 100nM Cas9 and 100nM (v1.0 or v2.1) sgRNA or 1000nM Cas9 and 1000nM (v1.0 or v2.1) sgRNA in Cas9 cleavage buffer (200mM HEPES (pH 7.5), 1.5M potassium chloride, 100mM magnesium chloride, 1mM EDTA, 5mM dithiothreitol). On-target cleavage reactions were purified using the QIAquick PCR purification kit (Qiagen), and off-target cleavage reactions were purified using the QIAquick nucleotide removal kit (Qiagen) before electrophoresis in a Criterion 5% polyacrylamide TBE gel (Bio-Rad).

[0123] In vitro selection: A 200 nM concatemer pre-selection library was incubated at 37°C for 10 min with 100 nM Cas9 and 100 nM sgRNA or 1000 nM Cas9 and 1000 nM sgRNA in Cas9 cleavage buffer (200 mM HEPES (pH 7.5), 1.5 M potassium chloride, 100 mM magnesium chloride, 1 mM EDTA, 5 mM dithiothreitol). The pre-selection libraries were also incubated separately at 37°C for 1 hour with 2 U of BspMI restriction endonuclease (NEB) in NEBuffer 3 (100 mM NaCl, 50 mM Tris-HCl, 10 mM MgCl2, 1 mM dithiothreitol (pH 7.9)). Library members selected after blunt ends or before sticky ends were purified using the QIAQuick PCR Purification Kit (Qiagen) and ligated with 10,000 U of T4 DNA ligase (NEB) in NEB T4 DNA ligase reaction buffer (50 mM Tris-HCl (pH 7.5), 10 mM magnesium chloride, 1 mM ATP, 10 mM dithiothreitol) with 10 pmol of adapter 1 / 2 (AACA) (Cas9:v2.1 sgRNA, 100 nM), adapter 1 / 2 (TTCA) (Cas9:v2.1 sgRNA, 1000 nM), adapter 1 / 2 (Cas9:v2.1 sgRNA, 1000 nM), or lib adapter 1 / CLTA(#) lib adapter 2 (pre-selection). Adapter-ligated DNA was purified using the QIAquick PCR purification kit and amplified for 10-13 cycles using Phusion Hot Start Flex DNA polymerase (NEB) and primers CLTA(#) sel PCR / PE2 short (after selection) or CLTA(#) lib seq PCR / lib fwd PCR (before selection) in Buffer HF (NEB).The amplified DNA was gel-purified, quantified using the KAPA Library Quantification Kit-Illumina (KAPA Biosystems), and subjected to single-read sequencing on Illumina MiSeq or Rapid Run single-read sequencing on Illumina HiSeq 2500 (Harvard University FAS Center for Systems Biology Core facility, Cambridge, MA).

[0124] Selection analysis. Pre-selection and post-selection sequencing data are as previously described. 21 The analysis was performed (with modifications using a script written in C++). Raw sequence data is not shown. See Table 2 for a curated summary. Specificity scores were calculated using the following formulas: positive specificity score = (base pair frequency at position [after selection] - base pair frequency at position [before selection]) / (1 - base pair frequency at position [before selection]) and negative specificity score = (base pair frequency at position [after selection] - base pair frequency at position [before selection]) / (base pair frequency at position [before selection]). Sequence logo normalization was performed as previously described. 22 .

[0125] Cell cutting assay. HEK293T cells 0.8 × 10⁶ per well before transcription. 5The cells were divided into 6-well plates and maintained in Dulbecco's Modified Eagle Medium (DMEM) supplemented with 10% fetal bovine serum (FBS) in a 37°C humidified incubator with 5% CO2. After 1 day, the cells were transiently transfected with Lipofectamine 2000 (Invitrogen) according to the manufacturer's protocol. HEK293T cells were transfected with 1.0 μg of Cas9 expression plasmid (Cas9-HA-2xNLS-GFP-NLS) and 2.5 μg of single-stranded RNA expression plasmid pSiliencer-CLTA (version 1.0 or 2.1) at 70% confluence in each well of the 6-well plate. Transfection efficiency was estimated to be approximately 70% based on the percentage of GFP-positive cells observed by fluorescence microscopy. 48 hours after transfection, the cells were washed with phosphate-buffered saline (PBS), pelletized, and frozen at -80°C. Genomic DNA was isolated from 200 μL of cell lysate using the DNeasy Blood and Tissue Kit (Qiagen) according to the manufacturer's protocol.

[0126] Off-target site sequencing. 100 ng of genomic DNA isolated from cells treated with Cas9 expression plasmid and single-stranded RNA expression plasmid (treated cells) or Cas9 expression plasmid alone (control cells) was amplified by PCR using Phusion Hot Start Flex DNA polymerase (NEB) in Buffer GC (NEB) supplemented with primers CLTA(#)-(#)-(#) fwd and CLTA(#)-(#)-(#) rev and 3% DMSO, with 35 cycles of 10 s 72°C extension. The relative amounts of crude PCR products were quantified by gel, and PCRs treated with Cas9 (control) and Cas9:sgRNA were pooled separately at equimolar concentrations before purification with the QIAquick PCR purification kit (Qiagen). The purified DNA was amplified by PCR using Phusion Hot Start Flex DNA polymerase (NEB) in Buffer HF (NEB) with 7 cycles of primers PE1-barcode# and PE2-barcode#. The amplified control and treated DNA pools were purified using the QIAquick PCR Purification Kit (Qiagen), followed by purification with Agencourt AMPure XP (Beckman Coulter). The purified control and treated DNA were quantified using the KAPA Library Quantification Kit-Illumina (KAPA Biosystems), pooled in a 1:1 ratio, and subjected to paired-end sequencing on Illumina MiSeq.

[0127] Statistical analysis. The statistical analysis was performed as previously described. 21 The p-values ​​in Tables 1 and 6 were calculated using a one-sided Fisher's exact test.

[0128] algorithm All scripts were written in C++. The algorithms used in this study are the same as those previously reported (see reference), with modifications.

[0129] Array Binning. 1) Specify the array pair starting with the barcode "AACA" or "TTCA" as the selected library member. 2) Regarding the selected library member (example provided), The lead in question:

number

number

[0130] NHEJ Array Call The lead in question:

number

number

number

[0131] Filter based on cutting site (arrangement after selection) 1) The cleavage sites are tallied through the recognition sites by identifying the first position within the fully sequenced recognition site (between the two steady sequences) which is identical to the first position in the sequencing read after the barcode (before the first steady sequence). 2) After aggregation, repeat step 1 and keep only sequences that have cleavage sites present in at least 5% of the sequencing reads.

[0132] result Broad-spectrum off-target DNA cleavage profiling reveals RNA-programmed Cas9 nuclease specificity. Sequence-specific endonucleases, including zinc finger nucleases (ZFNs) and transcription activator-like effector nucleases (TALENs), are used in induced pluripotent stem cells (iPSCs). 1~3 , multicellular organisms 4~8 , and ex vivo gene therapy clinical trials 9、10 In this context, ZFNs and TALENs have become important tools for gene modification. While ZFNs and TALENs have proven effective for such genetic manipulation, new ZFN or TALEN proteins must be generated for each DNA target site. In contrast, RNA-guided Cas9 endonucleases use RNA:DNA hybridization to determine the target DNA cleavage site, allowing a single monomeric protein to, in principle, cleave any sequence defined by the guide RNA. 11 .

[0133] From previous research 12~17 It has been demonstrated that Cas9 mediates genome editing at a site complementary to a 20-nucleotide sequence in the bound guide RNA. In addition, the target site must contain a protospacer facilitation motif (PAM) at its 3' end adjacent to the 20-nucleotide target site. For Streptococcus pyogenes Cas9, the PAM sequence is NGG. The specificity of Cas9-mediated DNA cleavage, both in vitro and in cells, has been previously inferred based on assays against small clusters of potential single-mutation off-target sites. These studies suggest that perfect complementarity between the guide RNA and target DNA is required within 7–12 base pairs adjacent to the PAM end of the target site (the 3' end of the guide RNA), while mismatches are acceptable at the non-PAM end (the 5' end of the guide RNA). 11、12、17~19 .

[0134] Cas9: The limited number of nucleotides that define the recognition of the guide RNA target is of medium to large size (>10). 7The Cas9: guide RNA complex will predict multiple sites of DNA breaks in the genome (bp), but the Cas9: guide RNA complex will predict multiple sites in the cell 12、13、15 and organisms 14 Both have been successfully used to modify zebrafish embryos. Studies using the Cas9:guideRNA complex to modify zebrafish embryos have observed toxicity rates similar to those of ZFNs and TALENs. 14 Recent extensive studies of the DNA binding (transcriptional repression) specificity of catalytically inactive Cas9 mutants in E. coli using high-throughput sequencing have not found any detectable off-target transcriptional repression in the relatively small E. coli transcriptome. 20 These studies have significantly advanced our fundamental understanding of Cas9. However, a systematic and comprehensive profile of Cas9:guideRNA-mediated DNA cleavage specificity, generated from measurements of Cas9 cleavage at numerous related mutant target sites, has not been described. Such a specificity profile is needed to understand and improve the potential of the Cas9:guideRNA complex as a research tool and future therapeutic agent.

[0135] The inventors of this invention previously published in vitro selection 21 The Cas9 single guide RNA (sgRNA) was modified and repurposed to process blunt-end cleavage products produced by Cas9 (compared to overhang-containing products of ZFN cleavage). 11 The off-target DNA cleavage profile of the complex was determined. Each selected experiment was approximately 10 12A DNA substrate library containing 100 sequences was used. This was large enough to provide 10 times the coverage of all sequences having 8 or 7 or fewer mutations relative to each 22-base-pair target sequence (including 2-base-pair PAMs) (Figure 1). The inventors used partially randomized nucleotide mixtures for all 22 target site base pairs to generate a binomial library of mutant target sites with an expected average of 4.62 mutations per target site. In addition, each target site library member was flanked by four fully randomized base pairs, testing for specificity patterns beyond those imposed by standard 20-base-pair target sites and PAMs.

[0136] 10 12Pre-selection libraries of potential off-target sites from individual organisms were generated for each of four different target sequences within the human clathrin light chain A (CLTA) gene (Figure 3). Synthetic 5' phosphorylated 53-base oligonucleotides were self-ligated in vitro to circular single-stranded DNA, and then converted to concatemer 53-base pair repeats by rolling circle amplification. The resulting pre-selection libraries were incubated with their corresponding Cas9:sgRNA complexes. Cleaved library members containing free 5' phosphate were isolated from intact library members by 5' phosphate-dependent ligation of an unphosphorylated double-stranded sequencing adapter. The ligated post-selection libraries were amplified by PCR. The PCR step produced a mixture of post-selection DNA fragments containing 0.5, 1.5, or 2.5 repeats of the Cas9-cleaved library members, resulting in amplification of adapter-ligated cleaved half-sites with or without one or more adjacent corresponding complete sites (Figure 1). Selected library members containing 1.5 target sequence repeats were isolated by gel purification and analyzed by high-throughput sequencing. In the final computer selection step to minimize the impact of errors during DNA amplification or sequencing, only sequences containing two identical copies of the cleaved half-region, which constitutes a repeat, were analyzed.

[0137] The pre-selection libraries were incubated under either enzyme restriction conditions (200 nM target site library, 100 nM Cas9:sgRNA v2.1) or enzyme saturation conditions (200 nM target site library, 1000 nM Cas9:sgRNA v2.1) for each of the four guide RNA targets tested (CLTA1, CLTA2, CLTA3, and CLTA4) (Figures 3C and 3D). The second guide RNA construct, sgRNA v1.0, was less active than sgRNA v2.1 and was assayed under enzyme saturation conditions alone for each of the four guide RNA targets tested (200 nM target site library, 1000 nM Cas9:sgRNA v1.0). The two guide RNA constructs differed in their length (Figure 3) and their DNA cleavage activity levels under selection conditions, which is consistent with previous reports. 15 (Figure 4). Both the pre-selection and post-selection libraries were characterized by high-throughput DNA sequencing and computer analysis. As expected, library members with fewer mutations were significantly enriched in the post-selection library compared to the pre-selection library (Figure 5).

[0138] Pre-selection and post-selection library compositions. The pre-selection libraries for CLTA1, CLTA2, CLTA3, and CLTA4 had observed mean mutation rates of 4.82 (n=1,129,593), 5.06 (n=847,618), 4.66 (n=692,997), and 5.00 (n=951,503) mutations per 22-base pair target site containing two-base pair PAMs, respectively. The post-selection libraries treated under enzyme restriction conditions with Cas9+CLTA1, CLTA2, CLTA3, or CLTA4 v.2.1 sgRNA contained mean mutation rates of 1.14 (n=1,206,268), 1.21 (n=668,312), 0.91 (n=1,138,568), and 1.82 (n=560,758) mutations per 22-base pair target site, respectively. Under enzyme overload conditions, the average number of mutations in the sequences that passed selection increased to 1.61 (n=640,391), 1.86 (n=399,560), 1.46 (n=936,414), and 2.24 (n=506,179) mutations per 22-base pair target site for CLTA1, CLTA2, CLTA3, and CLTA4 v2.1 sgRNAs, respectively. These results demonstrate that selection significantly enriched library members with fewer mutations for all Cas9:sgRNA complexes tested, and that enzyme overload conditions resulted in presumed cleavage of more highly mutated library members compared to enzyme restriction conditions (Figure 5).

[0139] The inventors calculated specificity scores to quantify the level of enrichment of each base pair at each position in the post-selection library relative to the pre-selection library, and normalized these scores to the maximum possible enrichment of that base pair. A positive specificity score indicates a base pair that was enriched in the post-selection library, while a negative specificity score indicates a base pair that was deenriched in the post-selection library. For example, a score of +0.5 indicates that the base pair was enriched to 50% of its maximum enrichment value, while a score of -0.5 indicates that the base pair was deenriched to 50% of its maximum deenrichment value.

[0140] In addition to the two base pairs defined by the PAM, all 20 base pairs targeted by the guide RNA were enriched in the sequences from the CLTA1 and CLTA2 selections (Figures 2, 6, and 9, and Table 2). For the CLTA3 and CLTA4 selections (Figures 7 and 8, and Table 2), the base pairs defined by the guide RNA were enriched at all positions except for the two most distal base pairs from the PAM (the 5' end of the guide RNA). At the undefined positions furthest from the PAM, at least two of the three alternative base pairs were enriched to approximately the same extent as the defined base pairs. Our finding that the overall 20 base pair target sites and the 2 base pair PAM may contribute to the DNA cleavage specificity of Cas9:sgRNA contrasts with results from previous single-substrate assays that suggested only 7–12 base pairs and the 2 base pair PAM were defined. 11、12、15 .

[0141] All pre-selection (n≧14,569) and post-selection (n≧103,660) library members of single mutants were computer-analyzed to provide selection enrichment values ​​for each possible single mutant sequence. The results of this analysis (Figures 2, 6, and 8) show that, when only single mutant sequences are considered, under enzyme restriction conditions, the 6-8 base pairs closest to the PAM are generally highly defined, while the non-PAM ends are poorly defined, which is consistent with previous findings. 11、12、17~19However, under enzyme saturation conditions, even single mutations of the 6–8 base pairs closest to the PAM are tolerated, suggesting that high specificity at the PAM terminus of the DNA target site may be impaired when the enzyme concentration is relatively high relative to the substrate (Figure 2). The observation of high specificity for single mutations close to the PAM applies only to sequences containing single mutations, and the selection results do not support a model in which any combination of mutations is tolerated in the target site region furthest from the PAM (Figures 10–15). Analysis of pre- and post-selection library composition is described elsewhere herein, position-dependent specificity patterns are illustrated in Figures 18–20, PAM nucleotide specificity is illustrated in Figures 21–24, and the more detailed effect of Cas9:sgRNA concentration on specificity is described in Figures 2G and 25).

[0142] Specificity at the non-PAM terminus of the target site. To evaluate the ability of Cas9:v2.1 sgRNA to tolerate multiple distal mutations in PAM under enzyme-over-enzyme conditions, the inventors calculated the maximum specificity score at each position of a sequence containing mutations only in the 1-12 base pair region at the terminal of the target site furthest from PAM (Figures 10-17).

[0143] The results of this analysis show that, at the molecular end furthest from the PAM when the rest of the sequence does not contain mutations, there is no selection for sequences with up to three mutations (maximum specificity score is approximately 0), depending on the target site. For example, when only the three base pairs furthest from the PAM can be changed within the CLTA2 target site (shown by the dark bars in Figure 11C), the maximum specificity score at each of the three variable positions is close to 0, indicating no selection for any of the four possible base pairs at each of the three variable positions. However, when the eight base pairs furthest from the PAM can be changed (Figure 11H), the maximum specificity scores at positions 4-8 are all greater than +0.4. This indicates that Cas9:sgRNA exhibits sequence preference at these positions even when the rest of the substrate contains preferred on-target base pairs.

[0144] The inventors also calculated the distribution of mutations in both pre-selection and post-selection libraries treated with v2.1 sgRNA under enzyme-over-enzyme conditions, when only the first 1–12 base pairs of the target site could be altered (Figures 15–17). For a subset of the data, there was considerable overlap between the pre-selection and post-selection libraries (Figures 15–17, a–c), demonstrating minimal to no selection in the post-selection library for sequences with mutations in only the first 3 base pairs of the target site. These results, together, indicate that Cas9:sgRNA can tolerate a small number of mutations (approximately 1–3) at the ends of sequences furthest from the PAM when the maximum sgRNA:DNA interaction is provided in the remainder of the target site.

[0145] Specificity of the PAM terminus of the target site. The inventors plotted positional specificity as the sum of the magnitudes of the specificity scores for all four base pairs at each position of each target site and normalized it to the same sum at the most highly specific position (Figures 18-20). Under both enzyme restriction and enzyme overload conditions, the PAM terminus of the target site is highly specific. Under enzyme restriction conditions, the PAM terminus of the molecule is almost absolutely specific by CLTA1, CTLA2, and CLTA3 guide RNAs (specificity score of base pairs specified by guide RNA ≥ +0.9) (Figures 2 and 6-9), and highly specific by CLTA4 guide RNA (specificity score of +0.7 to +0.9). Within this region of high specificity, specific single mutations that are consistent with the fluctuation pairing between the guide RNA and the target DNA are tolerated. For example, under enzyme restriction conditions for a single mutant sequence, dA:dT off-target base pairs and dG:dC base pairs defined by guide RNA are equally permissible at position 17 of 20 (relative to the non-PAM end of the target site) of the CLTA3 target site. At this position, rG:dT fluctuating RNA:DNA base pairs can be formed with minimal apparent loss of cleavage activity.

[0146] Importantly, the selection results also reveal that the selection of guide RNA hairpins influences specificity. When assayed under identical enzyme saturation conditions reflecting the substrate and relative enzyme excess in the cellular context, the shorter, less active sgRNA v1.0 construct was more specific than the longer, more active sgRNA v2.1 construct (Figures 2 and 5-8). The higher specificity of sgRNA v1.0 over sgRNA v2.1 was higher for CLTA1 and CLTA2 (approximately 40-90% difference) than for CLTA3 and CLTA4 (<40% difference). Interestingly, this difference in specificity localized to different regions of the target site for each target sequence (Figures 2H and 26). Together, these results suggest that different guide RNA configurations result in different DNA cleavage specificities, and that guide RNA-dependent changes in specificity do not equally affect all locations of the target site. Given the inverse relationship between Cas9:sgRNA concentration and specificity described above, the inventors surmise that the differences in specificity between guide RNA configurations stem from differences in their overall levels of DNA cleavage activity.

[0147] The effect of Cas9:sgRNA concentration on DNA cleavage specificity. To evaluate the effect of enzyme concentration on the specificity pattern for the four target sites tested, the inventors calculated the concentration-dependent difference in site specificity and compared it to the maximum possible change in site specificity (Figure 25). Generally, specificity was higher under enzyme restriction conditions than under enzyme overload conditions. The change from enzyme overload to enzyme restriction conditions generally increased specificity at the PAM terminus of the target by ≥80% of the maximum possible change in specificity. Decreasing the enzyme concentration generally induced a small increase (approximately 30%) in specificity at the end of the target site furthest from the PAM, but decreased concentration induced a considerably larger increase in specificity at the end of the target site closest to the PAM. For CLTA4, decreasing the enzyme concentration was accompanied by a small decrease (approximately 30%) in specificity at several base pairs near the end of the target site furthest from the PAM.

[0148] Specificity of PAM nucleotides. To evaluate the contribution of PAM to specificity, the inventors calculated the relative abundance of all 16 possible PAM dinucleotides in the pre-selection and post-selection libraries. All observed post-selection target site sequences were considered (Figure 21), or only post-selection target site sequences that did not contain mutations within the 20 base pairs defined by the guide RNA were considered (Figure 22). Considering all observed post-selection target site sequences, under enzyme restriction conditions, GG dinucleotides represented 99.8%, 99.9%, 99.8%, and 98.5% of the post-selection PAM dinucleotides for selection by CLTA1, CLTA2, CLTA3, and CLTA4 v2.1 sgRNA, respectively. In contrast, under enzyme-over-enzyme conditions, GG dinucleotides represented 97.7%, 98.3%, 95.7%, and 87.0% of the selected PAM dinucleotides, respectively, for selection by CLTA1, CLTA2, CLTA3, and CLTA4 v2.1 sgRNA. These data demonstrate that increasing enzyme concentration leads to increased cleavage of substrates containing non-standard PAM dinucleotides.

[0149] To describe the pre-selection library distribution of PAM dinucleotides, the inventors calculated specificity scores for PAM dinucleotides (Figure 23). When only on-target post-selection sequences are considered under enzyme overload conditions (Figure 24), non-standard PAM dinucleotides with a single G rather than two Gs are relatively tolerant. Under enzyme overload conditions, Cas9:CLTA4 sgRNA 2.1 showed the highest tolerance for non-standard PAM dinucleotides among all Cas9:sgRNA combinations tested. AG and GA dinucleotides were the most tolerant, followed by GT, TG, and CG PAM dinucleotides. In selection by Cas9:CLTA1, 2, or 3 sgRNA 2.1 under enzyme overload conditions, AG was the predominant non-standard PAM (Figures 23 and 24). The inventors' results are consistent with another recent study on PAM specificity (which shows that Cas9:sgRNA can recognize AG PAM dinucleotides). 23 In addition, the inventors' results indicate that under enzyme restriction conditions, GG PAM dinucleotides are highly restricted, while under enzyme overload conditions, non-standard PAM dinucleotides containing a single G may be tolerated, depending on the context of the guide RNA.

[0150] To confirm that the in vitro selection results accurately reflect the in vitro Cas9 cleavage behavior, the inventors performed individual cleavage assays of six CLTA4 off-target substrates containing 1 to 3 mutations within the target site. The inventors calculated the enrichment values ​​of all sequences in the post-selection library for Cas9:CLTA4 v2.1 sgRNA under enzyme saturation conditions by dividing the abundance of each sequence in the post-selection library by the calculated abundance in the pre-selection library. Under enzyme saturation conditions, single 1, 2, and 3-mutation sequences with the highest enrichment values ​​(27.5, 43.9, and 95.9) were cleaved to ≥71% completion (Figure 27). Two sequences with an enrichment value of 1.0 were cleaved to 35%, and two sequences with enrichment values ​​close to 0 (0.064) were not cleaved. Three sequences that were cleaved to 77% efficiency by CLTA4 v2.1 sgRNA were cleaved to a lower efficiency of 53% by CLTA4 v1.0 sgRNA (Figure 28). These results indicate that the selective enrichment value of individual sequences predicts in vitro cleavage efficiency.

[0151] To determine whether the results of in vitro selection and in vitro cleavage assays are directly related to Cas9:guide RNA activity in human cells, the inventors identified 51 off-target sites (19 in CLTA1 and 32 in CLTA4) containing up to eight mutations that were enriched in in vitro selection and present in the human genome (Tables 3-5). The inventors expressed Cas9:CLTA1 sgRNA v1.0, Cas9:CLTA1 sgRNA v2.1, Cas9:CLTA4 sgRNA v1.0, Cas9:CLTA4 sgRNA v2.1, or Cas9 without sgRNA in HEK293T cells by transient transfection, and used genomic PCR and high-throughput DNA sequencing to search for evidence of Cas9:sgRNA modification at 46 of the 51 off-target sites, as well as at on-target loci. Specific amplified DNA was not obtained for five of the 51 predicted off-target sites (three for CLTA1 and two for CLTA4).

[0152] Deep sequencing of genomic DNA isolated from HEK293T cells treated with Cas9:CLTA1 sgRNA or Cas9:CLTA4 sgRNA identified explicit non-homologous end joining (NHEJ) sequences at on-target sites and five of the 49 off-target sites tested (CLTA1-1-1, CLTA1-2-2, CLTA4-3-1, CLTA4-3-3, and CLTA4-4-8) (Tables 1 and 6-8). CLTA4 target sites were modified by Cas9:CLTA4 v2.1 sgRNA at a frequency of 76%, while off-target sites CLTA4-3-1, CLTA4-3-3, and CLTA4-4-8 were modified at frequencies of 24%, 0.47%, and 0.73%, respectively. The CLTA1 target site was modified by Cas9:CLTA1 v2.1 sgRNA at a frequency of 0.34%, while the off-target sites CLTA1-1-1 and CLTA1-2-2 were modified at frequencies of 0.09% and 0.16%, respectively.

[0153] Under v2.1 sgRNA-mediated enzyme saturation conditions, the two validated CLTA1 off-target sites (CLTA1-1-1 and CLTA1-2-2) were two of the three most highly enriched sequences identified in in vitro selection. CLTA4-3-1 and CLTA4-3-3 were the highest and third highest enriched sequences among the seven CLTA4 mutant sequences enriched in in vitro selection, which are also present in the genome. In vitro selection enrichment values ​​for the four mutant sequences were not calculated because 12 of the 14 CLTA4 sequences in the genome containing the four mutants (including CLTA4-4-8) were observed at the level of only one sequence count in the post-selection library. In summary, these results confirm that several off-target substrates identified in in vitro selection in the human genome are indeed cleaved by the Cas9:sgRNA complex in human cells, and also suggest that the most highly enriched genomic off-target sequences in selection are modified to the greatest extent possible in cells.

[0154] The intracellular off-target sites identified by the inventors were among the most highly enriched in the inventors' in vitro selection and contained up to four mutations relative to the target site. On the other hand, heterochromatin or covalent DNA modifications may impair the ability of the Cas9:guideRNA complex to access intracellular genomic off-target sites. From the identification of 5 off-target sites out of 49 cells tested in this study, rather than 0 or many, Cas9-mediated DNA breaks have been shown in recent studies. 11、12、19 As suggested by [the research], it is strongly suggested that the specific targeting is not limited to target sequences of 7-12 base pairs.

[0155] The cellular genome modification data are also consistent with the increased specificity of sgRNA v1.0 compared to sgRNA v2.1, observed in in vitro selection data and in individual assays. While the CLTA1-2-2, CLTA4-3-3, and CLTA4-4-8 sites were modified by the Cas9-sgRNA v2.1 complex, no evidence of modification at any of these three sites was detected in cells treated with Cas9:sgRNA v1.0. The CLTA4-3-1 site (modified at 32% of the on-target CLTA4 site modification frequencies in cells treated with Cas9:v2.1 sgRNA) was modified at only 0.5% of the on-target modification frequencies in cells treated with v1.0 sgRNA. This represents a 62-fold change in selectivity. In summary, these results demonstrate that guide RNA composition can have a significant impact on intracellular Cas9 specificity. The inventors' specificity profiling findings provide important advice to recent and current efforts to improve the overall DNA modification activity of the Cas9:guideRNA complex through the manipulation of guide RNA. 11、15 .

[0156] Overall, Cas9 off-target DNA cleavage profiling and subsequent analysis indicate that (i) Cas9:guide RNA recognition is extended to 18-20 defined target site base pairs and 2-base pair PAMs for the four target sites tested; (ii) increasing the Cas9:guide RNA concentration can decrease in vitro DNA cleavage specificity; (iii) using a more active sgRNA configuration can increase DNA cleavage specificity both in vitro and intracellularly, but can also impair it both in vitro and intracellularly; and (iv) as predicted by the inventors' in vitro results, Cas9:guide RNA can modify intracellular off-target sites by up to four mutations relative to the on-target site. The inventors' findings provide key knowledge to our understanding of RNA-programmed Cas9 specificity and reveal a previously unknown role of sgRNA configuration in DNA cleavage specificity. The principles revealed in this study may also be applicable to Cas9-based effectors engineered to mediate functions beyond DNA cleavage.

[0157] Equivalents and range Those skilled in the art will recognize, or be able to confirm by conventional experimental methods, many equivalents of the specific embodiments of the invention described herein. The scope of the invention is not intended to be limited to the foregoing, but is as set forth in the appended claims.

[0158] In a claim, unless otherwise indicated or otherwise obvious from the context, articles such as “a,” “an,” and “the” may mean one or more than one. Unless otherwise indicated or otherwise obvious from the context, any claim or statement containing “or” between one or more members of a group is considered satisfied if one, more than one, or all members of the group are present, used, or otherwise related in a given product or process. The present invention encompasses embodiments in which exactly one member of a group is present, used, or otherwise related in a given product or process. The present invention also encompasses embodiments in which more than one or all members of a group are present, used, or otherwise related in a given product or process.

[0159] Furthermore, it should be understood that the present invention encompasses all variations, combinations, and substitutions of one or more limitations, elements, clauses, explanatory terms, etc., from one or more claims or relevant parts thereof into another claim. For example, any claim that depends on another claim may be modified to encompass one or more limitations found in any other claim that depends on the same basic claim. Furthermore, where a claim describes a composition, it should be understood that, unless otherwise indicated or unless it is obvious to a person skilled in the art that such a modification would result in a contradiction or inconsistency, methods of using the composition for any of the purposes disclosed herein and methods of making the composition according to any of the methods disclosed herein or other methods known in the art are also encompassed.

[0160] Where elements are presented as a list (for example, in Markush group form), each subgroup of the element is also disclosed, and it should be understood that any element(s) may be removed from the group. The term “includes” is intended to be open, and it should also be noted that it allows for the inclusion of additional elements or steps. In general, where it is said that the present invention or an aspect of the present invention includes certain elements, features, steps, etc., it should be understood that certain aspects of the present invention or aspects of the present invention consist of or are essentially such elements, features, steps, etc. For the sake of simplicity, these aspects have not been specifically shown herein one by one. Therefore, for each aspect of the present invention that includes one or more elements, features, steps, etc., the present invention also provides aspects that consist of or are essentially such elements, features, steps, etc.

[0161] Where a range is given, the endpoints are encompassed. Furthermore, unless otherwise indicated or otherwise obvious from the context and / or the understanding of those skilled in the art, the values ​​expressed as a range may take any specific value within the range described in different embodiments of the invention, up to 1 / 10 of the lower limit unit of the range, unless the context explicitly states otherwise. Unless otherwise indicated or otherwise obvious from the context and / or the understanding of those skilled in the art, the values ​​expressed as a range may take any subrange within a given range, and the endpoints of subranges may be expressed with the same precision as 1 / 10 of the lower limit unit of the range.

[0162] In addition, it should be understood that any particular aspect of the present invention may be explicitly excluded from any one or more of the claims. Where a range is given, any value within that range may be explicitly excluded from any one or more of the claims. Any aspect, element, feature, use, or aspect of any composition and / or method of the present invention may be excluded from any one or more of the claims. For the sake of brevity, not all aspects in which one or more elements, features, uses, or aspects are excluded are explicitly shown herein.

[0163] table Table 1. Cellular modifications induced by Cas9:CLTA4 sgRNA. Thirty-three human genomic DNA sequences were identified that were enriched in in vitro selection with Cas9:CLTA4 v2.1 sgRNA under enzyme restriction or enzyme saturation conditions. Underlined sites contain insertions or deletions (indels) consistent with significant Cas9:sgRNA-mediated modifications in HEK293T cells. In vitro enrichment values ​​for selection with Cas9:CLTA4 v1.0 sgRNA or Cas9:CLTA4 v2.1 sgRNA are shown for sequences with three or fewer mutations. Due to the small number of in vitro selected sequence counts, enrichment values ​​were not calculated for sequences with four or five or more mutations. Modification frequency (number of sequences with indels divided by the total number of sequences) in HEK293T cells treated with Cas9 without sgRNA ("no sgRNA"), Cas9 with CLTA4 v1.0 sgRNA, or Cas9 with CLTA4 v2.1 sgRNA. Sites where the P-value shows significant modification in cells treated with v1.0 sgRNA or v2.1 sgRNA compared to cells treated with Cas9 without sgRNA are listed. "Not tested (nt)" indicates that PCR of the genome sequence did not provide specific amplification products.

[0164] Table 2: Raw selected sequence counts. Positions -4 to -1 are the four nucleotides immediately preceding the 20-base pair target site. PAM1, PAM2, and PAM3 are PAM positions immediately following the target site. Positions +4 to +7 are the four nucleotides immediately following the PAMs.

[0165] Table 3: CLTA1 genome off-target sequences. Twenty human genomic DNA sequences enriched in Cas9:CLTA1 v2.1 sgRNA in vitro selection under enzyme restriction or enzyme overload conditions were identified. "m" indicates the number of mutations from the on-target sequence that have mutations indicated in lowercase. Underlined sites contain insertions or deletions (indels) consistent with significant Cas9:sgRNA-mediated modifications in HEK293T cells. Human genomic coordinates are shown for each site (assembly GRCh37). CLTA1-0-1 is present at two loci, and sequence counts were pooled from both loci. Sequence counts are shown for amplified and sequenced DNA from HEK293T cells treated with Cas9 without sgRNA ("no sgRNA"), Cas9 with CLTA1 v1.0 sgRNA, or Cas9 with CLTA1 v2.1 sgRNA.

[0166] Table 4: CLTA4 genome off-target sequences. Thirty-three human genomic DNA sequences were identified that were enriched in Cas9:CLTA4 v2.1 sgRNA in vitro selection under enzyme restriction or enzyme overload conditions. "m" indicates the number of mutations from the on-target sequence that have mutations indicated in lowercase. Underlined sites contain insertions or deletions (indels) that are consistent with significant Cas9:sgRNA-mediated modifications in HEK293T cells. Human genomic coordinates are shown for each site (assembly GRCh37). Sequence counts are shown for each site for amplified and sequenced DNA from HEK293T cells treated with Cas9 without sgRNA ("no sgRNA"), Cas9 with CLTA4 v1.0 sgRNA, or Cas9 with CLTA4 v2.1 sgRNA.

[0167] Table 5: Genomic coordinates of CLTA1 and CLTA4 off-target sites. Fifty-four enriched human genomic DNA sequences were identified during in vitro selection of Cas9:CLTA1 v2.1 sgRNA and Cas9:CLTA4 v2.1 sgRNA under enzyme restriction or enzyme excess conditions. Human genomic coordinates are shown for each site (assembly GRCh37).

[0168] Table 6: Modification of cells induced by Cas9:CLTA1 sgRNA. Twenty human genomic DNA sequences enriched in the in vitro selection of Cas9:CLTA1 v2.1 sgRNA under enzyme restriction or enzyme excess conditions were identified. Sites indicated by underlines contain insertions or deletions (indels) that coincide with substantial Cas9:sgRNA-mediated modifications in HEK293T cells. In vitro enrichment values for selection by Cas9:CLTA1 v1.0 sgRNA or Cas9:CLTA1 v2.1 sgRNA are shown for sequences with three or fewer mutations. Due to the small number of in vitro selection sequences, enrichment values were not calculated for sequences with four or more mutations. Modification frequencies (number of sequences with indels divided by the total number of sequences) in HEK293T cells treated with Cas9 without sgRNA (“no sgRNA”), Cas9 with CLTA1 v1.0 sgRNA, or Cas9 with CLTA1 v2.1 sgRNA. P-values for sites showing substantial modification in cells treated with v1.0 sgRNA or v2.1 sgRNA compared to cells treated with Cas9 without sgRNA were 1.1E-05 (v1.0) and 6.9E-55 (v2.1) for CLTA1-0-1, 2.6E-03 (v1.0) and 2.0E-10 (v2.1) for CLTA1-1-1, and 4.6E-08 (v2.1) for CLTA1-2-2. P-values were calculated using a one-sided Fisher's exact test. “Not tested (n.t.)” indicates that the site was not tested or that PCR of the genomic sequence did not provide a specific amplification product.

[0169] Table 7: CLTA1 genomic off-target indel sequences. Insertion and deletion-containing sequences from treated HEK293T cells by DNA amplified and sequenced for the on-target genomic sequence (CLTA1-0-1) and each modified off-target site from cells treated with Cas9 without sgRNA (“no sgRNA”), Cas9 with CLTA1 v1.0 sgRNA, or Cas9 with CLTA1 v2.1 sgRNA. “ref” refers to the human genomic reference sequence for each site, and the modified sites are listed below. Mutations relative to the on-target genomic sequence are shown in lowercase. Insertions and deletions are indicated by underlined bold or dash, respectively. The percentage of modification is shown for conditions (v1.0 sgRNA or v2.1 sgRNA) that show a statistically significant enrichment of the modified sequences compared to the control (“no sgRNA”).

[0170] Table 8: CLTA4 genomic off-target indel sequences. Insertion and deletion-containing sequences from treated HEK293T cells by DNA amplified and sequenced for the on-target genomic sequence (CLTA4-0-1) and each modified off-target site from cells treated with Cas9 without sgRNA (“no sgRNA”), Cas9 with CLTA4 v1.0 sgRNA, or Cas9 with CLTA4 v2.1 sgRNA. “ref” refers to the human genomic reference sequence for each site, and the modified sites are listed below. Mutations relative to the on-target genomic sequence are shown in lowercase. Insertions and deletions are indicated by underlined bold or dash, respectively. The percentage of modification is shown for conditions (v1.0 sgRNA or v2.1 sgRNA) that show a statistically significant enrichment of the modified sequences compared to the control (“no sgRNA”).

[0171] Table 9: Oligonucleotides used in this study. All oligonucleotides were purchased from Integrated DNA Technologies. An asterisk (*) indicates that the preceding nucleotide was incorporated as a hand-mixed phosphoramidite consisting of 79 mol% of the phosphoramidite corresponding to the preceding nucleotide and 4 mol% each of three other standard phosphoramidites. " / 5Phos / " indicates the 5' phosphate group attached during synthesis. [Table 1] [Table 2-1] [Table 2-2] [Table 3] [Table 4-1] [Table 4-2] [Table 5] [Table 6-1] [Table 6-2] [Table 7-1] [Table 7-2] [Table 8-1] [Table 8-2]

Table 9-1

Table 9-2

Table 9-3

Table 9-4

Table 9-5

Table 10-1

Table 10-2

[0172] All publications, patents, and sequence database entries (including the items listed above) cited in this specification are hereby incorporated by reference in their entirety, as if each individual publication or patent were specifically and individually indicated to be incorporated by reference. In the case of a conflict, the present application (including any definitions in this specification) controls.

Claims

1. A library of nucleic acid molecules comprising multiple nucleic acid molecules, wherein each nucleic acid molecule comprises a concatemer of a repeat unit sequence having a candidate nuclease target site and a constant insertion sequence, wherein the repeat unit sequence has the structure: [(NZ)-(PAM)-(constant insertion sequence)], where N independently represents any nucleotide, Z is an integer from 10 to 50, and the candidate nuclease target site can be cleaved by a Cas9 nuclease.

2. The library according to claim 1, wherein the constant insertion sequence has at least 15 nucleotides and a nucleotide length of 70 or less.

3. The library according to claim 1, wherein the constant insertion sequence has a length of at least 10 nucleotides and 60 nucleotides or less.

4. A library according to any one of claims 1 to 3, comprising at least 105, at least 106, at least 107, at least 108, at least 109, at least 1010, at least 1011, or at least 1012 different candidate nuclease target sites.

5. A library according to any one of claims 1 to 4, comprising nucleic acid molecules having a molecular weight of at least 0.5 kDa, at least 1 kDa, at least 2 kDa, at least 3 kDa, at least 4 kDa, at least 5 kDa, at least 6 kDa, at least 7 kDa, at least 8 kDa, at least 9 kDa, at least 10 kDa, at least 12 kDa, or at least 15 kDa.

6. A library according to any one of claims 1 to 5, comprising candidate nuclease target sites which are variations of known target sites of Cas9.

7. The library according to claim 6, wherein the deformations of known target sites of Cas9 include 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, or 2 or fewer mutations compared to known target sites of Cas9.

8. The library according to claim 6 or 7, wherein the deformation differs on average from known target sites of Cas9 by more than 5%, more than 10%, more than 15%, more than 20%, more than 25%, or more than 30% (binomial distribution).

9. The library according to any one of claims 6 to 8, wherein the deformation differs on average from known target sites of Cas9 by 10% or less, 15% or less, 20% or less, 25% or less, 30% or less, 40% or less, or 50% or less (binomial distribution).

10. The library according to any one of claims 6 to 9, wherein the candidate nuclease target site is the target site of a Cas9 nuclease derivative.

11. The library according to any one of claims 1 to 10, wherein Z is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30.

12. The library according to claim 11, wherein the ZsgRNA complementary sequence is 20.

Citation Information

Patent Citations

  • Evaluation and improvement of nuclease cleavage specificity

    WO2013066438A2