Compositions and methods for improving specificity of genome engineering using rna-guided endonucleases

By optimizing guide RNA (gRNA) design and using nuclease-deficient Cas9 fusion protein, the off-target binding and cleavage problems of the CRISPR/Cas9 system are solved, the specificity and safety of genome editing are improved, and it is suitable for genome editing and gene expression regulation.

CN114634930BActive Publication Date: 2025-10-21DUKE UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202210083840.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-08-25
Filing Date
2016-08-25
Publication Date
2025-10-21
Estimated Expiration
2036-08-25

AI Technical Summary

Technical Problem

The existing CRISPR/Cas9 system has problems with off-target binding and cleavage in genome engineering, which affects its potential use in practical applications. It is necessary to improve the specificity of nucleases to reduce off-target binding and increase targeting specificity.

Method used

By optimizing the design method of guide RNA (gRNA), including identifying the target region, determining the full-length gRNA sequence, identifying off-target sites, calculating invasion kinetics and lifespan, randomizing linker nucleotides to improve gRNA binding specificity, and testing the optimized gRNA in vivo, it is combined with nuclease-deficient Cas9 fusion protein for genome editing and regulation.

Benefits of technology

It significantly improves the targeting specificity of the CRISPR/Cas9 system, reduces off-target binding and cleavage, and enhances the accuracy and safety of genome editing. It is suitable for genome editing, epigenome editing, and gene expression regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114634930B_ABST
    Figure CN114634930B_ABST
Patent Text Reader

Abstract

The present invention relates to compositions and methods for improving the specificity of genome engineering using RNA-guided endonucleases. Disclosed herein are optimized guide RNAs (gRNAs) and methods of designing and using the same with increased target binding specificity and reduced off-target binding.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 201680061639.8, filed on August 25, 2016, entitled "Compositions and methods for improving genome engineering specificity using RNA-guided endonucleases."

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No. 62 / 209,466, filed August 25, 2015, which is incorporated herein by reference in its entirety.

[0004] Statement of Government Interest

[0005] This invention was made with government support under Federal Grant Nos. MCB1244297 and CBET1151035 awarded by the National Science Foundation and F32GM11250201, R01DA036865, and DP2OD008586 awarded by the National Institutes of Health. The government has certain rights in this invention. Technical Field

[0006] The present disclosure relates to methods of optimizing guide RNAs (gRNAs) and designing and using such gRNAs with increased target binding specificity and reduced off-target binding. Background Art

[0007] RNA-guided endonucleases, particularly the protein Cas9, have been hailed as potential "perfect genome engineering tools" because they can be directed by a single 'guide RNA' molecule to cleave DNA of virtually any sequence. This capability has recently been exploited in many emerging biological and medical applications, generating tremendous excitement and hope for its future use. However, practical genome engineering requires extremely precise control of the ability to selectively target and cleave precise DNA sequences to avoid inadvertent damage and mutation of off-target DNA.

[0008] Cas9 is a prokaryotic type II CRISPR (clustered, regularly interspaced, short palindromic repeats) endonuclease—a CRISPR-associated (Cas) response to invading foreign DNA. During this response, Cas9 is first bound by a CRISPR RNA (crRNA): transactivating crRNA (tracrRNA) duplex, which then directs the cleavage of DNA containing a 20-base pair (bp) 'protospacer' site that is complementary to a variable 20-bp segment of the crRNA (Figure 1A). Following the binding of a single guide RNA (sgRNA), the Cas9-sgRNA complex binds to the 20-bp 'protospacer' sequence in the target DNA, provided that the protospacer is directly followed by a protospacer adjacent motif (PAM, here 'TGG'). Following binding, the Cas9 endonuclease creates a double-strand break (triangle) within the protospacer. Essentially, the only restriction on the sequences that Cas9 can target is that a short protospacer adjacent motif (PAM), such as 'NGG' in the case of S. pyogenes Cas9, must immediately follow the protospacer site in the foreign DNA molecule. Analysis of crystallographic and biochemical experiments suggests that specificity for protospacer binding and cleavage is conferred by first recognition of the PAM site by the Cas9 protein itself, followed by strand invasion by the bound RNA complex and direct Watson-Crick base pairing with the protospacer. Figure 1 A).

[0009] After the CRISPR-Cas9 system was redesignated for many heterologous biotechnology applications, Cas9's ability to be modularly 'programmed' to target almost any DNA by a single RNA hairpin has recently generated a huge stimulus. Notably, a single-guide RNA (sgRNA) hairpin was designed that combines the essential components of the crRNA:tracrRNA duplex into a single functional molecule. With this sgRNA, Cas9 can be introduced into various organisms to generate targeted double-strand breaks in vivo for significantly simpler genome engineering. Nuclease-free Cas9 (D10A / H840A, referred to as 'dCas9') and chimeric dCas9 derivatives have also been used to alter gene expression in vivo and introduce targeted epigenetic modifications via targeted binding at or near promoter sites.

[0010] Off-target binding and cleavage by Cas9 is of concern because it can adversely affect its potential use in practice. Significant efforts have been made to improve the specificity of Cas9 / dCas9 activity. First, the most extensive efforts were primarily accomplished by intelligently selecting target sequences without selecting similar other sequences in the genome, but recent investigations have found that these methods perform poorly in their ability to predict off-target cleavage. In addition, efforts have been made to directly engineer the protein itself by introducing point mutations that were found to modulate or increase PAM or protospacer binding specificity. Cas9 derivatives that only nick single strands of DNA without performing double-stranded DNA cleavage ('paired nickases') have also been used in pairs, assuming that the probability of off-target nicks at multiple sites being close enough to each other to produce double-strand breaks will be very low. Finally, some work has been done on preparing guide RNA variants themselves to try to achieve higher specificity. Early efforts to add a 5' extension region of the guide RNA to complement the additional nucleotides beyond the protospacer did not show increased Cas9 cleavage specificity in vivo. Instead, they were digested back to approximately their standard length ( Figure 1 A) For applications in genome engineering, especially therapeutic applications, extreme specificity of gene targeting is required to prevent off-target DNA from being destroyed and causing unauthorized mutations. However, there have been several reports of off-target binding and cleavage by Cas9, which can adversely affect its potential use in practice.

[0011] There remains a need to reduce off-target binding and increase nuclease specificity when using the CRISPR / Cas9 system. Summary of the Invention

[0012] The present invention relates to a method for generating optimized guide RNA (gRNA). The method comprises: a) identifying a target region of interest, wherein the target region of interest comprises a protospacer sequence; b) determining a polynucleotide sequence of a full-length gRNA targeting the target region of interest, wherein the full-length gRNA comprises a protospacer targeting sequence or segment; c) determining at least one or more off-target sites of the full-length gRNA; d) generating a polynucleotide sequence of a first gRNA, wherein the first gRNA comprises the polynucleotide sequence of the full-length gRNA

[0013] The present invention relates to a method for producing an optimized guide RNA (gRNA). The method comprises: a) identifying a target region of interest, the target region of interest comprising a protospacer sequence; b) determining a polynucleotide sequence of a full-length gRNA targeting the target region of interest, the full-length gRNA comprising a protospacer targeting sequence or segment; c) determining at least one or more off-target sites of the full-length gRNA; d) generating a polynucleotide sequence of a first gRNA, the first gRNA comprising the polynucleotide sequence of the full-length gRNA and an RNA segment, the RNA segment comprising a polynucleotide sequence having a length of M nucleotides, the polynucleotide sequence being complementary to a nucleotide segment of the protospacer targeting sequence or segment, the segment being located at the 3' end of the polynucleotide sequence of the full-length gRNA, the first gRNA optionally comprising a linker between the 3' end of the polynucleotide sequence of the full-length gRNA and the RNA segment, the linker comprising a polynucleotide sequence having a length of N nucleotides, the first gRNA being capable of invading the protospacer sequence and binding to a DNA sequence complementary to the protospacer sequence and forming a protospacer-duplex, and the first gRNA being capable of invading off-target sites. site and binds to a DNA sequence complementary to the off-target site and forms an off-target duplex; e) calculating an estimate of invasion kinetics and a lifetime that the first gRNA remains invaded in the protospacer and off-target site duplex or computationally simulating these invasion kinetics and lifetime, wherein the invasion dynamics are estimated nucleotide by nucleotide by determining the energy difference between further invasion by a different gRNA and reannealing of the first gRNA to the DNA sequence complementary to the protospacer sequence; f) comparing the estimated lifetime of the first gRNA at the protospacer and / or off-target site with the estimated lifetime of the full-length gRNA or truncated gRNA (tru-gRNA) at the protospacer and / or off-target site; g) randomizing 0 to N nucleotides in the linker and 0 to M nucleotides in the first gRNA and generating a second gRNA, and repeating step (e) with the second gRNA; h) identifying an optimized gRNA based on the gRNA sequence that meets the design criteria; and i) testing the optimized gRNA in vivo to determine binding specificity.

[0014] The present invention relates to an optimized gRNA produced by the above method.

[0015] The present invention relates to an isolated polynucleotide encoding the above-mentioned optimized gRNA.

[0016] The present invention relates to a vector comprising the isolated polynucleotide.

[0017] The present invention relates to a cell comprising the isolated polynucleotide or the vector.

[0018] The present invention relates to a kit comprising the isolated polynucleotide, the vector or the cell.

[0019] The present invention relates to a method for epigenomic editing in a target cell or subject. The method comprises contacting the cell or subject with an effective amount of the optimized gRNA molecule described above and a fusion protein comprising a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0020] The present invention relates to a method for performing site-specific DNA cleavage in a target cell or subject. The method comprises contacting the cell or subject with an effective amount of the above-described optimized gRNA molecule and a fusion protein or Cas9 protein, wherein the fusion protein comprises a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0021] The present invention relates to a method for performing genome editing in a cell. The method comprises administering to the cell an effective amount of the above-mentioned optimized gRNA molecule and contacting it with a fusion protein, wherein the fusion protein comprises a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0022] The present invention relates to a method for regulating gene expression in a cell. The method comprises contacting the cell with an effective amount of the above-described optimized gRNA and a fusion protein, wherein the fusion protein comprises a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 A shows a schematic diagram of Cas9 activity.

[0024] Figure 1 B shows an atomic force microscopy (AFM) image of dCas9-sgRNA bound at a protospacer sequence within a single streptavidin-labeled DNA molecule derived from the human AAVS1 locus.

[0025] Figure 1 C-1D shows the DNA substrate derived from AAVS1 that was designed to have a series of fully complementary and partially complementary protospacer sequences ( Figure 1 C) or the fraction of bound DNA occupied by Cas9 / dCas9-sgRNA for engineered DNA substrates (Figure 1D). The vertical lines represent the (23 bp) segment where each significant feature is located on the corresponding substrate.

[0026] Figures 2A-2D show the regulation of binding affinity and specificity of guide RNA variants. Figure 2A shows a schematic diagram of dCas9 binding to a single guide RNA with two nucleotide truncations at its 5' end (tru-gRNA, purple). Figure 2B shows a schematic diagram and proposed mechanism of dCas9 binding to a single guide RNA with a 5' extension region that forms a hairpin (hp-gRNA, blue) with a PAM distal binding segment of its targeting region. Figure 2C shows the unit site binding affinity (K) of dCas9 with tru-gRNA (purple, n=257) along an engineered DNA substrate. A )(See Figure 1 D). The dotted line shows the unit site affinity of dCas9-sgRNA for comparison. Figure 2D shows the unit site binding affinity (K) of dCas9 with guide RNA. A ), these guide RNAs have 5'-hairpins that overlap with nucleotides complementary to the last six (hp6-gRNA, blue) or ten (hp10-gRNA, green) PAM-distal nucleotides of the protospacer.

[0027] Figure 3A-3D shows that as it binds to sites that increasingly match protospacer sequences, Cas9 undergoes progressive conformational transitions. Figure 3A shows the fraction of bound DNA occupied by Cas9 / dCas9 along the DNA substrate, where the colors represent the Cas9 / dCas9 populations clustered according to their structure (by mean square error after alignment, see text). The different feature markers for site-specific analysis of Cas9 / Cas9 structural properties on DNA are: non-specific sequence (α; '20MM'), sites containing 10 PAM distal mismatches within the protospacer (β, '10MM'), sites containing 5 PAM distal mismatches within the protospacer (γ, '5MM'), or complete protospacer sites (δ or ε for dCas9 or Cas9, respectively; '0MM'). The overall average of the main clusters is shown in Figure 3C and color-coded according to the cluster structure it represents. Figure 3B shows the volume and height of the observed Cas9 / dCas9, color-coded according to the cluster to which each protein is assigned. Dashed lines delineate regions likely consisting of aggregates adsorbed near DNA (upper right) or streptavidin labels (lower left). For comparison, average height of streptavidin end labels: 0.92 nm ± 0.006 nm (SEM); average volume of streptavidin end labels: 0.110 x 10 4 nm 3 ±0.002x 10 4 nm 3 (SEM); n=1941. Figure 3D shows the average volume and height of Cas9 / dCas9 with sgRNA (red circles, red labels for Cas9 and blue labels for dCas9) or tru-gRNA (purple circles) bound to each feature of the substrate. It should be noted that dCas9 with tru-gRNA is only expected to interact with the first 3 or 8 (labeled herein as "3MM" and "8MM") PAM-distal mismatches at the 5MM and 10MM sites. For standard errors of the mean volume and height, see Table 2. For Cas9 / dCas9 with sgRNA, its structural properties at each feature are statistically different (δ-ε, α-ε: p < 0.05; α-β: p < 0.005; β-γ, γ-δ: p < < 0.0005. Hotelling's T 2 test).

[0028] Figures 4A-4D show Kinetic Monte Carlo (KMC) experiments that reveal differences in the stability of the R-loop, or structure formed by the protospacer duplex and the invading guide RNA, within stably bound Cas9 for different guide RNA variants. Figure 4A shows a schematic diagram of strand invasion of the protospacer (green) by a guide RNA (red) for the KMC experiment. The R-loop is highlighted. The transition rate of invasion (the rate v for m→m+1) is f , where m is the degree of strand invasion or equivalently the length of the R-loop) or the turnover rate of duplex reannealing (for a rate v of m → m-1 r ) is a function of the nearest neighbor DNA:DNA and RNA:DNA hybridization energies. See the main text and Supplementary Methods for details. Figure 4B shows the fractional time that the R loop has a size m for sg-RNA (red) or tru-gRNA (purple) derived from the 'equilibrium' KMC experiment (the simulation starts with m = 20 or 18, respectively). The simulation was run until t ≥ 10,000 (arbitrary units). Figure 4C shows the dynamic Monte Carlo time course of 'R loop breathing' for sgRNA (red) and tru-gRNA (purple) after complete invasion (the simulation starts with m = 20 or 18, respectively). The asterisk highlights the starting position of the simulation. (Inset) Histogram of the corresponding lifetimes of R loops ≥ 16 bp long. Figure 4D A proposed model for the mechanism controlling Cas9 / dCas9 specificity, based on results from AFM imaging and kinetic Monte Carlo (KMC) experiments (see text), is shown. Cas9 / dCas9 binds to the PAM and the guide RNA invades the PAM-adjacent protospacer duplex. During this strand invasion process, the guide RNA must displace the complementary strand of the protospacer. Competition between invasion of the duplex and reannealing results in a dynamic ('breathing') R-loop structure. The stability of sites 14-17 of the protospacer-guide RNA interaction is significantly increased by binding at sites 19 and 20, promoting conformational changes in Cas9 / dCas9, thereby authorizing DNA cleavage by Cas9.

[0029] Figures 5A-5C show that kinetic Monte Carlo (KMC) experiments reveal that the ability to cross a mismatch (MM) and invade a protospacer differs depending on the guide RNA structure. Figures 5A-5B The fractional occupancy of the R-loop length m of sgRNA ( FIG. 5A ) or tru-gRNA ( FIG. 5B ) during the invasion process (starting at m = 10, highlighted by asterisks) from a KMC experiment is shown over time. White Xs indicate the positions of mismatches. The simulation was run until t ≥ 10,000 (arbitrary units) and the results of 100 trials were averaged. Figure 5C A representative KMC time course of strand invasion is shown (starting at m = 10), with the mismatch site at m = 14 (arrows) for sgRNA (red) and tru-gRNA (purple). While sgRNAs invade robustly after bypassing mismatches, tru-gRNAs are repeatedly recaptured after mismatches due to the inherent volatility of their R-loops (see Figure 4).

[0030] Figures 6A-6B Experimental cleavage frequencies at target sites containing single rG·dG, rC·dC, rAdA, and rU·dT mismatches in the PAM distal region (≥10th protospacer site) (Hsu et al. (2013) Nature biotechnology, 31, 827-832) are shown to be correlated with R-loop stability determined from kinetic Monte Carlo experiments. Figure 6A shows the logarithm of the correlation between Cas9 cleavage frequency at site m and R-loop stability (the fraction of time the guide RNA remains bound to the protospacer at site m, see text) during strand invasion initiated at site m. 10 (p-value). (i) Stability at sites m=10 to m=14 is highly inversely correlated with the likelihood that the guide RNA will fall off the protospacer before crossing the mismatch ( Figure 15 ), while (ii) sites m = 14 to m = 17 are associated with conformational changes that induce cleavage activity (from AFM images). Colors correspond to correlation coefficients. Figure 6B shows that the experimental cleavage frequencies do not correlate with the estimated guide RNA-protospacer equilibrium binding free energy (ΔG° 37 (left) is significantly correlated with the stability of position m=14 during strand invasion (right). Error bars are the standard error of the mean occupancy time at position m=14. For these kinetic Monte Carlo experiments, max(t) = 100 (arbitrary units). Color bars are used to indicate the position of mismatch (MM) sites.

[0031] Figures 7A-7CThe structure of guide RNA affects the overview of the proposed mechanism of Cas9 / dCas9 specificity. Figure 7 A shows that for single guide RNA (sgRNA), the first few nucleotides of RNA (the 18-20 sites of its combination prototype spacer) stabilize the R ring breathing and the combination at the 14-17 sites of the prototype spacer, realizing efficient conformational conversion to active state to allow cracking. However, this stability increase given by these bases allows transient stability at the mismatch site and conformational changes that allow cracking. In many cases, if a mismatch is passed through, the R ring can remain stably invaded completely. Figure 7 B shows that for the guide RNA (tru- gRNA) of the first few (here 2) nucleotides of truncation, the stability reduction of the R ring (characterized by significant volatility) reduces the possibility of maintaining active conformation. When there is a mismatch site in the prototype spacer, the volatility of the R ring ensures that it will be quickly and repeatedly "recaptured" after the mismatch, and is greatly hindered at these sites. Figure 7C shows that although the 5'-end "simple" extension of the guide RNA used to target the protospacer and the adjacent site beyond the protospacer was found to be digested back to approximately the sgRNA length in vivo (Figure 7A), the guide RNA (hp-gRNA) with a 5'-hairpin complementary to the "PAM distal" targeting segment is expected to remain protected within the structure of Cas9 / dCas9 before invasion. After binding to the PAM site and initiating chain invasion by the hp-gRNA, the hairpin opens and complete chain invasion can occur after binding the complete protospacer. If there is a PAM distal mismatch at the target site, it is energetically more favorable for the hairpin to remain closed and chain invasion to be hindered. The ability of Cas9-hp-gRNA to cleave RNA remains to be verified.

[0032] Figures 8A-8B The purity of expressed Cas9 and dCas9 in SDS gels of purified Cas9 ( FIG. 8A ) and dCas9 ( FIG. 8B ) products (nominal molecular weight: 160 kDa) is shown. The eluted band shows a product purity of approximately 95%.

[0033] Figures 9A-9C Additional images of Cas9 / dCas9 bound to DNA are shown. A) Distribution of dCas9 binding to substrates that do not share homology with the AAVs1 protospacer sequence ( Figure 1 Comparison) (n = 443). Overlay is the cumulative distribution (CDF) of PAM sites (CDF PAM , black) and CDF of bases bound by dCas9 (red, CDF Cas9Comparisons were started at 100 bases at each end to avoid artifacts introduced by overlap with the streptavidin tag (a DNA selection criterion) and binding to exposed blunt ends of DNA (resulting in an expected increase in nonspecific binding). B) Absolute difference between protein-bound CDF and PAM site CDF D n The dashed line is the Kolmogorov-Smirnov criterion for goodness of fit of the two distributions. C) Comparison of the CDF of the combination of 100,000 randomly generated sequences with the same probabilities G, A, T, and C with the CDF of the PAM distribution using MATLAB. The vertical red line is the experimental Sup(D n ), indicating that the experimental dCas9 binding more closely matches the experimental PAM distribution than its 71.20% match to the generated sequence.

[0034] Figures 10A-10C Binding of "nonsense" substrates with no homology (>3 bp) to the protospacer sequence is shown. (A) Image of dCas9 alone. (B) Histogram of volume (left) and height (right) of dCas9 imaged alone (n=423), using a Gaussian fit to the original peaks. From the Gaussian fit, it can be seen that the average height is 1.746 nm (95% confidence: 1.689 nm-1.802 nm), with a standard deviation of 0.441 nm, and the average volume is 1302 nm 3 (95% confidence: 1266nm 3 -1337nm 3 ), with a standard deviation of 259.1 nm 3 (Note that because dCas9 here has no DNA within its binding channel, their recorded volumes may appear artificially low due to reduced mechanical resistance to the AFM probe.) Heights were measured relative to the median of a 10-pixel region surrounding each protein, and volumes were recorded as continuous features that were twice the standard deviation of the local background height. (C) Additional representative images of dCas9 bound to DNA that had been labeled at one end with monovalent streptavidin.

[0035] Figures 11A-11DRepresentative images of dCas9-sgRNA bound to RNA and examples of protein structural characterization are shown. Figure 11A shows a representative wide-field image of dCas9 bound to engineered DNA. Figure 11B shows a close-up of the framed area. The white arrow is monovalent streptavidin, and the red arrow is the dCas9 protein. Figures 11C-11D show examples extracted from the original image (Figure 11C) and examples of separation of Cas9 / dCas9 structures (Figure 11D). This extraction was repeated for each isolated protein bound to DNA, and then aligned in pairs by iterative translation, rotation, and reflection to minimize their mean squared topological differences. Based on these minimized mean squared errors, a distance matrix was constructed, and each protein was clustered according to the method of Laio and Rodriguez (2014) Science (New York, NY), 344, 1492-1496, and then the structural population was mapped back to its site on DNA according to the cluster (Figure 2A, Figures 10A-10C ).

[0036] Figures 12A-12B The properties of Cas9 / dCas9-sgRNAs mapped to corresponding binding sites are shown. Top: Stacked histograms of volume (left), maximum height (middle), and structure (clustered by mean square error) after alignment (right, see text) for all experimental conditions. As in the scatter plots below, the populations are colored according to the binned volume, height, or structure clusters. Binding distribution of extracted Cas9 / dCas9 molecules ( Figures 10A-10C ) and the entire dataset ( Figure 1 C-ID, Figures 8A-8B ), indicating that the selection procedure is unbiased and that the selected proteins are representative of the entire dataset. Below: Scatter plot of volume versus maximum height for all Cas9 / dCas9s, color-coded according to grouping by (left) volume, (middle) maximum height, and (right) structural clustering.

[0037] Figure 13Structural properties of Cas9 / dCas9 with tru-gRNA and hp-gRNA at corresponding binding sites are shown. The bound DNA fraction occupied by Cas9 / dCas9 along the engineered DNA substrate, where the color represents the Cas9 / dCas9 population clustered according to its structure (see Figure 3C). The protein structure is classified according to its closest similarity (by mean square error after alignment, see text) with sgRNA dCas9 / Cas9. For reference, on the engineered DNA substrate, the position of the complete prototype spacer site: 144-167bp; the position of the 10MM (8MM) site: 452-465bp; the position of the 5MM (3MM) site: 592-610bp. Similar trends were observed with dCas9 / Cas9 with sgRNAs: as dCas9 bound to sites with increasingly mismatched targets, the fraction of populations clustering with the largest (yellow) group increased, and while this effect was suppressed with tru-gRNAs, a substantial fraction of the population clustered with smaller (green and blue) populations even at full protospacer sites. The effect was particularly pronounced for hp10-gRNA, emphasizing its poor affinity for off-target sites.

[0038] Figures 14A-14C Figure 14A shows a schematic model of the chain invasion of guide RNA into DNA prototype spacers, and the estimated binding stability of RNA invasion into prototype spacers with PAM distal mismatches. Figure 14A shows a schematic model of the chain invasion of guide RNA into DNA prototype spacers. See also Figure 4A. It is assumed that the guide RNA dissociates at m=1. Figure 14B shows the calculated probability distribution of the dissociation time of guide RNAs for initial invasion of prototype spacers with different numbers of consecutive PAM distal mismatches to m=5. The length of these dissociation times can be regarded as an approximation of the dCas9 binding tendency at these sites. The asterisks highlight the dissociation time of the guide RNA population that was initially not fully invaded after the initial invasion to m=5. The invading RNA is highly unstable at the prototype spacer site with 15 PAM distal mismatches (15MM), and experimentally we rarely observe Cas9 / dCas9 bound at these sites ( Figure 1D). RNAs invading at protospacer sites with 10 or 5 PAM-distal mismatches (10MM and 5MM) were calculated to persist significantly longer (before dissociation) than RNAs invading at the 15MM site, but within orders of magnitude of each other; in AFM experiments we found that their binding propensities were approximately equal and lower than those at the full protospacer site (0MM). The sequence-specific transition rates (v) between the m states were calculated using the Q-matrix method as described previously (Sakmann et al. (1995) Single-channel recording, Springer; 2nd ed.). f and v r , see Supplementary Methods) to calculate the probability density function. Figure 14C shows that an examination of the estimated half-life of RNA-protospacer binding at protospacers with different numbers of PAM distal mismatches indicates that there are roughly three scenarios in which the stability of invading RNA is similar: those with >11 PAM distal mismatches (low stability); those with 3 to 11 PAM distal mismatches (medium stability); and those with <3 PAM distal mismatches (high stability). The results are qualitatively similar to the dCas9 distribution on engineered substrates observed via AFM (Figure 1D).

[0039] Figure 15 The average first passage time of sgRNA and tru-gRNA through the mismatch site during chain invasion is shown. The average first passage time of sgRNA (blue) and tru-gRNA (red) through the mismatch site during chain invasion is simulated (kinetic Monte Carlo method) for different positions of the mismatch site. The error bars are the standard deviations of the recorded first passage times. The sequence of the protospacer (AAVS1 site) is in the box.

[0040] Figures 16A-16BThe correlation between Cas9 cracking frequency (Hsu et al. (2013) Nature Biotechnology, 31,827-832) and the measured value of the R ring stability derived from the kinetic Monte Carlo method are shown. Figure 16 A shows R ring site stability (from kinetic Monte Carlo method, referring to the text) and from Hsu et al. (2013) Nature Science Technology, 31,827-832 The statistical power and intensity of the correlation between the experimental cracking frequency decrease with the increase of simulation length (max (t) = 100 to max (t) = 1000, arbitrary units). This result shows that the kinetics of chain invasion may be an important predictor of off-target cleavage rate. Figure 16 B shows that R ring has the time fraction of m size and the kinetic Monte Carlo test predicts the correlation between the possibility of invasion chain dissociation before passing through mispairing. Binding at sites 10 to 14-15 has a very strong inverse correlation (approximately 0.5-0.85) with the probability of dissociation before crossing the mismatch, while from AFM imaging experiments we found that binding at sites approximately ≥16 is associated with conformational changes in Cas9 / dCas9.

[0041] Figure 17 An overview of Deep-Seq data is shown, comparing on-target activity.

[0042] Figure 18 An overview of the deep sequencing data is shown, comparing increases in specificity.

[0043] Figure 19 Protospacer 1, dystrophin is shown; lane 1 shows GFP control; lane 2 shows full gRNA; lane 3 shows Tru-gRNA 19nt; lane 4 shows Tru-gRNA 18nt; lane 5 shows Tru-gRNA 17nt; lane 6 shows Tru-gRNA 16nt; lane 7 shows Hp-gRNA 4bp; lane 8 shows Hp-gRNA 5bp; lane 9 shows Hp-gRNA 6bp; lane 10 shows Hp-gRNA 7bp; lane 11 shows Hp-gRNA 8bp; and lane 12 shows Hp-gRNA 9bp, hairpin 1 (lane 12, 9nt hp)- (SEQ ID NO: 335), wherein a portion of the hairpin is italicized and the hairpin loop is underlined.

[0044] Figure 20 Protospacer 1, dystrophin, internal loop shown

[0045] Figure 21 shows the calculated secondary structure of the 5' end of the protospacer targeting segment of the hp-gRNA used for deep sequencing experiments (using the NuPack software suite). The color is the probability of each nucleotide being present in that secondary structure at equilibrium.

[0046] Figure 22 Shown are dystrophin, indel rates, all sites.

[0047] Figure 23 Dystrophin, on-target / total (off-target) is shown.

[0048] Figure 24 Protospacer 2, EMX1, is shown; lane 1 shows a GFP control; lane 2 shows a full gRNA; lane 3 shows a Tru-gRNA; lane 4 shows a 10-bp hp-gRNA; and lane 5 shows a 6-bp hp-gRNA, Hairpin 1. Switches - Surv_OT1 = DS_OT2; Surv_OT53 = DS_OT3.

[0049] Figure 25A and 25B Protospacer 2, EMX1, tru-hp, internal loop is shown.

[0050] Figures 26A-26C A hairpin structure is shown. Figure 26A Hairpin 1 is shown, which is a 6 bp 5' hairpin. Figure 26B Hairpin 2 is shown, which is a 5 bp 5' hairpin on the 18 nt (truncated) gRNA. Figure 26C Hairpin 3 is shown, which is a 3 bp 5' hairpin.

[0051] Figure 27 Shown are EMX1, indel rates, all loci.

[0052] Figure 28 Shown are EMX1, indel rate, and low off-target rate.

[0053] Figure 29 Shown are EMX1, on-target / total (off-target).

[0054] Figure 30 Protospacer 3, VEGFA1, is shown. Lane 1 shows a GFP control; lane 2 shows a full gRNA; lane 3 shows a Tru-gRNA; lane 4 shows a 10-bp hp-gRNA; and lane 5 shows a 6-bp hp-gRNA.

[0055] Figure 31 shows protospacer 3, VEGFA1: pam proximal hairpin. Lane 1 shows a GFP control; lane 2 shows a complete gRNA; lane 3 shows hp-gRNA1; lane 4 shows hp-gRNA2; lane 5 shows hp-gRNA3; lane 6 shows hp-gRNA4; lane 7 shows hp-gRNA5; and lane 8 shows hp-gRNA6.

[0056] Figure 32 Protospacer 3, VEGFA1:pam proximal hairpin, is shown.

[0057] Figure 33 Protospacer 3, VEGF1, internal loop are shown. Lane 1 shows control; Lane 2 shows full length; Lane 3 shows 2nt hp; Lane 4 shows 3nt hp, hairpin 5; and Lane 5 shows 4nt hp.

[0058] Figure 34A and 34B Deep sequencing experiments showing hairpins 1, 2, and 3 failed. Figure 25A Shown is hairpin 4 - a computationally derived hairpin designed to discriminate off-target site 2 while maintaining on-target activity. Figure 25B Hairpin 5-4 bp 5'-hairpin is shown (gRNA normally has significant 3' secondary structure).

[0059] Figure 35 VEGF1, indel rate, all sites are shown.

[0060] Figure 36 VEGF1, indel rate, and low off-target rate are shown.

[0061] Figure 37 VEGF1, on-target / total (off-target) are shown.

[0062] Figure 38 Protospacer 4, VEGFA3, is shown. Lane 1 shows a GFP control; lane 2 shows a full gRNA, lane 3 shows a Tru-gRNA; lane 4 shows a 3-bp hp-gRNA; lane 5 shows a 4-bp hp-gRNA; lane 6 shows a 5-bp hp-gRNA; lane 7 shows a 6-bp hp-gRNA; and lane 8 shows a 10-bp hp-gRNA.

[0063] Figure 39 shows gRNA4, VEGFA13: pam proximal hairpin. Lane 1 shows GFP control; Lane 2 shows complete gRNA; Lane 3 shows hp-gRNA1; Lane 4 shows hp-gRNA2; Lane 5 shows hp-gRNA3; Lane 6 shows hp-gRNA4; Lane 7 shows hp-gRNA5; and Lane 8 shows hp-gRNA6.

[0064] Figure 40A Hairpin 1 - a 4 bp hairpin targeting the 3'-region is shown.

[0065] Figure 40B Hairpin 2 - a 4 bp hairpin with a GU wobble base pair targeting the 3'-region is shown.

[0066] Figure 40C Shown is hairpin 3 - a 4 bp hairpin with a GU wobble base pair targeting the 3'-region (variant design).

[0067] Figure 41 VEGF3, indel rate, all sites are shown.

[0068] Figure 42 VEGF3, indel rate, and low off-target rate are shown.

[0069] Figure 43 VEGF3, on-target / total (off-target) are shown.

[0070] Figure 44A Shown is a hairpin designed to target the EMX1 gene.

[0071] Figure 44B Shown Figure 44A The hairpin sequence of EMX1-sg1.

[0072] Figure 44C The effect of decreasing the protospacer length and increasing the hairpin length on specificity is shown.

[0073] Figures 45A-45D DNA / RNA sequences are shown.

[0074] Figure 46 A diagram depicting the Surveyor assay is shown.

[0075] Figure 47Figure 2 shows the tolerance of AsCpf1 and LbCpf1 to mismatched or truncated crRNAs and the modification of endogenous genes of AsCpf1 and LbCpf1 using crRNAs containing a single mismatched base. Activity was determined by T7E1 assay; error bars, sem; n=3 (adapted from Kleinstiver et al., Nature Biotechnology 34:869-875).

[0076] Figure 48 Shown are the results of a surveyor assay for hp-gRNA used with the V-type CRISPR system, in which a hairpin was added to the 3' end of the full-length gRNA to eliminate off-target activity. DETAILED DESCRIPTION

[0077] Disclosed herein are compositions and methods for site-specific DNA targeting and epigenomic gene editing and / or transcriptional regulation, such as DNA cleavage and gene activation or repression. The present invention relates to a modular approach for designing and using optimized guide RNAs (hpgRNAs) with hairpin structures. These optimized guide RNAs can be easily incorporated into existing biotechnology infrastructure and result in controlled reduction of off-target activity while maintaining the ability to specifically target the correct DNA sequence. The methods described herein provide a novel approach for engineering optimized gRNAs that performs significantly better than other available methods and can be used in combination with other protein-specific means to improve specificity to highly improve performance.

[0078] The disclosed methods and optimized gRNAs have the great advantage of being easily adaptable to current methods and infrastructure already in place for RNA-guided genome engineering. In some embodiments, Cas9, dCas9, or Cpf1 are delivered to cells using viral vectors and vectors encoding optimized gRNAs for transcription in cells. The present invention requires only a few additional nucleotides for the vector encoding the optimized gRNA, which can be easily adjusted by current and standard practices. Similar to truncated guide RNAs (tru-gRNAs), optimized gRNAs or hpgRNAs can be used in combination with, for example, paired nickases, or other modifications of the endonuclease itself to further improve specificity. A series of experiments were performed in vitro that showed that the optimized gRNAs produced using the methods described herein increased the specificity of DNA binding relative to the best available gRNA options (see Figure 2). The use of optimized gRNAs eliminated or significantly impaired activity at targets containing only a small amount of mismatched DNA sequences, which are often sites where off-target activity of RNA-guided endonucleases occurs. The optimized gRNAs also provided specificity for cleavage activity at sites known to induce off-target activity in mammalian cells, even in the best known improvements to the guide RNA. The present invention is a general method for reducing the off-target activity of RNA-guided endonucleases, particularly Cas9, by engineering the structural design of guide RNA.

[0079] 1. Definition

[0080] As used herein, the terms "comprise," "include," "having," "has," "may," and "contain" are intended to be open transitional phrases, terms, or words that do not exclude the possibility of additional functions or structures. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates additional embodiments that "comprise," "consist of," and "consist essentially of" the embodiments or elements presented herein, whether or not explicitly stated.

[0081] For the recitation of numerical ranges herein, each number therebetween is expressly contemplated with the same degree of precision. For example, for the range 6-9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0-7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.

[0082] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those of ordinary skill in the art. In the case of conflict, this document (including definitions) will be used as the criterion. Although methods and materials similar to or equivalent to those described herein can be used in the practice or test of this disclosure, preferred methods and materials are described below. All publications, patent applications, patents and other references submitted herein are incorporated by reference in their entirety. These materials, methods and examples disclosed herein are only illustrative and are not intended to be restrictive.

[0083] "Adeno-associated virus" or "AAV," as used interchangeably herein, refers to a small virus belonging to the genus Dependovirus of the family Parvoviridae that infects humans and some other primate species. AAV is not currently known to cause disease and therefore the virus evokes a very mild immune response.

[0084] As used herein, "binding region" refers to a region within a nucleic acid target region that is recognized and bound by a nuclease, such as Cas9.

[0085] As used herein, "chromatin" refers to the organized complex of chromosomal DNA associated with histone proteins.

[0086] As used interchangeably herein, "cis-regulatory element" or "CRE" refers to a non-coding DNA region that regulates the transcription of nearby genes. CREs are present near one or more genes that they regulate. CREs typically regulate gene transcription by serving as binding sites for transcription factors. Examples of CREs include promoters, enhancers, super enhancers, silencers, insulators, and locus control regions.

[0087] "Clustered Regularly Interspaced Short Palindromic Repeats" and "CRISPR," as used interchangeably herein, refer to a locus containing multiple short direct repeats that are present in the genomes of approximately 40% of sequenced bacteria and 90% of sequenced archaea.

[0088] As used herein, "coding sequence" or "encoding nucleic acid" means a nucleic acid (RNA or DNA molecule) comprising a nucleotide sequence that encodes a protein. The coding sequence may further comprise a start signal and a stop signal operably linked to regulatory elements, including a promoter and a polyadenylation signal capable of directing expression of the nucleic acid in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence may be codon-optimized.

[0089] As used herein, "complementary" or "complementary" in reference to nucleic acids can refer to Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to a property shared between two nucleic acid sequences such that when they are aligned antiparallel to each other, the nucleotide bases at every position will be complementary.

[0090] As used herein, "correction", "genome editing" and "restoration" refer to changing a mutant gene that encodes a truncated protein or does not encode a protein at all, so as to obtain the expression of a full-length functional or partially full-length functional protein. Correcting or restoring a mutant gene can include replacing a gene region with a mutation or replacing the entire mutant gene with a gene copy that does not have a mutation using a repair mechanism such as homology-directed repair (HDR). Correcting or restoring a mutant gene can also include repairing a frameshift mutation that causes an early stop codon, an abnormal splicing acceptor site or an abnormal splicing donor site by: generating a double-strand break in the gene and then repairing it using non-homologous end joining (NHEJ). NHEJ can add or delete at least one base pair during the repair process, so that the correct reading frame can be restored and the early stop codon can be eliminated. Correcting or restoring a mutant gene can also include destroying an abnormal splicing acceptor site or a splicing donor sequence. Correcting or restoring a mutant gene can also include deleting a non-essential gene segment by the simultaneous action of two nucleases on the same DNA chain, so that the correct reading frame can be restored by removing the DNA between the two nuclease target sites and repairing the DNA break generated by NHEJ.

[0091] As used herein, "demethylase" refers to an enzyme that removes methyl (CH3-) groups from nucleic acids, proteins (particularly histones) and other molecules. Demethylases are very important in epigenetic modification mechanisms. Demethylase proteins change the transcriptional regulation of the genome by controlling the methylation levels that occur on DNA and histones, and then regulate the chromatin state of specific loci in the organism. "Histone demethylase" refers to a methylase that removes methyl groups from histones. There are several families of histone demethylases that act on different substrates and play different roles in cellular function. Fe (II)-dependent lysine demethylases can be JMJC demethylases. JMJC demethylases are histone demethylases that contain JumonjiC (JmjC) domains. JMJC demethylases can be members of KDM3, KDM4, KDM5 or KDM6 families of histone demethylases.

[0092] "DNase I hypersensitive site" or "DHS," as used interchangeably herein, refers to a docking site for transcription factors and chromatin modifiers, including p300, that coordinate the expression of distal target genes.

[0093] As used interchangeably herein, "donor DNA," "donor template," and "repair template" refer to a double-stranded DNA fragment or molecule that contains at least a portion of a gene of interest. The donor DNA can encode a fully functional protein or a partially functional protein.

[0094] As used herein, "endogenous gene" refers to a gene derived from an organism, tissue, or cell. Endogenous genes are natural to the cell, in a normal genomic and chromatin environment, and are not heterologous to the cell. Such cellular genes include, for example, animal genes, plant genes, bacterial genes, protozoan genes, fungal genes, mitochondrial genes, and chloroplast genes. As used herein, "endogenous target gene" refers to an endogenous gene targeted by an optimized gRNA and a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system.

[0095] As used herein, "enhancer" refers to a non-coding DNA sequence containing multiple activation binding sites and repressor binding sites. The length of the enhancer is in the range of from 50bp to 1500bp, and the enhancer can be at the proximal end, at the 5' upstream of the promoter, in any intron of the regulatory gene, or at the distal end, in the intron of the adjacent gene or in the intergenic region away from the locus, or in the district on different chromosomes. More than one enhancer can interact with the promoter. Similarly, enhancers can regulate more than one gene without connection restrictions, and can "jump" adjacent genes to regulate more distant genes. Transcriptional regulation can involve elements located on a chromosome different from the chromosome where the promoter is located. The proximal enhancer or promoter of an adjacent gene can serve as a platform for raising more distal elements.

[0096] "Duchenne muscular dystrophy" or "DMD," as used interchangeably herein, refers to a recessive, fatal, X-linked disorder that causes muscle degeneration and eventual death. DMD is a common, inherited, single-gene disorder that occurs in 1 in 3,500 males. DMD is caused by inherited or spontaneous mutations that result in nonsense or frameshift mutations in the dystrophin gene. Most dystrophin mutations that cause DMD are exon deletions that disrupt the reading frame and cause premature translation termination of the dystrophin gene. DMD patients typically lose the ability to physically support themselves during childhood, become progressively weaker during their teenage years, and die in their twenties.

[0097] As used herein, "dystrophin" refers to a rod-shaped cytoplasmic protein that is part of a protein complex that connects the muscle fiber cytoskeleton to the surrounding extracellular matrix through the cell membrane. Dystrophin provides structural stability to the cell membrane dystroglycan complex, which is responsible for regulating muscle cell integrity and function. The dystrophin gene, or "DMD gene," as used interchangeably herein, is a 2.2 megabase gene located at locus Xp21. The primary transcript measures approximately 2,400 kb, of which the mature mRNA is approximately 14 kb. 79 exons encode a protein of over 3,500 amino acids.

[0098] As used herein, "exon 51" refers to the 51st exon of the dystrophin gene. In DMD patients, exon 51 is often adjacent to deletions that disrupt the frame and has been targeted for oligonucleotide-based exon skipping in clinical trials. A clinical trial of the exon 51 skipping compound eteplirsen recently reported significant functional benefits across 48 weeks, with an average of 47% dystrophin-positive fibers compared to baseline. Mutations in exon 51 are ideally suited for permanent correction by NHEJ-based genome editing.

[0099] As used interchangeably herein, "frameshift" or "frameshift mutation" refers to a type of genetic mutation in which the addition or deletion of one or more nucleotides causes a shift in the reading frame of a codon in an mRNA. This shift in reading frame may result in changes in the amino acid sequence when the protein is translated, such as missense mutations or premature stop codons.

[0100] "Full-length gRNA" or "standard gRNA," as used interchangeably herein, refers to a gRNA that comprises a "backbone" and protospacer targeting sequence or segment, typically 20 nucleotides in length.

[0101] As used herein, "functional" and "fully functional" describe a protein that has biological activity. A "functional gene" refers to a gene that is transcribed into mRNA, which is translated into a functional protein.

[0102] As used herein, "fusion protein" refers to a chimeric protein produced by joining two or more genes that originally encoded separate proteins. Translation of the fusion gene produces a single polypeptide with functional properties derived from each original protein.

[0103] As used herein, "gene construct" refers to a DNA or RNA molecule comprising a nucleotide sequence encoding a protein. The coding sequence comprises a start signal and a stop signal operably linked to regulatory elements, including a promoter and a polyadenylation signal, which are capable of directing expression of the nucleic acid in the cells of an individual to whom the nucleic acid molecule is administered. As used herein, the term "expressible form" refers to a gene construct that contains the necessary regulatory elements operably linked to a coding sequence encoding a protein such that, when present in the cells of an individual, the coding sequence will be expressed.

[0104] As used herein, "hereditary disease" refers to a disease caused, in part or in whole, directly or indirectly, by one or more abnormalities in the genome, particularly a condition present from birth. Abnormalities can be mutations, insertions, or deletions. Abnormalities can affect the coding sequence of a gene or its regulatory sequence. Hereditary diseases can be, but are not limited to, DMD, hemophilia, cystic fibrosis, Huntington's chorea, familial hypercholesterolemia (LDL receptor deficiency), hepatoblastoma, Wilson's disease, congenital hepatic porphyria, inherited hepatic metabolic disorders, Lesch-Nyhan syndrome, sickle cell anemia, thalassemia, xeroderma pigmentosum, Fanconi's anemia, retinitis pigmentosa, ataxia telangiectasia, Bloom's syndrome, retinoblastoma, and Tay-Sachs disease.

[0105] As used herein, "genome" refers to the complete set of genes or genetic material present in a cell or organism. Genomes include DNA or RNA in RNA viruses. Genomes include genes (coding regions), non-coding DNA, and the genomes of mitochondria and chloroplasts.

[0106] "Guide RNA," "gRNA," "single-stranded gRNA," and "sgRNA," as used interchangeably herein, refer to a "backbone" sequence required for Cas9 binding or Cpf1 binding and a user-defined "spacer" or "targeting sequence" (also referred to herein as a protospacer targeting sequence or segment) that defines the genomic target to be modified. "hpgRNA," "hp-gRNA," and "optimized gRNA," as used interchangeably herein, refer to a gRNA with additional nucleotides at the 5' or 3' end that can form a secondary structure with all or part of the protospacer targeting sequence or segment.

[0107] "Histone acetyltransferase" or "HAT," as used interchangeably herein, refers to an enzyme that acetylates the conserved lysine amino acid on histones by transferring an acetyl group from acetyl-CoA to form ε-N-acetyllysine. DNA is wrapped around histones, and by transferring acetyl groups to histones, genes can be turned on and off. In general, histone acetylation increases gene expression because it is linked to transcriptional activation and is associated with euchromatin. Histone acetyltransferases can also acetylate non-histone proteins, such as nuclear receptors and other transcription factors, to promote gene expression.

[0108] As used interchangeably herein, "histone deacetylase" or "HDAC" refers to a class of enzymes that remove acetyl groups (O=C-CH3) from ε-N-acetyllysine amino acids on histones, thereby allowing histones to wrap more tightly around DNA. HDACs are also known as lysine deacetylases (KDACs) to describe their function rather than their targets, which also include non-histone proteins.

[0109] As used interchangeably herein, "histone methyltransferase" or "HMT" refers to a histone modifying enzyme (e.g., histone-lysine N-methyltransferase and histone-arginine N-methyltransferase) that catalyzes the transfer of one, two, or three methyl groups to lysine and arginine residues of histones. The attachment of methyl groups occurs primarily on specific lysine or arginine residues on histones H3 and H4.

[0110] "Homology-directed repair" or "HDR," as used interchangeably herein, refers to a mechanism by which double-stranded DNA damage is repaired in a cell when the same fragment of DNA is present in the nucleus (primarily in the G2 and S phases of the cell cycle). HDR uses a donor DNA template to guide repair and can be used to produce specific sequence changes in the genome, including targeted additions of entire genes. If the donor template is provided with a site-specific nuclease, such as a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, the cellular machinery will repair the break by homologous recombination, which is enhanced by several orders of magnitude in the presence of DNA cleavage. When there are no homologous DNA fragments, non-homologous end joining can occur instead.

[0111] As used herein, "genome" refers to the complete set of genes or genetic material present in a cell or organism. Genomes include DNA or RNA in RNA viruses. Genomes include genes (coding regions), non-coding DNA, and the genomes of mitochondria and chloroplasts.

[0112] As used herein, "genome editing" refers to changing a gene. Genome editing can include correcting or restoring a mutant gene. Genome editing can include knocking out a gene, such as a mutant gene or a normal gene. Genome editing can be used to treat disease or enhance muscle repair by changing a gene of interest.

[0113] As used herein, "identical" or "identity" in the context of two or more nucleic acid or polypeptide sequences means that the sequences have a specified percentage of identical residues over a specified region. The percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over a specified region, determining the number of positions at which identical residues occur in the two sequences to generate the number of matching positions, dividing the number of matching positions by the total number of positions in the specified region, and multiplying the result by 100 to obtain the percentage of sequence identity. Where the two sequences are of different lengths or the alignment produces one or more staggered ends and the specified region of comparison contains only a single sequence, the residues of the single sequence are included in the denominator of the calculation, but not in the numerator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered to be equivalent. Identity can be obtained manually or by using a computer sequence algorithm such as BLAST or BLAST 2.0.

[0114] As used herein, "insulator" refers to a genetic boundary element that blocks the interaction between an enhancer and a promoter. By being present between an enhancer and a promoter, an insulator can inhibit their subsequent interactions. An insulator can determine a group of genes that an enhancer can affect. When two adjacent genes on a chromosome have very different transcription patterns and the induction or repression mechanism of one does not interfere with the adjacent gene, an insulator is needed. Insulators have also been found to cluster at the boundaries of topologically associated domains (TADs) and can play a role in dividing the genome into "chromosome neighborhoods" (genomic regions where regulation occurs). It is believed that insulator activity occurs mainly through the DNA 3D structure mediated by proteins including CTCF. Insulators are likely to play a role through a variety of mechanisms. Many enhancers form DNA loops that bring enhancers very physically close to the promoter region during transcriptional activation. Insulators can promote the formation of DNA loops that prevent promoter-enhancer loops from forming. Barrier insulators can prevent heterochromatin from spreading from silent genes to transcriptionally active genes.

[0115] As used herein, "invasion" refers to the disruption of the DNA duplex at the protospacer region in the target region of the target gene, such as by a gRNA that binds to a DNA sequence complementary to the protospacer.

[0116] As used herein, "invasion kinetics" refers to the rate at which invasion proceeds. Invasion kinetics can refer to the rate at which a guide RNA invades a duplex to "full invasion" such that the protospacer is completely invaded, or the rate at which a segment of protospacer DNA that binds a guide RNA expands as it displaces from its complementary strand and binds the guide RNA nucleotide by nucleotide from its PAM proximal region to complete invasion.

[0117] As used herein, "lifespan" refers to the period of time that a gRNA remains invaded in the target region of a target gene.

[0118] As used herein, "locus control region" refers to a cis-regulatory element that enhances the expression of linked genes at distal chromatin sites. It acts in a copy number-dependent manner and is tissue-specific, as seen in the selective expression of β-globin genes in erythroid cells. The expression level of a gene can be altered by LCRs and gene proximal elements, such as promoters, enhancers, and silencers. LCRs act by recruiting chromatin modifications, coactivators, and transcriptional complexes. Its sequence is conserved in many vertebrates, and conservation of specific sites can indicate functional importance.

[0119] As used interchangeably herein, "mismatched" or "MM" refers to mismatched bases including G / T or A / C pairings. Mismatches are typically caused by structural interconversion of bases during G2. The damage is repaired by identifying the deformity caused by the mismatch, determining the template and non-template strands, and excising the incorrectly incorporated base and replacing it with the correct nucleotide.

[0120] As used herein, "modulate" may mean any change in an activity, such as regulating, downregulating, upregulating, decreasing, inhibiting, increasing, decreasing, inactivating, or activating.

[0121] As used interchangeably herein, "mutated gene" or "mutated gene" refers to a gene that has undergone a detectable mutation. A mutated gene has undergone a change, such as loss, gain, or exchange of genetic material, that affects the normal transmission and expression of the gene. As used herein, "disrupted gene" refers to a mutated gene that has a mutation that causes a premature stop codon. A disrupted gene product is truncated relative to the full-length, undisrupted gene product.

[0122] As used herein, "non-homologous end joining (NHEJ) pathway" refers to a pathway for repairing double-strand breaks by directly connecting the broken ends without the need for a homologous template. The template-independent reconnection of DNA ends by NHEJ is a random, error-prone repair process that introduces random micro-insertions and micro-deletions (indels) at the DNA breakpoints. This method can be used to intentionally destroy, delete, or change the reading frame of the target gene sequence. NHEJ typically uses short homologous DNA sequences (called microhomologs) to guide repair. These microhomologs are often present in single-stranded overhangs at the ends of double-strand breaks. When the overhangs are fully compatible, NHEJ typically accurately repairs the break, however, imprecise repair resulting in nucleotide loss may also occur, but imprecise repair is much more common when the overhangs are incompatible.

[0123] As used herein, "normal gene" refers to a gene that has not undergone changes, such as loss, gain, or exchange of genetic material. A normal gene undergoes normal gene transmission and gene expression.

[0124] As used herein, "nuclease-mediated NHEJ" refers to NHEJ that is initiated after a nuclease, such as Cas9, cleaves double-stranded DNA.

[0125] As used herein, "nucleic acid" or "oligonucleotide" or "polynucleotide" means at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary strand. Thus, nucleic acid also encompasses the complementary strand of the depicted single strand. Many variants of nucleic acids can be used for the same purpose as a given nucleic acid. Thus, nucleic acid also encompasses substantially identical nucleic acids and their complementary sequences. The single strand provides a probe that can hybridize to a target sequence under stringent hybridization conditions. Thus, nucleic acid also encompasses probes that hybridize under stringent hybridization conditions.

[0126] Nucleic acids can be single-stranded or double-stranded, or can contain portions of double-stranded and single-stranded sequences. Nucleic acids can be DNA, RNA, or hybrids of genomic and cDNA, wherein nucleic acids can contain a combination of deoxyribonucleotides and ribonucleotides, and a combination of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. Nucleic acids can be obtained by chemical synthesis or by recombinant methods.

[0127] As used herein, "on-target site" refers to the target region or sequence in the genome that the gRNA is intended to target. Ideally, the target site has complete homology (100% identity or homology) with the target DNA sequence, while there is no homology elsewhere in the genome.

[0128] As used herein, an "off-target site" refers to a region in the genome that has partial homology or partial identity to the on-target site or target region of a gRNA, but which is not the region that the gRNA is intended or designed to target.

[0129] As used herein, "operably linked" means that the expression of a gene is under the control of a promoter that is spatially linked thereto. A promoter can be located 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene can be approximately the same as the distance between the promoter and the gene it controls (in the gene from which the promoter is derived). As is known in the art, variations in this distance can be accommodated without loss of promoter function.

[0130] "p300 protein," "EP300," or "E1A binding protein p300," as used interchangeably herein, refers to the adenoviral E1A-associated cellular p300 transcriptional coactivator protein encoded by the EP300 gene. p300 is a highly conserved acetyltransferase involved in a wide range of cellular processes. p300 functions as a histone acetyltransferase, regulates transcription through chromatin remodeling, and is involved in cell proliferation and differentiation.

[0131] As used herein, "partially functional" describes a protein that is encoded by a mutant gene and that has biological activity that is lower than that of a functional protein but higher than that of a non-functional protein.

[0132] As used interchangeably herein, "premature stop codon" or "out-of-frame stop codon" refers to a nonsense mutation in a DNA sequence that results in a stop codon at a position not normally present in the wild-type gene. A premature stop codon may cause a protein to be truncated or shorter than the full-length version of the protein.

[0133] As used herein, "primary cells" refer to cells taken directly from living tissue (e.g., biopsy material). Primary cells can be established for in vitro growth. These cells undergo very few population doublings and are therefore more representative of the main functional components of the tissue from which they are derived than continuous (tumor or artificially immortalized) cell lines, thus representing a more representative in vivo state model. Primary cells can come from different species, such as mice or humans.

[0134] As used interchangeably herein, "protospacer sequence" or "protospacer segment" refers to a DNA sequence targeted by the Cas9 nuclease or Cpf1 nuclease in the CRISPR bacterial adaptive immune system. In the CRISPR / Cas9 system, the protospacer sequence is typically followed by a protospacer adjacent motif (PAM); the PAM is at the 5' end. In the CRISPR / Cpf1 system, the PAM is followed by the protospacer sequence; the PAM is at the 3' end.

[0135] "Protospacer targeting sequence" or "protospacer targeting segment" as used interchangeably herein refers to a nucleotide sequence in a gRNA that corresponds to a protospacer sequence and facilitates targeting of the protospacer sequence by a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system.

[0136] As used herein, "promoter" refers to a molecule of synthetic or natural origin that is capable of conferring, activating or enhancing expression of a nucleic acid in a cell. A promoter may comprise one or more specific transcriptional regulatory sequences to further enhance expression and / or alter the spatial expression and / or temporal expression of the sequence. A promoter may also comprise distal enhancer or repressor elements, which may be located up to several thousand base pairs from the transcription start site, or anywhere in the genome. Promoters may be derived from sources including viruses, bacteria, fungi, plants, insects and animals. Promoters may constitutively or differentially regulate the expression of a genomic component with respect to the cell, tissue or organ in which expression occurs, or with respect to the developmental stage at which expression occurs, or in response to external stimuli such as physiological stress, hormones, toxins, drugs, pathogens, metal ions or inducers. Representative examples of promoters include bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lactose operator promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter and CMV IE promoter.

[0137] As used herein, "protospacer adjacent motif" or "PAM" refers to a DNA sequence that immediately follows a DNA sequence targeted by Cas9 in the CRISPR bacterial adaptive immune system or immediately precedes a DNA sequence targeted by the Cpf1 nuclease in the CRISPR bacterial adaptive immune system. The PAM is a component of invading viruses or plasmids, but not a component of the bacterial CRISPR locus. Cas9 and Cpf1 will not be able to successfully bind to or cleave the target DNA sequence if they are not preceded or followed by the PAM sequence, respectively. The PAM is an essential targeting component (not present in the bacterial genome) that bacteria use to distinguish self DNA from non-self DNA, thereby preventing the CRISPR locus from being targeted and destroyed by nucleases.

[0138] The term "recombinant" when used to refer to, for example, a cell, nucleic acid, protein, or vector, indicates that the cell, nucleic acid, protein, or vector has been modified by the introduction of a heterologous nucleic acid or protein or the alteration of a native nucleic acid or protein, or that the cell is derived from a cell so modified. Thus, for example, these recombinant cells express genes not found in the native (naturally occurring) form of the cell or express a second copy of a native gene that is otherwise normally or abnormally expressed, underexpressed, or not expressed at all.

[0139] "Silencers" or "repressors" as used interchangeably herein refer to DNA sequences that can bind transcriptional regulatory factors and prevent genes from being expressed as proteins. Silencers are sequence-specific elements that induce a negative effect on the transcription of specific genes. Silencer elements can be located at many positions in DNA. The most common position is found upstream of the target gene, where it can help repress the transcription of the gene. This distance can vary greatly between approximately -20bp to -2000bp upstream of the gene. Some silencers can be found downstream of the promoter in the introns or exons of the gene itself. Silencers have also been found in the 3' untranslated region (3'UTR) of mRNA. There are two main types of silencers in DNA, namely, classical silencer elements and non-classical negative regulatory elements (NREs). In classical silencers, genes are actively repressed by silencer elements, mainly by interfering with the assembly of general transcription factors (GTFs). NREs typically passively repress genes by suppressing other elements upstream of the gene.

[0140] As used herein, "skeletal muscle" refers to a type of striated muscle that is under the control of the somatic nervous system and is connected to bone by bundles of collagen fibers called tendons. Skeletal muscle is composed of components called myocytes or "muscle cells" (sometimes referred to as "myofibers"). Myocytes are formed by the fusion of developing myoblasts (a type of embryonic progenitor cell that produces muscle cells) in a process called myogenesis. These long, cylindrical, multinuclear cells are also called myofibers.

[0141] As used herein, "skeletal muscle condition" refers to a condition associated with skeletal muscle, such as muscular dystrophy, aging, muscle degeneration, wound healing, and muscle weakness or atrophy.

[0142] As used interchangeably herein, "subject" and "patient" refer to any vertebrate, including but not limited to mammals (e.g., cows, pigs, camels, llamas, horses, goats, rabbits, sheep, hamsters, guinea pigs, cats, dogs, rats and mice, non-human primates (e.g., monkeys such as cynomolgus or macaques, chimpanzees, etc.), and humans). In some embodiments, the subject can be human or non-human. The subject or patient may be receiving other forms of treatment.

[0143] As used herein, "super enhancers" refer to regions in mammalian genomes that contain multiple enhancers that are collectively bound by an array of transcription factor proteins to drive transcription of genes involved in cell identity. Super enhancers are often identified near genes important for controlling and defining cell identity and can be used to rapidly identify key nodes that regulate cell identity. Enhancers have several quantifiable traits that have a range of values ​​and are typically elevated at super enhancers. Super enhancers are associated with higher levels of transcriptional regulatory proteins and are associated with more highly expressed genes. The expression of genes associated with super enhancers is particularly sensitive to perturbations, which can promote cell state transitions or explain the sensitivity of super enhancer-associated genes to small molecules that target transcription.

[0144] As used herein, "target enhancer" refers to an enhancer targeted by a gRNA and a CRISPR / Cas9-based system. The target enhancer can be within the target region.

[0145] As used herein, "target gene" refers to any nucleotide sequence encoding a known or putative gene product. The target gene can be a gene with a mutation involved in a genetic disease.

[0146] "Target region," "target sequence," "protospacer," or "protospacer sequence," as used interchangeably herein, refers to the region of a target gene that is targeted by a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system.

[0147] As used herein, "transcribed region" refers to the region of DNA that is transcribed into a single-stranded RNA molecule (called messenger RNA), resulting in the transfer of genetic information from the DNA molecule to the messenger RNA. During transcription, RNA polymerase reads the template strand in a 3' to 5' direction and synthesizes RNA from 5' to 3'. The mRNA sequence is complementary to the DNA strand.

[0148] As used herein, "target regulatory element" refers to a regulatory element targeted by a gRNA and a CRISPR / Cas9-based system. The target regulatory element can be within a target region.

[0149] As used herein, "transcribed region" refers to the region of DNA that is transcribed into a single-stranded RNA molecule (called messenger RNA), resulting in the transfer of genetic information from the DNA molecule to the messenger RNA. During transcription, RNA polymerase reads the template strand in a 3' to 5' direction and synthesizes RNA from 5' to 3'. The mRNA sequence is complementary to the DNA strand.

[0150] "Transcription start site" or "TSS," as used interchangeably herein, refers to the first nucleotide of a transcribed DNA sequence where RNA polymerase begins synthesizing an RNA transcript.

[0151] As used herein, "transgene" refers to a gene or genetic material containing a gene sequence isolated from one organism and introduced into a different organism. This non-native segment of DNA may retain the ability to produce RNA or protein in the transgenic organism or may alter the normal function of the genetic code of the transgenic organism. The introduction of a transgene has the potential to alter the phenotype of an organism.

[0152] As used herein, "tru gRNA" refers to a full-length guide RNA with nucleotides truncated from its 5' end, typically by 2 nucleotides.

[0153] As used herein, "trans-regulatory elements" refer to regions of non-coding DNA that regulate the transcription of genes distant from the gene from which they are transcribed. Transcriptional regulatory elements can be located on the same or different chromosomes as the target gene. Examples of trans-regulatory elements include enhancers, super enhancers, silencers, insulators, and locus control regions.

[0154] "Variant" as used herein with respect to nucleic acids means (i) a portion or fragment of a reference nucleotide sequence (including a nucleotide sequence having insertions or deletions compared to a reference nucleotide sequence); (ii) a complementary sequence of a reference nucleotide sequence or a portion thereof; (iii) a nucleic acid substantially identical to a reference nucleic acid or its complement; or (iv) a nucleic acid that hybridizes under stringent conditions to a reference nucleic acid, its complement, or a sequence substantially identical thereto.

[0155] A "variant" of a peptide or polypeptide differs in amino acid sequence by insertion, deletion, or conservative substitution of amino acids but retains at least one biological activity. Variant may also refer to a protein having an amino acid sequence that is substantially identical to a reference protein having an amino acid sequence that retains at least one biological activity. Conservative substitution of amino acids, i.e., replacing an amino acid with a different amino acid having similar properties (e.g., hydrophilicity, degree and distribution of charged regions), is considered in the art to typically involve minor changes. As is understood in the art, these minor changes can be identified, in part, by considering the hydropathic index of amino acids. Kyte et al., J. Mol. Biol. 157:105-132 (1982). The hydropathic index of an amino acid is based on considerations of its hydrophobicity and charge. It is known in the art that amino acids with similar hydropathic indices can be substituted and still retain protein function. In one aspect, amino acids with hydropathic indices of ±2 are substituted. The hydropathicity of amino acids can also be used to reveal substitutions that result in proteins that retain biological function. Considering the hydropathicity of amino acids in the context of a peptide allows calculation of the maximum local average hydropathicity of the peptide. Substitutions can be made with amino acids whose hydrophilicity values ​​are within ±2 of each other. Both the hydrophobicity index and the hydrophilicity value of an amino acid are influenced by the particular side chain of the amino acid. It will be appreciated that, consistent with this observation, amino acid substitutions that are compatible with biological function depend on the relative similarity of the amino acids, and in particular the side chains of those amino acids, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.

[0156] As used herein, "vector" means a nucleic acid sequence containing an origin of replication. The vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. The vector can be a DNA or RNA vector. The vector can be a self-replicating extrachromosomal vector and is preferably a DNA plasmid. For example, the vector can encode Cas9 and at least one optimized gRNA nucleotide sequence of any one of SEQ ID NOs: 149-315, 321-323, and 326-329.

[0157] Unless otherwise defined herein, scientific and technical terms used in conjunction with the present disclosure shall have the meanings commonly understood by those of ordinary skill in the art. For example, any nomenclature used in conjunction with cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein is well known and commonly used in the art. The meaning and scope of the terms should be clear; however, in the event of any potential ambiguity, the definitions provided herein take precedence over any dictionary or external definition. Furthermore, unless the context otherwise requires, singular terms shall include the plural meaning and plural terms shall include the singular meaning.

[0158] 2. CRISPR system

[0159] The CRISPR system is a microbial nuclease system that participates in the defense of invading phages and plasmids that provide acquired immunity. The CRISPR locus in the microbial host contains a combination of CRISPR-related (Cas) genes and non-coding RNA elements that can program the specific nucleic acid cleavage of CRISPR mediation. A short foreign DNA segment (called a spacer) is integrated into the genome between the CRISPR repeats and used as a "memory" after exposure. Cas9 forms a complex with the 3' end of a single guide RNA (sgRNA), and the protein-RNA pair recognizes its genomic target by complementary base pairing between the 5' end of the sgRNA sequence and a predetermined 20 bp DNA sequence (called a protospacer). The complex is guided to the homologous locus of the pathogen DNA via the region (i.e., protospacer) encoded in the CRISPR RNA ("crRNA") and the protospacer adjacent motif (PAM) in the pathogen genome. The non-coding CRISPR array in the direct repeat sequence is transcribed and cleaved into a short crRNA containing a separate spacer sequence, which guides the Cas nuclease to the target site (protospacer). By simply exchanging the 20-bp recognition sequence of an expressed chimeric sgRNA, the Cas9 nuclease can be directed to a new genomic target. CRISPR spacers are used to recognize and silence exogenous genetic elements in a manner similar to RNAi in eukaryotes.

[0160] 3 categories (I, II and III type effector systems) of known CRISPR systems. Type II effector systems use single effector enzyme Cas9 to carry out targeted DNA double-strand breaks with cracking dsDNA in four successive steps. Compared with the Type I and Type III effector systems that require multiple different effectors to play the role of complex, Type II effector systems can work in alternative situations (such as eukaryotic cells). Type I effector systems are composed of long pre-crRNA (which is transcribed from the CRISPR locus containing introns), Cas9 protein and the tracrRNA involved in pre-crRNA processing. TracrRNA hybridizes with the repeat region of the intron separating pre-crRNA, and therefore dsRNA cracking is started by endogenous RNA enzyme III. This cracking is followed by the second cleavage event by Cas9 in each intron, producing a mature crRNA retained and associated with tracrRNA and Cas9, thereby forming Cas9:crRNA-tracrRNA complex.

[0161] An engineered version of the Type II effector system has been shown to function in human cells for genome engineering. In this system, the Cas9 protein is directed to the genomic target site by a synthetically reconstructed "guide RNA" ("gRNA," also interchangeably referred to herein as a chimeric sgRNA, which is a crRNA-tracrRNA fusion that generally avoids the need for RNase III and crRNA processing).

[0162] The Cas9:crRNA-tracrRNA complex unwinds the DNA duplex and searches for a matching crRNA sequence for cleavage. Target recognition occurs when the complementarity between the "protospacer" sequence in the target DNA and the remaining spacer sequences in the crRNA is detected. If the corrective protospacer adjacent motif (PAM) is also present at the 3' end of the protospacer, Cas9 mediates the cleavage of the target DNA. For protospacer targeting, the sequence must be immediately followed by the protospacer adjacent motif (PAM), a short sequence recognized by the Cas9 nuclease necessary for DNA cleavage. Different type II systems have different PAM requirements. The Streptococcus pyogenes CRISPR system may have a PAM sequence of this Cas9 (SpCas9) as 5'-NRG-3', where R is A or G, and is characterized by the specificity of the system in human cells. The unique ability of the CRISPR / Cas9 system is the direct ability to simultaneously target multiple different loci by co-expressing a single Cas9 protein with two or more sgRNAs. For example, the Streptococcus pyogenes type II system naturally prefers to use the "NGG" sequence, where "N" can be any nucleotide, but other PAM sequences such as "NAG" are also acceptable in engineered systems (Hsu et al. (2013) Nature Biotechnology, 31, 827-832). Similarly, Cas9 from Neisseria meningitidis (NmCas9) typically has a natural PAM of NNNNGATT, but has activity across a variety of PAMs, including the highly degenerate NNNNGNNN PAM (Esvelt et al. Nature Methods (2013) doi: 10.1038 / nmeth.2681).

[0163] 3. CRISPR / Cas9-based systems

[0164] Provided herein are CRISPR / Cas9 systems comprising optimized gRNAs, such as hairpin gRNAs (also referred to herein as "hpgRNAs" or "hp-gRNAs"), which allow for improved DNA targeting for epigenome editing and transcriptional regulation, such as specific cleavage of a target region of interest, such as a target gene, or activation or repression of gene expression of a target gene. The optimized gRNAs provide increased target binding specificity by modulating the lifetime at off-target positions to reduce activity at those off-target sites, while having reduced off-target binding and off-target activity of CRISPR / Cas9-based systems and CRISPR / Cpf1-based systems.

[0165] Optimizing gRNA can regulate the activity of Cas9 fusion protein by regulating the Cas9 life span at these positions and regulating the overall invasion dynamics without considering the activity of the second domain. In addition, the gRNA bound to the prototype spacer at the 5' end of the prototype spacer targeting segment can also participate in Cas9 cracking. Reducing the combination with the off-target site will limit the possibility of complete invasion / cracking at these off-target sites. A kind of engineering form of type II effector system is shown in human cells and acts on genome engineering. In this system, by synthesizing " guide RNA " (" gRNA ", also can be used interchangeably as chimeric single guide RNA (" sgRNA ") herein) (for Cas9, this guide RNA is the crRNA-tracrRNA fusion that generally avoids the needs of RNA enzyme III and crRNA processing) by guiding Cas9 protein to genomic target site. Provided herein is an engineering system based on CRISPR / Cas9 for genome editing and treating hereditary diseases. The engineering system based on CRISPR / Cas9 can be designed to target any gene, including genes related to hereditary diseases, aging, tissue regeneration or wound healing. A CRISPR / Cas9-based system can include a Cas9 protein or a Cas9 fusion protein and at least one optimized gRNA, as described below. The Cas9 fusion protein can include, for example, a domain with a different activity that is endogenous to Cas9, such as a transactivation domain.

[0166] The target gene may have a mutation such as a frameshift mutation or a nonsense mutation. If the target gene has a mutation that causes an early termination codon, an abnormal splicing acceptor site or an abnormal splicing donor site, the system based on CRISPR / Cas9 can be designed to recognize and combine the nucleotide sequence upstream or downstream of the early termination codon, the abnormal splicing acceptor site or the abnormal splicing donor site. The system based on CRISPR-Cas9 can also be used to destroy normal gene splicing by targeting splicing acceptors and donors to induce early termination codon skipping or restore a destroyed reading frame. The system based on CRISPR / Cas9 may or may not mediate off-target changes to become the protein coding region of the genome.

[0167] i.Cas9

[0168] The system based on CRISPR / Cas9 can include Cas9 protein or Cas9 fusion protein. Cas9 Protein is an endonuclease that cleaves nucleic acids and is encoded by the CRISPR locus and participates in the type II CRISPR system. Cas9 protein can be from any bacterial or archaeal species, such as Streptococcus pyogenes. The Cas9 protein can be mutated to inactivate the nuclease activity. An inactivated Cas9 protein (iCas9, also referred to as "dCas9") from Streptococcus pyogenes without endonuclease activity has recently been used to target genes in bacteria, yeast, and human cells via gRNA to silence gene expression by steric hindrance. As used herein, "iCas9" and "dCas9" both refer to Cas9 proteins with amino acid substitutions D10A and H840A and inactivated nuclease activity. In some embodiments, an inactivated Cas9 protein from Neisseria meningitidis, such as NmCas9, can be used. For example, a system based on CRISPR / Cas9 can include iCas9 of SEQ ID NO: 1.

[0169] ii. Cas9 fusion protein

[0170] The system based on CRISPR / Cas9 may include a fusion protein of Cas9 proteins such as dCas9 and a second domain without nuclease activity.The second domain may include a transcriptional activation domain (such as a VP64 domain or a p300 domain), a transcription repression domain (such as a KRAB domain), a nuclease domain, a transcription release factor domain, a histone modification domain, a nucleic acid association domain, an acetylase domain, a deacetylase domain, a methylase domain (such as a DNA methylase domain), a demethylase domain, a phosphorylation domain, a ubiquitination domain or a sumoylation domain.The second domain can be a modification factor of DNA methylation or chromatin looping.

[0171] In some embodiments, the fusion protein may comprise a dCas9 domain and a transcriptional activator. For example, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 2. In other embodiments, the fusion protein may comprise a dCas9 domain and a transcriptional repressor. For example, the fusion protein comprises the amino acid sequence of SEQ ID NO: 3. In other aspects, the fusion protein may comprise a dCas9 domain and a site-specific nuclease that has a different nuclease activity than Cas9.

[0172] The fusion protein can include two heterologous polypeptide domains, wherein the first polypeptide domain includes a Cas protein and the second polypeptide domain does not have nuclease activity. As described above, the fusion protein can include a Cas9 protein or a mutated Cas9 protein fused to a second polypeptide domain with nuclease activity. The second polypeptide domain may have a nuclease activity different from the nuclease activity of the Cas9 protein. Nucleases or proteins with nuclease activity are enzymes that can cleave the phosphodiester bonds between the nucleotide subunits of nucleic acids. Nucleases are generally further divided into endonucleases and exonucleases, but some enzymes can fall into two categories. Well-known nucleases are deoxyribonucleases and ribonucleases.

[0173] (1) CRISPR / Cas9-based gene activation system

[0174] The system based on CRISPR / Cas9 can be the gene activation system based on CRISPR / Cas9, and this system can activate regulatory element function and have the exceptional specificity to epigenome editing.The gene activation system based on CRISPR / Cas9 can be used for screening and can be targeted to increase or decrease enhancer, insulator, silencer and locus control region of target gene expression.This technology can be used for function assignment to the putative regulatory element identified by genome research (such as ENCODE and epigenomics roadmap plan (Roadmap Epigenomics project)).

[0175] CRISPR / Cas9-based gene activation systems can activate gene expression by modifying DNA methylation, chromatin looping, or catalyzing the acetylation of histone H3 lysine 27 at its target site, leading to robust transcriptional activation of the target gene from the promoter and proximal and distal enhancers. CRISPR / Cas9-based gene activation systems are highly specific and can be directed to the target gene using as little as a single guide RNA. CRISPR / Cas9-based gene activation systems can activate the expression of a gene or gene family by targeting enhancers at distal locations in the genome.

[0176] (a) Histone acetyltransferase (HAT) proteins

[0177] The gene activation system based on CRISPR / Cas9 can include histone acetyltransferase protein, Such as p300 protein, CREB binding protein (CBP; and analogs of p300), GCN5 or PCAF or its fragment. Using programmable fusion protein based on CRISPR / Cas9 by the histone acetylation in regulatory elements is an effective strategy to increase target gene expression. The histone acetyltransferase based on CRISPR / Cas9 of any site in the target genome can uniquely activate distal regulatory elements. Histone acetyltransferase protein can include people p300 protein or its fragment. Histone acetyltransferase protein can include wild-type people p300 protein or mutant people p300 protein or its fragment. Histone acetyltransferase protein can include the core lysine-acetyltransferase domain of people p300 protein, i.e. p300 HAT core (also referred to as "p300 core").

[0178] (b) CRISPR / dCas9 p300核心 Activate the system

[0179] The p300 protein regulates the activity of many genes in tissues throughout the body. It plays a role in regulating cell growth and division, enabling cells to mature and take on specialized functions (differentiation), and preventing the growth of cancerous tumors. The p300 protein can activate transcription by connecting transcription factors to protein complexes that carry out transcription in the cell nucleus. The p300 protein also functions as a histone acetyltransferase, regulating transcription through chromatin remodeling.

[0180] dCas9 p300核心 The fusion protein is an efficient and easily programmable tool for synthetically manipulating acetylation at targeted endogenous loci, resulting in regulation of proximal and distal enhancer-regulated genes. The p300 core acetylates lysine 27 (H3K27ac) on histone H3 and can provide H3K27ac enrichment. Despite robust protein expression, fusion of the catalytic core domain of p300 to dCas9 can result in significantly higher transactivation of downstream genes compared to direct fusion of the full-length p300 protein. dCas9 p300核心 Fusion proteins can also be expressed relative to dCas9 VP64 Demonstrated increased transactivation capacity, including in the context of the Nm-dCas9 backbone, particularly at distal enhancer regions where dCas9 VP64 showed little, if any, measurable downstream transcriptional activity. p300核心 Demonstrates precise and robust genome-wide transcriptional specificity. dCas9p300核心 It may enable efficient transcriptional activation and co-enrich acetylation at promoters targeted by epigenetically modified enhancers.

[0181] dCas9 p300核心 Gene expression can be activated by targeting and binding to single-stranded gRNA of promoters and / or characterized enhancers. This technology also provides the ability to synthesize transactivating distal genes from putative and known regulatory regions, and simplifies transactivation by applying a single programmable effector and a single target site. These capabilities allow multiplexing to simultaneously target multiple promoters and / or enhancers. The mammalian origin of p300 can provide advantages over viral-derived effector domains for in vivo applications by minimizing potential immunogenicity.

[0182] dCas9 p300核心 The gene activation is highly specific for the target gene. In some embodiments, the p300 core comprises amino acids 1048-1664 of SEQ ID NO: 2 (i.e., SEQ ID NO: 4). In some embodiments, the CRISPR / Cas9-based gene activation system comprises dCas9 of SEQ ID NO: 2 p300核心 Fusion protein or Nm-dCas9 of SEQ ID NO: 5 p300核心 fusion protein.

[0183] (2) CRISPR / Cas9-based gene repression system

[0184] The CRISPR / Cas9-based system can be a CRISPR / Cas9-based gene repression system that can inhibit regulatory element function and has exceptional specificity for epigenome editing. In some embodiments, the CRISPR / Cas9-based gene repression system (such as a CRISPR / Cas9-based gene repression system) can be a CRISPR / Cas9-based gene repression system that can inhibit regulatory element function and has exceptional specificity for epigenome editing. KRAB Gene repression systems (e.g., genomic loci) can interfere with distal enhancer activity by highly specific remodeling of the epigenetic state of targeted genetic loci.

[0185] (a) CRISPR / dCas9 KRAB Gene repression system

[0186] dCas9 KRAB Repressors are highly specific epigenome editing tools that can be used in loss-of-function screens to study gene function and discover targets for drug development. KRAB It has exceptional specificity for targeting specific enhancers, silencing only the target gene of that enhancer and creating a repressive heterochromatin environment at that site. KRABIt can be used to screen for novel regulatory elements within the endogenous genomic environment by silencing proximal or distal regulatory elements and corresponding gene targets. The specificity of the dCas9-KRAB repressor allows its use for genome-wide specificity in silencing endogenous genes. Epigenetic mechanisms such as histone methylation are disrupted at the target locus.

[0187] The KRAB domain, a common heterochromatin-forming motif in naturally occurring zinc-finger transcription factors, has been genetically linked to dCas9 to generate the RNA-guided synthetic repressor dCas9. KRAB The Kruppel-associated box ("KRAB") recruits heterochromatin-forming factors: Kap1, HP1, SETDB1, and NuRD. It induces H3K0 trimethylation and histone deacetylation. KRAB-based synthetic repressors can effectively silence the expression of individual genes and have been used to repress oncogenes, inhibit viral replication, and treat dominant-negative diseases.

[0188] 4. CRISPR / Cpf1-based systems

[0189] The disclosed optimized gRNAs can be used with the clustered regularly interspaced short palindromic repeats or ("CRISPR / Cpf1") system from Prevotella and Francisella 1. The CRISPR / Cpf1 system, a DNA editing technology similar to the CRISPR / Cas9 system, is found in Prevotella and Francisella bacteria and prevents genetic damage from viruses. Cpf1 is an RNA-guided endonuclease of the class II CRISPR / Cas system that contains a 1,300 amino acid protein. The Cpf1 gene is associated with the CRISPR locus and encodes an endonuclease that uses a guide RNA to find and cleave viral DNA. Cpf1 is a smaller and simpler endonuclease than Cas9 and has a smaller sgRNA molecule (approximately half the nucleotides of Cas9) because functional Cpf1 does not require tracrRNA and only requires crRNA. Examples of Cpf1 that can be used with the optimized gRNA include Cpf1 from Acidaminococcus and Lachnospiraceae bacteria.

[0190] The Cpf1 locus, which encodes Cas1, Cas2, and Cas4 proteins, is more similar to Type I and Type III systems than to those from Type II systems. The Cpf1 locus contains a mixed α / β domain, RuvC-I, followed by a helical region, RuvC-II, and zinc finger domains. The Cpf1 protein possesses a RuvC-like endonuclease domain similar to the RuvC domain of Cas9. Cpf1 lacks the HNH endonuclease domain, and the N-terminus of Cpf1 lacks the α-helical recognition lobe of Cas9. The Cpf1 CRISPR-Cas domain architecture reveals that Cpf1 is functionally unique and is classified as a Class 2, Type V CRISPR system.

[0191] The CRISPR / Cpf1 system consists of the Cpf1 enzyme and a guide RNA that finds the correct position on the duplex and positions the complex there to cleave the target DNA. The CRISPR / Cpf1 system has three phases of activity: adaptation, crRNA formation, and interference. During the adaptation phase, the Cas1 and Cas2 proteins facilitate the adaptation of small fragments of DNA into the CRISPR array. The crRNA phase involves processing pre-cr-RNA to produce mature crRNA to guide the Cas proteins. During the interference phase, Cpf1 binds to the crRNA to form a binary complex that identifies and cleaves the target DNA sequence.

[0192] The Cpf1-crRNA complex cleaves target DNA or RNA by identifying the protospacer-adjacent motif 5'-YTN-3' (where "Y" is a pyrimidine and "N" is any nucleobase) or 5'-TTN-3' (which is different from the G-rich PAM targeted by Cas9). Unlike the PAM targeted by Cas9, which is located on the 3' side of the guide RNA, the PAM targeted by Cpf1 is located on the 5' side of the guide RNA. After identifying the PAM, unlike the blunt-end cleavage of Cas9, Cpf1 introduces a sticky-end-like DNA double-strand break with a 4 or 5 nucleotide overhang, thereby enhancing the efficiency and specificity of gene insertion during NHEJ or HDR. TTN PAM sites are more useful for human genome engineering than GGN PAM sites because the human genome is rich in T than G. For Cpf1, the protospacer-targeting segment of the gRNA is at its 3' end, while the Cas9 gRNA is at its 5' end.

[0193] 5. gRNA

[0194] The system based on CRISPR / Cas9 or the system based on CRISPR / Cpf1 can include at least one gRNA of the targeting nucleic acid sequence, such as the optimized gRNA described herein.gRNA provides the specific targeting of the target region or target gene by the system based on CRISPR / Cas9 or the system based on CRISPR / Cpf1.For the system based on CRISPR / Cas9, gRNA is a fusion of two non-coding RNAs, i.e., crRNA and tracrRNA.gRNA or sgRNA can target any desired DNA sequence by exchanging the sequence encoding the 20bp prototype spacer (the sequence is given targeting specificity by pairing with the complementary base of the desired DNA target).gRNA simulates the naturally occurring crRNA:tracrRNA duplex involved in the II type effector system.This duplex, such as 42 nucleotide crRNA and 75 nucleotide tracrRNA, can be included to guide Cas9 to crack the target nucleic acid.gRNA can target and bind to the target region of the target gene.For the system based on CRISPR / Cpf1, gRNA is crRNA.

[0195] A system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 can include at least one gRNA, such as an optimized gRNA as described herein, wherein the gRNA targets different DNA sequences. The target DNA sequences can be overlapping. The target sequence or protospacer is followed by a PAM sequence at the 3' end of the protospacer. Different type II systems have different PAM requirements. For example, the Streptococcus pyogenes type II system uses an "NGG" sequence, where "N" can be any nucleotide.

[0196] 6. Methods for Generating Optimized Guide RNA (gRNA)

[0197] The present disclosure relates to methods for generating optimized gRNAs, such as hairpin gRNAs (also referred to herein as "hpgRNAs" and "hp-gRNAs"). The optimized gRNA comprises the nucleotide sequence of the full-length gRNA and nucleotides added to the 5' end or 3' end of the full-length gRNA. In some embodiments, the full-length gRNA can be designed using programs such as SgRNA designer, CRISPRMultiTargeter, or SSFinder. The nucleotides added to the 5' end of the CRISPR / Cas9 system or the 3' end of the CRISPR / Cpf1 system of the full-length gRNA can form a secondary structure by hybridizing or partially hybridizing to the nucleotides in the protospacer targeting sequence of the full-length gRNA. The secondary structure regulates DNA binding or cleavage by disrupting the invasion of the gRNA into the DNA duplex. The secondary structure affects the invasion kinetics of the gRNA, rather than the binding energy of the gRNA to the complementary DNA chain. As described in the examples below, the guide RNA of the type II CRISPR-Cas system binds to the protospacer through a Cas9-facilitated process called "strand invasion," in which the Cas9 protein itself first binds and melts the protospacer adjacent motif (PAM) through direct interaction, after which the 3' end of the gRNA base pairs with the PAM adjacent nucleotides (the "seed" region), and then proceeds nucleotide by nucleotide from the 3' end of the gRNA to the 5' end where it base pairs with the protospacer. A similar mechanism is used for the CRISPR / Cpf1 system.

[0198] The nucleotides added to the 5' end or 3' end of full-length gRNA are not only added to hybridize with the prototype spacer targeting segment of guide RNA (hairpin) to block the approach to the prototype spacer under thermodynamic equilibrium. As described in the examples, the equilibrium thermodynamic secondary structure properties (such as the melting temperature of gRNA secondary structure) are completely unrelated to the specificity of guide RNA. On the contrary, in the case of cracking and in the subsequent calculation work combined with Cas9 (such as by ChIP-Seq measurement in cells (see doi:10.1038 / nbt.2916; doi:10.1038 / nbt.2889)), these and estimated chain invasion kinetics are adjusted to the structure, design and function of the guide RNA for the chain invasion in the prototype spacer. There is a significant and substantial correlation, these guide RNAs are necessarily different from the hairpin designed to thermodynamically compete with the target site and the combination of the off-target site under equilibrium. For example, secondary structural elements designed to be stable at equilibrium (e.g., RNAs that form hairpin-like structures containing internal rG-rU wobble base pairs within the stem) can rapidly destabilize during strand invasion (e.g., because the rG-rU wobble base pair becomes the terminal base pair of the stem when the adjacent nucleotide invades the protospacer), resulting in a significant energy penalty for the RNA secondary structure, regulating strand invasion and binding kinetics by completely independent mechanisms rather than simply by blocking access to the protospacer at thermodynamic equilibrium. Secondary structures that are stable at equilibrium but rapidly destabilize during strand invasion can be designed using the methods described herein in a manner that allows for differentiation between on-target and off-target sites with minimal thermodynamic energy differences between the sites (e.g., as a result of a single internal mismatch), which are virtually indistinguishable by cis-blocking or thermodynamic competition. Invasion of the on-target site destabilizes the hairpin containing the GU wobble base pair, and these sites are differentiated by invasion kinetics. For example, the use of computationally designed secondary structures responsible for strand invasion enables the VEGFA1 site described in the examples below (target site is GGGTGGGGGGAGTTTGCTCC and off-target site 2 is GGGTGGGGGAGTTTGCTCC) to be modified compared to standard or full-length guide RNA or truncated guide RNA, respectively. A TGG A GGGAGTTTGCTCC; mismatches are underlined) were reduced by 93% and 98% in off-target cleavage.

[0199] Additionally, nucleotides can be added to the 5' or 3' end of the full-length gRNA to disrupt "naturally occurring" secondary structures on the protospacer-targeting segment of the gRNA in the "seed" region to enhance initiation of strand invasion by the guide RNA. Thus, the addition of these nucleotides that form secondary structures that alter strand invasion to modulate DNA binding or cleavage by hybridizing or partially hybridizing to nucleotides in the protospacer-targeting sequence represents a distinct class of guide RNA modifications.

[0200] The optimized gRNA is designed to minimize binding at off-target sites and to allow binding to the protospacer sequence. In some embodiments, the off-target site is a known or predicted off-target site. In some embodiments, the method includes: identifying a target region of interest, the target region of interest comprising a protospacer sequence; determining the polynucleotide sequence of a full-length gRNA targeting the target region of interest, the full-length gRNA comprising a protospacer targeting sequence or segment; determining at least one or more off-target sites of the full-length gRNA; generating a polynucleotide sequence of a first gRNA, the first gRNA comprising the polynucleotide sequence of the full-length gRNA and an RNA segment, the RNA segment comprising a polynucleotide sequence having a length of M nucleotides, the polynucleotide sequence being complementary to a nucleotide segment of the protospacer targeting sequence or segment, the RNA segment being located at the 5' end of the polynucleotide sequence of the full-length gRNA, the first gRNA optionally comprising a linker between the 5' end of the polynucleotide sequence of the full-length gRNA and the RNA segment, the linker comprising a polynucleotide sequence having a length of N nucleotides, the first gRNA The invention relates to a method for synthesizing a first gRNA that is capable of invading the protospacer sequence and binding to a DNA sequence complementary to the protospacer sequence and forming a protospacer-duplex, and the first gRNA is capable of invading an off-target site and binding to a DNA sequence complementary to the off-target site and forming an off-target duplex; calculating an estimate of invasion kinetics and a lifetime that the first gRNA remains invaded in the protospacer and off-target site duplex or computationally simulating these invasion kinetics and lifetime, wherein the invasion kinetics are estimated on a nucleotide-by-nucleotide basis by determining the energy difference between further invasion by a different gRNA and reannealing the first gRNA to the DNA sequence complementary to the protospacer sequence; comparing the estimated lifetime of the first gRNA at the protospacer and / or off-target site with the estimated lifetime of the full-length gRNA or truncated gRNA (tru-gRNA) at the protospacer and / or off-target site; randomizing 0 to N nucleotides in the linker and 0 to M nucleotides in the first gRNA and generating a second gRNA, and using the second gRNA Repeat step (e); identify optimized gRNAs based on gRNA sequences that meet the design criteria; and test the optimized gRNAs in vivo to determine binding specificity.

[0201] In some embodiments, the method comprises: identifying a target region of interest, the target region of interest comprising a protospacer sequence; determining a polynucleotide sequence of a full-length gRNA targeting the target region of interest, the full-length gRNA comprising the protospacer targeting sequence or segment; determining at least one or more off-target sites of the full-length gRNA; generating a polynucleotide sequence of a first gRNA, the first gRNA comprising the polynucleotide sequence of the full-length gRNA and an RNA segment, the RNA segment comprising a polynucleotide sequence having a length of M nucleotides, the polynucleotide sequence being complementary to a nucleotide segment of the protospacer targeting sequence or segment, the RNA segment being located at the 3' end of the polynucleotide sequence of the full-length gRNA, the first gRNA optionally comprising a linker between the 3' end of the polynucleotide sequence of the full-length gRNA and the RNA segment, the linker comprising a polynucleotide sequence having a length of N nucleotides, the first gRNA being capable of invading the protospacer sequence and binding to a DNA sequence complementary to the protospacer sequence and forming a protospacer-duplex, and and the first gRNA is capable of invading the off-target site and binding to a DNA sequence complementary to the off-target site and forming an off-target duplex; calculating an estimate of the invasion kinetics and the lifetime that the first gRNA remains invaded in the duplex of the protospacer and off-target site or computationally simulating these invasion kinetics and lifetime, wherein the invasion dynamics are estimated nucleotide by nucleotide by determining the energy difference between further invasion by a different gRNA and reannealing the first gRNA to the DNA sequence complementary to the protospacer sequence; comparing the estimated lifetime of the first gRNA at the protospacer and / or off-target site with the estimated lifetime of the full-length gRNA or truncated gRNA (tru-gRNA) at the protospacer and / or off-target site; randomizing 0 to N nucleotides in the linker and 0 to M nucleotides in the first gRNA and generating a second gRNA, and repeating step (e) with the second gRNA; identifying an optimized gRNA based on the gRNA sequence that meets the design criteria; and testing the optimized gRNA in vivo to determine binding specificity.

[0202] In some embodiments, the energetics of further invasion of a different gRNA is determined by determining the energetics of at least one of the following: (I) disruption of DNA-DNA base pairing, (II) formation of RNA-DNA base pairs, (III) the energy difference resulting from disruption or formation of a different secondary structure within the uninvaded guide RNA, and (IV) formation or disruption of interactions between the displaced DNA strand complementary to the protospacer and any unpaired guide RNA nucleotides that are not involved in the secondary structure. In some embodiments, the energetics of reannealing of a first gRNA to a DNA sequence complementary to the protospacer sequence is determined by determining the energetics of at least one of the following: (I) formation of DNA-DNA base pairing, (II) disruption of RNA-DNA base pairs, (III) the energy difference resulting from disruption or formation of a different secondary structure within the newly uninvaded guide RNA, and (IV) formation or disruption of interactions between the displaced DNA strand complementary to the protospacer and any unpaired guide RNA nucleotides that are not involved in the secondary structure. In some embodiments, the method further comprises determining an energetic consideration based on at least one of: (V) base pairing across mismatches, (VI) interaction with the Cas9 protein, and / or (VII) an additional heuristic, wherein the additional heuristic relates to binding lifetime, extent of invasion, stability of the invading guide RNA, or other computational / simulated properties of gRNA invasion to Cas9 cleavage activity.

[0203] The system based on CRISPR / Cas9 or the system based on CRISPR / Cpf1 can use the gRNA of sequence and length variation, such as optimization gRNA as described herein.In certain embodiments, full-length gRNA can include the prototype spacer targeting section corresponding to the polynucleotide sequence of target DNA sequence (i.e., prototype spacer).In certain embodiments, the prototype spacer targeting section can have at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 30 nucleotides or at least 35 nucleotides.GRNA can target at least one of the following items: promoter region, enhancer region, repressor region, insulator region, silencer region, relate to the transcription region of the region looped with promoter region DNA, gene splicing region or target gene. In some embodiments, the full-length gRNA comprises a protospacer targeting segment of about 15 to 20 nucleotides.

[0204] In some embodiments, the RNA segment comprises 2 to 20 nucleotides, 3 to 10 nucleotides, or 5 to 8 nucleotides. In some embodiments, the RNA segment comprises 2 to 20 nucleotides, 3 to 10 nucleotides, or 5 to 8 nucleotides that are complementary to the protospacer targeting sequence. In some embodiments, M is between 1 and 20, between 1 and 19, between 1 and 18, between 1 and 17, between 1 and 16, between 1 and 15, between 1 and 14, between 1 and 13, between 1 and 12, between 1 and 11, between 1 and 10, between 1 and 9, between 1 and 8, between 1 and 7, between 1 and 6, between 1 and 5, between 2 and 20, between 2 and 19, between 2 and 18, between 2 and 17, between 2 and 16, between 2 and 15, between 2 and 14, between 2 and 13, between 2 and 12, between 2 and 11, between 2 and 10, between 2 and 9, between 2 and 8, between 2 and 7, between 2 and 6, between 2 and 5, between 3 and 20, between 3 and 19, between 3 and 18, between 3 and 17, between 3 and 16, between 3 and 15, between 3 and 14 Between, between 3 and 13, between 3 and 12, between 3 and 11, between 3 and 10, between 3 and 9, between 3 and 8, between 3 and 7, between 3 and 6, between 3 and 5, between 4 and 20, between 4 and 19, between 4 and 18, between 4 and 17, between 4 and 16, between 4 and 15, between 4 and 14, 4 Between 4 and 13, Between 4 and 12, Between 4 and 11, Between 4 and 10, Between 4 and 9, Between 4 and 8, Between 4 and 7, Between 4 and 6, Between 4 and 5, Between 5 and 20, Between 5 and 19, Between 5 and 18, Between 5 and 17, Between 5 and 16, Between 5 and 15, Between 5 and 14, Between 5 and 13, Between 5 and 12, Between 5 and 11, Between 5 and 10, Between 5 and 9, Between 5 and 8, Between 5 and 7, Between 5 and 6, Between 6 and 20, Between 6 and 19, Between 6 and 18, Between 6 and 17, Between 6 and 16, Between 6 and 15, Between 6 and 14, Between 6 and 13, Between 6 and 12, 17, 8 and 16, 8 and 14, 8 and 13, 8 and 13, 8 and 12, 8 and 11, 8 and 10, 8 and 9, 9 and 20, 9 and 19, 9 and 18, 9 and 17, 9 and 16, 9 and 15, 9 and 14, 9 and 13, 9 and 12, 9 and 11, or 9 and 10.For example, M can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20. In some embodiments, the RNA segments can have 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15 , 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14 , 5 to 13, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8, 5 to 7, 5 to 6, 6 to 20, 6 to 19, 6 to 18, 6 to 17, 6 to 16, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 19, 7 to 18, 7 to 17, 7 to 16, 7 to 15, 7 to 14, 7 to 13, 7 to 12, 7 to 11, 7 to 10 In some embodiments, the RNA segment may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides.

[0205] In some embodiments, N is between 1 and 20, between 1 and 19, between 1 and 18, between 1 and 17, between 1 and 16, between 1 and 15, between 1 and 14, between 1 and 13, between 1 and 12, between 1 and 11, between 1 and 10, between 1 and 9, between 1 and 8, between 1 and 7, between 1 and 6, between 1 and 5, between 2 and 20, between 2 and 19, between 2 and 18, between 2 and 17, between 2 and 16, between 2 and 15, between 2 and 14, between 2 and 13, between 2 and 12, and between 2 and 11. between 3 and 10, between 2 and 9, between 2 and 8, between 2 and 7, between 2 and 6, between 2 and 5, between 3 and 20, between 3 and 19, between 3 and 18, between 3 and 17, between 3 and 16, between 3 and 15, between 3 and 14, between 3 and 13, between 3 and 12, between 3 and 11, between 3 and 10, between 3 and 9, between 3 and 8, between 3 and 7, between 3 and 6, between 3 and 5, between 4 and 20, between 4 and 19, between 4 and 18, between 4 and 17, between 4 and 16, between 4 and 15 Between, between 4 and 14, between 4 and 13, between 4 and 12, between 4 and 11, between 4 and 10, between 4 and 9, between 4 and 8, between 4 and 7, between 4 and 6, between 4 and 5, between 5 and 20, between 5 and 19, between 5 and 18, between 5 and 17, between 5 and 16, between 5 and 15, 5 Between 5 and 14, Between 5 and 13, Between 5 and 12, Between 5 and 11, Between 5 and 10, Between 5 and 9, Between 5 and 8, Between 5 and 7, Between 5 and 6, Between 6 and 20, Between 6 and 19, Between 6 and 18, Between 6 and 17, Between 6 and 16, Between 6 and 15, Between 6 and 14, Between 6 and 13, Between 6 and 12, Between 6 and 11, Between 6 and 10, Between 6 and 9, Between 6 and 8, Between 6 and 7, Between 7 and 20, Between 7 and 19, Between 7 and 18, Between 7 and 17, Between 7 and 16, Between 7 and 15, Between 7 and 14, Between 7 and 13, Between 7 and 12, Between 7 and 11, 17, 9 and 16, 8 and 14, 8 and 13, 8 and 13, 8 and 12, 8 and 11, 8 and 10, 8 and 9, 9 and 20, 9 and 19, 9 and 18, 9 and 17, 9 and 16, 9 and 15, 9 and 14, 9 and 13, 9 and 12, 9 and 11, or 9 and 10. For example, N can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.In some embodiments, the linker comprises 1 to 20 nucleotides, 3 to 10 nucleotides, or 5 to 8 nucleotides. For example, a linker can have 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 2 to 14, 2 to 13, 2 to 12, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8 , 5 to 7, 5 to 6, 6 to 20, 6 to 19, 6 to 18, 6 to 17, 6 to 16, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 19, 7 to 18, 7 to 17, 7 to 16, 7 to 15, 7 to 14, 7 to 13, 7 to 12, 7 to 11 17, 8, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the linker may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides. In some embodiments, the linker may include a stable linker such as a tetraloop. Examples of tetranucleotide loops include, but are not limited to, ANYA, CUYG, GNRA, UMAC, and UNCG.

[0206] In some embodiments, the RNA segment and / or the protospacer targeting sequence provide a secondary structure. In some embodiments, the secondary structure is formed by partially hybridizing the protospacer targeting sequence with the RNA segment. In some embodiments, the secondary structure regulates DNA binding or Cas9 cleavage by destroying the invasion of the optimized gRNA into the protospacer duplex or off-target duplex. In some embodiments, the secondary structure stabilizes the 5' end of the gRNA within the protein and protects the optimized gRNA within Cas9 to prevent degradation.

[0207] In some embodiments, the secondary structure is formed by hybridizing all or part of the RNA segment to nucleotides in the 5' end of the protospacer targeting sequence or segment, nucleotides in the middle of the protospacer targeting sequence or segment, and / or nucleotides at the 3' end of the protospacer targeting sequence or segment. In some embodiments, a continuous segment of the RNA segment hybridizes to the protospacer targeting sequence or segment. In some embodiments, a non-continuous segment of the RNA segment hybridizes to the protospacer targeting sequence or segment. In some embodiments, the secondary structure is a hairpin.

[0208] In some embodiments, the secondary structure is stable at room temperature or 37° C. In some embodiments, the total equilibrium free energy of the secondary structure is less than about 2 kcal / mol at a temperature between about 4° C. and about 50° C., such as at room temperature or 37° C. For example, the total equilibrium free energy of the secondary structure can be less than about 10 kcal / mol, less than about 5 kcal / mol, less than about 4 kcal / mol, less than about 3 kcal / mol, less than about 2 kcal / mol, less than about 1 kcal / mol, or less than about 0.5 kcal / mol at temperatures between about 4° C. and about 50° C., between about 4° C. and about 40° C., between about 4° C. and about 37° C., between about 4° C. and about 30° C., between about 4° C. and about 25° C., between about 4° C. and about 20° C., between about 4° C. and about 10° C., between about 5° C. and about 50° C., between about 5° C. and about 40° C., between about 5° C. and about 37° C. In some embodiments, the RNA segment hybridizes or forms a non-canonical base pair with at least two nucleotides of the protospacer targeting sequence or segment. In some embodiments, the non-canonical base pair is rU-rG.

[0209] In some embodiments, 1 to 20 nucleotides are randomized in the linker. For example, 1 to 20, 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 20, 2 to 15, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 3 to 20, 3 to 15, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 to 4, 4 to 20, 4 to 15, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 15, 5 to 10, 5 to 9, 5 to 8 In some embodiments, 5 to 7, 5 to 6, 6 to 20, 6 to 15, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 15, 7 to 10, 7 to 9, 7 to 8, 8 to 20, 8 to 15, 8 to 10, 8 to 9, 9 to 20, 9 to 15, or 9 to 10, 10 to 20, 10 to 15, or 15 to 20 nucleotides can be randomized in the linker.

[0210] In some embodiments, 1 to 20 nucleotides are randomized in the RNA segment. For example, 1 to 20, 1 to 15, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 1 to 4, 1 to 3, 1 to 2, 2 to 20, 2 to 15, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 3 to 20, 3 to 15, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 to 4, 4 to 20, 4 to 15, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 15, 5 to 10, 5 to 9 In some embodiments, 5 to 8, 5 to 7, 5 to 6, 6 to 20, 6 to 15, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 15, 7 to 10, 7 to 9, 7 to 8, 8 to 20, 8 to 15, 8 to 10, 8 to 9, 9 to 20, 9 to 15, or 9 to 10, 10 to 20, 10 to 15, or 15 to 20 nucleotides can be randomized in the RNA segment.

[0211] In some embodiments, step (g) is repeated X times to produce X number of gRNAs, and step (e) is repeated for each X number of gRNAs, where X is between 0 and 20. In some embodiments, X can be between 1 and 20, between 1 and 19, between 1 and 18, between 1 and 17, between 1 and 16, between 1 and 15, between 1 and 14, between 1 and 13, between 1 and 12, between 1 and 11, between 1 and 10, between 1 and 9, between 1 and 8, between 1 and 7, between 1 and 6, between 1 and 5, between 2 and 20, between 2 and 19, between 2 and 18, between 2 and 17, between 2 and 16, between 2 and 15, between 2 and 14, between 2 and 13, between 2 and 12, between 2 and 11, between 1 and 10, between 2 and 9, between 2 and 8, between 2 and 7, between 2 and 6, between 2 and 5, between 3 and 20, between 3 and 19, between 3 and 18, between 3 and 17, between 3 and 16, between 3 and 15, Between 3 and 14, between 3 and 13, between 3 and 12, between 3 and 11, between 3 and 10, between 3 and 9, between 3 and 8, between 3 and 7, between 3 and 6, between 3 and 5, between 4 and 20, between 4 and 19, between 4 and 18, between 4 and 17, between 4 and 16, between 4 and 15, between 4 and 14, between 4 and 13, between 4 and 12, between 4 and 11, between 4 and 10, between 4 and 9, between 4 and 8, between 4 and 7, between 4 and 6, between 4 and 5, between 5 and 20, between 5 and 19, between 5 and 18, between 5 and 17, between 5 and 16, between 5 and 15, between 5 and 14, Between 5 and 13, between 5 and 12, between 5 and 11, between 5 and 10, between 5 and 9, between 5 and 8, between 5 and 7, between 5 and 6, between 6 and 20, between 6 and 19, between 6 and 18, between 6 and 17, between 6 and 16, between 6 and 15, between 6 and 14, between 6 and 13, between 6 and 12, between 6 and 11, between 6 and 10, between 6 and 9, between 6 and 8, between 6 and 7, between 7 and 20, between 7 and 19, between 7 and 18, between 7 and 17, between 7 and 16, between 7 and 15, between 7 and 14, between 7 and 13, between 7 and 12, between 7 and 11, between 7 and 10, between 8 and 9, between 7 and 8, between 8 and 20, between 8 and 19, between 8 and 18, between 8 and 17, between 8 and 16, between 8 and 14, between 8 and 13, between 8 and 13, between 8 and 12, between 8 and 11, between 8 and 10, between 8 and 9, between 9 and 20, between 9 and 19, between 9 and 18, between 9 and 17, between 9 and 16, between 9 and 15, between 9 and 14, between 9 and 13, between 9 and 12, between 9 and 11, or between 9 and 10.For example, X can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.

[0212] In some embodiments, invasion kinetics and lifespan are calculated using the kinetic Monte Carlo method or the Gillespie algorithm. In some embodiments, invasion kinetics and lifespan can be determined using "deterministic" methods known to those skilled in the art, such as differential equations for chain invasion modeling. The kinetic Monte Carlo (KMC) method is a Monte Carlo computer simulation of the time evolution of certain processes intended to simulate natural occurrences. These processes are typically processes that occur under known transition rates between states. These known transition rates are input values ​​for the KMC algorithm. The Gillespie algorithm (also known as the Doob-Gillespie algorithm) generates statistically correct trajectories (possible solutions) of random equations. The Gillespie algorithm can be used to simulate increasingly complex systems. The algorithm is particularly useful for simulating reactions within cells, where the number of reactants is typically tens of molecules (or less). Mathematically, it is a dynamic Monte Carlo method and is similar to the kinetic Monte Carlo method. The Gillespie algorithm allows for discrete and random simulation of systems using a small number of reactants because each reaction is explicitly simulated. The trajectory corresponding to a single Gillespie simulation represents an exact sample from the probability mass function, which is the solution to the master equation.

[0213] In certain embodiments, design criteria can be specificity, binding lifespan regulation and / or estimated cracking specificity.For example, optimizing gRNA can be designed to have the binding lifespan of the binding lifespan of being greater than or equal to complete gRNA at the target site, and / or the binding lifespan of the binding lifespan of being less than or equal to the binding lifespan of the binding lifespan of the binding lifespan of the binding lifespan of the missing target site.In certain embodiments, optimizing gRNA is selected as the binding lifespan of the binding lifespan of the binding lifespan of at least three missing target sites being less than or equal to full-length gRNA, wherein the missing target site is predicted to be the closest missing target site or is predicted to have the highest identity with the target site.In certain embodiments, design criteria is included in the lifespan at missing target site or the cracking rate being less than or equal to the lifespan of full-length gRNA or brachymemma gRNA at missing target site or the cracking rate and / or prediction in target activity rate being greater than 10% of the prediction in target activity rate of full-length gRNA or brachymemma gRNA.

[0214] In some embodiments, the optimized gRNA is tested in step i) using a mismatch-sensitive nuclease to determine CRISPR activity, such as using a surveyor assay or T7 endonuclease 1 (T7E1) assay, or a next-generation sequencing technology such as Illumina MiSeq or GUIDE-Seq. In some embodiments, the optimized gRNA is tested in step i) using a reporter assay in which Cas9-fusion protein activity alters expression of a reporter protein such as GFP. GUIDE-Seq is an assay designed to measure off-target cleavage.

[0215] In some embodiments, programs such as CRISPR Design (Ran et al. Nature Protocols (2013) 8: 2281-2308) and CCTop (Stemmer, PLoS One (2015) 10: e0124633) tools can be used to identify target regions based on sequences adjacent to the PAM sequence. In some embodiments, target sites can include promoters, DNase I hypersensitive sites, transposase accessible chromatin sites, DNA methylation sites, transcription factor binding sites, epigenetic markers, expression quantitative trait loci, and / or regions associated with human traits or phenotypes in genetic association studies. Target sites can be identified by DNA enzyme sequencing (DNase-seq), assay for transposase-accessible chromatin with high-throughput sequencing (ATAC-seq), ChIP sequencing, self-transcriptionally active regulatory region sequencing (STARR-seq), single-molecule real-time sequencing (SMRT), formaldehyde-assisted isolation of regulatory elements sequencing (FAIRE-seq), micrococcal nuclease sequencing (MNase-seq), reduced representation bisulfite sequencing (RRBS-seq), whole-genome bisulfite sequencing, methyl-incorporating DNA immunoprecipitation (MEDIP-seq), or genetic association studies. In some embodiments, off-target sites can be determined using CasOT (PKU Zebrafish Functional Genomics group, Peking University), CHOPCHOP (Harvard University), CRISPR design (MIT), CRISPR design tool (The Broad Institute of Harvard and MIT), CRISPR / Cas9 gRNA finder (University of Colorado), CRISPRfinder (Université Paris-Sud), E-CRISP (DKFZ German Cancer Research Center), CRISPR gRNA design tool (DNA 2.0), PROGNOS (Emory University / Georgia Institute of Technology), ZiFiT (Massachusetts General Hospital). Examples of tools that can be used to determine target regions and off-target sites are described in International Patent Application No. WO2016109255, which is incorporated herein by reference in its entirety.

[0216] 7. Target Gene

[0217] As disclosed herein, CRISPR / Cas9-based systems or CRISPR / Cpf1-based systems can be designed to target and cleave any target gene. For example, gRNA, such as the optimized gRNA described herein, can target and bind to a target region in a target gene. The target gene can be an endogenous gene, a transgenic gene, or a viral gene in a cell line. In some embodiments, the target gene can be a known gene. In some embodiments, the target gene is an unknown gene. The gRNA can target any nucleic acid sequence. The nucleic acid sequence target can be DNA. The DNA can be any gene. For example, the gRNA can target genes such as DMD, EMX1, or VEGFA.

[0218] In some aspects, the target gene is a disease-related gene. In some embodiments, the target cell is a mammalian cell. In some embodiments, the genome includes the human genome. In some embodiments, the target gene can be a prokaryotic gene or a eukaryotic gene, such as a mammalian gene. For example, a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 can target mammalian genes, such as DMD (dystrophin gene), EMX1, VEGFA, IL1RN, MYOD1, OCT4, HBE, HBG, HBD, HBB, MYOCD (myocardial protein), PAX7 (pairing box protein Pax-7), FGF1 (fibroblast growth factor-1) gene, such as FGF1A, FGF1B and FGF1C.Other target genes include, but are not limited to, Atf3, Axud1, Btg2, c-Fos, c-Jun, Cxcl1, Cxcl2, Edn1, Ereg, Fos, Gadd45b, Ier2, Ier3, Ifrd1, Il1b, Il6, Irf1, Junb, Lif, Nfkbia, Nfkbiz, Ptgs2, S1c25a25, Sqstm1, Tieg, Tnf, Tnfaip3, Zfp36, Birc2, Ccl2, Ccl20, Ccl7, Cebpd, Ch25h, CSF1, Cx3cl1, Cxcl10, Cxcl5, Gch, Icam1, Ifi47, Ifngr2, Mmp10, Nfkbie, Npal1, p21, Relb, Ripk2, Rnd1, S1pr3, Stx11, Tgtp, Tlr2, Tmem140, Tnfaip2, Tnfrsf6, Vcam1, 1110004C05Rik (GenBank accession number BC010291), Abca1, AI561871 (GenBank accession number BI143915), AI882074 (GenBank accession number BB730912), Arts1, AW049765 (GenBank accession number BC026642.1), C3, Casp4, Ccl5, Ccl9, Cdsn, Enpp2, Gbp2, H2-D1, H2-K, H2-L, Ifit1, Ii, Il13ra1, Il1rl1, Lcn2, Lhfpl2, LOC677168 (GenBank accession number AK019325), Mmp13, Mmp3, Mt2, Naf1, Ppicap, Prnd, Psmb10, Saa3, Serpina3g, Serpinf1, Sod3, Stat1, Tapbp, U90926 (GenBank accession number NM_020562), Ubd, A2AR (adenosine A2A receptor), B7-H3 (also known as CD276), B7-H4 (also known as VTCN1), BTLA (B and T lymphocyte attenuator; also known as CD272), CTLA-4 (cytotoxic T lymphocyte-associated protein 4; also known as CD152), IDO (indoleamine 2,3 dioxygenase), KIR (killer cell immunoglobulin-like receptor), LAG3 (lymphocyte activation gene-3), PD-1 (programmed death 1 (PD-1) receptor), TIM-3 (T cell immunoglobulin domain and mucin domain 3) and VISTA (V-domain Ig inhibitor of T cell activation). In some embodiments, the target gene is DMD (dystrophin), EMX1 or VEGFA gene.

[0219] 8. Compositions for genome editing

[0220] The present invention relates to compositions for genome editing, genome alteration, or altering the gene expression of a target gene. The compositions comprise optimized gRNAs produced by the disclosed methods and a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system. In some embodiments, the gRNAs can distinguish between on-target and off-target sites with minimal thermodynamic energy differences between the sites, and provide increased specificity. In some embodiments, the optimized gRNAs regulate strand invasion into protospacers.

[0221] Increased specificity is achieved by adding an extension to the 5' or 3' end of the full-length or standard gRNA such that it forms a "hairpin" structure that is self-complementary to the segment of the full-length or standard gRNA that targets the protospacer (e.g., the prototargeting sequence). Figure 1 B and Figure 2B. The hairpin acts as a kinetic barrier to protospacer strand invasion, but the hairpin is displaced during strand invasion across the target site, allowing full invasion to occur.

[0222] As shown in Figure 2D, binding by dCas9 preferentially occurs to the complete protospacer, strongly suggesting that the hairpin is actually displaced during the invasion process. The disclosed optimized gRNA, which is a hairpin structure, is designed to increase specificity of binding to the target site by inhibiting invasion when there is a mismatch between the target and the PAM-distal targeting region of the guide RNA. In these cases, it is energetically more favorable for the hairpin to remain closed, and the presence of the hairpin likely promotes unwinding and disengagement of Cas9 / dCas9 from these sites.

[0223] Compared to standard guide RNA and best guide RNA variants (see Examples), optimized gRNA (hpgRNA) with 5'-hairpin or 3'-hairpin is significantly enhanced in terms of binding specificity, and eliminates or significantly weakens the binding at the protospacer site containing mismatch. Increasing the hairpin length increases the specificity of dCas9 binding. Optimized gRNA and hpgRNA can be used to adjust Cas9 / dCas9 or Cpf1 binding affinity and specificity. Based on the size and structure of the hairpin, the hairpin of the hpgRNA can be accommodated in the DNA binding channel of the Cas9 / dCas9 molecule and protect it from degradation. In certain embodiments, the hairpin length, loop length and loop composition can be changed to allow for more delicate control of these characteristics. In certain embodiments, the hairpin length can be about 1 to about 20 nucleotides, or about 3 to about 10 nucleotides.For example, the hairpins can be 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10, 5 to 9, 5 to 8 , 5 to 7, 5 to 6, 6 to 20, 6 to 19, 6 to 18, 6 to 17, 6 to 16, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 19, 7 to 18, 7 to 17, 7 to 16, 7 to 15, 7 to 14, 7 to 13, 7 to 12, 7 to 11, 7 to 10, 7 to 9, 7 to 8, 8 to 20, 8 to 19, 8 to 18, 8 to 17, 8 to 16, 8 to 15, 8 to 14, 8 to 13, 8 to 12, 8 to 11 In some embodiments, the hairpin can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, or about 5 to about 8 nucleotides in length.

[0224] In some embodiments, the loop length can be about 1 to about 20 nucleotides, about 3 to about 10 nucleotides, or about 5 to about 8 nucleotides. For example, the loop length can be 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10 , 5 to 9, 5 to 8, 5 to 7, 5 to 6, 6 to 20, 6 to 19, 6 to 18, 6 to 17, 6 to 16, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 19, 7 to 18, 7 to 17, 7 to 16, 7 to 15, 7 to 14, 7 to 13, 7 to 12, 7 to 11, 7 to 10, 7 to 9, 7 to 8, 8 to 20, 8 to 19, 8 to 18, 8 to 17, 8 to 16, 8 to 15, 8 to 14, 8 to 13 In some embodiments, the loop may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, or about 5 to about 8 nucleotides in length.

[0225] In some embodiments, the loop composition can be about 1 to about 20 nucleotides, about 3 to about 10 nucleotides, or about 5 to about 8 nucleotides. For example, the loop composition can be 1 to 20, 1 to 19, 1 to 18, 1 to 17, 1 to 16, 1 to 15, 1 to 14, 1 to 13, 1 to 12, 1 to 11, 1 to 10, 1 to 9, 1 to 8, 1 to 7, 1 to 6, 1 to 5, 2 to 20, 2 to 19, 2 to 18, 2 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 3 to 20, 3 to 19, 3 to 18, 3 to 17, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 4 to 20, 4 to 19, 4 to 18, 4 to 17, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 20, 5 to 19, 5 to 18, 5 to 17, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10 , 5 to 9, 5 to 8, 5 to 7, 5 to 6, 6 to 20, 6 to 19, 6 to 18, 6 to 17, 6 to 16, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 20, 7 to 19, 7 to 18, 7 to 17, 7 to 16, 7 to 15, 7 to 14, 7 to 13, 7 to 12, 7 to 11, 7 to 10, 7 to 9, 7 to 8, 8 to 20, 8 to 19, 8 to 18, 8 to 17, 8 to 16, 8 to 15, 8 to 14, 8 to 13 In some embodiments, the loop may be comprised of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, or about 5 to about 8 nucleotides.

[0226] The composition can include a viral vector and a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as the optimized gRNA described herein). In certain embodiments, the composition includes a modified AAV vector and a nucleotide sequence encoding a CRISPR / Cas9-based system with at least one gRNA (such as the optimized gRNA described herein). The composition can further include donor DNA or transgenic. These compositions can be used for genome editing, genome engineering, and correcting or reducing the effects of gene mutations involved in genetic diseases.

[0227] The target gene may participate in the differentiation of cells or any other process that may need to activate, repress or destroy genes, or may have mutations such as deletions, frameshift mutations or nonsense mutations. If the target gene has a mutation that causes premature termination codons, abnormal splicing acceptor sites or abnormal splicing donor sites, There is at least one gRNA (such as optimized gRNA as described herein) based on CRISPR / Cas9 System or a system based on CRISPR / Cpf1 can be designed to recognize and combine the nucleotide sequence upstream or downstream of premature termination codons, abnormal splicing acceptor sites or abnormal splicing donor sites. There is at least one gRNA (such as optimized gRNA as described herein) based on CRISPR / Cas9 System or a system based on CRISPR / Cpf1 system can also be used to destroy normal gene splicing by targeting splicing acceptor and donor to induce premature termination codon skipping or restore destroyed reading frame. There is at least one gRNA (such as optimized gRNA as described herein) based on CRISPR / Cas9 System or a system based on CRISPR / Cpf1 system can or may not mediate the off-target changes for the protein coding region of genome.

[0228] In some embodiments, the CRISPR / Cas9-based system induces or represses gene expression of a target gene by at least about 1-fold, at least about 2-fold, at least about 3-fold, at least about 4-fold, at least about 5-fold, at least about 6-fold, at least about 7-fold, at least about 8-fold, at least about 9-fold, at least about 10-fold, at least 15-fold, at least 20-fold, at least 30-fold, at least 40-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, at least 100-fold, at least about 110-fold, at least 120-fold, at least 130-fold, at least 140-fold, at least 150-fold, at least 160-fold, at least 170-fold, at least 180-fold, at least 190-fold, at least 200-fold, at least about 300-fold, at least 400-fold, at least 500-fold, at least 600-fold, at least 700-fold, at least 800-fold, at least 90 ... The control level of gene expression of the target gene can be the gene expression level of the target gene in cells that have not been treated with any CRISPR / Cas9-based system.

[0229] a. Modified lentiviral vector

[0230] The composition for genome editing, genome alteration or altering the gene expression of a target gene may comprise a modified lentiviral vector. The modified lentiviral vector comprises a first polynucleotide sequence encoding a DNA targeting system and a second polynucleotide sequence encoding at least one sgRNA. The first polynucleotide sequence may be operably linked to a promoter. The promoter may be a constitutive promoter, an inducible promoter, a repressible promoter or a regulatable promoter.

[0231] The second polynucleotide sequence encodes at least one gRNA, such as an optimized gRNA as described herein. For example, the second polynucleotide sequence can encode at least one gRNA, at least two gRNAs, at least three gRNAs, at least four gRNAs, at least five gRNAs, at least six gRNAs, at least seven gRNAs, at least eight gRNAs, at least nine gRNAs, at least ten ... The second polynucleotide sequence can encode 1 to 50 gRNAs, 1 to 45 gRNAs, 1 to 40 gRNAs, 1 to 35 gRNAs, 1 to 30 gRNAs, 1 to 25 different gRNAs, 1 to 20 gRNAs, 1 to 16 gRNAs, 1 to 8 different gRNAs, 4 different gRNAs to 50 different gRNAs, 4 different gRNAs to 45 different gRNAs, 4 different gRNAs to 40 different gRNAs, 4 different gRNAs to 35 different gRNAs, 4 different gRNAs to 30 different gRNAs, 4 different gRNAs to 25 different gRNAs, 4 different gRNAs to 20 different gRNAs, 4 different gRNAs to 16 different gRNAs, 4 different gRNAs to 8 different gRNAs, 8 different gRNAs to 50 different gRNAs. NA, 8 different gRNAs to 45 different gRNAs, 8 different gRNAs to 40 different gRNAs, 8 different gRNAs to 35 different gRNAs, 8 different gRNAs to 30 different gRNAs, 8 different gRNAs to 25 different gRNAs, 8 different gRNAs to 20 different gRNAs, 8 different gRNAs to 16 different gRNAs, 16 different gRNAs to 50 different gRNAs, 16 different gRNAs to 45 different gRNAs, 16 different gRNAs to 40 different gRNAs, 16 different gRNAs to 35 different gRNAs, 16 different gRNAs to 30 different gRNAs, 16 different gRNAs to 25 different gRNAs, 16 different gRNAs to 20 different gRNAs.Each polynucleotide sequence encoding different gRNAs can be operably connected to a promoter. The promoters operably connected to different gRNAs can be the same promoter. The promoters operably connected to different gRNAs can be different promoters. The promoter can be a constitutive promoter, an inducible promoter, a repressible promoter or a regulatable promoter. At least one gRNA can bind to a target gene or locus. If more than one gRNA is included, each gRNA binds to different target regions within a target locus, or each gRNA binds to different target regions within different loci.

[0232] b. Adeno-associated virus vector

[0233] AAV can be used to deliver compositions to cells using a variety of construct configurations. For example, AAV can deliver a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system and a gRNA expression cassette on separate vectors. Alternatively, if a small Cas9 protein derived from a species such as Staphylococcus aureus or Neisseria meningitidis is used, expression cassettes for Cas9 and up to two gRNAs can be combined in a single AAV vector within the 4.7 kb packaging limit.

[0234] As described above, the composition comprises a modified adeno-associated virus (AAV) vector. The modified AAV vector can be capable of delivering a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system to a mammalian cell and expressing the system in the cell. For example, the modified AAV vector can be an AAV-SASTG vector (Piacentino et al. (2012) Human Gene Therapy [Human Gene Therapy] 23: 635-646). The modified AAV vector can be based on one or more of several capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, and AAV9. Modified AAV vectors can be based on AAV2 pseudotypes with alternative muscle-tropic AAV capsids, such as AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, and AAV / SASTG vectors, which efficiently transduce skeletal or cardiac muscle by systemic and local delivery (Seto et al. Current Gene Ther (2012) 12: 139-151).

[0235] 9. Target cells

[0236] As disclosed herein, gRNAs, such as the optimized gRNAs described herein, can be used with any type of cell in conjunction with the CRISPR / Cas9 system. In some embodiments, the cell is a bacterial cell, a fungal cell, an archaeal cell, a plant cell, or an animal cell, such as a mammalian cell. In some embodiments, this can be an organ or an animal organism. In some embodiments, the cells can be any cell type or cell line, including but not limited to 293-T cells, 3T3 cells, 721 cells, 9L cells, A2780 cells, A2780ADR cells, A2780cis cells, A172 cells, A20 cells, A253 cells, A431 cells, A-549 cells, ALC cells, B16 cells, B35 cells, BCP-1 cells, BEAS-2B cells, bEnd.3 cells, BHK-21 cells, BR 293 cells, BxPC3 cells, C2C12 cells, C3H-10T1 / 2 cells, C6 / 36 cells, Cal-27 cells, CHO cells, COR-L23 cells, COR-L23 / CPR cells, COR-L23 / 5010 cells, COR-L23 / R23 cells, COS-7 cells, COV-434 cells, CML cells, T1 cells, CMT cells, CT26 cells, D17 cells, DH82 cells, DU145 cells, DuCaP cells, EL4 cells, EM2 cells, EM3 cells, EMT6 / AR1 cells, EMT6 / AR10.0 cells, FM3 cells, H1299 cells, H69 cells, HB54 cells, HB55 cells, HCA2 cells, HEK-293 cells, HeLa cells, Hepa1c1c7 cells, HL-60 cells, HMEC cells, HT-29 cells, Jurkat cells, J558L cells, JY cells, K562 cells, Ku812 cells, KCL22 cells, KG1 cells, KYO1 cells, LNCap cells, Ma-MeI cells 1, 2, 3...48 cells, MC-38 cells, MCF-7 cells, MCF-10A cells, MDA-MB-231 cells, MDA-MB-468 cells, MDA-MB-435 cells, MDCK II cells, MDCK II cells, MG63 cells, MOR / 0.2R cells, MONO-MAC 6 cells, MRC5 cells, MTD-1A cells, MyEnd cells, NCI-H69 / CPR cells, NCI-H69 / LX10 cells, NCI-H69 / LX20 cells, NCI-H69 / LX4 cells, NIH-3T3 cells, NALM-1 cells, NW-145 cells, OPCN / OPCT cells, Peer cells, PNT-1A / PNT2 cells, Raji cells, RBL cells, RenCa cells, RIN-5F cells, RMA / RMAS cells, Saos-2 cells, Sf-9 cells, SiHa cells, SkBr3 cells, T2 cells, T-47D cells, T84 cells, THP1 cells, U373 cells, U87 cells, U937 cells, VCaP cells, Vero cells, WM39 cells, WT-49 cells, X63 cells, YAC-1 cells, YAR cells, GM12878, K562, H1 human embryonic stem cells, HeLa- S3, HepG2, HUVEC, SK-N-SH, IMR90, A549, MCF7, HMEC or LHCM, CD14+, CD20+, primary cardiac or hepatic cells, differentiated H1 cells, 8988T, Adult_CD4_naive, Adult_CD4_Th0, Adult_CD4_Th1, AG04449, AG04450, AG09309, AG09319, AG10803, AoAF, AoSMC, BC_Adipose_UHN00001, BC_Adrenal_Gland_H12803N, BC_Bladder_01- 11002, BC_Brain_H11058N, BC_Breast_02-03015, BC_Colon_01-11002, BC_Colon_H12817N, BC_Esophagus_01-11002, BC_Esophagus_H12817N, BC_Jejunum_H12817N, BC_Kidney_01-11002, BC_Kidney_H12817N, BC_Left_Ventricle_N41, BC_Leukocyte_UHN00204, BC_Liver_01-11002, BC_Lung_01-11002, BC_Lung_H12817N, BC_Pancreas_H12817N, BC_Penis_H12817N, BC_Pericardium_H12529N, BC_Placenta_UHN00189, BC_Prostate_Gland_H12817N, BC_Rectum_N29, BC_Skeletal_Muscle_01- 11002, BC_Skeletal_Muscle_H12817N, BC_Skin_01-11002, BC_Small_Intestine_01-11002, BC_Spleen_H12817N, BC_Stomach_01-11002、BC_Stomach_H12817N、BC_Testis_N30、BC_Uterus_BN0765、 BE2_C、BG02ES、BG02ES-EBD、BJ、bone_marrow_HS27a、 bone_marrow_HS5、bone_marrow_MSC、Breast_OC、Caco-2、 CD20+_RO01778、CD20+_RO01794、CD34+_Mobilized、 CD4+_Naive_Wb11970640、CD4+_Naive_Wb78495824、Cerebellum_OC、 Cerebrum_frontal_OC、Chorion、CLL、CMK、Colo829、Colon_BC、Colon_OC、Cord_CD4_naive、Cord_CD4_Th0、Cord_CD4_Th1、Decidua、 Dnd41、ECC-1、Endometrium_OC、Esophagus_BC、Fibrobl、Fibrobl_GM03348、FibroP、FibroP_AG08395、FibroP_AG08396、 FibroP_AG20443、Frontal_cortex_OC、GC_B_cell、Gliobla、GM04503、 GM04504、GM06990、GM08714、GM10248、GM10266、GM10847、GM12801、GM12812、GM12813、GM12864、GM12865、GM12866、 GM12867、GM12868、GM12869、GM12870、GM12871、GM12872、 GM12873、GM12874、GM12875、GM12878-XiMat、GM12891、GM12892、 GM13976、GM13977、GM15510、GM18505、GM18507、GM18526、 GM18951、GM19099、GM19193、GM19238、GM19239、GM19240、 GM20000、H0287、H1-neurons、H7-hESC、H9ES、H9ES-AFP-、H9ES- AFP+、H9ES-CM、H9ES-E、H9ES-EB、H9ES-EBD、HAc、HAEpiC、HA- h、HAL、HAoAF、HAoAF_6090101.11、HAoAF_6111301.9、HAoEC、HAoEC_7071706.1、HAoEC_8061102.1、HA-sp、HBMEC、HBVP、 HBVSMC、HCF、HCFaa、HCH、HCH_0011308.2P、HCH_8100808.2、 HCM、HConF、HCPEpiC、HCT-116、Heart_OC、Heart_STL003、HEEpiC、 HEK293、HEK293T、HEK293-T-REx、Hepatocytes、HFDPC、 HFDPC_0100503.2、HFDPC_0102703.3、HFF、HFF-Myc、HFL11W、HFL24W、HGF、HHSEC、HIPEpiC、HL-60、HMEpC、 HMEpC_6022801.3、HMF、hMNC-CB、hMNC-CB_8072802.6、hMNC- CB_9111701.6、hMNC-PB、hMNC-PB_0022330.9、hMNC-PB_0082430.9、hMSC-AT、hMSC-AT_0102604.12、hMSC-AT_9061601.12、hMSC-BM、 hMSC-BM_0050602.11、hMSC-BM_0051105.11、hMSC-UC、hMSC- UC_0052501.7、hMSC-UC_0081101.7、HMVEC-dAd、HMVEC-dBl-Ad、 HMVEC-dBl-Neo、HMVEC-dLy-Ad、HMVEC-dLy-Neo、HMVEC-dNeo、 HMVEC-LBl、HMVEC-LLy、HNPCEpiC、HOB、HOB_0090202.1、HOB_0091301、HPAEC、HPAEpiC、HPAF、HPC-PL、HPC- PL_0032601.13、HPC-PL_0101504.13、HPDE6-E6E7、HPdLF、HPF、 HPIEpC、HPIEpC_9012801.2、HPIEpC_9041503.2、HRCEpiC、HRE、 HRGEC、HRPEpiC、HSaVEC、HSaVEC_0022202.16、 HSaVEC_9100101.15、HSMM、HSMM_emb、HSMM_FSHD、HSMMtube、 HSMMtube_emb、HSMMtube_FSHD、HT-1080、HTR8svn、Huh-7、Huh-7.5、HVMF、HVMF_6091203.3、HVMF_6100401.3、HWP、 HWP_0092205、HWP_8120201.5、iPS、iPS_CWRU1、iPS_hFib2_iPS4、 iPS_hFib2_iPS5、iPS_NIHi11、iPS_NIHi7、Ishikawa、Jurkat、Kidney_BC、 Kidney_OC、LHCN-M2、LHSR、Liver_OC、Liver_STL004、Liver_STL011、 LNCaP、Loucy、Lung_BC、Lung_OC、Lymphoblastoid_cell_line、M059J、 MCF10A-Er-Src、MCF-7、MDA-MB-231、Medullo、Medullo_D341、 Mel_2183、Melano、Monocytes-CD14+、Monocytes-CD14+_RO01746、Monocytes-CD14+_RO01826、MRT_A204、MRT_G401、MRT_TTC549、 Myometr、Naive_B_cell、NB4、NH-A、NHBE、NHBE_RA、NHDF、 NHDF_0060801.3、NHDF_7071701.2、NHDF-Ad、NHDF-neo、NHEK、NHEM.f_M2、NHEM.f_M2_5071302.2、NHEM.f_M2_6022001、 NHEM_M2、NHEM_M2_7011001.2、NHEM_M2_7012303、NHLF、NT2- D1、Olf_neurosphere、Osteobl、ovcar-3、PANC-1、Pancreas_OC、 PanIsletD、PanIslets、PBDE、PBDEFetal、PBMC、PFSK-1、pHTE、 Pons_OC、PrEC、ProgFib、Prostate、Prostate_OC、Psoas_muscle_OC、Raji、 RCC_7860、RPMI-7951、RPTEC、RWPE1、SAEC、SH-SY5Y、 Skeletal_Muscle_BC、SkMC、SKMC、SkMC_8121902.17、SkMC_9011302、SK-N-MC, SK-N-SH_RA, Small_intestine_OC, Spleen_OC, Stellate, Stomach_BC, T_cells_CD4+, T-47D, T98G, TBEC, Th1, Th1_Wb33676984, Th1_Wb54553204, Th17, Th2, Th2_Wb33676984, Th2_Wb54553204, Treg_Wb78495824, Treg_Wb83319432, U2OS, U87, UCH-1, Urothelia, WERI-Rb-1 and WI-38. In some embodiments, the target cell can be any cell, such as a primary cell, HEK293 cell, 293Ts cell, SKBR3 cell, A431 cell, K562 cell, HCT116 cell, HepG2 cell, or K-Ras-dependent and K-Ras-independent cell populations.

[0237] 10. Methods of epigenome editing

[0238] The present disclosure relates to a method for epigenome editing in a target cell or subject using a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1. The method can be used to activate or repress a target gene. The method includes contacting the cell or subject with an effective amount of an optimized gRNA molecule as described herein and a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1. In some embodiments, the optimized gRNA is encoded by a polynucleotide sequence and packaged into a lentiviral vector. In some embodiments, the lentiviral vector comprises an expression cassette comprising a promoter operably connected to a polynucleotide sequence encoding sgRNA. In some embodiments, the promoter operably connected to the polynucleotide encoding the optimized gRNA is inducible.

[0239] 11. Site-specific DNA cleavage method

[0240] The present disclosure relates to methods for site-specific DNA cleavage in target cells or subjects using a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system. The method comprises contacting the cell or subject with an effective amount of an optimized gRNA molecule as described herein and a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system. In some embodiments, the optimized gRNA is encoded by a polynucleotide sequence and packaged into a lentiviral vector. In some embodiments, the lentiviral vector comprises an expression cassette comprising a promoter operably linked to a polynucleotide sequence encoding the sgRNA. In some embodiments, the promoter operably linked to the polynucleotide encoding the optimized gRNA is inducible.

[0241] The number of gRNAs administered to a cell or sample can be at least 1 gRNA, at least 2 different gRNAs, at least 3 different gRNAs, at least 4 different gRNAs, at least 5 different gRNAs, at least 6 different gRNAs, at least 7 different gRNAs, at least 8 different gRNAs, at least 9 different gRNAs, at least 10 different gRNAs, at least 11 different gRNAs, at least 12 different gRNAs, at least 13 different gRNAs, at least 14 different gRNAs, at least 15 different gRNAs, at least 16 different gRNAs, at least 17 different gRNAs, at least 18 different gRNAs, at least 18 different gRNAs, at least 20 different gRNAs, at least 25 different gRNAs, at least 30 different gRNAs, at least 35 different gRNAs, at least 40 different gRNAs, at least 45 different gRNAs, or at least 50 different gRNAs.The number of gRNAs administered to a cell can be from at least 1 gRNA to at least 50 different gRNAs, from at least 1 gRNA to at least 45 different gRNAs, from at least 1 gRNA to at least 40 different gRNAs, from at least 1 gRNA to at least 35 different gRNAs, from at least 1 gRNA to at least 30 different gRNAs, from at least 1 gRNA to at least 25 different gRNAs, from at least 1 gRNA to at least 20 different gRNAs, from at least 1 gRNA to at least 16 different gRNAs, from at least 1 gRNA to at least 12 different gRNAs, from at least 1 gRNA to at least 8 different gRNAs, from at least 1 gRNA to at least 4 different gRNAs, from at least 4 gRNAs to at least 50 different gRNAs, from at least 4 different gRNAs to at least 45 different gRNAs, from at least 4 different gRNAs to at least 40 different gRNAs, from at least 4 different gRNAs to at least 35 different gRNAs, from at least 4 different gRNAs to at least 30 In some embodiments, the present invention provides a plurality of different gRNAs, at least 4 different gRNAs to at least 25 different gRNAs, at least 4 different gRNAs to at least 20 different gRNAs, at least 4 different gRNAs to at least 16 different gRNAs, at least 4 different gRNAs to at least 12 different gRNAs, at least 4 different gRNAs to at least 8 different gRNAs, at least 8 different gRNAs to at least 50 different gRNAs, at least 8 different gRNAs to at least 45 different gRNAs, at least 8 different gRNAs to at least 40 different gRNAs, at least 8 different gRNAs to at least 35 different gRNAs, 8 different gRNAs to at least 30 different gRNAs, at least 8 different gRNAs to at least 25 different gRNAs, 8 different gRNAs to at least 20 different gRNAs, at least 8 different gRNAs to at least 16 different gRNAs, or 8 different gRNAs to at least 12 different gRNAs.

[0242] The gRNA may comprise a complementary polynucleotide sequence to the target DNA sequence followed by a PAM sequence. The gRNA may comprise a "G" at the 5' end of the complementary polynucleotide sequence. The gRNA may comprise a complementary nucleotide sequence of at least 10 base pairs, at least 11 base pairs, at least 12 base pairs, at least 13 base pairs, at least 14 base pairs, at least 15 base pairs, at least 16 base pairs, at least 17 base pairs, at least 18 base pairs, at least 19 base pairs, at least 20 base pairs, at least 21 base pairs, at least 22 base pairs, at least 23 base pairs, at least 24 base pairs, at least 25 base pairs, at least 30 base pairs, or at least 35 base pairs to the target DNA sequence followed by a PAM sequence. The PAM sequence may be "NGG," wherein "N" may be any nucleotide. The gRNA may target at least one of the promoter region, enhancer region, or transcribed region of the target gene. In some embodiments, the gRNA targets a nucleic acid sequence having a polynucleotide sequence of at least one of SEQ ID NOs: 13-148, 316, 317, or 320. The gRNA may comprise a nucleic acid sequence of at least one of SEQ ID NOs: 149-315, 321-323, or 326-329.

[0243] 12. Methods for correcting mutated genes and treating subjects

[0244] The present disclosure also relates to a method for correcting a mutant gene in a subject. The method includes administering a composition as described above to a cell of a subject. Compositions will have a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as optimized gRNA as described herein) delivered to the cell for purposes of restoring the expression of fully functional or partially functional proteins with a repair template or donor DNA, thereby replacing the entire gene or the region containing a mutation. A CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as optimized gRNA as described herein) can be used to introduce site-specific double-strand breaks at the targeted locus. When a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as optimized gRNA as described herein) binds to a target DNA sequence, site-specific double-strand breaks are produced, thereby allowing the target DNA to be cracked. This DNA cracking can stimulate natural DNA repair mechanisms, resulting in one of two possible repair pathways: homology-directed repair (HDR) or non-homologous end joining (NHEJ) pathway.

[0245] The present disclosure relates to genome editing with a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as optimized gRNA as described herein) in the absence of a repair template, which can efficiently correct the reading frame and restore the expression of functional proteins involved in genetic diseases. The disclosed CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as optimized gRNA as described herein) can involve the use of homology-directed repair or nuclease-mediated non-homologous end joining (NHEJ) correction methods, which makes it possible to efficiently correct in primary cell lines that may not be suitable for homologous recombination or gene correction based on selection that is restricted in proliferation. This strategy integrates the rapid and robust organization of an active CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with an efficient gene editing method for treating genetic diseases caused by mutations in non-essential coding regions that cause frameshifts, premature stop codons, abnormal splicing donor sites, or abnormal splicing acceptor sites.

[0246] a. Nuclease-mediated non-homologous end joining

[0247] Restoring protein expression from endogenous mutant genes can be achieved through template-free NHEJ-mediated DNA repair. In contrast to transient methods targeting target gene RNA, correcting the target gene reading frame in the genome by transiently expressing a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as the optimized gRNA described herein) can result in permanent restoration of target gene expression in each modified cell and all its progeny.

[0248] Nuclease-mediated NHEJ gene correction can correct the target gene of mutation and provide several potential advantages over HDR approach.For example, NHEJ does not need the donor template that may cause non-specific insertion mutagenesis.Compared with HDR, NHEJ is efficiently operated in all stages of the cell cycle, Therefore it can be effectively used in both cells in the cycle and post-mitotic cells (such as muscle fibers).Relative to the pharmacological forced reading of oligonucleotide-based exon skipping or stop codons, this provides alternative robust permanent gene recovery, and in theory can be required to be as little as one drug treatment.In addition to the plasmid electroporation method described herein, the gene correction based on NHEJ using a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 and other engineered nucleases including meganucleases and zinc finger nucleases can be combined with other existing in vitro and in vivo platforms for cell-based and gene-based therapies. For example, CRISPR / Cas9-based systems or CRISPR / Cpf1-based systems could enable DNA-free genome editing approaches through mRNA-based gene transfer or delivery as purified cell-permeable proteins, which would overcome any potential for insertional mutagenesis.

[0249] b. Homology-directed repair

[0250] Restoring protein expression from an endogenous mutant gene can involve homology-directed repair. The above method further comprises administering a donor template to the cell. The donor template can include a nucleotide sequence encoding a fully functional protein or a partially functional protein. For example, the donor template can include a miniaturized dystrophin construct, referred to as a minidystrophin ("minidys"), a fully functional dystrophin construct for restoring a mutant dystrophin gene or a fragment of a dystrophin gene, which, after homology-directed repair, results in restoration of the mutant dystrophin gene.

[0251] 13. Genome Editing Methods

[0252] The present disclosure also relates to genome editing using the above-mentioned CRISPR / Cas9-based system or the CRISPR / Cpf1-based system to restore the expression of fully functional or partially functional proteins with a repair template or donor DNA, thereby replacing the entire gene or the region containing the mutation. The CRISPR / Cas9-based system or the CRISPR / Cpf1-based system can be used to introduce site-specific double-strand breaks at the targeted locus. When gRNA is used to make the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system bind to the target DNA sequence, site-specific double-strand breaks are produced, thereby allowing the target DNA to be cracked. The CRISPR / Cas9-based system and the CRISPR / Cpf1-based system have the advantage of advanced genome editing due to their high rate of success and efficient genetic modification. This DNA cracking can stimulate natural DNA repair mechanisms, leading to one of two possible repair pathways: homology-directed repair (HDR) or non-homologous end joining (NHEJ) pathway.

[0253] The present disclosure relates to genome editing with a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system in the absence of a repair template, which can efficiently correct the reading frame and restore the expression of functional proteins involved in genetic diseases. The disclosed CRISPR / Cas9-based system or the CRISPR / Cpf1-based system and method may involve the use of homology-directed repair or nuclease-mediated non-homologous end joining (NHEJ) correction methods, which make it possible to efficiently correct in primary cell lines that may not be suitable for homologous recombination or gene correction based on selection that is restricted in proliferation. This strategy integrates the rapid and robust organization of active CRISPR / Cas9-based systems or CRISPR / Cpf1-based systems with efficient gene editing methods for the treatment of genetic diseases caused by mutations in non-essential coding regions that cause frameshifts, premature stop codons, abnormal splicing donor sites, or abnormal splicing acceptor sites.

[0254] The present disclosure provides methods for correcting mutant genes in cells and treating subjects suffering from genetic diseases, such as DMD. The method may include administering to cells or subjects a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 as described above, a polynucleotide or vector encoding the system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1, or the system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1. The method may include administering a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1, such as administering a Cas9 protein, a Cpf1 protein, a Cas9 fusion protein containing a second domain, a nucleotide sequence encoding the Cas9 protein, the Cpf1 protein or the Cas9 fusion protein, and / or at least one gRNA, wherein the gRNA targets different DNA sequences. The target DNA sequences may be overlapping. As described above, the number of gRNAs administered to the cell can be at least 1 gRNA, at least 2 different gRNAs, at least 3 different gRNAs, at least 4 different gRNAs, at least 5 different gRNAs, at least 6 different gRNAs, at least 7 different gRNAs, at least 8 different gRNAs, at least 9 different gRNAs, at least 10 different gRNAs, at least 15 different gRNAs, at least 20 different gRNAs, at least 30 different gRNAs, or at least 50 different gRNAs. The gRNA may comprise a nucleic acid sequence of at least one of SEQ ID NOs: 149-315, 321-323, or 326-329. The method may involve homology-directed repair or non-homologous end joining.

[0255] 14. Constructs and Plasmids

[0256] The composition as described above may include a gene construct encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 as disclosed herein. A gene construct (such as a plasmid) may include a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 such as Cas9 protein, Cpf1 protein and Cas9 fusion protein and / or at least one nucleic acid encoding an optimized gRNA as described herein. The composition as described above may include a gene construct encoding a modified AAV vector and a nucleic acid sequence encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 with at least one gRNA (such as optimized gRNA as described herein). A gene construct (such as a plasmid) may include a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 with at least one gRNA (such as optimized gRNA as described herein). A gene construct (such as a plasmid) may include a nucleic acid encoding a modified lentiviral vector as disclosed herein. A gene construct (such as a plasmid) may include a nucleic acid encoding Cas9-fusion protein and at least one sgRNA. The genetic construct may be present in the cell as a functional extrachromosomal molecule. The genetic construct may be a linear minichromosome, including a centromere, a telomere, or a plasmid or a cosmid.

[0257] The gene construct can also be part of the genome of a recombinant viral vector, including recombinant lentivirus, recombinant adenovirus, and recombinant adeno-associated virus. The gene construct can be part of the genetic material of a live attenuated microorganism or recombinant microbial vector living in a cell. The gene construct can contain regulatory elements for gene expression of the nucleic acid coding sequence. The regulatory elements can be promoters, enhancers, initiation codons, termination codons, or polyadenylation signals.

[0258] The nucleic acid sequence may constitute a genetic construct that may be a vector. The vector may be capable of expressing a fusion protein, such as a Cas9 fusion protein, in mammalian cells. The vector may be recombinant. The vector may contain a heterologous nucleic acid encoding a Cas9 fusion protein. The vector may be a plasmid. The vector may be used to transfect cells with a nucleic acid encoding a Cas9 fusion protein, and culture and maintain the transformed host cells under conditions in which the Cas9 fusion protein is systematically expressed.

[0259] Coding sequences can be optimized for stability and high levels of expression. In some cases, codons are selected to reduce secondary structure formation in the RNA, such as secondary structure formed by intramolecular bonding.

[0260] The vector may comprise a heterologous nucleic acid encoding a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, and may further comprise a start codon (which may be upstream of the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence) and a stop codon (which may be downstream of the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence). The start codon and the stop codon may be in the reading frame with the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence. The vector may also comprise a promoter operably linked to the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence. The promoter operably linked to the CRISPR / Cas9-based system or CRISPR / Cpf1-based system coding sequence can be a promoter from Simian Virus 40 (SV40), Mouse Mammary Tumor Virus (MMTV) promoter, Human Immunodeficiency Virus (HIV) promoter such as Bovine Immunodeficiency Virus (BIV) long terminal repeat (LTR) promoter, Moloney virus promoter, Avian Leukosis Virus (ALV) promoter, Cytomegalovirus (CMV) promoter such as CMV immediate early promoter, Epstein-Barr virus (EBV) promoter or Rous sarcoma virus (RSV) promoter. The promoter can also be a promoter from a human gene such as human ubiquitin C (hUbC), human actin, human myosin, human hemoglobin, human muscle creatine or human metallothionein. The promoter can also be a natural or synthetic tissue-specific promoter, such as a muscle or skin-specific promoter. Examples of such promoters are described in US Patent Application Publication No. US 20040175727, the disclosure of which is incorporated herein by reference in its entirety.

[0261] The vector may further comprise a polyadenylation signal, which may be downstream of a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system. The polyadenylation signal may be an SV40 polyadenylation signal, an LTR polyadenylation signal, a bovine growth hormone (bGH) polyadenylation signal, a human growth hormone (hGH) polyadenylation signal, or a human β-globin polyadenylation signal. The SV40 polyadenylation signal may be a polyadenylation signal from a pCEP4 vector (Invitrogen, San Diego, CA).

[0262] The vector may also be included in a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system (i.e., Cas9 protein, Cpf1 protein or Cas9 fusion protein coding sequence or sgRNA (such as the optimized gRNA described herein)). An enhancer may be necessary for DNA expression. The enhancer may be human actin, human myosin, human hemoglobin, human muscle creatine or a viral enhancer, such as an enhancer from CMV, HA, RSV or EBV. Polynucleotide functional enhancers are described in U.S. Patent Nos. 5,593,972, 5,962,428 and WO 94 / 016737, the contents of which are fully incorporated by reference. The vector may also include a mammalian origin of replication to maintain the vector outside the chromosome and to produce multiple copies of the vector in the cell. The vector may also include a regulatory sequence that is very suitable for gene expression in mammalian or human cells to which the vector is administered. The vector may also include a reporter gene, such as green fluorescent protein ("GFP") and / or a selective marker, such as hygromycin ("Hygro").

[0263] The vector can be an expression vector or a system for producing proteins by conventional techniques and readily available starting materials, including Sambrook et al., Molecular Cloning and Laboratory Manual, 2nd ed., Cold Spring Harbor (1989), which is incorporated by reference in its entirety. In some embodiments, the vector can comprise a nucleic acid sequence encoding a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, including a nucleic acid sequence encoding a Cas9 protein, a Cpf1 protein, or a Cas9 fusion protein and a nucleic acid sequence encoding at least one gRNA comprising a nucleic acid sequence of at least one of SEQ ID NOs: 149-315, 321-323, or 326-329.

[0264] 15. Pharmaceutical compositions

[0265] The composition may be in the form of a pharmaceutical composition. The pharmaceutical composition may include about 1 ng to about 10 mg of DNA encoding a CRISPR / Cas9-based system, a CRISPR / Cpf1-based system, or a CRISPR / Cas9-based system protein component (i.e., Cas9 protein, Cpf1 protein, or Cas9 fusion protein). The pharmaceutical composition may include about 1 ng to about 10 mg of a modified AAV vector and DNA encoding a nucleotide sequence of a CRISPR / Cas9-based system with at least one gRNA (such as the optimized gRNA described herein). The pharmaceutical composition may include about 1 ng to about 10 mg of modified lentiviral vector DNA. The pharmaceutical composition according to the present invention is formulated according to the mode of administration to be used. In the case where the pharmaceutical composition is an injectable pharmaceutical composition, they are sterilized, pyrogen-free, and particle-free. It is preferred to use isotonic formulations. Generally speaking, additives for isotonicity may include sodium chloride, dextrose, mannitol, sorbitol, and lactose. In some cases, isotonic solutions are preferred, such as phosphate-buffered saline. Stabilizers include gelatin and albumin. In some embodiments, a vasoconstrictor is added to the formulation.

[0266] Compositions can further comprise a pharmaceutically acceptable excipient. A pharmaceutically acceptable excipient can be a functional molecule as a vehicle, adjuvant, carrier or diluent. A pharmaceutically acceptable excipient can be a transfection facilitator, which can include surfactants (such as immunostimulatory complexes (ISCOMS)), Freund's incomplete adjuvant, LPS analogs (including monophosphoryl lipid A), muramyl peptides, quinone analogs, vesicles (such as squalene and squalene), hyaluronic acid, lipids, liposomes, calcium ions, viral proteins, polyanions, polycations or nanoparticles or other known transfection facilitators.

[0267] Transfection facilitator is polyanion, polycation, including poly-L-glutamic acid (LGS) or lipid. Transfection facilitator is poly-L-glutamic acid, and more preferably poly-L-glutamic acid is present in the composition for genome editing with a concentration of less than 6mg / ml. Transfection facilitator may also include surfactant (such as immunostimulatory complex (ISCOMS)), Freund's incomplete adjuvant, LPS analogs (including monophosphoryl lipid A), muramyl peptide, quinone analogs and vesicles (such as squalene and squalene), and hyaluronic acid may also be used to be administered in combination with gene constructs. In certain embodiments, the DNA vector encoding the composition may also include transfection facilitator such as lipid, liposome (including lecithin liposomes or other liposomes known in the art, as DNA-liposome mixture (see, for example, WO9324640)), calcium ion, viral protein, polyanion, polycation or nanoparticle or other known transfection facilitator. Preferably, transfection facilitator is polyanion, polycation, including poly-L-glutamic acid (LGS) or lipid.

[0268] 16. Constructs and Plasmids

[0269] Compositions as described above may include a gene construct encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 as disclosed herein. Gene constructs (such as plasmids or expression vectors) may include nucleic acids encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 and / or at least one gRNA (such as optimized gRNA as described herein). Compositions as described above may include a gene construct encoding a modified lentiviral vector and a nucleic acid sequence encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 as disclosed herein. Gene constructs (such as plasmids) may include nucleic acids encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1. Compositions as described above may include a gene construct encoding a modified lentiviral vector. Gene constructs (such as plasmids) may include nucleic acids encoding a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 and at least one sgRNA (such as optimized gRNA as described herein). Gene constructs may be present in cells as extrachromosomal molecules that function. The genetic construct may be a linear minichromosome including a centromere, telomeres or a plasmid or cosmid.

[0270] The gene construct can also be part of the genome of a recombinant viral vector, including recombinant lentivirus, recombinant adenovirus, and recombinant adeno-associated virus. The gene construct can be part of the genetic material of a live attenuated microorganism or recombinant microbial vector living in a cell. The gene construct can contain regulatory elements for gene expression of the nucleic acid coding sequence. The regulatory elements can be promoters, enhancers, initiation codons, termination codons, or polyadenylation signals.

[0271] The nucleic acid sequence may constitute a genetic construct that may be a vector. The vector will be capable of expressing a fusion protein, such as a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, in mammalian cells. The vector may be recombinant. The vector may contain a heterologous nucleic acid encoding a fusion protein, such as a CRISPR / Cas9-based system. The vector may be a plasmid. The vector may be used to transfect cells with a nucleic acid encoding a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, and to culture and maintain the transformed host cells under conditions where expression of the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system occurs.

[0272] Coding sequences can be optimized for stability and high levels of expression. In some cases, codons are selected to reduce secondary structure formation in the RNA, such as secondary structure formed by intramolecular bonding.

[0273] The vector may include a heterologous nucleic acid encoding a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system, and may further include a start codon (which may be upstream of a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system coding sequence) and a stop codon (which may be downstream of a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system coding sequence). The start codon and the stop codon may be in a reading frame together with the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence. The vector may also include a promoter operably connected to the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system coding sequence. The CRISPR / Cas9-based system or the CRISPR / Cpf1-based system can be under light-inducible or chemical-inducible control to enable dynamic control of time and space. The promoter operably linked to the CRISPR / Cas9-based system or CRISPR / Cpf1-based system coding sequence can be a promoter from Simian Virus 40 (SV40), Mouse Mammary Tumor Virus (MMTV) promoter, Human Immunodeficiency Virus (HIV) promoter such as Bovine Immunodeficiency Virus (BIV) long terminal repeat (LTR) promoter, Moloney virus promoter, Avian Leukosis Virus (ALV) promoter, Cytomegalovirus (CMV) promoter such as CMV immediate early promoter, Epstein-Barr virus (EBV) promoter or Rous sarcoma virus (RSV) promoter. The promoter can also be a promoter from a human gene such as human ubiquitin C (hUbC), human actin, human myosin, human hemoglobin, human muscle creatine or human metallothionein. The promoter can also be a natural or synthetic tissue-specific promoter, such as a muscle or skin-specific promoter. Examples of such promoters are described in US Patent Application Publication No. US20040175727, the disclosure of which is incorporated herein by reference in its entirety.

[0274] The vector may further comprise a polyadenylation signal, which may be downstream of a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system. The polyadenylation signal may be an SV40 polyadenylation signal, an LTR polyadenylation signal, a bovine growth hormone (bGH) polyadenylation signal, a human growth hormone (hGH) polyadenylation signal, or a human β-globin polyadenylation signal. The SV40 polyadenylation signal may be a polyadenylation signal from a pCEP4 vector (Invitrogen, San Diego, CA).

[0275] The vector may also be included in a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system and / or an enhancer upstream of sgRNA (such as the optimized gRNA described herein). The enhancer may be necessary for DNA expression. The enhancer may be human actin, human myosin, human hemoglobin, human muscle creatine or a viral enhancer, such as an enhancer from CMV, HA, RSV or EBV. Polynucleotide functional enhancers are described in U.S. Patent Nos. 5,593,972, 5,962,428 and WO94 / 016737, the contents of which are fully incorporated by reference. The vector may also include a mammalian origin of replication to maintain the vector outside the chromosome and to produce multiple copies of the vector in the cell. The vector may also include a regulatory sequence that is very suitable for gene expression in mammalian or human cells to which the vector is administered. The vector may also include a reporter gene, such as green fluorescent protein ("GFP") and / or a selective marker, such as hygromycin ("Hygro").

[0276] The vector can be an expression vector or a system for producing proteins by conventional techniques and readily available starting materials, including Sambrook et al., Molecular Cloning and Laboratory Manual, 2nd ed., Cold Spring Harbor (1989), which is incorporated by reference in its entirety. In some embodiments, the vector can comprise a nucleic acid sequence encoding a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system and a nucleic acid sequence encoding at least one gRNA (such as an optimized gRNA described herein).

[0277] In some embodiments, the gRNA (such as the optimized gRNA described herein) is encoded by a polynucleotide sequence and packaged into a lentiviral vector. In some embodiments, the lentiviral vector comprises an expression cassette. The expression cassette may comprise a promoter operably linked to a polynucleotide sequence encoding the gRNA (such as the optimized gRNA described herein). In some embodiments, the promoter operably linked to the polynucleotide encoding the gRNA is inducible.

[0278] i. Adeno-associated virus vector

[0279] As described above, the composition comprises a modified adeno-associated virus (AAV) vector. The modified AAV vector can have enhanced myocardial and skeletal muscle tissue tropism. The modified AAV vector can deliver a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system with at least one gRNA (such as the optimized gRNA described herein) into mammalian cells and express the system in the cells. For example, the modified AAV vector can be an AAV-SASTG vector (Piacentino et al. (2012) Human Gene Therapy 23: 635-646). The modified AAV vector can deliver nucleases to skeletal and myocardial muscles in vivo. The modified AAV vector can be based on one or more of several capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8 and AAV9. Modified AAV vectors can be based on AAV2 pseudotypes with alternative muscle-tropic AAV capsids, such as AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, and AAV / SASTG vectors, which efficiently transduce skeletal or cardiac muscle by systemic and local delivery (Seto et al. Current Gene Ther (2012) 12: 139-151).

[0280] 17. Delivery Method

[0281] Provided herein are systems for delivering CRISPR / Cas9 or systems based on CRISPR / Cpf1 and optimization gRNA as described herein to provide gene constructs and / or protein methods based on CRISPR / Cas9 or systems based on CRISPR / Cpf1. The delivery of a system based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 and optimization gRNA as described herein can be transfection or electroporation of one or more nucleic acid molecules expressed in a cell based on CRISPR / Cas9 or a system based on CRISPR / Cpf1 and optimization gRNA as described herein and delivered to the cell surface. The system based on CRISPR / Cas9 or the system based on CRISPR / Cpf1 can be delivered to cells. The nucleic acid molecules can be electroporated using BioRad Gene Pulser Xcell or Amaxa Nucleofector IIb devices or other electroporation devices. Several different buffers can be used, including BioRad electroporation solution, Sigma phosphate buffered saline product #D8537 (PBS), Invitrogen OptiMEM I (OM), or Amaxa Nucleofector solution V (NV). Transfection can include a transfection reagent such as Lipofectamine 2000.

[0282] Vectors encoding CRISPR / Cas9-based systems or CRISPR / Cpf1-based system proteins can be delivered to modified target cells in tissues or subjects by DNA injection (also referred to as DNA vaccination) with or without in vivo electroporation, liposome mediation, nanoparticle promotion, and / or recombinant vectors. The recombinant vector can be delivered by any viral means. The viral means can be a recombinant lentivirus, a recombinant adenovirus, and / or a recombinant adeno-associated virus.

[0283] The nucleotides encoding the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system protein can be introduced into the cell to induce the gene expression of the target gene. For example, one or more nucleotide sequences encoding the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system for the target gene can be introduced into a mammalian cell. When the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system is delivered to the cell, the vector enters the mammalian cell, and the transfected cell will express the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system. The CRISPR / Cas9-based system or the CRISPR / Cpf1-based system can be given to a mammal to induce or regulate the gene expression of the target gene in the mammal. The mammal can be a human, a non-human primate, a cow, a pig, a sheep, a goat, an antelope, a bison, a buffalo, a bovine, a deer, a thorn, an elephant, a llama, an alpaca, a mouse, a rat or a chicken, and preferably a human, a cow, a pig or a chicken.

[0284] Methods for introducing nucleic acid into host cells are known in the art, and any known method can be used to introduce nucleic acid (e.g., expression construct) into cells. Suitable methods include, for example, viral or phage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran-mediated transfection, liposome-mediated transfection, gene gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, etc. In certain embodiments, compositions can be delivered by mRNA delivery and ribonucleoprotein (RNP) complex delivery.

[0285] 18. Giving Channels

[0286] The composition can be administered to the subject by different approaches, including oral, parenteral, sublingual, transdermal, rectal, transmucosal, topical, via inhalation, buccal administration, intrapleural, intravenous, intraarterial, intraperitoneal, subcutaneous, intramuscular, intranasal, intrathecal and intraarticular or a combination thereof. For veterinary use, the composition is given as a suitable acceptable formulation according to normal veterinary practice. Veterinarians can easily determine the dosage regimen and route of administration that is most suitable for a specific animal. The composition can be administered by traditional syringes, needleless injection devices, "microprojectile bombardment gene guns" or other physical methods such as electroporation ("EP"), "hydrodynamic methods" or ultrasound.

[0287] The composition can be delivered to a mammal by several techniques including DNA injection (also referred to as DNA vaccination), with or without in vivo electroporation, liposome mediation, nanoparticle promotion, recombinant vectors such as recombinant lentivirus, recombinant adenovirus, and recombinant adeno-associated virus. The composition can be injected into skeletal muscle or cardiac muscle. For example, the composition can be injected into the tibialis anterior muscle.

[0288] 19. Test kit

[0289] Provided herein is a kit that can be used for site-specific DNA binding. The kit includes a composition as described above and instructions for use of the composition. The instructions contained in the kit can be affixed to packaging material, or can be included as a packaging insert. Although instructions are typically written or printed materials, they are not limited to this. Any medium that can store such instructions and transmit them to end users is encompassed by this disclosure. Such media include, but are not limited to, electronic storage media (e.g., disks, tapes, cylinders, chips), optical media (e.g., CDROMs), etc. As used herein, the term "instructions" may include the address of an internet website that provides instructions.

[0290] As described above, the composition can include a modified lentiviral vector and a nucleotide sequence encoding a CRISPR / Cas9-based system and an optimized gRNA. The CRISPR / Cas9-based system as described above can be included in the kit to specifically bind to and target a specific regulatory region of a target gene.

[0291] 20. Examples

[0292] The foregoing may be better understood by reference to the following examples, which are presented for purposes of illustration and are not intended to limit the scope of the invention.

[0293] Example 1

[0294] Materials and Methods

[0295] Materials: Tris-HCl (pH 7.6) buffer was obtained from Corning Life Sciences. L-glutamic acid monopotassium salt monohydrate, dithiothreitol (DTT), and magnesium chloride were obtained from Sigma Aldrich Co., LLC.

[0296] Cloning of Cas9, dCas9, and sgRNA expression plasmids; Standard techniques were used to clone, express, and purify plasmids encoding Cas9, dCas9, and sgRNA targeting the AAVS1 locus of human chromosome 19. Standard techniques were also used to generate DNA substrates for imaging—(i) an 1198bp substrate derived from the AAVS1 locus segment of human chromosome 19; (ii) an "engineered" 989bp DNA substrate containing a series of six complete, partial, or mismatched target sites; and (iii) a 1078bp "nonsense" substrate that does not have homology (>3bp) to the protospacer. Plasmids encoding wild-type Cas9 and dCas9 were obtained from Addgene (plasmid 39312 and plasmid 47106). Gateway cloning (Life Technologies) was used to clone plasmids for expressing Cas9 and dCas9 in bacteria. In short, PCR was used to amplify the Cas9 and dCas9 genes and add flanking attL1 and attL2 sites. BP recombination was performed to transfer these genes into a shuttle vector, followed by LP recombination to transfer these genes into pDest17, which had an N-terminal hexa-histidine tag added (Life Technologies). Plasmids encoding chimeric sgRNAs and sgRNA variants (described below) were cloned as previously described (Perez-Pinera et al., (2013) Nature Methods, 10, 973-976).

[0297] Expression and purification of Cas9 and dCas9. The plasmid encoding Cas9 or dCas9 was transformed into SoluBL21 competent cells (Genlantis) according to standard techniques (Sambrook, J., Fritsch, EF and Maniatis, T. (1989) Molecular cloning [Molecular cloning]. Cold Spring Harbor Laboratory Press, New York). A single colony was used to inoculate a 25 mL starter culture. The 25 mL starter culture was grown overnight and used to inoculate a 1 L culture. The inoculated 1 L culture was grown at 25 ° C for 5 hours, after which the temperature was lowered to 16 ° C, and protein expression was induced by adding 0.1 mM IPTG. The induced culture was grown for another 12 hours at 16 ° C. The cells were harvested by centrifugation at 4000 x g and stored at -80 ° C for long-term storage.

[0298] The cell pellet was resuspended in 30 mL of lysis buffer (50 mM Tris-HCl, 500 mM NaCl, 10 mM MgCl2, 10% v / v glycerol, 0.2% Triton-1000, and 1 mM PMSF). The cell suspension was lysed by sonication at a 30% duty cycle for 5 minutes. The suspension was then centrifuged at 12,000 x g for 30 minutes. The supernatant was then taken and incubated with Ni-NTA resin (Qiagen) for 30 minutes under gentle agitation. The resin was then loaded onto a column, washed with wash buffer (35 mM imidazole, 50 mM Tris-HCl, 500 mM NaCl, 10 mM MgCl2, 10% v / v glycerol), and eluted with elution buffer (120 mM imidazole, 50 mM Tris-HCl, 500 mM NaCl, 10 mM MgCl2, 10% v / v glycerol). Ultracel-30k centrifugal filters were then used to exchange the solvent into storage buffer (50 mM Tris-HCl, 500 mM NaCl, 10 mM MgCl2, 10% v / v glycerol). The samples were then aliquoted and frozen at -80°C. Representative polyacrylamide SDS gels of purified Cas9 and dCas9 are presented in Figure S1, indicating approximately >95% purity.

[0299] Expression and purification of sgRNA and guide RNA variants. Guide RNA was transcribed in vitro using the MEGAshortscript T7 transcription kit (Life Technologies). DNA templates with T7 promoters were generated from guide RNA plasmids via PCR and the reaction was set up according to the manufacturer's instructions. T7 templates for guide RNA (tru-gRNA) with 2 nucleotides truncated from its 5' end and guide RNA (hp-gRNA) with a 5' extension region forming a hairpin were generated from standard gRNA plasmids by PCR. RNA was then purified using phenol-chloroform extraction using standard techniques (Sambrook et al. (1989) Molecular Cloning. Cold Spring Harbor Laboratory Press, New York).

[0300] Generation of DNA substrates.Use DNeasy test kit (Qiagen) to extract and purify genomic DNA according to the scheme HEK293T cell line of manufacturer.Then use PCR amplification AAVS1 locus.From genomic DNA, use the following primers from integrated DNA technology company (Integrated DNA Technologies) (IDT) via direct PCR to construct the substrate of 1198bpAAVS1 source: 5'-\Bt\-CCAGGATCAGTGAAACGCAC-3' and 5'-GAGCTCTACTGGCTTCTGCG-3', wherein\Bt\ represents that the primer is biotinylated at the 5' end.The "through engineering approaches" DNA substrate containing a series of PAM and complete or partial prototype spacer sites is sorted into two gBlock fragments, each gBlock fragment contains EcoRI restriction site at one end.Substrate is digested, linked together, and then enriched via PCR using the following primers (Integrated DNA Technologies, IDT): 5'-\Bt\-CATGACGTGCAGCAAGC-3' and 5'-CGACGATGCGCTGAATC-3'. To construct a "nonsense" substrate that does not contain sites showing homology (greater than 3 bp) to the protospacer: a 690 bp DNA construct containing a series of restriction sites was synthesized (GeneScript, Inc.), and an additional DNA length from λ DNA (New England Biolabs) was subcloned into the construct; the 1078 bp substrate was then PCR amplified using the following primers (IDT): 5'-\Bt\- GACCTGCAGGCATGCAAGCTTGG-3' and 5'-CAGCGTCCCCGGTTGTGAATCT-3'. All DNA was gel purified, diluted to 25 nM in working buffer (20 mM Tris-HCl (pH 7.6), 100 mM potassium glutamate, 5 mM MgCl2 and 0.4 mM DTT), and incubated with 40x excess monomeric streptavidin (Howarth et al., (2006) Nature Methods, 3, 267-273) for 10 minutes before incubation with Cas9 / dCas9.

[0301] Sodium dodecyl sulfate-polyacrylamide gels of purified Cas9 and dCas9 are shown in Figures 8A-8B This indicates a purity of approximately >95%.

[0302] Atomic force microscopy. Atomic force microscopy (AFM) was performed in air using a Bruker Nanoscope V Multimode with an RTSEP (Bruker) probe (nominal spring constant 40 N / m, resonant frequency, 300 kHz). Prior to the experiment, the protein and guide RNA were mixed in a 1:1.5 ratio for 10 minutes. The protein and DNA were mixed in a working buffer solution at room temperature for at least 10 minutes (maximum 35 minutes), deposited on freshly cleaved mica (Ted Pella, Inc.) treated with 3-aminopropylsiloxane (prepared as described previously (24)) for 8 seconds, rinsed with ultrapure (>17 MΩ) water, and dried in air. The protein was briefly centrifuged before incubation with the DNA. When using standard sgRNA, at least 4 preparations were imaged for each experimental condition, and at least 2 preparations were imaged for experiments using other guide RNA variants. Typically, for each sample, images were acquired at 1-1.5 lines / s at a pixel resolution of 1024 x 1024 over a 2.75 μm square area or 2048 x 2048 over a 5.5 μm square area. For each experimental condition, images of several thousand (approximately 2500-6000) DNA molecules were resolved.

[0303] DNA tracing and thinning using sub-pixel resolution. The AFM images obtained are smoothed and leveled (planar state, line by line and leveled by 3rd order polynomial) using the open source image analysis software for scanning probe microscope Gwyddion, and then exported to MATLAB (Mathworks). For the presence of clearly identifiable streptavidin labels, at least one combined Cas9 / dCas9 molecule and ensuring that there is no clear end-to-end path with other DNA molecules that is aggregated or overlapped, 151 × 151 pixels (405nm × 405nm) areas containing each DNA molecule are sorted. The outline of DNA is manually traced, and the estimated boundaries of streptavidin and Cas9 / dCas9 are marked. Then the method based on Wiggins et al. (2006) Nature nanotechnology [Natural Nanotechnology], 1,137-141) is used to carry out algorithmic refinement of track. Starting from the weighted centroid of streptavidin (x1), the position of the next element of the backbone (x2) is estimated by stepping 2.5 nm toward the nearest manually traced point outside the estimated boundary of streptavidin. An 11-pixel line is drawn at x2 with two times linear interpolation of the DNA image perpendicular to the (x1-x2) line segment. x2 is repositioned to the position with the maximum feature height on the normal line and then adjusted to 2.5 nm from x1 on the new (x1-x2) line. x3...x is then iteratively estimated using the nearest manually traced points. n The position of , to produce an initial guess for the next skeleton position, is then corrected as before, and the correction process continues until point x n Less than 2.5 mm from the end of the traced DNA molecule. i When the estimated boundary of the Cas9 / dCas9 molecule is entered, the position of the DNA is estimated as a cubic Hermite spline (using the point x i-1 、x i 、x j and x j+1 , where x j is the first point of the hand-drawn trajectory outside the estimated Cas9 / dCas9 boundary) located at a distance from x i Point at 2.5nm.

[0304] Once the tracing is complete, the height of the DNA along the contour is extracted (relative to the median pixel height of the local area). The estimated boundaries of streptavidin and Cas9 / dCas9 are iteratively expanded or retracted around the original estimate until they are expanded to greater than (μ d +σ d) continuous region, where μ d and σ d are the mean and standard deviation of the height of the traced DNA outside the estimated position of the bound protein, and the estimates converge.

[0305] To account for any instrumental hysteresis that may distort the apparent length of the DNA, the length of the DNA was normalized and only 20% of the DNA molecules that were initially measured to be their expected length (given the known number of base pairs, 0.33 nm / base pair) were used for further analysis (for AAVS1 substrate number - trace number: 804; nominal length: 1198 bp, recorded average length: 1283 bp, standard deviation: 154 bp; for engineered substrates - trace number: 1520, nominal length: 986 bp, recorded average length: 1071 bp, standard deviation: 124 bp; for "nonsense" substrates - trace number: 616, nominal length: 1078 bp, recorded average length: 1217 bp, standard deviation: 135 bp). This step prevented us from improperly analyzing, for example, two DNA molecules that appeared collinear, DNA that may have been fragmented, or DNA that may have been cleaved and separated by Cas9 (which is rare, see the text).

[0306] Figure 1 C-1D, Figures 2C-2D and Figures 9A-9C Binding histograms were generated by mapping the relative position of each bound protein to the bases of protein overlap (nearest neighbor interpolation) and summing the total number of proteins bound to each site (if a single Cas9 / dCas9 can be interpreted as contacting multiple (k) sites, each contact region is weighted by 1 / k in the binding histogram). Peaks in the binding histograms were fitted to an empirical Gaussian exp(-((x-μ) / w)) using MATLAB. 2 ), where μ is the average peak position, and w is the peak width parameter (w = √2σ, where σ is the standard deviation).

[0307] Determination of dCas9 apparent dissociation constant. The apparent dissociation constant of dCas9 with different guide RNA variants was determined as previously described (Yang et al. (2005) Nucleic Acids Res. [Nucleic Acids Research], 33, 4322-4334). In short, at known solution concentrations of dCas9-guide RNA ([dCas9] 0) and DNA molecules ([DNA] 0), the corresponding number of "engineered" DNA molecules with or without bound protein was counted (protein-bound DNA fraction θ). dCas9After profiling the DNA with bound proteins (see above), the average number of proteins bound per DNA molecule (n dCas9 The overall dissociation constant was calculated as K d,DNA =[DNA][dCas9] / [DNA·dCas9]=(1-Θ dCas9 )([dCas9]0-n dCas9 [DNA]0) / (Θ dCas9 )

[0308] Alternatively, the fraction Θ of DNA with bound dCas9 within one peak width of the Gaussian fit of its corresponding binding histogram is used. dCas9,原型间隔子 The protospacer-specific dissociation constant K was calculated similarly d,原型间隔子 (ie, see Table 1), as in the case of using a fraction Θ for each site on the DNA with bound dCas9 dCas9,ss Calculation of site-specific dissociation constant K a,ss =K d,ss -1 Same.

[0309] Protein alignment and clustering. Images of Cas9 and dCas9 proteins were extracted that were isolated and appeared to contact DNA at only a single location. These features were selected as having a protein size greater than (μm) that was completely within a 134 nm × 134 nm bounding box. d +2σ d ) features, where μ d and σ d is the mean and standard deviation of the DNA height of the bound protein; this step essentially has the effect of removing from the set most of the aggregated / dense Cas9 / dCas9 as well as those proteins from images with large extrinsic noise. After four-fold nearest neighbor interpolation, proteins with a height greater than (μ d +σ d) to minimize the mean squared error between their topographic heights. A distance matrix was constructed from these minimized mean squared errors, and proteins with standard sgRNAs were clustered according to this criterion using the method of Rodriguez and Laio (27); proteins with guide RNA variants were clustered according to the closest Cas9 / dCas9 structure with the standard sgRNA. The overall average structure was extracted by performing a reference-free alignment between each member of each cluster according to the method of Penczek, Radermacher, and Frank (28). The characteristics of the Cas9 / dCas9 population at each feature on the DNA (e.g., protospacer site) were determined using proteins that bound within one peak width of a Gaussian distribution fitted to the binding histogram (i.e., see Table 1).

[0310] Kinetic Monte Carlo (KMC) of guide RNA strand invasion and R-loop "breathing". Kinetic Monte Carlo (KMC) experiments simulating strand invasion of guide RNAs at protospacer sites were performed using a Gillespie-type (continuous time, discrete state) ((Gillespie (1976) Journal of computational physics, 22, 403-434) implemented in MATLAB. Strand invasion was modeled as a one-dimensional random walk in a position-dependent potential determined by the relative nearest-neighbor-dependent DNA:DNA and RNA:DNA binding free energies. See, e.g., FIG4A . That is, the guide RNA base pairs with the protospacer up to m protospacer sites (1≥m≥20 for sgRNA1 and 1≥m≥18 for truncated sgRNA (tru-gRNA), and, to first order, the forward rate (additional guide RNA invasion rate) v f Use exp(-(ΔG°(m+1) RNA:DNA - ΔG°(m+1) DNA:DNA ) / 2RT), where R is the Boltzmann constant, T is the temperature (here 37°C to correspond to the parameter settings used), ΔG°(m+1) RNA:DNA is the base pairing free energy between RNA and the protospacer at site m+1, and ΔG°(m+1) DNA:DNA is the base pairing free energy between the protospacer and its complementary DNA strand (including a 1 / 2 correction term to satisfy careful balance). For sgRNA or tru-gRNA, v in state m = 20 or 18 f Set to 0. Reverse rate (rate of re-hybridization between the protospacer and its complementary DNA strand) v rProportional to exp(-(ΔG°(m) DNA:DNA -ΔG°(m) RNA:DNA ) / 2RT); if state m=1, the simulation stops (indicating guide RNA-protospacer dissociation). Starting from time t=0 (in arbitrary time units), for each iteration of the algorithm, the m-dependent rate is determined and two random numbers r1 and r2 are generated from a uniform distribution between 0 and 1. t is calculated with Δt=log(r1) / (v f +v r ) is incremented. If r2≥v f / (v f +v r ), state m increases to m+1, otherwise decreases to m-1. For measurements of "equilibrium" of R-loop breathing, m starts at m=20 (or 18 in the case of tru-gRNA) and the algorithm is iterated until t≥10,000. To measure "invasion" kinetics (e.g., in the presence of mismatched base pairs), m starts at m=10 (up to t=1000).

[0311] The free energy parameters are obtained from experimental literature at 37°C in 1 M NaCl. Sequence-dependent DNA:DNA hybridization free energy ΔG°(x) DNA:DNA Adapted from SantaLucia et al. (1996) Biochemistry, 35, 3555-3562; Sequence-dependent RNA:DNA hybridization free energy ΔG°(x) RNA:DNA Obtained from Sugimoto et al. (1995) Biochemistry, 34, 11211-11216; and in the case of introduced point mismatches rG·dG, rC·dC, rA·dA and rU·dT ΔG°(x) RNA:DNA Values ​​are obtained from Watkins et al. (2011) Nucleic Acids Res., 39, 1894-1902 (under slightly higher salt conditions). The sequence of the protospacer used was "ATCCTGTCCCTAGTGGCCCC' (SEQ ID NO: 336), as in the AAVS1 target site in the AFM experiment; the sequence of the protospacer complementary DNA was "GGGGCCACTAGGGACAGGAT" (SEQ ID NO: 337), and the sequence of the guide RNA was "GGGGCCACUAGGGACAGGAU" (SEQ ID NO: 338) for sgRNA or "GGCCACUAGGGACAGGAU" (SEQ ID NO: 339) for truncated RNA.

[0312] The correlation between the R-loop stability and experimental Cas9 cleavage rate derived from KMC.In order to analyze the correlation between the interaction of guide RNA-protospacer and Cas9 cleavage rate in vivo, Extracted the sequence of guide RNA and target DNA from Hsu et al. (2013) Nature Biotechnology, 31,827-832 and its experimentally determined maximum likelihood estimate (MLE) Cas9 cleavage frequency.By from Hsu et al. (2013) Nature Biotechnology, 31,827-832 with single nucleotide PAM far-end (away from PAM site ≥ 10bp) mismatch (rG dG, rC dC, rA dA and rU dT type) guide RNA and target DNA sequence and the experimentally determined maximum likelihood estimate (MLE) Cas9 cleavage frequency at these sites input (n=136) into KMC script. For each sequence, the strand invasion simulation starting at m = 10 was repeated 1000 times (maximum t = 100) to obtain an average time fraction with m ≥ 16 and correlated with the empirical cleavage rate. Significance was determined by bootstrapping the average occupancy fraction with the MLE cleavage frequency via 100,000 permutations and then recalculating the correlation coefficient and p-value. The guide RNA-protospacer binding free energy was estimated by summing the nearest neighbor energies using the parameter set listed above and using −3.1 kcal mol -1 Starting factor correction.

[0313] dCas9-tru-gRNA and dCas9-hp-gRNA data for comparison with dCas9-sgRNA structural properties. When comparing height and volume measurements of proteins between experiments, AFM imaging conditions should be kept roughly consistent to avoid introducing artifacts. For example, this is generally not a problem when comparing the height and volume of dCas9 bound to different sites on an engineered DNA molecule, but it poses a challenge when comparing the structural properties of dCas9 / Cas9 using different guide RNAs or DNA substrates. As a control, the height and volume of streptavidin used to label the ends of the traced DNA molecules were used for different experiments and should remain constant under all experimental conditions. For experiments with sgRNAs, the average height of the streptavidin differed by less than 0.1 nm (mean difference: 0.087 nm; standard deviation of the difference: 0.052 nm) and their average volume (1098 nm) was 0.1 nm. 3 ) is less than 15nm 3 (Average difference: 14.461nm 3 ; Standard deviation of the difference: 10.419nm 3However, the average height and volume between experiments performed with tru-gRNA and hp-gRNA differed from those performed with sgRNA, with differences of up to 0.14 nm and 225 nm, respectively. 3 To directly compare the results of these experiments, the height of dCas9 with tru-gRNA and hp-gRNA on engineered DNA was scaled according to the average height difference relative to the height of dCas9 with sgRNA and the volume was scaled according to the percentage difference in average volume.

[0314] Example 2

[0315] Atomic force microscopy captures specific and nonspecific binding of Cas9 / dCas9 along engineered DNA substrates at high resolution

[0316] Analysis of crystallographic and biochemical experiments suggests that specificity for protospacer binding and cleavage is conferred first by PAM site recognition by Cas9 itself, followed by strand invasion of the bound RNA complex and direct Watson-Crick base pairing with the protospacer ( Figure 1 A), but a complete mechanistic picture has yet to emerge. To directly probe the relative propensity for binding to protospacers and off-target sites at single-molecule resolution, 50 nM Cas9-sgRNA or dCas9-sgRNA complexes targeting the AAVS1 locus on human chromosome 19 were imaged in air by AFM after incubation with one of three DNA substrates (2.5 nM):

[0317] (i) An 1198 bp segment of the AAVS1 locus containing the complete target site following the PAM (hre "TGG") Figure 1 C);

[0318] (ii) A 989 bp engineered DNA substrate containing a series of six complete, partial, or mismatched target sites, each separated by approximately 150 bp ( Figure 1 D). Mismatches at these sites may span the "seed" (PAM proximal, approximately 12 bp) and "non-seed" (PAM distal) regions of the protospacer. The only PAM sites in this engineered substrate are at these explicitly designed positions; and

[0319] (iii) 1078 bp of “nonsense” DNA with no homology to the target sequence (beyond 3 bp sequence) Figures 9A-9C ).

[0320] Figure 1C shows that dCas9 and Cas9 exhibit almost identical binding distributions on the AAVS1 substrate (n = 404 and n = 250, respectively). Figure 1 D shows that on engineered substrates (n=536), dCas9 binds with the highest affinity to intact protospacers without mismatch (MM) sites (peak 1, hereafter referred to as complete or "0MM" sites), and also binds to sites with 5 or 10 mismatched bases located distal to the PAM site (the third and fourth features from the streptavidin label, hereafter referred to as "5MM" or "10MM" sites, respectively), although with reduced affinity. Sites containing a large number of mismatches (the second and fifth features) or having two PAM-proximal mismatched nucleotides (the sixth feature) bind at significantly lower rates. (Below) Distribution of PAM ("TGG") sites in each substrate.

[0321] Structurally, Streptococcus pyogenes Cas9 is a 160 kDa monomeric protein of approximately 10 nm × 10 nm × 5 nm (from the crystal structure), roughly divided into two leaf-shaped halves, each containing a nuclease domain. Consistent with the x-ray structure, dCas9-sgRNA images by AFM appear as a large egg-shaped structure ( Figures 10A-10C ), after incubating Cas9 or dCas9 with DNA, these structures bound to DNA were observed and designated as Cas9 or dCas9 ( Figure 1 B. Figures 10A-10C To clearly determine the sequences of sites bound by Cas9 and dCas9, biotinylated DNA molecules were labeled at one end with a monovalent streptavidin tag before AFM imaging. DNA molecules observed to have bound Cas9 or dCas9 protein were selected for further analysis and traced with subpixel resolution according to a modified protocol adapted from Wiggins et al. (25), and sites bound by Cas9 / dCas9 were extracted (see Supplementary Methods for details).

[0322] The method proved to be very robust (Table 1): on DNA bound by either Cas9 or dCas9, a concentration of 100 bp was observed exactly on sites with adjacent PAMs (within the expected 23 bp, Figure 1 Significant protein enrichment at the protospacer site of CD was observed and presented as a sharp peak. No such obvious peak was observed in the DNA substrate without the target site ( Figures 9A-9C). The standard deviation of the peak widths was in the range of 36-60 bp, which is a significant improvement compared to binding experiments using single-molecule fluorescence, which resulted in a standard deviation σ of peak widths of approximately 1000 bp. The average apparent Cas9 / dCas9 "footprint" on DNA covered 78.1 bp ± 37.9 bp; the enlargement of the apparent footprint relative to the approximately 20 bp Cas9 footprint on DNA determined by biochemical and crystallographic methods is a well-known result of imaging convolution with AFM tip width. Previously, it was observed in vitro that Cas9 remained bound to the target DNA for long periods of time (>10 min) after putative DNA cleavage as a single-turnover endonuclease and was unable to displace from the cleaved strand without harsh chemical treatment. It was observed that the majority of DNA molecules bound to Cas9 behaved as full-length AAVS1-derived substrates, with only a small (approximately 5%) percentage of the substrate having been cleaved and dissociated. After tracing these DNA molecules, Cas9 was observed to bind these "full-length" substrates with almost the same distribution as dCas9 (two-sided Koch-Stokes test, significance level 5%) ( Figure 1 C).

[0323] Table 1 Targeting Cas9 / dCas9-sgRNA Figure 1 Peaks recorded in the CD binding histogram and in the binding histogram of Figure 2C for dCas9 with a sgRNA with a 2 nt truncation at the 5' end (tru-gRNA) are based on the Gaussian ∝exp(-((x-μ) / w) 2 )

[0324]

[0325]

[0326] a Standard single guide RNA (sgRNA)

[0327] b Single guide RNA shortened by 2 nt from the 5' end (tru-gRNA)

[0328] c Subsequent tracings revealed the number of DNA molecules that were monovalently streptavidin-labeled and bound to the protein (see Supporting Methods for details).

[0329] d Target site with 10 PAM-distal mismatches

[0330] e Target sites with 5 PAM-distal mismatches

[0331] fOn the engineered DNA substrate, the tru-gRNA is expected to interact only with the first 8 of the 10 PAM-distal mismatched nucleotides at the 10MM site.

[0332] g On the engineered DNA substrate, the tru-gRNA is expected to interact only with the first three of the five PAM-distal mismatched nucleotides at the 5MM site.

[0333] h bp from the streptavidin-tagged end (from PAM to the end of the site)

[0334] i Combined with the peak maximum in the histogram (from the Gaussian fit)

[0335] j The peak width is √2σ, where σ is the standard deviation

[0336] k The number of dCas9 molecules observed within one peak width (√2σ) of the binding site. If Cas9 / dCas9 appears to contact DNA at n sites, the molecule is weighted 1 / n. If the molecule overlaps 10MM and 5MM sites, the # is further weighted 1 / 2.

[0337] By examining the occupancy of dCas9 bound to different positions along the engineered substrate, the relative binding propensity of dCas9 to various mismatched and partial target sites can be determined (Figure ID, Table 1). The overall dissociation constant between dCas9 and the entire DNA substrate was estimated to be 2.70nM (±1.58nM, 95% confidence, Table 2). In particular, the dCas9 dissociation constant at the site of the prototype spacer (within one peak width of the binding histogram) that was completely (perfectly matched) on the substrate was 44.67nM (±1.04nM, 95% confidence). Early electrophoretic mobility shift assays (EMSA) estimated that the binding of dCas9-sgRNA to the prototype spacer site on short DNA molecules (approximately 50bp) was between 0.5nM and 2nM. While the observed increase in the dissociation constant at the protospacer site may be related to the presence of multiple off-target sites on the engineered DNA substrate, the dissociation constants determined by AFM are typically nearly an order of magnitude higher than those determined by conventional assays (26). This discrepancy is generally attributed to nonspecific interactions between the protein and the blunt ends of the short DNA strands, which are not accounted for in EMSA.

[0338] Table 2 dCas9 with different guide RNA variants and 989 bp "engineered" DNA substrates containing a series of fully and partially complementary protospacer sites (e.g., Figure 1D, 2C, and 2D)

[0339]

[0340]

[0341] a Full-length single guide RNA (sgRNA)

[0342] b Truncated sgRNA (truncated the first two nt at the 5' end)

[0343] c sgRNA with an additional 5'-hairpin that overlaps six PAM-distal targeting nts (see text)

[0344] d sgRNA with an additional 5'-hairpin that overlaps ten PAM-distal targeting nts (see text)

[0345] On engineered substrates, dCas9 is relatively tolerant of distal mismatches (exhibiting a 50%-60% tendency relative to the intact target site, Figure 1 D and Table 1) and have the same apparent affinity for target sites containing 5 and 10 distal mismatches (MM) (within the apparent confidence). However, the tendency to bind to protospacer sites containing only two PAM adjacent mismatches is similar to the tendency to bind to sites with 15 or even 20 (PAM site only) distal mismatches (approximately 5%-10% binding tendency relative to the complete target, which is close to the background binding signal), a finding consistent with previous biochemical studies. Although there are no PAM sites other than those adjacent to the protospacer sites on the engineered substrates, there is a clear "shoulder" of enhanced Cas9 and dCas9 binding near the AAVS1 target on the AAVS1-derived substrate, which is particularly enriched in the PAM sites. The slight enrichment of dCas9 on nonsense and AAVS1-derived substrates at regions distal to the target site closely mirrored the distribution of PAM sites (two-sided Koch-Smith test, 5% significance level), and the dCas9 distribution on nonsense substrates more closely mirrored the experimental PAM distribution compared to dCas9 matching 71.20% of 100,000 randomly generated sequences with the same dA, dT, dC, and dG distribution ( Figures 9A-9C). Since dCas9 binding along the "nonsense" substrate (with 879 PAM sites in 1079bp) corresponds very well to the PAM site distribution, this is understood to be a measure of the true dCas9-PAM interaction. The average unit-site dissociation constant of dCas9 binding along the "nonspecific" substrate is estimated to be approximately 867nM (standard deviation ± 209nM). This can be understood as an estimate of the dCas9 binding dissociation constant on DNA that does not have protospacer homology.

[0346] Example 3

[0347] sgRNA with two nucleotide truncations at the 5' end (tru-gRNA) does not increase dCas9 binding specificity in vitro

[0348] It was found that Cas9 still exhibited cleavage activity even when the guide (protospacer targeting) segment of the sgRNA or crRNA was truncated by up to four nucleotides from its 5' end, and Fu et al. (21) recently showed that the use of sgRNAs with these 5'-truncations (optimally truncated by 2-3 nucleotides) can actually lead to an order of magnitude increase in the fidelity of Cas9 cleavage in vivo. It has been proposed that the increased sensitivity to mismatch sites (MMs) using these truncated sgRNAs (referred to as "tru-gRNAs," Figure 2A) is a result of reduced binding energy between the guide RNA and the in situ spacer site. This suggests that the binding energy conferred by the additional 5'-nucleotides on the sgRNA can compensate for any mismatched nucleotides and stabilize Cas9 at the incorrect site, whereas tru-gRNAs would be relatively less stable on DNA if a mismatch is present.

[0349] As a test of this proposed mechanism, dCas9 was imaged in the presence of a tru-gRNA that was truncated two nucleotides 5' relative to the previously used sgRNA. The dCas9-tru-gRNA complex was incubated with an engineered substrate containing a series of complete and partial protospacer sites. Again, a peak was found precisely at the complete protospacer site (Figure 2C and Table 1), although the apparent binding constant at this site was significantly reduced relative to dCas9 with a complete sgRNA (i.e., the dissociation constant increased, see Table 2). However, relative to binding at the complete protospacer site, off-target binding of dCas9 with a tru-gRNA having a PAM distal mismatch at the protospacer site was actually increased compared to dCas9 with an sgRNA (Figure 2C and Table 1). Similar to dCas9 with sgRNA, dCas9 with tru-gRNA bound with approximately equal propensity to protospacers with 10 or 5 PAM-distal mismatch sites (note: tru-gRNA is only expected to interact with the first 8 and first 3 mismatches at those sites, respectively). These results suggest that the increased cleavage fidelity using tru-gRNA is not necessarily conferred by a relative decrease in binding propensity or relative stability at off-target sites in the presence of mismatches. Rather, there is some "threshold" effect where the binding constant decreases to approximately 4-5 × 10 6 M effectively abolished cleavage activity in vivo, but these and additional results presented below suggest that the increased specificity exhibited by tru-gRNAs may be influenced by differences in the cleavage mechanism itself. Furthermore, these findings suggest that while tru-gRNAs can improve the cleavage specificity of active Cas9, they may not improve the specificity of binding activity for in vivo applications involving dCas9 (or chimeric derivatives).

[0350] In addition, previous reports have shown that tru-gRNA with 5'-truncation (optimal truncation of 2-3 nucleotides) in its protospacer targeting fragment can lead to an order of magnitude increase in the fidelity of Cas9 cleavage in vivo (Fig. 2A), and the results shown in the embodiment show that the truncated gRNA does not improve the specificity of dCas9 binding (Fig. 2C). Figure 2C shows a comparison of the binding affinity of dCas9 with tru-gRNA (trugRNA, purple line) on DNA molecules containing a complete protospacer (site i) and a protospacer site with 5 and 10 PAM distal mismatches (sites ii and iii, respectively). Figure 2C shows that standard guide RNA retains a significant ability to bind to these off-target sites (containing mismatches), and trugRNA does not show relative enhancement in terms of binding specificity at the mismatch site at 510 nucleotides distal to the protospacer PAM. The binding distribution of dCas9 with tru-gRNA showed a clear peak of affinity at the protospacer sites with 10 PAM distal mismatches and 5 PAM distal mismatches, which demonstrated that tru-gRNA did not increase binding specificity relative to the full sgRNA (see Table 1). The "peaks" in the binding histogram indicate specific, stable binding at these off-target sites. In fact, compared with the standard guide RNA, the binding of dCas9-trugRNA at off-target sites is actually increased relative to binding to the protospacer. This promiscuous binding may limit their utility for dCas9 and chimeric dCas9 derivatives. It may also reflect the off-target cleavage reported by this system, which, although improved relative to the standard guide RNA, is still significant at some off-target sites. For comparison, we did not find specific binding of hpgRNA at these sites with mismatches (Figure 2D). hpgRNA binds to these sites with approximately the same affinity as they nonspecifically bind to DNA that does not have homology to the protospacer, with the observed off-target binding affinity being reduced by approximately 22% relative to the truncated gRNA. Additionally, based on the narrow geometry of the Cas9 DNA binding channel, we anticipate that the presence of an unopened hairpin within the mismatched protospacer may inhibit the conformational changes in Cas9 necessary for cleavage ( Figure 1 B).

[0351] Extensive work has been done to characterize this off-target activity—and to improve the specificity of Cas9 / dCas9 by: intelligent choice of protospacer target sequences; optimization of sgRNA structure by, for example, truncation of the first two 5′-nucleotides in the sgRNA; and use of a “double-nicking” Cas9 enzyme—but a clear understanding of the exact mechanism of RNA-guided cleavage (as it relates to the structural biology of Cas9) is necessary to develop Cas9 derivatives and guide RNAs with increased fidelity for its novel medical and biological applications.

[0352] In line with this goal, here we used atomic force microscopy (AFM) to resolve individual Streptococcus pyogenes Cas9 and dCas9 proteins as they bound targets along engineered DNA substrates after incubation with different sgRNA variants. This technique allowed us to directly resolve the binding sites and structures of individual Cas9 / dCas9 proteins simultaneously, providing a wealth of mechanistic information about Cas9 / dCas9 specificity at single-molecule resolution. Consistent with traditional biochemical studies, we found that significant binding of Cas9 / dCas9 with sgRNAs occurred at sites containing up to 10 mismatched base pairs in the target sequence. However, while the use of guide RNAs truncated by two nucleotides from their 5' end (tru-gRNAs) has previously been shown to result in up to 5000-fold reduction in off-target mutagenesis by Cas9 in vivo, we found that the in vitro binding specificity of dCas9 with tru-gRNAs to mismatched targets was similar to that of standard sgRNAs. Adding a hairpin to the 5' end of the sgRNA that partially overlaps the target binding region of the guide RNA increased dCas9 specificity, but at the expense of an overall reduced propensity to bind to DNA. Our results suggest that the overall stability of the guide RNA-DNA binding does not necessarily control Cas9 cleavage specificity when the mismatch is located 10 bp away from the PAM.

[0353] Example 4

[0354] Absolute binding propensity and distribution profile of dCas9 binding to DNA with mismatched protospacers modulated in vitro by guide RNAs with 5′-hairpins complementary to the “PAM-distal” targeting segment (hp-gRNA)

[0355] dCas9 specificity can be increased by extending the 5' end of sgRNA so that it forms a hairpin structure that overlaps the "PAM distal" targeting (or "non-seed") segment of gRNA (Fig. 2B). After the PAM site is bound and the chain invasion of the guide RNA to DNA has been initiated, the hairpin opens after binding to the complete prototype spacer and complete chain invasion can occur. If there is a PAM distal mismatch at the target site, it is energetically more conducive to the hairpin remaining closed and chain invasion is hindered. Similar topological structures have recently been used in "dynamic DNA loops" driven by chain invasion. In these systems, Hairpins act as kinetic barriers to invasion, and the oligonucleotide invasion rate slows down by several orders of magnitude when attempting to invade targets with mismatches. The hairpins here can shift during the invasion of the complete target site, but if there is a mismatch between the target and the non-seed targeting region of the guide RNA, invasion is suppressed (Fig. 2B). In those cases, it is energetically more conducive to the hairpin remaining closed. Although previous work has added 5'-extensions to sgRNAs to complement additional nucleotides beyond the protospacer, these guide RNAs did not show increased Cas9 cleavage specificity in vivo. Instead, they were digested back to approximately their standard length in living cells. Based on the size and structure of the hairpin, it can fit within the DNA binding channel of the Cas9 / dCas9 molecule and protect it from degradation.

[0356] sgRNAs (hp-gRNAs) were generated with a 5′-hairpin overlapping nucleotides complementary to the last six (hp6-gRNA) or ten (hp10-gRNA) PAM-distal sites of the protospacer. By mapping the observed binding positions of dCas9-hp-gRNAs on the engineered DNA substrates ( FIG. 2D ), a sharp peak was observed precisely at the protospacer site (PAM and protospacer located at positions 144–167, with binding peaks at position 154.0 for dCas9-hp6-gRNA (95% confidence: 153.3–154.8) and at 158.3 for dCas9-hp10-gRNA (95% confidence: 157.6–158.9). The specificity peaks at sites with 5 and 10 distal mismatches were significantly flattened, with dCas9 and hp10-gRNA exhibiting significantly reduced off-target site affinity (down 22% relative to dCas9 with tru-gRNA). The affinity peak at the full protospacer site suggests that the hairpin is indeed opened after full invasion. For hp6-gRNA n=243 and for hp10-gRNA n= 212. dCas9 with hp-gRNA showed a similar decrease in target site affinity as dCas9 with tru-gRNA; however, in contrast to dCas9 with tru-gRNA, dCas9 with hp-RNA did not exhibit any sharp binding peaks at off-target sites, which would otherwise indicate strong specific binding. For hp6-gRNA, there was binding enrichment around protospacer sites with 5 or 10 mismatches at the PAM distal site. Because they lacked the sharp binding peaks observed with sgRNA and tru-gRNA, these enrichments are unlikely to indicate specific binding, but rather may indicate that dCas9 dissociates from these sites after adsorption to the surface. This would indicate that binding at these off-target sites is very weak in the case of hp6-gRNA.

[0357] In the case of hp10-gRNA, binding to these mismatch sites was roughly below the level of nonspecific binding elsewhere on the substrate, representing a maximum reduction of 22% in off-target binding affinity relative to that observed for tru-gRNA (the observed association constant decreased from 3.18 × 10 6 M to 2.48×10 6 This increase in hp10-gRNA specificity was also reflected by a similar binding dissociation constant to the protospacer site as hp6-gRNA, but a significant increase in the overall dissociation constant for the entire (specific + nonspecific) engineered substrate (Table 2).

[0358] The pronounced enrichment at the exact protospacer site suggests that the hairpin in the hp-gRNA is actually opened after invasion of the perfect protospacer site, as nucleotides bound to the PAM-distal site of the protospacer are otherwise trapped within the hairpin. A possible mechanism for improved binding specificity is that the presence of the hairpin promotes unwinding of the guide RNA from these off-target sites when protospacer sites with PAM-distal mismatches are not opened. The results suggest that hp-gRNAs can be used to modulate Cas9 / dCas9 binding affinity and specificity, and further manipulation of hairpin length, loop length, and loop composition may allow for finer control of these properties.

[0359] Example 5

[0360] Cas9 and dCas9 undergo progressive structural transitions as they bind to DNA sites that increasingly match the target protospacer sequence.

[0361] Using negative stain transmission electron microscopy (TEM), it was observed that upon binding to sgRNA, the dCas9 structure is compacted and rotated to open a putative DNA binding channel between its two lobes. After binding to DNA containing a PAM and protospacer sequence, dCas9 undergoes a second structural reorientation to an unfolded conformation. It is proposed that the role of this second transition is related to sgRNA strand invasion, or the alignment of the two main Cas9 nuclease sites to the two separate DNA strands. However, these studies were performed only in the presence or absence of DNA containing a fully matched protospacer sequence, and examining the transitions between these conformations at partially matched protospacer sites can provide insights into off-target binding and cleavage mechanisms. Therefore, in addition to determining relative binding propensities, AFM imaging was used to capture these putative conformational transitions performed by Cas9 and dCas9 as they bind DNA at various sites complementary to the protospacer. We extracted the volume and maximum topographic height of Cas9 and dCas9 proteins with sgRNA that appeared to be isolated on DNA (n=839) and mapped these values ​​to their corresponding binding sites on DNA (Figure 3, Figures 11A-11D and Figures 12A-12B The binding site distribution is almost identical to that of the full dataset, indicating that this selection is unbiased and representative. Recorded images of each of these proteins were extracted (Figures 11C-11D) and aligned pairwise by iterative rotation, reflection, and translation. The protein structures were clustered according to their pairwise mean variance topographic differences ( Figures 12A-12Band Table 3). A significant advantage of this technique is that it naturally clusters any monovalent streptavidin or any aggregated Cas9 / dCas9 proteins that colocalize with DNA on the surface independently of those assigned as single Cas9 / dCas9 molecules, thereby allowing unbiased analysis of the structural properties of these proteins on DNA. Analysis of the distribution of binding sites by putative streptavidin molecules or aggregated proteins revealed that they are all rare and evenly distributed along the DNA and therefore do not interfere with analysis of the distribution of binding sites ( Figures 12A-12B ).

[0362] At sites with no homology to the target, such as on “nonsense” DNA substrates, dCas9 molecules with sgRNA were mostly small and oval (Fig. 3C(iii) and Table 3). However, as the dCas9 protein bound to increasingly complementary target sequences (Fig. 3(α-δ)), their height and volume increased significantly relative to nonspecific binding (Fig. 3D and Figures 12A-12B , Table 2), reaching a maximum size at the protospacer sequence. This increase was also accompanied by the dCas9 population (Figures 3A and Figures 12A-12B , Table 2) shifted from structures clustered in relatively flat and ovoid conformations (Figures 3C(ii) and C3(iii), blue and green) to those increasingly clustered into slightly rounder structures with a large central bulge (Figure 3C(i), yellow). This latter observed conformation is likely the unfolded conformation previously observed via TEM and more recently by size exclusion chromatography, and is likely the active state, in which the nuclease domain of Cas9 is correctly positioned around the DNA so that cleavage can occur most efficiently.

[0363] The catalytically active Cas9 undergoes a significant size increase as it also binds to the protospacer sequence (Figure 3(ε)); however, there is a small, but statistically significant, decrease in size relative to dCas9, and the conformation of Cas9 at the complete protospacer site tends to cluster in a flatter (green) structure. Because we did not simultaneously monitor whether the DNA was cleaved during imaging, it is unclear whether this represents another conformational change after DNA cleavage or is the result of mutational differences between Cas9 and dCas9; however, because binding and strand invasion have been previously identified as rate-limiting steps, it is likely that the DNA within Cas9 was cleaved during these measurements.

[0364] Table 3 Properties of dCas9 / Cas9 with different guide RNA variants at full- and partial- and non-complementary protospacer sites

[0365]

[0366]

[0367]

[0368] a Total molecules observed within two standard deviations of these sites. Below: Population fraction (±95% binomial confidence) in the main three structural clusters as colored in Figure 2 in the main text (Y = yellow cluster, G = green cluster, B = light blue cluster). Figures 12A-12B According to the complete distribution characteristics of clustering.

[0369] b Standard mean error

[0370] c Standard mean error

[0371] d The null hypothesis that the height-volume distributions were different was rejected (p > 0.05; Hotelling's T 2 test)

[0372] e On the engineered DNA substrate, tru-gRNA is expected to interact only with the first 8 of the 10 PAM-distal mismatched nucleotides at the 10 MM site (labeled “8 MM” in FIG3D ).

[0373] f On the engineered DNA substrate, tru-gRNA is expected to interact only with the first three of the five PAM-distal mismatched nucleotides at the 5MM site (labeled "3MM" in Figure 3D).

[0374] g See Supplementary Note 1 in the Supporting Information regarding correction of height and volume of proteins with tru-gRNA and hp-gRNA so that they can be compared with proteins with sgRNA.

[0375] Example 6

[0376] Interactions between guide RNA and target DNA at or near the 16th protospacer site stabilize Cas9 / dCas9 conformational changes

[0377] AFM imaging directly revealed that while dCas9 / Cas9 retained a significant preference for binding to protospacer sites with up to ten distal mismatches, binding to DNA sites that were increasingly complementary to the protospacer drove an increased shift in the dCas9 / Cas9 protein population toward the conformation that represents the active site. Notably, similar structural shifts between off-target and perfectly matched sites were also observed for dCas9 with hp-gRNA (Tables 2 and 3). Figure 13 ). The presence of complementary PAM-distal sequences is known to correlate with increased stability of Cas9 on DNA. It has also been recently discovered that Cas9 binds to single-stranded DNA and that increasing PAM-distal complementarity to the protospacer (from 10 to 20 sites) leads to an increased change in protein size. This is then also associated with a switch in Cas9 activity from nicking behavior to full cleavage. Here, we can directly determine the volume of Cas9 / dCas9 bound to double-stranded DNA sites. Analysis of the structural properties of individual Cas9 / dCas9 proteins on double-stranded DNA reveals stable conformational transitions with increasingly matched target sequences, consistent with a “conformational gating” mechanism, where sgRNAs base-paired to these distal sites also stabilize the active conformation, allowing efficient cleavage to occur, while binding to sites with multiple distal mismatches shifts the equilibrium away from the active structure (i.e., see ). Figure 4D ).

[0378] Accordingly, we found that this effect was substantially attenuated for dCas9 with tru-gRNA (Figure 3D and Table 3), with less movement between structural populations within the protein clusters (Figure 13). In addition, although we found statistical differences between the height-volume properties of nonspecifically bound dCas9-tru-gRNAs and those bound at full or partial protospacer sites (p < 0.05; Hotelling T 2 3D and Table 3). It has been recently hypothesized that although invasion of the first 10 bp of the protospacer triggers a conformational change in Cas9, complete invasion of the protospacer by the guide RNA helps drive further movement to a fully active state. Therefore, we hypothesize that the suppression of conformational changes at increasingly matched protospacer sites observed for dCas9 with tru-gRNAs (relative to those with sgRNAs) is a result of reduced stability of these guide RNAs at PAM-distal sites.

[0379] To investigate the relative stabilities of sgRNAs and tru-gRNAs at these sites, we performed kinetic Monte Carlo (KMC) studies of the dynamic structure of R-loops during and after strand invasion—structures formed by an invading guide RNA bound to a segment of continuous DNA, exposing a single-stranded loop of its complementary DNA (Fig. 4A). See Supplementary Methods for further details. Briefly, using a Gillespie-type algorithm, we modeled strand invasion of up to m protospacer sites by bound guide RNAs as a nucleotide-by-nucleotide sequential competition between invasion (disruption of base pairing between the protospacer and its complementary DNA strand followed by replacement with a protospacer-guide RNA base pair) and reannealing (the reverse), with invasion sequence-dependent rates v and v, respectively. f and the reannealing sequence-dependent rate v r (Figure 4A). For the first order, we define the transition rate v from state m to m+1 as f Approximately the same as exp(-(ΔG°(m+1) RNA:DNA -ΔG°(m+1) DNA:DNA ) / 2RT), where ΔG°(m+1) RNA:DNA is the free energy of base pairing between RNA and the protospacer at position m+1, and ΔG°(m+1) DNA:DNA is the free energy of base pairing between the protospacer and its complementary DNA strand at position m+1 (R is the ideal gas constant, T is the temperature, and the 1 / 2 term is added to satisfy careful balance). r Similarly, with exp(-(ΔG°(m) DNA:DNA -ΔG°(m) RNA:DNA ) / 2RT). This type of transition rate has been previously used in computational studies of nucleotide base pairing and stability, and here they allow us to capture the general dynamics of R-loops in a sequence-dependent manner.

[0380] In general, RNA:DNA base pairs are energetically stronger than DNA:DNA base pairs, and as expected, at equilibrium, we see from the KMC trajectory that the guide RNA stably binds to the protospacer (Figure 4C). However, while the sgRNA is quite stable and maintains almost complete invasion—the strand still invades to the 19th protospacer site during 95% of the simulation time course (Figure 4B)—the tru-gRNA exhibits significant fluctuations in protospacer re-annealing at PAM-distal sites (Figures 4B and 4C). Because the only difference between dCas9-sgRNA and dCas9-tru-gRNA is the truncation of only two 5'-nucleotides from the guide RNA, and because we found that the conformational change of dCas9-sgRNA is suppressed at sites containing 5 PAM-distal mismatches, these results suggest that the conformational change to the fully active state is stabilized by the interaction between the guide RNA and the protospacer near the 16th site of the protospacer, and this stabilization is disrupted by the instability of the tru-gRNA in this region. Indeed, KMC experiments showed that when sgRNA was replaced by tru-gRNA, the average lifetime between complete DNA invasion and reannealing to position 16 was reduced by two orders of magnitude (Figure 4C inset). This result is consistent with previous findings: in the case of Cas9 activity modulated by tru-gRNA variants with 2 or 3 nucleotide (nt) truncations depending on the sequence context, cleavage in all tested cases was significantly reduced by about 90%-100% by 4nt truncation and abolished after 5nt truncation. The conformational change to the protein's activated state is stabilized by these interactions at or near position 16 of the protospacer. This finding is supported by the stability of gRNAs at the 14th-17th protospacer positions, which was estimated by additional KMC experiments described below and correlated with experimental off-target cleavage in vivo (see below), while the stability of guide RNAs at protospacer positions 18-20 was not the case.

[0381] Example 7

[0382] Fluctuations in the guide RNA-protospacer R-loop suggest a mechanism for mismatch tolerance of Cas9 / dCas9 and increased cleavage specificity by tru-gRNA

[0383] To investigate the mechanism by which Cas9 or dCas9 can tolerate mismatches in the protospacer or become sensitive to these mismatches, we performed a series of KMC experiments using the AAVS1 protospacer site, introducing one or two PAM distal (≥10 bp from PAM) mismatches in the AAVS1 protospacer site (Figure 5). Cas9 is generally more tolerant to PAM distal mismatches than to PAM proximal mismatches. However, Hsu et al. (2013) Nature Biotechnology, 31, 827-832 identified significant and variable differences in the Cas9 cleavage rate estimated at the protospacer containing PAM distal mismatches based on sequence context, mismatch type, and mismatch site. Based on our AFM and early KMC experiments, we hypothesized that the difference in cleavage rate may similarly be a result of the different stability of the guide RNA near the 16th site of the protospacer. For these simulations, we only examined sequences with protospacer-guide RNA pairs with rG·dG, rC·dC, rA·dA, and rU·dT mismatches that would result in separation, for which the sequence context-dependent thermodynamic data are the most complete and applicable to our KMC model. The effects of these mismatched base pairs are not expected to significantly reduce the overall binding energy between the sgRNA and the protospacer (Table 4); for example, single rG·dG, rC·dC, rA·dA, and rU·dT mismatches reduce the RNA:DNA melting temperature by an average of 1.7°C. Instead, their effects are expected to be kinetic rather than thermodynamic by hindering strand displacement at the mismatch. Therefore, we started the kinetic Monte Carlo experiment from the 10th protospacer site (initial R loop length m = 10) (such as would occur during chain invasion).

[0384] Table 4 Sequences from Hsu et al. (2013) Nature Biotechnology, 31, 827-832 used for correlation analysis and maximum likelihood estimation (MLE) cleavage frequencies (mismatch sites in the target sequence are bolded).

[0385]

[0386]

[0387]

[0388]

[0389]

[0390]

[0391] KMC experiments were then performed to investigate the kinetics of strand invasion in the presence of PAM-distal mismatches. In all cases (1000 trials per condition), the guide RNA remained fairly stably bound even in the presence of mismatches (i.e., no complete melting was observed) and was often able to rapidly bypass these sites to complete invasion ( Figure 5C and Figures 14A-14C ), but the average first-pass time of overall chain invasion varies significantly, depending on the location of the mismatch site ( Figures 14A-14C ). R-loops are quite stable during invasion (Figure 5A), as sgRNAs are often able to remain fully invaded even in the presence of multiple mismatches. Qualitatively, the results are similar to those of previous in vitro studies of dCas9 / Cas9 binding and cleavage of mismatched targets. However, in the case of tru-gRNAs (Figure 5B), R-loops are often trapped behind the mismatch site. The average first passage time across the mismatch is similar for sgRNAs and tru-gRNAs ( Figures 14A-14C ), but examination of the time course of KMC revealed that tru-gRNAs are often “recaptured” quickly after mismatches due to the inherent volatility of their R-loops ( Figure 5C For sgRNAs, this recapture frequency is much lower. Therefore, combined with AFM imaging, the results of the KMC experiment suggest that the origin of the increased specificity of tru-gRNA lies not in differences in the binding process, but in the volatility of its R loop ( Figure 4D ), causing it to be repeatedly trapped following a mismatch even after initially bypassing it, making it less likely that Cas9 will assume an active conformation. For the sgRNA, once it bypasses a mismatch, it can remain fully invasive with relatively minor perturbations, suggesting a mechanism for mismatch tolerance.

[0392] Example 8

[0393] The stability of the guide RNA interaction with positions 14–17 of the protospacer correlates with experimental off-target Cas9 cleavage rates, whereas the overall guide RNA–protospacer binding energy does not.

[0394] To verify whether the stability of the R loop at or near position 16 of the protospacer in vivo (which was suggested by AFM studies to be linked to conformational changes in Cas9) is associated with Cas9 activity, we performed a kinetic Monte Carlo (KMC) analysis of the R loop stability on the sequences used by Hsu et al. (2013) Nature Biotechnology, 31, 827-832. The data set of Hsu et al. (2013) Nature Biotechnology, 31, 827-832 consists of measurements of the cleavage frequency at 15 different protospacer targets containing various point mutations relative to the guide RNAs performed to study the cleavage specificity of Cas9. This data set contains 136 protospacer-guide RNA pairs with single isolated mismatches of the type rG·dG, rC·dC, rA·dA, and rU·dT in the PAM distal region (Table 4). We used the KMC method to study the protospacer-guide RNA pairs starting at an R loop size of m = 10 to simulate invasion. Inclusion of a single mismatch site from this set reduced the magnitude of their overall guide RNA–protospacer binding free energy by an average of approximately 6% relative to the perfectly matched target, but as mentioned above, a wide distribution of Cas9 cleavage frequencies was observed for these guide RNA–protospacer pairs whose origin was not obvious.

[0395] The average fraction of time that the RNA was stably bound to each site of the protospacer was determined for each guide RNA over 1000 trials and then correlated with the maximum likelihood estimate of the cleavage activity of Cas9 (Table 4, Figures 6A-6B and Figure 15 A moderate (0.433) but statistically significant correlation (p < 1 × 10) was found between guide RNA stability at the 16th protospacer position and reported off-target cleavage activity. -6 ) correlation. Notably, no statistically significant correlation was found between cleavage rate and the predicted individual DNA:RNA binding energy (0.0786; p = 0.3631) (Figures 6A and 6B). In addition to R-loop stability at position 16, stability at the 17th protospacer site was also found to be significantly correlated with reported cleavage (Table 5), but this was not the case for sites ≥ position 18 ( Figures 6A-6B Although the kinetic Monte Carlo model presented here is based on a relatively simple strand invasion model, these results further suggest that the stability of positions 16–17 of the protospacer, and therefore the accompanying conformational changes we observe, is correlated with Cas9 cleavage activity in vivo ( Figure 4D ).

[0396] Table 5 In the PAM distal region (≥10th protospacer site) aCorrelation between experimental cleavage frequencies at target sites containing single rG·dG, rC·dC, rA dA, and rU·dT mismatches (Hsu et al. (2013) Nature Biotechnology, 31, 827-832) and measures of guide RNA-protospacer stability

[0397]

[0398] a n=136.

[0399] b See Table 4 for details.

[0400] c See the text for details. Max(t)=100.

[0401] Due to the observed structural differences between dCas9 and tru-gRNAs and sgRNAs, we restricted most of our analyses to interactions with nucleotides 16–18 of the protospacer. However, we also observed an increase in the strength and statistical significance of the correlation between cleavage and stability at nucleotides 16–18 of the protospacer. Figures 6A-6B ), among which the correlation at position 14 is the most significant. Because the R loop is a dynamic structure ( Figure 4D ), so it is possible that interactions with these sites are those thought to be key interactions responsible for DNA cleavage. Truncating the guide RNA by 4 or 5 nucleotides could abolish cleavage activity by substantially destabilizing the R-loop at positions 14 or 15 in much the same way that tru-gRNA destabilizes the R-loop at positions 16-17. However, because in our model positions 14 and 15 are necessarily invaded whenever position 16 is bound by an sgRNA, it is likely that these positions have additional information because they are more strongly anticorrelated with the probability of the sgRNA dissociating from the duplex before bypassing the mismatch site (Figure 6Ai and Figure 16B), an alternative mechanism by which cleavage would not occur. Currently, there is no crystallographic evidence directly linking strand invasion to the observed conformational changes thought to authorize cleavage. However, based on the evidence provided by the AFM experiments presented here and the results of kinetic Monte Carlo simulations, we conclude that the stability of the guide RNA at positions 14-17 of the protospacer during invasion is critical for this conformational change and ultimately the specificity of Cas9 cleavage.

[0402] In addition, the R-loop, as a dynamic structure in competition between chain invasion and DNA reannealing, can be used to understand the mechanisms of off-target cleavage and mismatch tolerance. No statistically significant correlation was found between the cleavage rate and the predicted DNA-RNA binding energy alone (Figure 6B), suggesting that the kinetics of chain invasion can be considered when attempting to determine Cas9 activity at the detachment site. Although cleavage is eliminated when 4 or 5 nucleotides are truncated from the guide RNA, Cas9 is still able to cleave DNA with up to 6 distal mismatch sites. Transient nonspecific interactions at these PAM-distal sites can fully stabilize the conformational movement required for cleavage. Because we see a small fraction of the dCas9-sgRNA population at partial protospacer sites with similar structures to those at the full protospacer (yellow, Figure 3C (i)), this population may represent the fraction of Cas9 in a transiently stable active conformation. Therefore, this population may be responsible for off-target cleavage.

[0403] While Cas9 / dCas9 binding specificity is primarily determined by interactions with the PAM-proximal region, DNA cleavage specificity is likely controlled by a conformational change to an activated structure stabilized by the interaction of the guide RNA at the 14-17 bp region of the protospacer ( Figure 4D Kinetic Monte Carlo experiments revealed that R-loops formed during strand invasion of guide RNAs can be quite dynamic structures even when the guide RNA remains stably bound, suggesting a mechanism for improved tru-gRNA specificity and the origin of off-target cleavage via transient stabilization of the guide RNA-protospacer at a critical region around the mismatch site. The proposed mechanisms for the effects of each sgRNA variant on Cas9 / dCas9 specificity are summarized in Figures 7A-7C middle.

[0404] Using AFM, we found that HP-gRNAs significantly weakened or abolished specific binding at the cognate target. hp-gRNAs may be valuable for modulating dCas9 binding affinity and specificity in potential biological and medical applications. Specifically, based on the narrow geometry of the Cas9 binding channel, the presence of an unopened hairpin within the mismatched protospacer may inhibit Cas9 from undergoing conformational changes to its active state. The opening of the hairpin in hp-gRNAs upon binding could also serve as a binding-dependent signal in vivo, for example nucleating dynamic DNA / RNA structures only after binding to a specific site.

[0405] Previous studies of guide RNA truncation raised the question of why the native Cas9 system employed crRNAs targeting 20 bp protospacer sites when only 16 nucleotides of the guide sequence were required for cleavage and additional nucleotides (>18) did not improve in vivo cleavage specificity. These results suggest that the presence of the “extra” 5′ nucleotides binding to the 19th and 20th protospacer sites buffers transient reannealing at the critical 14th-17th positions of the protospacer, allowing efficient conformational changes to the active state and subsequent cleavage to occur. Results from AFM and KMC experiments indicate that the stability of the guide RNA at these sites after complete invasion shifts the equilibrium structure of Cas9 toward the active conformation ( FIG4A ), while the volatility of the R loop of the “truncated” guide RNA reduces the pressure to shift the equilibrium toward the active state. Cas9 with sgRNAs may also have an evolutionary advantage in its role as an adaptive immune agent against invasive DNA in prokaryotes compared to the promiscuous activity of tru-gRNAs, because the DNA of invading phages undergoes rapid point mutations at the sites targeted by Cas9 in order to avoid cleavage.

[0406] The design of guide RNA sequences for in vivo Cas9 / dCas9 applications has primarily focused on avoiding targets in the genome that contain multiple sites with similar sequences. However, recent studies exploring off-target cleavage have found that current methods for predicting off-target activity are largely ineffective. The stability of the R-loop during invasion correlates significantly better with off-target cleavage rates than does the guide RNA-protospacer binding energy or the location of mismatches alone (another important criterion used in guide RNA design; Table 3). The stability of the R-loop shortly after invasion initiation correlates much better with experimental cleavage rates than does the long-term stability in KMC experiments (Figure 16A), suggesting that the kinetics of strand invasion are a factor in predicting off-target activity.

[0407] Example 9

[0408] In vivo testing

[0409] The optimized gRNA activity was tested in living cells to investigate dCas9 binding specificity. Several hairpin gRNAs (hp-gRNAs) were designed for each of the four target positions (protospacers) in the human genome ( Figure 17 and Figure 18 One target is in the dystrophin gene ( Figure 19-23 ), another target is in the EMX1 gene ( Figure 24-29 and Figure 44), and two targets were in the VEGFA gene, labeled VEGFA1 ( Figures 30-37 ) and VEGFA3( Figure 38All experiments were performed in HEK293T cells.

[0410] Additional nucleotides (nt) are added to the 5' end of the full guide RNA (gRNA, full length 20 nt) and designed to form hairpins and secondary structures by hybridizing with the 5'-protospacer targeting nucleotides or nucleotides in the middle or 3' end of the protospacer targeting region to regulate the binding and cleavage activity of Cas9 towards the protospacer.

[0411] A secondary structure of the VEGFA1-targeting hp-gRNA was computationally designed using the methods described herein to prevent binding at known off-target sites while allowing binding to the full protospacer ( Figures 44A-44C ). hp-gRNAs were selected to have a binding lifetime greater than or equal to that of the full gRNA at the on-target site, and less than or equal to that of the full-length gRNA at the first three off-target sites. Other 5' structures were designed to include dG-rU wobble base pairs to adjust the energetics of the hp-gRNA's secondary structure, or added to the ends of truncated gRNAs (tru-gRNAs, <20 nt), which themselves have been shown to promote greater specificity of Cas9 activity.

[0412] Cell work. For deep sequencing analysis, 293T cells were transfected with plasmids expressing Cas9 and the gRNA of interest. The cells were incubated for 4 days to allow Cas9 and gRNA to exert their maximum activity. The cells were then harvested and their genomic DNA was purified. gRNAs that have been well characterized in the literature (i.e., their on-target and off-target sites are known) were used.

[0413] Surveyor determination. Compared to deep sequencing, the surveyor assay has a lower throughput and lower sensitivity. However, the surveyor assay is faster and less technical in data analysis, providing gel images. Therefore, the surveyor assay was performed as the first lane, and the optimal conditions were analyzed in triplicate using deep sequencing. Both deep sequencing and the surveyor assay are methods for quantifying mutation events caused by Cas9+gRNA.

[0414] The cell work for Surveyor is the same as described above. After genomic DNA purification, primers are designed to amplify target sites. In this experiment, 200k cell pools were used, and because DNA repair is random, each of them has different mutations. The sites on 200k cells were amplified to produce heterologous PCR products: because each cell randomly (i.e., randomly, fallible) repairs the Cas9 cleavage site, some amplicons have deletions, some amplicons have insertions, and some amplicons are wild type and unmodified.

[0415] The heterologous PCR pool is heated and repaired, and in some cases the different strands anneal to each other: the wild-type DNA strand may bind to the DNA with the insertion, or the insertion may bind to the deletion. When this happens, a small "bubble" is formed, and the structure is called a DNA heteroduplex (see Figure 46 ).

[0416] Surveyor nuclease was used to detect these heteroduplexes by digesting and cleaving them. DNA cleavage was then a proxy for Cas9 mutation activity. PCR pools were separated on gels, and the intensity of these digestion bands was used to quantify the rate of Cas9 activity.

[0417] Deep sequencing. Primers were designed to amplify these known targets / off-targets. A high-fidelity polymerase was used in the PCR. Illumina adapters were also present on these primers so that they could be barcoded and loaded onto the Illumina Mi-Seq platform. The number of hairpins, number of targets, number of off-targets, sequencing coverage, etc. are described in the accompanying drawings and the accompanying description. Good coverage was obtained on the samples used in the analysis. The average number of reads / sample was 20,000. The sample with the fewest reads was 1,700. A very small number of targets did not generate enough aligned reads and were not included in the analysis.

[0418] The resulting sequencing data were analyzed using CRISPResso software (Pinello et al. Nature Biotechnology (2016) 34(7): 695-697), which aligns deep sequencing reads to specific sites known to be off-target or on-target. The results from this software were compared with an in-house script that globally aligned deep sequencing reads to the human genome and correlated well. Mutation rates were quantified using CRISPResso, and the resulting data are presented in histograms for each target gene.

[0419] The Surveyor assay was first designed to test indels at known target sites and off-target sites after expressing Cas9 and hp-gRNA in HEK cells using standard gRNAs (see Table 6). The activity at these sites was compared with standard gRNAs and truncated gRNAs (tru-gRNAs). These are shown below as gels showing the cleavage of genomic DNA by Surveyor nuclease, where cleavage indicates mutagenesis of Cas9.

[0420] Table 6

[0421]

[0422] The most promising hp-gRNA designs were selected for additional quantitative analysis using next-generation sequencing to evaluate Cas9 activity at target and off-target sites in HEK cells. Specificity was defined as the number of on-target hits / total (number of off-target hits).

[0423] Although Cas9 activity was generally equal or slightly decreased when hp-gRNAs were used, each hp-gRNA selected for deep sequencing experiments showed enhanced specificity compared to truncate gRNAs and, in most cases, was equal to or greater than tru-gRNAs in terms of specificity.

[0424] In one case, an hp-gRNA hairpin targeting EMX1 exhibited a >6000-fold improvement in specificity compared to a full gRNA (compared to a ≈100 improvement for a tru-gRNA). A VEGFA1-targeting hp-gRNA with a secondary structure computationally designed using an in-house algorithm significantly outperformed tru-gRNA activity in terms of specificity (3-fold improvement compared to an 18-fold improvement for the gRNA). These hp-gRNAs were tested in combination with Streptococcus pyogenes Cas9. Figures 44A-44C Shown is a Surveyor assay using an EMX1-targeting hp-gRNA of Cas9 from Streptococcus pyogenes that exhibited on-target activity and no detectable off-target activity, in contrast to a tru-gRNA that showed significant off-target activity.

[0425] Example 10

[0426] hp-gRNA for CRISPR / Cpf1 system

[0427] Experiments were designed to reproduce the results of Kleinstiver et al., Nature Biotechnology (2016) 34:869-874. Kleinstiver et al. used full-length gRNAs to show that Cpf1 of the Lachnospiraceae family readily cleaves off-target sites with mismatches at 8-9 nucleotides beyond the PAM distal site. By using gRNAs with mismatches at different positions in the target site ( Figure 47 ). In this example, hairpin guide RNAs for use with the type V CRISPR-Cas system CRISPR-Cpf1 were designed and tested as described above using the methods of the invention.

[0428] To test the off-target activity of Cpf1 with and without additional secondary structural elements, the DNMT1 gene (TTTC CTGATGGGTCCATGTCTGTTACTC (SEQ ID NO: 330)) was targeted for Cpf1 cleavage. "Off-target activity" was tested by using a guide RNA with a mismatch nucleotide at position 9, e.g., CTGATGGTgCATGTCT GTTA( SEQ ID NO: 331), using a 20-nucleotide full-length guide RNA or a 17-nucleotide truncated gRNA, CTGATGGTgCATG TCTG (SEQ ID NO: 332). A 9 nucleotide long secondary structure element was added to the 3' end of the Cpf1 guide RNA to hybridize to the segment of the guide RNA surrounding the mismatched nucleotide, where in this case the "linker" element includes the 4 3' nt of the protospacer targeting segment, i.e., CTGATGGTgCATGTCT GTTA AGACATGcACCA (SEQ ID NO: 333) and CTGATGGTgCATG TCTG CATGcACCA (SEQ ID NO: 334). Surveyor assays showed that inclusion of these additional 3' elements reduced or eliminated off-target activity at the DNMT1 site exhibited by full or truncated gRNAs.

[0429] HP-gRNA was designed with an "internal" hairpin design, in which the 4 nucleotides distal to the PAM act as a loop. The hairpin was added to the 3' end of the gRNA. Table 7 shows the sequence of the hp-gRNA with spaces separating this region. Mismatches are shown in lowercase letters.

[0430] The survey results of these HP-gRNAs are shown in Figure 48The results show that adding a hairpin to the 3' end eliminates off-target activity. Lane 1 shows a control; Lane 2 shows a full-length gRNA containing a mismatched nucleotide at position 9; Lane 3 shows a full-length gRNA containing a mismatched nucleotide at position 9 and an additional 3' hairpin structure; Lane 4 shows a truncated gRNA containing a mismatched nucleotide at position 9; and Lane 5 shows a truncated gRNA containing a mismatched nucleotide at position 9 and an additional 3' hairpin structure. The Surveyor primers used are also shown in Table 7.

[0431] When using normal guide RNA, Cpf1 tolerates mismatches at 8–10 nucleotides and cleaves DNA at these off-target sites ( Figure 47 ).like Figure 48 As shown, Cpf1 hp-gRNA was able to eliminate the off-target activity displayed in Kleinstiver, while the truncated gRNA was not.

[0432] Table 7

[0433]

[0434] It should be understood that the foregoing detailed description and accompanying examples are illustrative only and should not be taken as limiting the scope of the present invention, which is defined only by the appended claims and their equivalents.

[0435] Various changes and modifications to the disclosed examples will be apparent to those skilled in the art. Such changes and modifications, including but not limited to those related to the chemical structures, substituents, derivatives, intermediates, syntheses, compositions, formulations or methods of use of the present invention, may be made without departing from the spirit and scope of the present invention.

[0436] For reasons of completeness, various aspects of the invention will be set out in the following numbered clauses:

[0437] Item 1. A method for producing an optimized guide RNA (gRNA), the method comprising: a) Identifying a target region of interest, the target region of interest comprising a protospacer sequence; b) determining a polynucleotide sequence of a full-length gRNA targeting the target region of interest, the full-length gRNA comprising a protospacer targeting sequence or segment; c) determining at least one or more off-target sites of the full-length gRNA; d) generating a polynucleotide sequence of a first gRNA, the first gRNA comprising the polynucleotide sequence of the full-length gRNA and an RNA segment, the RNA segment comprising a polynucleotide sequence having a length of M nucleotides, the polynucleotide sequence being complementary to a nucleotide segment of the protospacer targeting sequence or segment, the RNA segment being located at the 5' end of the polynucleotide sequence of the full-length gRNA, the first gRNA optionally comprising a linker between the 5' end of the polynucleotide sequence of the full-length gRNA and the RNA segment, the linker comprising a polynucleotide sequence having a length of N nucleotides, the first gRNA being capable of invading the protospacer sequence and binding to a DNA sequence complementary to the protospacer sequence and forming a protospacer-duplex, and the first gRNA being capable of invading an off-target site and binding to a DNA sequence complementary to the off-target site. sequence and form an off-target duplex; e) calculating an estimate of invasion kinetics and a lifetime that the first gRNA remains invaded in the protospacer and off-target site duplex or computationally simulating these invasion kinetics and lifetime, wherein invasion dynamics are estimated nucleotide by nucleotide by determining the energy difference between further invasion by a different gRNA and reannealing of the first gRNA to the DNA sequence complementary to the protospacer sequence; f) comparing the estimated lifetime of the first gRNA at the protospacer and / or off-target site with the estimated lifetime of the full-length gRNA or truncated gRNA (tru-gRNA) at the protospacer and / or off-target site; g) randomizing 0 to N nucleotides in the linker and 0 to M nucleotides in the first gRNA and generating a second gRNA, and repeating step (e) with the second gRNA; h) identifying an optimized gRNA based on the gRNA sequence that meets the design criteria; and i) testing the optimized gRNA in vivo to determine binding specificity.

[0438] Item 2. A method for producing an optimized guide RNA (gRNA), the method comprising: a) Identifying a target region of interest, the target region of interest comprising a protospacer sequence; b) determining a polynucleotide sequence of a full-length gRNA targeting the target region of interest, the full-length gRNA comprising a protospacer targeting sequence or segment; c) determining at least one or more off-target sites of the full-length gRNA; d) generating a polynucleotide sequence of a first gRNA, the first gRNA comprising the polynucleotide sequence of the full-length gRNA and an RNA segment, the RNA segment comprising a polynucleotide sequence having a length of M nucleotides, the polynucleotide sequence being complementary to a nucleotide segment of the protospacer targeting sequence or segment, the RNA segment being located at the 3' end of the polynucleotide sequence of the full-length gRNA, the first gRNA optionally comprising a linker between the 3' end of the polynucleotide sequence of the full-length gRNA and the RNA segment, the linker comprising a polynucleotide sequence having a length of N nucleotides, the first gRNA being capable of invading the protospacer sequence and binding to a DNA sequence complementary to the protospacer sequence and forming a protospacer-duplex, and the first gRNA being capable of invading an off-target site and binding to a DNA sequence complementary to the off-target site. sequence and form an off-target duplex; e) calculating an estimate of invasion kinetics and a lifetime that the first gRNA remains invaded in the protospacer and off-target site duplex or computationally simulating these invasion kinetics and lifetime, wherein invasion dynamics are estimated nucleotide by nucleotide by determining the energy difference between further invasion by a different gRNA and reannealing of the first gRNA to the DNA sequence complementary to the protospacer sequence; f) comparing the estimated lifetime of the first gRNA at the protospacer and / or off-target site with the estimated lifetime of the full-length gRNA or truncated gRNA (tru-gRNA) at the protospacer and / or off-target site; g) randomizing 0 to N nucleotides in the linker and 0 to M nucleotides in the first gRNA and generating a second gRNA, and repeating step (e) with the second gRNA; h) identifying an optimized gRNA based on the gRNA sequence that meets the design criteria; and i) testing the optimized gRNA in vivo to determine binding specificity.

[0439] Clause 3. The method of clause 1 or 2, wherein the energetics of further invasion of a different gRNA is determined by determining the energetics of at least one of: (I) disruption of DNA-DNA base pairing, (II) formation of RNA-DNA base pairs, (III) the energy difference resulting from disruption or formation of a different secondary structure within the non-invaded guide RNA, and (IV) formation or disruption of an interaction between a displacing DNA strand complementary to the protospacer and any unpaired guide RNA nucleotides that are not involved in a secondary structure.

[0440] Clause 4. The method of any of clauses 1-3, wherein the energetics of reannealing of the first gRNA to the DNA sequence complementary to the protospacer sequence is determined by determining the energetics of at least one of: (I) formation of DNA-DNA base pairing, (II) disruption of RNA-DNA base pairs, (III) the energy difference resulting from disruption or formation of a different secondary structure within the newly uninvaded guide RNA, and (IV) formation or disruption of an interaction between the displaced DNA strand complementary to the protospacer and any unpaired guide RNA nucleotides that are not involved in secondary structure.

[0441] Item 5. The method of item 3 or 4, further comprising determining an energy consideration based on at least one of: (V) base pairing across mismatches, (VI) interaction with the Cas9 protein, and / or (VII) additional heuristics, wherein the additional heuristics relate to binding lifetime, extent of invasion, stability of the invading guide RNA, or other computational / simulated properties of gRNA invasion to Cas9 cleavage activity.

[0442] Item 6. The method of any one of items 1-5, wherein the full-length gRNA comprises about 15 to 20 nucleotides.

[0443] Clause 7. The method of any one of clauses 1-5, wherein M is between 1 and 20.

[0444] Clause 8. The method of clause 7, wherein M is between 4 and 10.

[0445] Item 9. The method of any one of items 1-8, wherein the RNA segment comprises 2 to 15 nucleotides complementary to the protospacer targeting segment.

[0446] Clause 10. The method of any one of clauses 1-9, wherein N is between 1 and 20.

[0447] Clause 11. The method of clause 10, wherein N is between 3 and 10.

[0448] Item 12. The method of any one of items 1-11, wherein the RNA segment and / or protospacer targeting sequence provides secondary structure.

[0449] Item 13. The method of Item 12, wherein the secondary structure is formed by partially hybridizing the protospacer targeting sequence to the RNA segment.

[0450] Item 14. The method of Item 13, wherein the secondary structure modulates DNA binding or Cas9 cleavage by disrupting invasion of the optimized gRNA into the protospacer duplex or the off-target duplex.

[0451] Item 15. The method of any of items 12-14, wherein the secondary structure is formed by hybridizing all or part of the RNA segment to nucleotides in the 5' end of the protospacer targeting sequence or segment, nucleotides in the middle of the protospacer targeting sequence or segment, and / or nucleotides at the 3' end of the protospacer targeting sequence or segment.

[0452] Clause 16. The method of any one of clauses 12-15, wherein the secondary structure is a hairpin.

[0453] Clause 17. The method of any one of clauses 12-16, wherein the secondary structure is stable at room temperature or 37°C.

[0454] Clause 18. The method of any one of Clauses 12-17, wherein the total equilibrium free energy of the secondary structure is less than about 2 kcal / mol at room temperature or 37°C.

[0455] Clause 19. The method of any one of clauses 1-18, wherein the RNA segment hybridizes or forms non-canonical base pairs with at least two nucleotides of the protospacer targeting sequence or segment.

[0456] Item 20. The method of Item 19, wherein the non-canonical base pair is rU-rG.

[0457] Item 21. The method of any one of items 1-20, wherein the optimized gRNA is used in a cell together with a CRISPR / Cas9-based system or a CRISPR / Cpf1-based system.

[0458] Item 22. The method of any one of items 1-21, wherein the secondary structure protects the optimized gRNA within the CRISPR / Cas9-based system or the CRISPR / Cpf1-based system to prevent degradation in the cell.

[0459] Clause 23. The method of any one of clauses 1-22, wherein 1-20 nucleotides are randomized in the linker.

[0460] Item 24. The method of any one of items 1-23, wherein 1-20 nucleotides are randomized in the RNA segment.

[0461] Clause 25. The method of any one of clauses 1-24, wherein step (g) is repeated X times, thereby producing X number of gRNAs, and step (e) is repeated for each X number of gRNAs, wherein X is between 0 and 20.

[0462] Clause 26. The method of any one of Clauses 1-25, wherein the invasion kinetics and the lifetime are calculated using a kinetic Monte Carlo method or a Gillespie algorithm.

[0463] Item 27. A method as described in any of items 1-26, wherein the invasion kinetics is the rate at which the guide RNA invades the protospacer duplex to complete invasion so that the protospacer is completely invaded, or the rate at which the segment of the protospacer DNA that binds the gRNA expands as it shifts from its complementary strand and binds the guide RNA nucleotide by nucleotide from its PAM-proximal region to complete invasion.

[0464] Clause 28. The method of any one of clauses 1-27, wherein the design criteria comprises specificity, binding lifetime modulation and / or estimated cleavage specificity.

[0465] Item 29. A method as described in Item 28, wherein the design criteria include optimizing the gRNA to have a binding lifetime that is greater than or equal to the binding lifetime of the full-length gRNA to the target site, or a binding lifetime that is less than or equal to the binding lifetime of the full-length gRNA to the off-target site.

[0466] Item 30. A method as described in Item 29, wherein the design criteria include optimizing the gRNA to have a binding lifetime that is less than or equal to the binding lifetime of the full-length gRNA for at least three off-target sites, wherein these off-target sites are predicted to be the closest off-target sites or are predicted to have the highest identity with these on-target sites.

[0467] Item 31. A method as described in Item 28, wherein the design criteria include a lifetime or cleavage rate at off-target sites that is less than or equal to the lifetime or cleavage rate of the full-length gRNA or truncated gRNA at these off-target sites and / or a predicted on-target activity rate that is greater than 10% of the predicted on-target activity rate of the full-length gRNA or truncated gRNA.

[0468] Item 32. The method of any one of items 1-31, wherein the optimized gRNA is tested in step i) using a surveyor assay, next generation sequencing technology, or GUIDE-Seq.

[0469] Item 33. The method of any one of items 1-32, wherein the optimized gRNA is designed to minimize binding at off-target sites and to allow binding to protospacer sequences.

[0470] Item 34. The method of any one of items 1-33, wherein the off-target site is a known or predicted off-target site.

[0471] Item 35. The method of any one of items 1-34, wherein the full-length gRNA targets a mammalian gene.

[0472] Item 36. The method of any one of items 1-35, wherein the target gene comprises an endogenous target gene or a transgene.

[0473] Item 37. The method of any one of items 1-36, wherein the target gene comprises a disease-associated gene.

[0474] Item 38. The method of any one of items 1-37, wherein the target gene is the DMD, EMX1 or VEGFA gene.

[0475] Item 39. The method of Item 38, wherein the VEGFA gene is VEGFA1 or VEGFA3.

[0476] Item 40. An optimized gRNA produced by the method of any one of items 1-39.

[0477] Item 41. The optimized gRNA of Item 40, wherein the gRNA can distinguish between on-target and off-target sites with minimal thermodynamic energy difference between these sites.

[0478] Item 42. The optimized gRNA of Item 40 or 41, wherein the optimized gRNA regulates strand invasion into the protospacer.

[0479] Item 43. The optimized gRNA of any one of items 40-42, wherein the optimized gRNA comprises a nucleotide sequence of at least one of SEQ ID NOs: 149-315, 321-323, and 326-329.

[0480] Item 44. An isolated polynucleotide encoding the optimized gRNA of any one of items 40-43.

[0481] Item 45. A vector comprising the isolated polynucleotide of Item 44.

[0482] Item 46. A cell comprising the isolated polynucleotide of Item 44 or the vector of Item 45.

[0483] Item 47. A kit comprising the isolated polynucleotide of Item 44, the vector of Item 45, or the cell of Item 46.

[0484] Item 48. A method for epigenome editing in a target cell or subject, the method comprising contacting the cell or subject with an effective amount of an optimized gRNA molecule as described in any of items 40-43 or an isolated polynucleotide as described in item 44 and a fusion protein, the fusion protein comprising a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0485] Item 49. A method for performing site-specific DNA cleavage in a target cell or subject, the method comprising contacting the cell or subject with an effective amount of an optimized gRNA molecule as described in any of items 40-43 or an isolated polynucleotide as described in item 44 and a fusion protein or Cas9 protein, wherein the fusion protein comprises a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0486] Item 50. A method for performing genome editing in a cell, the method comprising administering to the cell an effective amount of an optimized gRNA molecule as described in any of items 40-43 or an isolated polynucleotide as described in item 44 and a fusion protein, the fusion protein comprising a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0487] Item 51. A method as described in Item 50, wherein the genome editing includes correcting a mutant gene or inserting a transgene.

[0488] Item 52. The method of Item 51, wherein correcting the mutant gene comprises deleting, rearranging or replacing the mutant gene.

[0489] Item 53. A method as described in any of items 51 or 52, wherein correcting the mutant gene comprises nuclease-mediated non-homologous end joining or homology-directed repair.

[0490] Item 54. A method for regulating gene expression in a cell, the method comprising contacting the cell with an effective amount of an optimized gRNA molecule as described in any of items 40-43 or an isolated polynucleotide as described in item 44 and a fusion protein, the fusion protein comprising a first polypeptide domain comprising a nuclease-deficient Cas9 and a second polypeptide domain having an activity selected from the group consisting of: transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

[0491] Item 55. The method of Item 54, wherein the gene expression of the at least one target gene is regulated when the gene expression of the at least one target gene is increased or decreased compared to the normal gene expression level of the at least one target gene.

[0492] Item 56. A method as described in item 54 or 55, wherein the fusion protein comprises a dCas9 domain and a transcriptional activator.

[0493] Item 57. The method of Item 56, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO: 2.

[0494] Item 58. A method as described in item 54 or 55, wherein the fusion protein comprises a dCas9 domain and a transcriptional repressor.

[0495] Item 59. The method of Item 58, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO: 3.

[0496] Item 60. A method as described in item 54 or 55, wherein the fusion protein comprises a dCas9 domain and a site-specific nuclease.

[0497] Item 61. The method of any one of items 48-60, wherein the optimized gRNA is encoded by a polynucleotide sequence and packaged into a lentiviral vector.

[0498] Item 62. A method as described in Item 61, wherein the lentiviral vector comprises an expression cassette comprising a promoter operably linked to the polynucleotide sequence encoding the gRNA.

[0499] Item 63. A method as described in Item 62, wherein the promoter operably linked to the polynucleotide encoding the optimized gRNA is inducible.

[0500] Item 64. The method of any one of items 61-63, wherein the lentiviral vector further comprises a polynucleotide sequence encoding the Cas9 protein or fusion protein.

[0501] Item 65. The method of any one of items 48-64, wherein the at least one target gene is a disease-associated gene.

[0502] Item 66. The method of any one of items 48-65, wherein the target cell is a eukaryotic cell.

[0503] Item 67. The method of any one of items 48-66, wherein the target cell is a mammalian cell.

[0504] The method of any one of clauses 48-67, wherein the target cell is a HEK293T cell.

[0505] Appendix - Sequence

[0506] Streptococcus pyogenes Cas9 (with D10A, H840A) (SEQ ID NO: 1)

[0507]

[0508] dCas9 p300核心 :(Addgene plasmid 61357) amino acid sequence; 3X "Flag" epitope, nuclear localization sequence, Streptococcus pyogenes Cas9 ( D10A 、 H840A ), p300 core effector, "HA" epitope (SEQ ID NO: 2)

[0509]

[0510]

[0511] dCas9 KRAB (SEQ ID NO:3)

[0512]

[0513]

[0514] Nm-dCas9 p300核心 :(Addgene plasmid 61365) amino acid sequence; Neisseria meningitidis Cas9 ( D16A 、 D587A 、 H588A 、 N611A ), nuclear localization sequence, p300 core effector, " HA ” Epitope (SEQ ID NO: 5)

[0515]

[0516] Sequence Listing <110> Duke University <120> Compositions and methods for improving specificity of genome engineering using RNA-guided endonucleases <130> 028193-9240-WO00 <140> PCT / US2016 / 048798 <141> 2016-08-25 <150> 62 / 209,466 <151> 2015-08-25 <160> 348 <170> PatentIn version 3.5 <210> 1 <211> 1368 <212> PRT <213> Streptococcus pyogenes <400> 1 Met Asp Lys Lys Tyr Ser Ile Gly Leu Ala Ile Gly Thr Asn Ser Val 1 5 10 15 Gly Trp Ala Val Ile Thr Asp Glu Tyr Lys Val Pro Ser Lys Lys Phe 20 25 30 Lys Val Leu Gly Asn Thr Asp Arg His Ser Ile Lys Lys Asn Leu Ile 35 40 45 Gly Ala Leu Leu Phe Asp Ser Gly Glu Thr Ala Glu Ala Thr Arg Leu 50 55 60 Lys Arg Thr Ala Arg Arg Arg Tyr Thr Arg Arg Lys Asn Arg Ile Cys 65 70 75 80 Tyr Leu Gln Glu Ile Phe Ser Asn Glu Met Ala Lys Val Asp Asp Ser 85 90 95 Phe Phe His Arg Leu Glu Glu Ser Phe Leu Val Glu Glu Asp Lys Lys 100 105 110 His Glu Arg His Pro Ile Phe Gly Asn Ile Val Asp Glu Val Ala Tyr 115 120 125 His Glu Lys Tyr Pro Thr Ile Tyr His Leu Arg Lys Lys Leu Val Asp 130 135 140 Ser Thr Asp Lys Ala Asp Leu Arg Leu Ile Tyr Leu Ala Leu Ala His 145 150 155 160 Met Ile Lys Phe Arg Gly His Phe Leu Ile Glu Gly Asp Leu Asn Pro 165 170 175 Asp Asn Ser Asp Val Asp Lys Leu Phe Ile Gln Leu Val Gln Thr Tyr 180 185 190 Asn Gln Leu Phe Glu Glu Asn Pro Ile Asn Ala Ser Gly Val Asp Ala 195 200 205 Lys Ala Ile Leu Ser Ala Arg Leu Ser Lys Ser Arg Arg Leu Glu Asn 210 215 220 Leu Ile Ala Gln Leu Pro Gly Glu Lys Lys Asn Gly Leu Phe Gly Asn 225 230 235 240 Leu Ile Ala Leu Ser Leu Gly Leu Thr Pro Asn Phe Lys Ser Asn Phe 245 250 255 Asp Leu Ala Glu Asp Ala Lys Leu Gln Leu Ser Lys Asp Thr Tyr Asp 260 265 270 Asp Asp Leu Asp Asn Leu Leu Ala Gln Ile Gly Asp Gln Tyr Ala Asp 275 280 285 Leu Phe Leu Ala Ala Lys Asn Leu Ser Asp Ala Ile Leu Leu Ser Asp 290 295 300 Ile Leu Arg Val Asn Thr Glu Ile Thr Lys Ala Pro Leu Ser Ala Ser 305 310 315 320 Met Ile Lys Arg Tyr Asp Glu His His Gln Asp Leu Thr Leu Leu Lys 325 330 335 Ala Leu Val Arg Gln Gln Leu Pro Glu Lys Tyr Lys Glu Ile Phe Phe 340 345 350 Asp Gln Ser Lys Asn Gly Tyr Ala Gly Tyr Ile Asp Gly Gly Ala Ser 355 360 365 Gln Glu Glu Phe Tyr Lys Phe Ile Lys Pro Ile Leu Glu Lys Met Asp 370 375 380 Gly Thr Glu Glu Leu Leu Val Lys Leu Asn Arg Glu Asp Leu Leu Arg 385 390 395 400 Lys Gln Arg Thr Phe Asp Asn Gly Ser Ile Pro His Gln Ile His Leu 405 410 415 Gly Glu Leu His Ala Ile Leu Arg Arg Gln Glu Asp Phe Tyr Pro Phe 420 425 430 Leu Lys Asp Asn Arg Glu Lys Ile Glu Lys Ile Leu Thr Phe Arg Ile 435 440 445 Pro Tyr Tyr Val Gly Pro Leu Ala Arg Gly Asn Ser Arg Phe Ala Trp 450 455 460 Met Thr Arg Lys Ser Glu Glu Thr Ile Thr Pro Trp Asn Phe Glu Glu 465 470 475 480 Val Val Asp Lys Gly Ala Ser Ala Gln Ser Phe Ile Glu Arg Met Thr 485 490 495 Asn Phe Asp Lys Asn Leu Pro Asn Glu Lys Val Leu Pro Lys His Ser 500 505 510 Leu Leu Tyr Glu Tyr Phe Thr Val Tyr Asn Glu Leu Thr Lys Val Lys 515 520 525 Tyr Val Thr Glu Gly Met Arg Lys Pro Ala Phe Leu Ser Gly Glu Gln 530 535 540 Lys Lys Ala Ile Val Asp Leu Leu Phe Lys Thr Asn Arg Lys Val Thr 545 550 555 560 Val Lys Gln Leu Lys Glu Asp Tyr Phe Lys Lys Ile Glu Cys Phe Asp 565 570 575 Ser Val Glu Ile Ser Gly Val Glu Asp Arg Phe Asn Ala Ser Leu Gly 580 585 590 Thr Tyr His Asp Leu Leu Lys Ile Ile Lys Asp Lys Asp Phe Leu Asp 595 600 605 Asn Glu Glu Asn Glu Asp Ile Leu Glu Asp Ile Val Leu Thr Leu Thr 610 615 620 Leu Phe Glu Asp Arg Glu Met Ile Glu Glu Arg Leu Lys Thr Tyr Ala 625 630 635 640 His Leu Phe Asp Asp Lys Val Met Lys Gln Leu Lys Arg Arg Arg Tyr 645 650 655 Thr Gly Trp Gly Arg Leu Ser Arg Lys Leu Ile Asn Gly Ile Arg Asp 660 665 670 Lys Gln Ser Gly Lys Thr Ile Leu Asp Phe Leu Lys Ser Asp Gly Phe 675 680 685 Ala Asn Arg Asn Phe Met Gln Leu Ile His Asp Asp Ser Leu Thr Phe 690 695 700 Lys Glu Asp Ile Gln Lys Ala Gln Val Ser Gly Gln Gly Asp Ser Leu 705 710 715 720 His Glu His Ile Ala Asn Leu Ala Gly Ser Pro Ala Ile Lys Lys Gly 725 730 735 Ile Leu Gln Thr Val Lys Val Val Asp Glu Leu Val Lys Val Met Gly 740 745 750 Arg His Lys Pro Glu Asn Ile Val Ile Glu Met Ala Arg Glu Asn Gln 755 760 765 Thr Thr Gln Lys Gly Gln Lys Asn Ser Arg Glu Arg Met Lys Arg Ile 770 775 780 Glu Glu Gly Ile Lys Glu Leu Gly Ser Gln Ile Leu Lys Glu His Pro 785 790 795 800 Val Glu Asn Thr Gln Leu Gln Asn Glu Lys Leu Tyr Leu Tyr Tyr Leu 805 810 815 Gln Asn Gly Arg Asp Met Tyr Val Asp Gln Glu Leu Asp Ile Asn Arg 820 825 830 Leu Ser Asp Tyr Asp Val Asp Ala Ile Val Pro Gln Ser Phe Leu Lys 835 840 845 Asp Asp Ser Ile Asp Asn Lys Val Leu Thr Arg Ser Asp Lys Asn Arg 850 855 860 Gly Lys Ser Asp Asn Val Pro Ser Glu Glu Val Val Lys Lys Met Lys 865 870 875 880 Asn Tyr Trp Arg Gln Leu Leu Asn Ala Lys Leu Ile Thr Gln Arg Lys 885 890 895 Phe Asp Asn Leu Thr Lys Ala Glu Arg Gly Gly Leu Ser Glu Leu Asp 900 905 910 Lys Ala Gly Phe Ile Lys Arg Gln Leu Val Glu Thr Arg Gln Ile Thr 915 920 925 Lys His Val Ala Gln Ile Leu Asp Ser Arg Met Asn Thr Lys Tyr Asp 930 935 940 Glu Asn Asp Lys Leu Ile Arg Glu Val Lys Val Ile Thr Leu Lys Ser 945 950 955 960 Lys Leu Val Ser Asp Phe Arg Lys Asp Phe Gln Phe Tyr Lys Val Arg 965 970 975 Glu Ile Asn Asn Tyr His His Ala His Asp Ala Tyr Leu Asn Ala Val 980 985 990 Val Gly Thr Ala Leu Ile Lys Lys Tyr Pro Lys Leu Glu Ser Glu Phe 995 1000 1005 Val Tyr Gly Asp Tyr Lys Val Tyr Asp Val Arg Lys Met Ile Ala 1010 1015 1020 Lys Ser Glu Gln Glu Ile Gly Lys Ala Thr Ala Lys Tyr Phe Phe 1025 1030 1035 Tyr Ser Asn Ile Met Asn Phe Phe Lys Thr Glu Ile Thr Leu Ala 1040 1045 1050 Asn Gly Glu Ile Arg Lys Arg Pro Leu Ile Glu Thr Asn Gly Glu 1055 1060 1065 Thr Gly Glu Ile Val Trp Asp Lys Gly Arg Asp Phe Ala Thr Val 1070 1075 1080 Arg Lys Val Leu Ser Met Pro Gln Val Asn Ile Val Lys Lys Thr 1085 1090 1095 Glu Val Gln Thr Gly Gly Phe Ser Lys Glu Ser Ile Leu Pro Lys 1100 1105 1110 Arg Asn Ser Asp Lys Leu Ile Ala Arg Lys Lys Asp Trp Asp Pro 1115 1120 1125 Light Light Tyr Gly Gly Phe Asp Ser Pro Thr Val Ala Tyr Ser Val 1130 1135 1140 Leu Val Val Ala Lys Val Glu Lys Gly Lys Ser Lys Lys Leu Lys 1145 1150 1155 Ser Val Lys Glu Leu Leu Gly Ile Thr Ile Met Glu Arg Ser Ser 1160 1165 1170 Phe Glu Lys Asn Pro Ile Asp Phe Leu Glu Ala Lys Gly Tyr Lys 1175 1180 1185 Glu Val Lys Lys Asp Leu Ile Ile Lys Leu Pro Lys Tyr Ser Leu 1190 1195 1200 Phe Glu Leu Glu Asn Gly Arg Lys Arg Met Leu Ala Ser Ala Gly 1205 1210 1215 Glu Leu Gln Lys Gly Asn Glu Leu Ala Leu Pro Ser Lys Tyr Val 1220 1225 1230 Asn Phe Leu Tyr Leu Ala Ser His Tyr Glu Lys Leu Lys Gly Ser 1235 1240 1245 Pro Glu Asp Asn Glu Gln Lys Gln Leu Phe Val Glu Gln His Lys 1250 1255 1260 His Tyr Leu Asp Glu Ile Ile Glu Gln Ile Ser Glu ...

Claims

1. A Cpf1 gRNA comprising a protospacer targeting sequence, wherein the protospacer targeting sequence consists of a polynucleotide encoded by the sequence of SEQ ID NO:

328.

2. An isolated polynucleotide encoding the gRNA according to claim 1.

3. The isolated polynucleotide of claim 2, further encoding Cpf1.

4. A vector comprising the isolated polynucleotide according to claim 2.

5. A kit comprising the isolated polynucleotide according to claim 2.

6. Use of the gRNA according to claim 1 in preparing a kit for a method for epigenome editing a target gene in a target cell or subject, the method comprising contacting a cell or subject with an effective amount of the gRNA and a fusion protein or expressing an effective amount of the gRNA and the fusion protein in a cell or subject, the fusion protein comprising a first polypeptide domain comprising nuclease-deficient Cpf1 and a second polypeptide domain having an activity selected from the group consisting of: Transcriptional activation activity, transcriptional repression activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity, wherein the gRNA targets a target region of a target gene, and wherein the target gene is DNMT1.

7. Use of the gRNA according to claim 1 in the preparation of a kit for a method for site-specific DNA cleavage of a target gene in a target cell or subject, the method comprising contacting the cell or subject with an effective amount of the gRNA and Cpf1 RNA-guided endonuclease or expressing an effective amount of the gRNA and Cpf1 RNA-guided endonuclease in the cell or subject, wherein the gRNA targets a target region of the target gene, and wherein the target gene is DNMT1.

8. The use according to claim 7, wherein the Cpf1 RNA-guided endonuclease comprises Cpf1 or a fusion protein comprising a first polypeptide domain containing nuclease-deficient Cpf1 and a second polypeptide domain having nuclease activity.

9. Use of the gRNA and Cpf1 RNA-guided endonuclease according to claim 1 in preparing a kit for a method for performing genome editing in a cell, the method comprising administering an effective amount of the gRNA and Cpf1 RNA-guided endonuclease according to claim 1 to the cell or expressing an effective amount of the gRNA and Cpf1 RNA-guided endonuclease according to claim 1 in the cell, wherein the gRNA targets a target region of a target gene, and wherein the target gene is DNMT1.

10. The use according to claim 9, wherein the Cpf1 RNA-guided endonuclease comprises a fusion protein comprising a first polypeptide domain comprising nuclease-deficient Cpf1 and a second polypeptide domain having an activity selected from the group consisting of: Transcriptional activation activity, transcriptional repressor activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

11. The use of claim 9, wherein the genome editing comprises correcting a mutation in a target gene or inserting a transgene.

12. The method of claim 11, wherein correcting a mutation in the target gene comprises deleting, rearranging or replacing the mutation.

13. The use of claim 11, wherein correcting a mutation in a target gene comprises nuclease-mediated non-homologous end joining or homology-directed repair.

14. Use of the gRNA and Cpf1 RNA-guided endonuclease according to claim 1 in preparing a kit for a method of regulating gene expression of a target gene in a cell, the method comprising contacting the cell with an effective amount of the gRNA and Cpf1 RNA-guided endonuclease or expressing an effective amount of the gRNA and Cpf1 RNA-guided endonuclease in the cell, wherein the gRNA targets a target region of the target gene, and wherein the target gene is DNMT1.

15. The use according to claim 14, wherein the Cpf1 RNA-guided endonuclease comprises Cpf1 or a fusion protein comprising a first polypeptide domain comprising nuclease-deficient Cpf1 and a second polypeptide domain having an activity selected from the group consisting of: Transcriptional activation activity, transcriptional repressor activity, nuclease activity, transcriptional release factor activity, histone modification activity, nucleic acid association activity, DNA methylase activity, and direct or indirect DNA demethylase activity.

16. The use according to claim 15, wherein the gene expression of the target gene is regulated when the gene expression level of the target gene is increased or decreased compared to the gene expression level of the target gene in a control. The method according to claim 15 , wherein the fusion protein comprises a transcriptional activator. The use according to claim 15 , wherein the fusion protein comprises the amino acid sequence of SEQ ID NO:

2.

19. The use of claim 15, wherein the fusion protein comprises a transcriptional repressor.

20. The use according to claim 19, wherein the fusion protein comprises the amino acid sequence of SEQ ID NO:

3.

21. The use of claim 15, wherein the fusion protein comprises a site-specific nuclease.

22. The use according to any one of claims 6 to 21, wherein the gRNA is encoded by a polynucleotide sequence and packaged into a recombinant lentiviral vector, a recombinant adenovirus and / or a recombinant adeno-associated virus.

23. The method of any one of claims 6 to 21, wherein contact of the gRNA with the cell is nanoparticle-facilitated.

24. The use according to claim 22, wherein the recombinant lentiviral vector, recombinant adenovirus and / or recombinant adeno-associated virus comprises an expression cassette comprising a promoter operably linked to the polynucleotide sequence encoding the gRNA.

25. The use of claim 24, wherein the promoter is inducible. 26 . The use according to claim 22 , wherein the recombinant lentiviral vector, recombinant adenovirus and / or recombinant adeno-associated virus further comprises a polynucleotide sequence encoding the Cpf1 or fusion protein.

27. The use according to any one of claims 6 to 21, wherein the target gene is a disease-related gene.

28. The use of any one of claims 6-21, wherein the target cell is a eukaryotic cell.

29. The use of any one of claims 6-21, wherein the target cell is a mammalian cell.

Citation Information

Patent Citations

  • Synthetic muscle promoters with activities exceeding naturally occurring regulatory sequences in cardiac cells

    US20040175727A1

  • Genetic immunization

    US5593972A

  • Compositions and methods for delivery of genetic material

    US5962428A

  • Compositions and methods for delivery of genetic material

    WO1994016737A1

  • Methods and compositions for cloning into large vectors

    WO2016109255A1