In silico prediction tool for meganuclease off-target sites

The development of an in silico prediction tool for meganucleases, using a position weighted matrix refined through iterations, addresses the lack of tools for predicting off-target sites, achieving high accuracy in identifying potential off-target sites for safe genome editing.

WO2025122552A1PCT designated stage expired Publication Date: 2025-06-12THE TRUSTEES OF THE UNIV OF PENNSYLVANIA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058367
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-04
Filing Date
2024-12-04
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

There are no available in silico prediction tools for the M2PCSK9 meganuclease or other meganucleases, which hinders the prediction of off-target sites essential for safe and effective genome editing.

Method used

A method is developed to predict off-target sites for meganucleases by generating a position weighted matrix (PWM) from experimentally identified off-target sites, refining it through iterations, and using it to identify potential off-target sites in the genome.

Benefits of technology

The method effectively predicts off-target sites with high accuracy, as demonstrated by the prediction of 25 out of 18,729 unique off-target sites identified in multiple studies, with only two experimentally identified sites not predicted by the tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058367_12062025_PF_FP_ABST
    Figure US2024058367_12062025_PF_FP_ABST
Patent Text Reader

Abstract

An in silico method of providing a predicted off-target site of a nuclease having an on-target site in a genome is provided. The method includes generating a consensus sequence, as described herein. In certain embodiments, the nuclease is a meganuclease that targets PCSK9. In certain embodiments, the nuclease is ARCUS.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] IN SILICO PREDICTION TOOL FOR MEGANUCLEASE OFF -TARGET SITES

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to US Provisional Patent Application 63 / 605,657, filed December 4, 2023, which application is incorporated herein by reference in its entirety.

[0004] INCORPORATION BY REFERENCE OF SEQUENCE LISTING

[0005] The sequence listing file, entitled “24-10536PCT_Seq-Listing”, created December 3, 2024, being 11,348 bytes is incorporated by reference herein in its entirety.

[0006] BACKGROUND OF THE INVENTION

[0007] Whereas CRISPR-based genome editing approaches include a guide RNA (gRNA), allowing for prediction of off-target sites based on target site accessibility and the affinity of the gRNA to the target sequence (including both overall nucleotide usage, position-specific nucleotides, and specific mismatch-containing motifs), genome editing by the M2PCSK9 meganuclease is based on a protein-DNA interaction. Substantial protein engineering, with a limited ability to model this interaction, is required to change the DNA sequences that this meganuclease recognizes, with some of this previously performed in the optimization of the M1PCSK9 meganuclease which preceded the version utilized in ECUR-506 (Wang et al., 2018). While there are a variety of available in silico prediction tools for CRISPR nucleases, there are no available in silico prediction tools for the M2PCSK9 meganuclease or other meganucleases.

[0008] What are needed are improved compositions and methods for prediction of off-target sites for the M2PCSK9 meganuclease and other meganucleases.

[0009] SUMMARY OF THE INVENTION

[0010] Provided herein is a method of providing a predicted off-target site of a nuclease having an on-target site in a genome. The method includes providing a plurality of experimentally identified off-target sites and identifying the best match to the on-target site in each experimentally identified off-target site, and generating a first position weighted matrix (1-PWM) from the best matches. The method further includes identifying the best match to the fPWM in each experimentally identified off-target site, and generating a preliminary position weighted matrix (pPWM) from the best matches. The method further includes performing at least 5 iterations of (b), wherein each iteration comprises finding the best match in each of the experimentally identified off-target sites to the pPWM generated in the iteration prior. The method further includes identifying a final PWM when the pPWM is identical in two or more iterations; and locating the final PWM or a sequence sharing at least 85% identity with the final PWM in the genome, thereby identifying a predicted off-target site in the genome. In certain embodiments, the nuclease is a meganuclease. In certain embodiments, the meganuclease is the ARCUS meganuclease.

[0011] Other aspects and advantages of the invention will be apparent from the following detailed description of the invention.

[0012] BRIEF DESCRIPTION OF THE DRAWINGS

[0013] FIG. 1 is a venn diagram showing how commonly the predicted motifs from the described in silico prediction tool for the M2PCSK9 meganuclease (ARCUS) occur within the rhesus macaque genome in comparison to the experimentally identified off-target sites from previous studies in rhesus macaques, iPSC-derived human hepatocytes, and human hepatocytes in FRG mice.

[0014] FIGS. 2A-2C show the results of the experiment described in Example 1. Molecular analysis on the Day 84 liver biopsy samples from each animal were performed to measure transgene copy numbers per diploid genome (FIG. 2A), mRNA expression levels (FIG. 2B), and off-target editing (FIG. 2C).

[0015] DETAILED DESCRIPTION OF THE INVENTION

[0016] Due to the predictive characteristics of gRNA-target interactions, a variety of in silico tools currently exist for prediction of potential CRISPR nuclease off-target sites, while no such tools exist for meganucleases, including M2PCSK9. As such, motif analysis was performed utilizing all previously generated off-target data sets with the M2PCSK9 meganuclease to establish a mechanism for in silico prediction of potential M2PCSK9 off- target sites.

[0017] The data described herein confirm that off-target analysis in rhesus macaques is predictive of events that may occur in the human genome through comparison of identified off-target sites within liver samples from a GLP-compliant toxicology study in infant rhesus macaques to studies in human cells (iPSC-derived hepatocytes and chimeric liver-humanized FRG mouse model). This was taken even further through the use of a newly developed in silico prediction tool; out of a total 18,729 unique off-target sites identified in all three studies and by our in silico prediction tool, only two sites identified experimentally were not predicted.

[0018] Provided herein are methods which provide in silico prediction of off-target sites for meganucleases, including the M2PCSK9 meganuclease, also known as ARCUS.

[0019] PCSK9

[0020] Proprotein convertase subtilisin kexin 9 (PCSK9) is a serine protease that reduces both hepatic and extrahepatic low-density lipoprotein (LDL) receptor (LDLR; 606945) levels and increases plasma LDL cholesterol. PCSK9 is critical in the regulation of plasma cholesterol homeostasis. PCSK9 binds to the low-density lipid receptor family members low density lipoprotein receptor (LDLR), very low-density lipoprotein receptor (VLDLR), apolipoprotein E receptor (LRP1 / APOER) and apolipoprotein receptor 2 (LRP8 / APOER2), and promotes their degradation in intracellular acidic compartments. Human PCSK9 has a protein sequence of NP_777596.2, as shown in SEQ ID NO: 1, with the coding sequence shown in SEQ ID NO: 2.

[0021] While the PCSK9 gene has been targeted for treatment of cholesterol related diseases, it has been demonstrated that the PSCK9 gene locus is a safe harbor for gene targeting for insertion of other, non-PCSK9 transgenes. However, as noted above, tools are lacking for prediction of off-target effects of the PCSK9-targeting nucleases that target the PCSK9 gene locus.

[0022] As used herein, the “target PCSK9 locus” or “PCSK9 gene locus” is any site in the PCSK9 coding region where insertion of a heterologous transgene is desired. In certain embodiments, the target PCSK9 locus is in Exon 7 of the PCSK9 coding sequence. In certain embodiments, the nuclease is a meganuclease that targets PCSK9. Meganucleases are endodeoxyribonucleases characterized by a large recognition site (double-stranded DNA sequences of 12 to 40 base pairs), for example, I-Scel. When combined with a nuclease, DNA can be cut at a specific location. The restriction enzymes can be introduced into cells, for use in gene editing or for genome editing in situ. In certain embodiments, the nuclease is a member of the LAGLID ADG (SEQ ID NO: 3) family of homing endonucleases. In certain embodiments, the nuclease is a member of the LCrel family of homing endonucleases which recognizes and cuts a 22 base pair recognition sequence. See, e.g., WO 2009 / 059195. Methods for rationally-designing mono-LAGLIDADG (SEQ ID NO: 3) homing endonucleases were described which are capable of comprehensively redesigning I-Crel and other homing endonucleases to target widely-divergent DNA sites, including sites in mammalian, yeast, plant, bacterial, and viral genomes (WO 2007 / 047859). The term “homing endonuclease” is synonymous with the term “meganuclease.” See, WO 2018 / 195449, describing certain PCSK9 meganucleases, which is incorporated herein in its entirety.

[0023] In certain embodiments, the PCSK9 meganuclease is that described in WO 2022 / 232232, sometimes referred to herein as M2PCSK9 or ARCUS. The sequence of the ARCUS meganuclease is reproduced in SEQ ID Nos: 6 and 7. In certain embodiments, the PCSK9 meganuclease is used in conjunction with a heterologous transgene that does not encode PCSK9 to provide a therapeutic composition for treatment of any number of conditions. One such condition is human ornithine transcarbamylase (OTC) deficiency. Compositions and methods for treating OTC deficiency are described in WO 2023 / 140971, which is incorporated herein by reference.

[0024] In Silico Methods

[0025] In one aspect, an in silico method of predicting off-target sites for a nuclease, is provided. The method includes generating a consensus sequence using a position weighted matrix, as further described herein.

[0026] 90 previously generated ITR-seq libraries were collated from rhesus macaque samples following IV administration of an AAV vector expressing M2PCSK9, including those previously published (Wang et al., 2018; Wang et al., 2021; Breton et al., 2020) and those detailed in the non-clinical proof-of-concept (POC) studies. Upon review of these previously sequenced libraries through the updated ITR-seq bioinformatics analysis, the number of off-target sites identified from each library ranged from 0 to 261 (mean=25.3). When these data sets were combined there were 2275 total sites, which consisted of 1492 unique sites (i.e., there were sites identified in more than one sample evaluated). Each identified site is an “experimentally identified off-target site”. The number of libraries from which each unique site was identified ranged from 1 to 37 (mean=1.5). With this compiled data set, we trained a M2PCSK9 motif using the experimentally identified off-target sites. We searched for the best matches to the 22 bases at the on-target site of the M2PCSK9 meganuclease (SEQ ID NO: 4 - TCCCCTGGGGCAAAGAGGTCCA) within the local sequences of the 1492 unique experimentally identified off-target sites identified in previous data sets. The on-target site is specific for each meganuclease, and can be used as the ARCUS on-target site exemplified herein.

[0027] The best match at each site is the 22-base subsequence having the highest matching score to the on-target site of the M2PCSK9 meganuclease. The sequences of these matches were used to generate a first position weighted matrix (1-PWM) for the M2PCSK9 motif, as performed previously (Wasserman, W. W., & Sandelin, A. (2004). Applied bioinformatics for the identification of regulatory elements. Nature Reviews Genetics, 5(4), 276-287, which is incorporated herein by reference).

[0028] As used herein, a “position weighted matrix” or “PWM” refers to a scoring matrix composed of the log likelihood of each nucleotide in a target sequence. Based on an alignment of all known sites, the total number of observations of each nucleotide is recorded for each position, producing a position frequency matrix (PFM). A normalized PFM, in which each column adds up to a total of one, is a table of probabilities for observing each nucleotide at each position. For efficient computational analysis, the PFM must be converted to a log-scale, referred to as the PWM. Various iterations of position weighted matrices are utilized (e.g., first position weighted matrix, preliminary position weighted matrix), prior to arriving at a final PWM that is used to generate a consensus sequence which is used to predict off-target sites in a particular genome.

[0029] Multiple iterations are performed to refine the first PWM. In the example provided herein, 100 iterations were completed to refine the first PWM. During each iteration, local sequences at the 1492 off-target sites are searched to find the best match to the PWM, while the best match at each site is the 22-base subsequence having the highest matching score to the PWM (Wasserman, W. W., & Sandelin, A. (2004)). The best matches are used to generate a preliminary PWM (pPWM) for the next iteration. The pPWM stabilized after ~10 iterations, generating the final PWM. Any number of iterations can be performed, where the desired outcome is stabilization of the pPWM, i.e., an identical pPWM in two or more consecutive iterations. In certain embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,

[0030] 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,

[0031] 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65,

[0032] 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more iterations are performed. The final pPWM is identified as the stabilized pPWM.

[0033] The matching scores of all sites to the PWM had a bimodal distribution with about one third of all sites having scores higher than 0.8 in one peak. We performed another 100 iterations using only high score matches to generate a PWM at each round. The final version of the PWM was used for all follow-up analyses. The consensus sequence was generated from the pPWM using the nucleotide with the highest score at each position. The consensus sequence was TGCCCTGGGGAAAAAAGGGCCA (SEQ ID NO: 5).

[0034] As used herein, “on-target site” refers to the recognition sequence for which a given nuclease has specificity. With respect to meganucleases, which have recognition sites of 12 to 40 base pairs, the recognition site generally occurs only once in a given genome. For example, as described herein, in certain embodiments, the meganuclease is the ARCUS meganuclease that targets a region in the PCSK9 locus having the sequence TCCCCTGGGGCAAAGAGGTCCA (SEQ ID NO: 4).

[0035] Due to the length of the recognition sequence and high cleavage specificity meganucleases are believed to have lower rates of off-targets than other nucleases. However, off-target activities are affected by the structure of meganucleases and the methods of delivery. As used herein, “experimentally identified off-target site” refers to a sequence of a site in the genome that has been observed to have been cleaved by the meganuclease under experimental conditions, as determined using various techniques, such as ITR-seq.

[0036] As used herein, “predicted off-target site” refers to a sequence of a potential site in the genome that may be cleaved by a nuclease, e.g., a meganuclease. In certain embodiments, the predicted off-target site has the sequence of the consensus sequence as described herein. In certain embodiments, the predicted off-target site has a sequence at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% identical to the consensus sequence. In certain embodiments, the predicted off-target site has a sequence at least 85% identical to SEQ ID NO: 5.

[0037] As noted above, in certain embodiments, the meganuclease is that described in WO 2022 / 232232, sometimes referred to herein as M2PCSK9 or ARCUS. In other embodiments, the meganuclease targets the human beta-2 microglobulin gene, such as those described in WO2017112859A1. In other embodiments, the meganuclease targets the human T cell receptor alpha constant region gene, such as those described in WO2017062439A1. In other embodiments, the meganuclease targets the human dystrophin gene, such as those described in WO 2011 / 141820A1. In other embodiments, the meganuclease targets the human C-C chemokine receptor type 5 gene (CCR5), such as those described in WO 2014191525A1. In other embodiments, the meganuclease targets recognition sequences present in a mutant Rhodopsin P23H allele, such as those described in US 10,758,595.

[0038] In certain embodiments, the nuclease is utilized with a donor vector, that provides the coding sequence for a therapeutic transgene. In certain embodiments, the donor vector contains an expression cassette comprising a nucleic acid sequence encoding a transgene, and regulatory sequences that direct expression of the transgene in the target cell. In certain embodiments, the transgene encodes a protein that is aberrantly expressed in a liver metabolic disorder or other genetic disorder. The transgene encodes a protein other than PCSK9. Such proteins include, but are not limited to OTC, low density lipoprotein receptor (LDLr), Factor IX, and Factor VIII.

[0039] Further illustrative genes which may be delivered via the donor vector include, without limitation, glucose-6-phosphatase, associated with glycogen storage disease or deficiency type 1A (GSD1), phosphoenolpyruvate-carboxykinase (PEPCK), associated with PEPCK deficiency; cyclin-dependent kinase-like 5 (CDKL5), also known as serine / threonine kinase 9 (STK9) associated with seizures and severe neurodevelopmental impairment; galactose- 1 phosphate uridyl transferase, associated with galactosemia; phenylalanine hydroxylase (PAH), associated with phenylketonuria (PKU); gene products associated with Primary Hyperoxaluria Type 1 including Hydroxyacid Oxidase 1 (GO / HAO1), AGXT, branched chain alpha-ketoacid dehydrogenase, including BCKDH, BCKDH-E2, BAKDH- Ela, and BAKDH-Elb, associated with Maple syrup urine disease; fumarylacetoacetate hydrolase, associated with tyrosinemia type 1 ; methylmalonyl-CoA mutase, associated with methylmalonic acidemia; medium chain acyl CoA dehydrogenase, associated with medium chain acetyl CoA deficiency; ornithine transcarbamylase (OTC), associated with ornithine transcarbamylase deficiency; argininosuccinic acid synthetase (ASS1), associated with citrullinemia; lecithin-cholesterol acyltransferase (LCAT) deficiency; amethylmalonic acidemia (MMA); NPC1 associated with Niemann-Pick disease, type Cl); propionic academia (PA); low density lipoprotein receptor (LDLR) protein, associated with familial hypercholesterolemia (FH), LDLR variant, such as those described in WO 2015 / 164778; ApoE and ApoC proteins, associated with dementia; lipoprotein lipase (LPL) (Lipoprotein Lipase Deficiency), UDP-glucouronosyltransferase, associated with Crigler-Najjar disease; adenosine deaminase, associated with severe combined immunodeficiency disease; hypoxanthine guanine phosphoribosyl transferase, associated with Gout and Lesch-Nyan syndrome; biotimidase, associated with biotimidase deficiency; alpha-galactosidase A (a-Gal A) associated with Fabry disease; beta-galactosidase (GLB1) associated with GM1 gangliosidosis; ATP7B associated with Wilson’s Disease; beta-glucocerebrosidase, associated with Gaucher disease type 2 and 3; peroxisome membrane protein 70 kDa, associated with Zellweger syndrome; arylsulfatase A (ARSA) associated with metachromatic leukodystrophy, galactocerebrosidase (GALC) enzyme associated with Krabbe disease, alpha-glucosidase (GAA) associated with Pompe disease; sphingomyelinase (SMPD1) gene associated with Nieman Pick disease type A; camosinase (CN1); hypoxanthine-guanine phosphoribosyltransferase (HGPRT); erythropoietin (EPO); Carbamyl Phosphate Synthetase (CPS1), N-Acetylglutamate Synthetase (NAGS); Argininosuccinate Lyase (ASL) (Argininosuccinic Aciduria); and Arginase (AG); argininosuccsinate synthase associated with adult onset type II citrullinemia (CTLN2) (WO 2018 / 144709, which is incorporated herein by reference); carbamoyl-phosphate synthase 1 (CPS1) associated with urea cycle disorders; survival motor neuron (SMN) protein, associated with spinal muscular atrophy; ceramidase associated with Farber lipogramilomatosis; b-hexosaminidase associated with GM2 gangliosidosis and Tay-Sachs and Sandhoff diseases; aspartylglucosaminidase associated with aspartyl-glucosaminuria; a-fucosidase associated with fucosidosis; a- mannosidase associated with alpha-mannosidosis; porphobilinogen deaminase, associated with acute intermittent porphyria (AIP); alpha- 1 antitrypsin for treatment of alpha- 1 antitrypsin deficiency (emphysema); erythropoietin for treatment of anemia due to thalassemia or to renal failure; vascular endothelial growth factor, angiopoietin-1, and fibroblast growth factor for the treatment of ischemic diseases; thrombomodulin and tissue factor pathway inhibitor for the treatment of occluded blood vessels as seen in, for example, atherosclerosis, thrombosis, or embolisms; aromatic amino acid decarboxylase (AADC), and tyrosine hydroxylase (TH) for the treatment of Parkinson's disease; the beta adrenergic receptor, anti-sense to, or a mutant form of, phospholamban, the sarco(endo)plasmic reticulum adenosine triphosphatase-2 (SERCA2), and the cardiac adenylyl cyclase for the treatment of congestive heart failure; a tumor suppressor gene such as p53 for the treatment of various cancers; a cytokine such as one of the various interleukins for the treatment of inflammatory and immune disorders and cancers; dystrophin or mini dystrophin and utrophin or miniutrophin for the treatment of muscular dystrophies; and, insulin or GLP-1 for the treatment of diabetes.

[0040] Examples of suitable transgenes for delivery include, e.g., those associated with familial hypercholesterolemia (e.g., VLDLr, LDLr, ApoE, see, e.g., WO 2020 / 132155, WO 2018 / 152485, WO 2017 / 100682, which are incorporated herein by reference), muscular dystrophy, cystic fibrosis, and rare or orphan diseases. Examples of such rare disease may include spinal muscular atrophy (SMA), Huntingdon’s Disease, Rett Syndrome (e.g., methyl-CpG-binding protein 2 (MeCP2); UniProtKB - P51608), Amyotrophic Lateral Sclerosis (ALS), Duchenne Type Muscular dystrophy, Friedrichs Ataxia (e.g., frataxin), progranulin (PRGN) (associated with non- Alzheimer’s cerebral degenerations, including, frontotemporal dementia (FTD), progressive non-fluent aphasia (PNFA) and semantic dementia), among others. Other useful gene products include, carbamoyl synthetase I, ornithine transcarbamylase (OTC), arginosuccinate synthetase, arginosuccinate lyase (ASL) for treatment of arginosuccinate lyase deficiency, arginase, fumaryl acetate hydrolase, phenylalanine hydroxylase, alpha- 1 antitrypsin, rhesus alpha- fetoprotein (AFP), rhesus chorionic gonadotrophin (CG), glucose-6-phosphatase, plasma protease Cl inhibitor (SERPING1) associated with hereditary angioedema, porphobilinogen deaminase, cystathione beta-synthase associated with homocystinuria, branched chain ketoacid decarboxylase, albumin, isovaleryl-coA dehydrogenase, propionyl CoA carboxylase, methyl malonyl CoA mutase, glutaryl CoA dehydrogenase, insulin, beta-glucosidase, pyruvate carboxylate, hepatic phosphorylase, phosphorylase kinase, glycine decarboxylase, H-protein, T -protein, a cystic fibrosis transmembrane regulator (CFTR) sequence, and a dystrophin gene product [e.g., a mini- or micro-dystrophin]. Still other useful gene products include enzymes such as may be useful in enzyme replacement therapy, which is useful in a variety of conditions resulting from deficient activity of enzyme. For example, enzymes that contain mannose-6-phosphate may be utilized in therapies for lysosomal storage diseases (e.g., a suitable gene includes that encoding P-glucuronidase (GUSB)). Examples of suitable transgene for delivery may include human frataxin delivered in an AAV vector as described, e.g., PCT / US20 / 66167, December 18, 2020, US Provisional Patent Application No. 62 / 950,834, filed December 19, 2019, and US Provisional Application No. 63 / 136,059 filed on January 11, 2021 which are incorporated herein by reference. Another example of suitable transgene for delivery may include human acid-a-glucosidase (GAA) delivered in an AAV vector as described, e.g., PCT / US20 / 30493, April 30, 2020, now published as WO2020 / 223362A1, PCT7US20 / 30484, April 20, 2020, now published as WO 2020 / 223356 Al, US Provisional Patent Application No. 62 / 840,911, filed April 30, 2019, US Provisional Application No. 62.913,401, filed October 10, 2019, US Provisional Patent Application No. 63 / 024,941, filed May 14, 2020, and US Provisional Patent Application No. 63 / 109,677, filed November 4, 2020 which are incorporated herein by reference. Also, another example of suitable transgene for delivery may include human a-L-iduronidase (IDUA) delivered in an AAV vector as described, e.g., PCT / US2014 / 025509, March 13, 2014, now published as WO 2014 / 151341, and US Provisional Patent Application No. 61 / 788,724, filed March 15, 2013 which are incorporated herein by reference.

[0041] Other useful therapeutic products include those expressed in muscle, including heart muscle. Other useful therapeutic products encoded by the transgene include hormones and growth and differentiation factors including, without limitation, insulin, glucagon, glucagon- like peptide 1 (GLP-1), growth hormone (GH), parathyroid hormone (PTH), growth hormone releasing factor (GRF), follicle stimulating hormone (FSH), luteinizing hormone (LH), human chorionic gonadotropin (hCG), vascular endothelial growth factor (VEGF), angiopoietins, angiostatin, granulocyte colony stimulating factor (GCSF), erythropoietin (EPO), connective tissue growth factor (CTGF), basic fibroblast growth factor (bFGF), acidic fibroblast growth factor (aFGF), epidermal growth factor (EGF), transforming growth factor a (TGFa), platelet-derived growth factor (PDGF), insulin growth factors I and II (IGF- I and IGF-II), any one of the transforming growth factor 0 superfamily, including TGF 0, activins, inhibins, or any of the bone morphogenic proteins (BMP) BMPs 1-15, any one of the heregluin / neuregulin / ARIA / neu differentiation factor (NDF) family of growth factors, nerve growth factor (NGF), brain-derived neurotrophic factor (BDNF), neurotrophins NT-3 and NT-4 / 5, ciliary neurotrophic factor (CNTF), glial cell line derived neurotrophic factor (GDNF), neurturin, agrin, any one of the family of semaphorins / collapsins, netrin-1 and netrin-2, hepatocyte growth factor (HGF), ephrins, noggin, sonic hedgehog and tyrosine hydroxylase. Other transgenes useful herein include those for treating mucopolysaccharidosis type I-VII (IDUA, IDS, GNA, HGSNAT, NAGLU, SGSH, GALNS, GLB1, ARSB, GUSB). Exemplary sequences for useful for treating MPSI can be found in WO 2019 / 010335, which is incorporated herein by reference. Exemplary sequences for useful for treating MPSII can be found in WO 2019 / 060662, which is incorporated herein by reference. Exemplary sequences for useful for treating MPSIIIa can be found in WO

[0042] 2019 / 108857, which is incorporated herein by reference. Exemplary sequences for useful for treating MPSIIIb can be found in WO 2019 / 108856 which is incorporated herein by reference.

[0043] As used herein, “a,” “an,” or “the” can mean one or more than one. For example, “a” cell can mean a single cell or a multiplicity of cells.

[0044] As used herein, the term “specificity” means the ability of a nuclease to recognize and cleave double-stranded DNA molecules only at a particular sequence of base pairs referred to as the recognition sequence, or only at a particular set of recognition sequences. The set of recognition sequences will share certain conserved positions or sequence motifs, but may be degenerate at one or more positions. A highly-specific nuclease is capable of cleaving only one or a very few recognition sequences. Specificity can be determined by any method known in the art.

[0045] The term “exogenous” as used to describe a nucleic acid sequence or protein means that the nucleic acid or protein does not naturally occur in the position in which it exists in a chromosome, or host cell. An exogenous nucleic acid sequence also refers to a sequence derived from and inserted into the same expression cassette or host cell, but which is present in a non-natural state, e.g., a different copy number, or under the control of different regulatory elements.

[0046] The term “heterologous” when used with reference to a protein or a nucleic acid indicates that the protein or the nucleic acid comprises two or more sequences or subsequences which are not found in the same relationship to each other in nature. For instance, the nucleic acid is typically recombinantly produced, having two or more sequences from unrelated genes arranged to make a new functional nucleic acid. For example, in one embodiment, the nucleic acid has a promoter from one gene arranged to direct the expression of a coding sequence from a different gene.

[0047] As used herein, the term “host cell” may refer to the packaging cell line in which a vector (e.g., a recombinant AAV) is produced from a production plasmid. In the alternative, the term “host cell” may refer to any target cell in which expression of the transgene is desired. Thus, a “host cell,” refers to a prokaryotic or eukaryotic cell that contains a exogenous or heterologous nucleic acid sequence that has been introduced into the cell by any means, e.g., electroporation, calcium phosphate precipitation, microinjection, transformation, viral infection, transfection, liposome delivery, membrane fusion techniques, high velocity DNA-coated pellets, viral infection and protoplast fusion. In certain embodiments herein, the term “host cell” refers to cultures of cells of various mammalian species for in vitro assessment of the compositions described herein. In other embodiments herein, the term “host cell” refers to the cells employed to generate and package the viral vector or recombinant virus. Still in other embodiment, the term “host cell” is intended to reference the target cells of the subject being treated in vivo for the diseases or conditions as described herein. In certain embodiments, the term “host cell” is a liver cell or hepatocyte.

[0048] A “subject” is a mammal, e.g., a human, mouse, rat, guinea pig, dog, cat, horse, cow, pig, or non-human primate, such as a monkey, chimpanzee, baboon or gorilla. A patient refers to a human. A veterinary subject refers to a non-human mammal. In certain embodiments, the subject does not have a defect in their PCSK9 gene.

[0049] A “replication-defective virus” or “viral vector” refers to a synthetic or artificial viral particle in which an expression cassette containing a gene of interest is packaged in a viral capsid or envelope, where any viral genomic sequences also packaged within the viral capsid or envelope are replication-deficient; i.e., they cannot generate progeny virions but retain the ability to infect target cells. In one embodiment, the genome of the viral vector does not include genes encoding the enzymes required to replicate (the genome can be engineered to be “gutless” - containing only the gene of interest flanked by the signals required for amplification and packaging of the artificial genome), but these genes may be supplied during production. Therefore, it is deemed safe for use in gene therapy since replication and infection by progeny virions cannot occur except in the presence of the viral enzyme required for replication.

[0050] The terms “sequence identity” “percent sequence identity” or “percent identical” in the context of nucleic acid sequences refers to the residues in the two sequences which are the same when aligned for maximum correspondence. The length of sequence identity comparison may be over the full-length of the genome, the full-length of a gene coding sequence, or a fragment of at least about 500 to 5000 nucleotides, is desired. However, identity among smaller fragments, e.g. of at least about nine nucleotides, usually at least about 20 to 24 nucleotides, at least about 28 to 32 nucleotides, at least about 36 or more nucleotides, may also be desired. Similarly, “percent sequence identity” may be readily determined for amino acid sequences, over the full-length of a protein, or a fragment thereof. Suitably, a fragment is at least about 8 amino acids in length and may be up to about 700 amino acids. Examples of suitable fragments are described herein. The term “substantial homology” or “substantial similarity,” when referring to amino acids or fragments thereof, indicates that, when optimally aligned with appropriate amino acid insertions or deletions with another amino acid (or its complementary strand), there is amino acid sequence identity in at least about 95 to 99% of the aligned sequences. Preferably, the homology is over full-length sequence, or a protein thereof, e.g., a cap protein, a rep protein, or a fragment thereof which is at least 8 amino acids, or more desirably, at least 15 amino acids in length. Examples of suitable fragments are described herein.

[0051] By the term “highly conserved” is meant at least 80% identity, preferably at least 90% identity, and more preferably, over 97% identity. Identity is readily determined by one of skill in the art by resort to algorithms and computer programs known by those of skill in the art.

[0052] Generally, when referring to “identity”, “homology”, or “similarity” between two different adeno-associated viruses, “identity”, “homology” or “similarity” is determined in reference to “aligned” sequences. “Aligned” sequences or “alignments” refer to multiple nucleic acid sequences or protein (amino acids) sequences, often containing corrections for missing or additional bases or amino acids as compared to a reference sequence. In the examples, AAV alignments are performed using the published AAV9 sequences as a reference point. Alignments are performed using any of a variety of publicly or commercially available Multiple Sequence Alignment Programs. Examples of such programs include, “Clustal Omega”, “Clustal W”, “CAP Sequence Assembly”, “MAP”, and “MEME”, which are accessible through Web Servers on the internet. Other sources for such programs are known to those of skill in the art. Alternatively, Vector NTI utilities are also used. There are also a number of algorithms known in the art that can be used to measure nucleotide sequence identity, including those contained in the programs described above. As another example, polynucleotide sequences can be compared using Fasta™, a program in GCG Version 6.1. Fasta™ provides alignments and percent sequence identity of the regions of the best overlap between the query and search sequences. For instance, percent sequence identity between nucleic acid sequences can be determined using Fasta™ with its default parameters (a word size of 6 and the NOPAM factor for the scoring matrix) as provided in GCG Version 6.1, herein incorporated by reference. Multiple sequence alignment programs are also available for amino acid sequences, e.g., the “Clustal Omega”, “Clustal X”, “MAP”,

[0053] “PIMA”, “MSA”, “BLOCKMAKER”, “MEME”, and “Match-Box” programs. Generally, any of these programs are used at default settings, although one of skill in the art can alter these settings as needed. Alternatively, one of skill in the art can utilize another algorithm or computer program which provides at least the level of identity or alignment as that provided by the referenced algorithms and programs. See, e.g., J. D. Thomson et al, Nucl. Acids. Res., “A comprehensive comparison of multiple sequence alignments”, 27(13):2682-2690 (1999).

[0054] As used herein, the term “about” refers to a variant of ±10% from the reference integer and values therebetween. For example, “about” 40 base pairs, includes ±4 (i.e., 36 - 44, which includes the integers 36, 37, 38, 39, 40, 41, 42, 43, 44). For other values, particularly when reference is to a percentage (e.g., 90% identity, about 10% variance, or about 36% mismatches), the term “about” is inclusive of all values within the range including both the integer and fractions.

[0055] Throughout this disclosure, various aspects of the invention can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the invention. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range such as from 1 to 6 should be considered to have specifically disclosed subranges such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6 etc., as well as individual numbers within that range, for example, 1, 2, 2.7, 3, 4, 5, 5.3, and 6. This applies regardless of the breadth of the range.

[0056] As used throughout this specification and the claims, the terms “comprising”, “containing”, “including”, and its variants are inclusive of other components, elements, integers, steps and the like. Conversely, the term “consisting” and its variants are exclusive of other components, elements, integers, steps and the like.

[0057] Unless defined otherwise in this specification, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art and by reference to published texts, which provide one skilled in the art with a general guide to many of the terms used in the present application.

[0058] EXAMPLES

[0059] Ornithine transcarbamylase (OTC) deficiency is an X-linked urea cycle disorder associated with high mortality. Although a promising treatment for late-onset OTC deficiency, adeno-associated virus (AAV) neonatal gene therapy would only provide shortterm therapeutic effects as the non-integrated genome gets lost during hepatocyte proliferation. Nuclease-mediated, site-specific integration of an OTC mini gene cassette in a safe harbor in the genome would provide long-term therapeutic benefits to patients with OTC deficiency.

[0060] EXAMPLE 1 - ARCUS2-MEDIATED HOTC GENE TARGETING IN NEWBORN NHPS

[0061] Newborn (1-16 days old) or infant (3-4 months old) rhesus macaques were used in non-GLP-compliant POC pharmacology studies. The M2PCSK9 meganuclease targets a 22- bp sequence present in the human and rhesus macaque PCSK9 gene. Thus, rhesus macaques can be used to evaluate on-target editing (pharmacology) and safety / toxicology. Furthermore, newborn and infant rhesus macaques have similar anatomical and physiological features as human infants and will allow for the use of the intended clinical ROA (IV). It is anticipated that the similarity in anatomy and ROA will result in representative vector distribution and transduction profiles, which will enable more accurate assessment of the pharmacology and toxicity of the test article, including on- and off-target editing, and clinical pathology, which is not possible in newborn mice.

[0062] In this study newborn NHPs were administered ARCUS2 nuclease vectors, and donor vectors having HDR arms of varying length - 500bp arm or short HDR arm. All 14 newborn macaques tolerated vector infusions well (i.e., no apparent clinical sequelae) and gained weight over time. Liver enzyme levels were within the normal range except for transient and modest elevation of ALT levels in some animals on Day 14.

[0063] Analysis on the Day 0 plasma samples collected from the newborns prior to dosing showed 3 animals (21-111, 21-113, 21-122) had high levels (> 400) of binding antibodies to AAVrh79. These pre-existing anti-AAVrh79 antibodies would block AAV gene transfer.

[0064] PCSK9 levels were followed in all newborn animals including the donor-only control animals over time. PCSK9 levels on Day 0 varied between the newborns. Nine animals showed a trend of reduced PCSK9 levels post vector administration including one donor-only control animal, while the while the remaining five animals showed persistent or transient elevation of PCSK9 levels post dosing.

[0065] On Day 84, a liver biopsy via laparotomy was performed. Transduction efficiencies of hOTC in liver were evaluated by dual ISH with hOTC- and M2PCSK9-specific probes to detect transgene mRNA, and by OTC immunofluorescence to detect human OTC protein, followed by quantification on scanned slides. The three animals (21-111, 21-113, and 21- 122) with pre-existing anti-AAVrh79 binding antibodies at the time of dosing did show any OTC-positive hepatocytes by both methods. The two donor-only control animals showed low level (< 1%) of hOTC transduction. The highest transduction efficiencies (11.9 and 18.6% by OTC immunofluorescence) were detected in the two animals that received a co-admimstration of AAVrh79.TBG.PI.M2PCSK9.WPRE.bGH and AAVrh79.rhHDR.TBG.hOTCco.bGH donor vectors (G6). Positive hOTC-expressing hepatocytes were also found to be present in clusters. These levels are above the threshold for substantially benefitting patients, which is ~5% OTC-expressing cells.

[0066] Molecular analysis on the Day 84 liver biopsy samples from each animal were performed to measure transgene copy numbers per diploid genome, mRNA expression levels, on-target editing, and off-target editing (FIG. 2A-2C). Consistent with the transduction efficiency analyses, the two animals (21-157 and 21-175) in Group 6 had the highest hOTC vector GC (FIG. 2A), hOTC mRNA (FIG. 2B), and on-target indel% (FIG. 2C). The M2PCSK9 vector GC in animals were 2-fold to 7-fold lower than the hOTC vector GC, while M2PCSK9 mRNA levels were 23-fold and 765-fold lower than the hOTC mRNA levels (FIG. 2A and 2B).

[0067] Off-target activity evaluated by ITR-seq identified 2 to 40 potential off-targets in the Day 84 liver biopsy samples in this study. Off-target editing will be further characterized by amplicon-seq on the potential off-target sites.

[0068] EXAMPLE 2 - IN SILICO PREDICTION OF OFF-TARGET SITES

[0069] As part of the non-clinical assessment of off-target sites by the M2PCSK9 meganuclease, we developed an in silico tool to predict potential off-target sites. Whereas CRISPR-based genome editing approaches include a guide RNA (gRNA), allowing for prediction of off-target sites based on target site accessibility and the affinity of the gRNA to the target sequence (including both overall nucleotide usage, position-specific nucleotides, and specific mismatch-containing motifs), genome editing by the M2PCSK9 meganuclease is based on a protein-DNA interaction. Substantial protein engineering, with a limited ability to model this interaction, is required to change the DNA sequences that this meganuclease recognizes, with some of this previously performed in the optimization of the M1PCSK9 meganuclease which preceded the version utilized in ECUR-506 (Wang et al., 2018).

[0070] Due to the predictive characteristics of gRNA-target interactions, a variety of in silico tools currently exist for prediction of potential CRISPR nuclease off-target sites, while no such tools exist for meganucleases, including M2PCSK9. As such, motif analysis was performed utilizing all previously generated off-target data sets with the M2PCSK9 meganuclease to establish a mechanism for in silico prediction of potential M2PCSK9 off- target sites.

[0071] 90 previously generated ITR-seq libraries were collated from rhesus macaque samples following IV administration of an AAV vector expressing M2PCSK9, including those previously published (Wang et al., 2018; Wang et al., 2021; Breton et al., 2020) and those detailed in the non-clinical proof-of-concept (POC) studies. Upon review of these previously sequenced libraries through the updated ITR-seq bioinformatics analysis, the number of off-target sites identified from each library ranged from 0 to 261 (mean=25.3). When these data sets were combined there were 2275 total sites, which consisted of 1492 unique sites (i.e., there were sites identified in more than one sample evaluated). The number of libraries from which each unique site was identified ranged from 1 to 37 (mean=1.5). With this compiled data set, we were then able to train a M2PCSK9 motif using these numerous identified sites by first searching for the best matches to the 22 bases at the on- target site of the M2PCSK9 meganuclease (TCCCCTGGGGCAAAGAGGTCCA - SEQ ID NO: 4) within the local sequences of the 1492 unique sites identified in our previous data sets. The sequences of these matches were used to generate a position weighted matrix (PWM) for the M2PCSK9 motif. We then ran 100 iterations to refine this PWM. During each iteration, local sequences at the 1492 off-target sites were searched to find the best match to the PWM, while the best match at each site is the 22-base subsequence having the highest matching score to the PWM (Wasserman and Sandelin, 2004). The best matches were used to generate another PWM for the next iteration. The PWM stabilized after ~10 iterations.

[0072] The matching scores of all sites to the PWM had a bimodal distribution with about one third of all sites having scores higher than 0.8 in one peak. We performed another 100 iterations using only high score matches to generate a PWM at each round. The final version of the PWM was used for all follow-up analyses. The consensus sequence generated was TGCCCTGGGGAAAAAAGGGCCA (SEQ ID NO: 5). 585 out of the 1492 off-target sites had high matching scores to the final PWM. In general, these sites were identified in more ITR-seq libraries than sites with lower matching scores (mean=2.06 vs. 1.18).

[0073] To predict more off-target sites for the M2PCSK9 meganuclease, we used MEME (version 5.5.0) to search for matches to the PWM in both the rhesus (Mmul lO) and human (GRCh38) reference genomes. 14,274 and 17,393 sites were identified from Mmul_10 and GRCh38, respectively, using a MEME false discovery rate (FDR) of 0.1 as cutoff. The minimum matching score to the PWM is 0.86. According to the alignment of 100-base local sequences at the rhesus sites to the human reference genome, 13,219 of the sites are common to rhesus and human.

[0074] This novel in silico prediction tool was then used to determine whether there were overlapping sites between different data sets, including those identified in the liver samples obtained in the GLP-compliant toxicology study in infant NHPs administered with ECUR- 506 (Non-clinical Study 2020-007) and those identified in the human genome in human hepatocytes derived from iPSCs and a chimeric liver-humanized mouse model. While the prediction of off-target sites in the human genome could be validated based on a number of overlapping off-target sites identified, the inherent differences between proteimDNA and gRNA:DNA interactions mean additional caveats to in silico prediction for meganucleases need to be considered when interpreting these data compared to similar predictive methods for CRISPR nucleases.

[0075] In addition to the comparison between off-target sites identified across the GLP- compliant toxicology study in rhesus macaques with ECUR-506 and those identified with the M2PCSK9 meganuclease in human cells (iPSC-derived human hepatocytes and FRG mice), we also added in off-target sites predicted to occur based on our newly developed inhouse prediction tool.

[0076] Whereas CRISPR-based genome editing approaches include a gRNA allowing for prediction of off-target sites based on target site accessibility due to genomic DNA structure and epigenetic modifications, and the affinity of the gRNA to the target sequence (including both overall nucleotide usage, position-specific nucleotides, and motifs with some nucleases allowing for up to 3 bp mismatches), genome editing by the M2PCSK9 meganuclease is based on a protein-DNA interaction.

[0077] As there are no available in silico prediction tools for the M2PCSK9 or other meganucleases, compared to the variety available for CRISPR nucleases, motif analysis was performed utilizing all previously generated data sets with the M2PCSK9 meganuclease. 90 previously generated ITR-seq libraries were collated from rhesus macaque samples collected by either biopsy or at necropsy following IV administration of an AAV vector expressing M2PCSK9, including those previously published (Wang et al., 2018; Wang et al., 2021; Breton et al., 2020) and those detailed in the non-clinical POC studies above.

[0078] Evaluation of off-target sites looked for motifs of 22 bp or less that commonly occur in identified off-targets (the M2PCSK9 meganuclease has a 22-nucleotide target site within the PCSK9 gene). The resulting 17,165 predicted off-target sites were then compared to identified potential M2PCSK9 meganuclease off-target sites compiled from the totality of unique off-target sites identified across three different studies using different methods, including sites identified in the rhesus genome in previous studies with the M2PCSK9 meganuclease, identified off-target sites in the human genome in iPSC-derived hepatocytes, and sites identified in the chimeric liver-humanized FRG mouse model, as further described below:

[0079] 90 previously generated ITR-seq libraries were collated from rhesus macaque samples collected by either biopsy or at necropsy following IV administration of an AAV vector expressing M2PCSK9, including the off-target sites identified in the GLP-compliant toxicology study in infant rhesus macaques (Non-clinical Study 2020-007). There was a total of 982 unique off-target sites identified in the rhesus macaque reference genome.

[0080] 160 unique off-target sites were predicted and identified in iPSC-derived human hepatocytes using GUIDE-seq. These off-target sites were then mapped to their homologous loci in the rhesus reference genome to allow for direct comparison to the off-target sites identified in the human cells to all identified off-target sites in the rhesus reference genome.

[0081] Following reanalysis of the ITR-seq data obtained from the FRG mouse studies, there were 1093 unique off-target sites identified across the 10 FRG mouse samples. Again, the data on off-target genome editing obtained in human hepatocytes in FRG mice should be considered a worst-case scenario as this is an immune-deficient model with persistent expression of the M2PCSK9 meganuclease that may have led to a high-level of off-target editing in human hepatocytes. However, as we sought to include all off-target sites in this subsequent analysis, no sites were excluded here. Again, these off-target sites were then mapped to their homologous loci in the rhesus reference genome to allow for direct comparison between the off-target sites identified in the human cells and all identified off- target sites in the rhesus reference genome.

[0082] As noted above, for this analysis we mapped all identified sites in human cells to their homologous loci in the rhesus reference genome. This ensured that no off-target sites were excluded from this analysis due to the lack of a human homologue. We then evaluated how commonly these predicted motifs from our in silico prediction tool for the M2PCSK9 meganuclease occur within the rhesus macaque genome in comparison to the experimentally identified off-target sites from previous studies in rhesus macaques, iPSC-derived human hepatocytes, and human hepatocytes in FRGmice (Figure 1).

[0083] As shown in Figure 1, 25 out of these total 18,729 unique off-target sites were identified in all three studies and by our in silico prediction tool, and only two sites identified experimentally were not predicted. While other identified off-target sites could be of high importance, we focused additional evaluation on these 25 off-target sites due to the frequency of their occurrence.

[0084] The location of the 25 common off-target sites identified in the human genome were then annotated, including the gene and genetic region closest to the off-target site. Table 1 details the gene name, gene description, closest gene region, and distance to the transcriptional start site for that gene as a surrogate for whether regulatory elements could be impacted by vector integration at this off-target site. As seen previously with the M2PCSK9 meganuclease, most of the off-target sites were determined to be intragenic and those sites found to be intergenic mostly did not affect the coding sequence (CDS, Table 1).

[0085] Of the 17 overlapping off-target sites in the human reference genome resulting from the comparison of the frequency of their occurrence across studies in rhesus macaques and human hepatocytes both in vitro (iPSC-derived) and in vivo (FRG mouse model) and the 25 overlapping off-target sites resulting following inclusion of our in silico prediction tool, 14 of these sites were found in both analyses. These included genes ACSF3, ALG8, ATP11A, CLIP3, COMMD9, COX7A2L, DRD3, HERPUD1, MTA2, NPBWR2, PLCG2, PRKCE, RBM15, and TBC1D22A.

[0086] Table 1 Additional exploratory analysis confirmed that off-target analysis in rhesus macaques is predictive of events that may occur in the human genome through comparison of identified off-target sites within liver samples from the GLP-compliant toxicology study in infant rhesus macaques to studies in human cells (iPSC-derived hepatocytes and chimeric liver-humanized FRG mouse model). This was taken even further through the use of a newly developed in silico prediction tool; out of a total 18,729 unique off-target sites identified in all three studies and by our in silico prediction tool, only two sites identified experimentally were not predicted.

[0087] In summary, identified off-target sites in macaque liver samples are located predominantly in the intragenic regions of the genome and, therefore, would not affect expression of endogenous genes. The majority of potential off-target sites identified by GUIDE-seq on M2PCSK9-transduced human iPSC cells or by ITR-seq on the humanized FRG mice liver samples also reside in intronic, non-coding regions of the human genome.

[0088] Therefore, the data presented here demonstrate sustained and therapeutically beneficial editing in non-human primate liver with an engineered meganuclease and provide a data set on in vivo genome editing efficacy, immune toxicity, and safety. We believe that potential for off-target editing has been thoroughly assessed in NHPs, which provide a relevant model for translation to events that may occur in patients. While we cannot predict whether patient-specific single nucleotide polymorphisms throughout the genome could increase off-target editing at a particular site, we do not see areas of specific concern.

[0089] Ultimately the risks and benefits of in vivo genome editing should be considered in the context of the unmet need of the target patient populations.

[0090] All documents cited in this specification are incorporated herein by reference, as are sequences and text of the Sequence Listing filed herewith are incorporated by reference. While the invention has been described with reference to particular embodiments, it will be appreciated that modifications can be made without departing from the spirit of the invention. Such modifications are intended to fall within the scope of the appended claims.

Claims

WHAT IS CLAIMED IS:

1. A method of providing a predicted off-target site of a nuclease having an on-target site in a genome, the method comprising:(a) providing a plurality of experimentally identified off-target sites and identifying the best match to the on-target site in each experimentally identified off-target site, and generating a first position weighted matrix (1-PWM) from the best matches;(b) identifying the best match to the fPWM in each experimentally identified off- target site, and generating a preliminary position weighted matrix (pPWM) from the best matches;(c) performing at least 5 iterations of (b), wherein each iteration comprises finding the best match in each of the experimentally identified off-target sites to the pPWM generated in the iteration prior;(d) identifying a final PWM when the pPWM is identical in two or more iterations; and(e) generating a consensus sequence from the final PWM;(f) locating the consensus sequence or a sequence sharing at least 85% identity with the consensus sequence in the genome, thereby identifying a predicted off-target site in the genome.

2. The method of claim 1, wherein the nuclease on-target site is TCCCCTGGGGCAAAGAGGTCCA (SEQ ID NO: 4).

3. The method of claim 1, wherein the consensus sequence is TGCCCTGGGGAAAAAAGGGCCA (SEQ ID NO: 5).

4. The method of claim 2, wherein the consensus sequence is TGCCCTGGGGAAAAAAGGGCCA (SEQ ID NO: 5).

5. The method of any one of claims 1 to 4, wherein step (c) comprises at least 10 iterations.

6. The method of any of claims 1 to 5, wherein the nuclease is a meganuclease.

7. The method of claim 6, wherein the meganuclease is ARCUS.

Citation Information

Patent Citations

  • Engineered meganucleases specific for recognition sequences in the PCSK9 gene

    US11680254B2

  • Soy nucleic acid molecules and other molecules associated with transcription plants and uses thereof for plant improvement

    US20040031072A1

  • Methods and systems for identifying crispr / cas off-target sites

    US20190295689A1

  • Stable proteins and methods for designing same

    WO2017017673A2