Compositions for regulating and self-inactivating enzyme expression and methods for modulating off-target activity of enzymes

HK40069051BActive Publication Date: 2026-07-17THE TRUSTEES OF THE UNIV OF PENNSYLVANIA

Patent Information

Authority / Receiving Office
HK · HK
Patent Type
Patents
Current Assignee / Owner
THE TRUSTEES OF THE UNIV OF PENNSYLVANIA
Filing Date
2022-08-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, AAV-mediated nuclease delivery leads to persistent expression in target tissues, triggering immune responses and cytotoxicity, and also exhibits off-target activity, resulting in insertions and deletions in other regions of the genome.

Method used

Design a self-regulating gene-editing nuclease expression cassette containing a nuclease coding sequence and a protein degradation signal. Guide nuclease expression in host cells through the regulatory sequence and reduce off-target activity by fusing the protein degradation signal, for example, using a PEST sequence as the protein degradation signal.

Benefits of technology

It reduces the off-target activity of nucleases, improves the safety of delivery enzymes, reduces immune responses and cytotoxicity during gene editing, and lowers the risk of genome insertions and deletions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides a gene editing nuclease expression cassette comprising a nucleic acid sequence comprising a nuclease-encoding sequence operably linked to a regulatory sequence that directs expression of the nuclease upon delivery to a host cell having a sequence targeted by the nuclease, and at least one nuclease modulating sequence selected from a target sequence of the nuclease or a mutated target sequence recognized by the nuclease upon its expression. A vector comprising the gene editing nuclease expression cassette is provided. Compositions containing the vector and methods of use are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] The use of engineered nucleases to edit malfunctioning genes has been described. AAV-mediated delivery of such nucleases has also been described. However, while AAV-mediated delivery of nucleases avoids the need for repeated re-administration, the resulting nucleases persist in the target tissue after vector transduction, which can induce an immune response and cellular toxicity.

[0002] Furthermore, both in vitro and in vivo studies have shown that nucleases generate indels (insertions and deletions) in other regions of the genome, regardless of the delivery vehicle, suggesting off-target activity.

[0003] There is a need for improved compositions and methods for gene editing. SUMMARY

[0004] A delivery system for enzymes is provided that provides for a reduction in off-target activity, thus improving the safety of the delivered enzyme. The system is particularly well suited for use in gene editing therapies and / or enzymes delivered by systems in which the gene persists in the cell, such as AAV-mediated delivery.

[0005] In one aspect, a self-regulating gene editing nuclease expression cassette is provided. The expression cassette includes (a) a nucleic acid sequence comprising a nuclease-encoding sequence operably linked to a regulatory sequence that directs expression of the nuclease upon delivery to a host cell having a sequence targeted by the nuclease; and (b) at least one nuclease regulatory sequence selected from a target sequence of the nuclease or a mutated target sequence recognized by the nuclease upon its expression. Optionally, the coding sequence encodes a fusion protein comprising at least one peptide degradation signal in-frame with the nuclease-encoding sequence. In certain embodiments, the protein degradation signal is 10 amino acids to 50 amino acids in length. In certain embodiments, the protein degradation sequence is a PEST sequence of about 42 amino acids in length. In certain embodiments, the expression cassette includes more than one protein degradation signal. In certain embodiments, the expression cassette includes a mutated target sequence, which can be located upstream of, downstream of, or within the nuclease-encoding sequence. In embodiments in which the expression cassette has two or more nuclease regulatory sequences, they can be independently positioned and can be the same or different sequences. In certain embodiments, the self-regulating nuclease expression cassette can have more than one protein degradation signal.

[0006] In certain embodiments, a self-inactivating meganuclease expression cassette is provided, comprising a nucleic acid sequence comprising a sequence encoding a meganuclease fused to at least one protein degradation sequence, and a regulatory sequence that directs expression of the meganuclease and protein degradation signal as a fusion protein in a host cell. In certain embodiments, the protein degradation signal is 10 amino acids to 50 amino acids in length. In certain embodiments, the protein degradation sequence is a PEST sequence that is about 42 amino acids in length. In certain embodiments, the expression cassette comprises more than one protein degradation signal.

[0007] In certain embodiments, a recombinant AAV suitable for gene editing comprises an AAV capsid and a vector genome packaged in the AAV capsid, the vector genome comprising an expression cassette as described in the preceding paragraph and AAV inverted terminal repeat (ITR) sequences required for packaging the expression cassette into the capsid.

[0008] A pharmaceutical composition is provided comprising an expression cassette as described herein, and one or more of a carrier, diluent, and / or excipient. In certain embodiments, the expression cassette is in a non-vector based delivery system, such as a lipid nanoparticle (LNP). In certain embodiments, the expression cassette is engineered into a non-viral vector (e.g., a plasmid or other genetic element), the recombinant vector is admixed with a carrier, diluent, and / or excipient to form a pharmaceutical composition. In certain embodiments, the expression cassette is engineered into a viral vector, the recombinant vector is admixed with a carrier, diluent, and / or excipient to form a pharmaceutical composition. In certain embodiments, the expression cassette is engineered into a vector genome, and packaged into an AAV capsid to form a recombinant adeno-associated virus (rAAV), the rAAV is admixed with a carrier, diluent, and / or excipient to form a pharmaceutical composition.

[0009] In certain embodiments, a nuclease expression cassette, non-viral vector (e.g., DNA or RNA), or viral vector (e.g., rAAV or lentivirus) as described herein can be administered for gene editing in a patient. In certain embodiments, the method is suitable for non-embryonic gene editing. In certain embodiments, the patient is an infant (e.g., birth to about 9 months). In certain embodiments, the patient is older than an infant, such as 12 months or older.

[0010] In certain embodiments, a pharmaceutical composition as described herein can be administered for gene editing in a patient. In certain embodiments, the method is suitable for non-embryonic gene editing. In certain embodiments, the patient is an infant (e.g., birth to about 9 months). In certain embodiments, the patient is older than an infant, such as 12 months or older.

[0011] Other aspects and advantages of the invention will become apparent from the following detailed description. Attached Figure Description

[0012] Figure 1 This is a schematic diagram of the AAV "suicide" vector. The AAV vector is constructed by inserting a combination of a large-scale nuclease (M2PCSK9) with a target sequence (black bars marked with *), a mutant target sequence containing 8 mismatches (black bars marked with *mut), and a PEST sequence (black bars with lines).

[0013] Figure 2A Showing the timeline of in vivo experiments. Mouse studies. RAG KO mice were intravenously (IV) injected with AAV expressing human PCSK9 (hPCSK9). Two weeks later, treated mice were injected with a single dose of the AAV suicide vector or the corresponding control (AAV8.M2PCSK9) IV. Non-human primate studies. Rhesus macaques were administered AAV.mutant target.M2PCSK9-PEST or AAV.M2PCSK9. Liver biopsies were collected on day 18 or day 128 post-treatment.

[0014] Figure 2B This shows insertions and deletions in the target region of the purified AAV vector.

[0015] Figure 3A (Left side) and Figure 3B (Right) Shows extensive nuclease-induced insertions and deletions in the hPCSK9 gene and the AAV suicide vector. Mice were treated with the indicated AAV vector and euthanized at the indicated time following AAV9.hPCSK9 administration. The percentage of insertions / deletions for each target is shown as a percentage of total readings.

[0016] Figure 4A (Left side) and Figure 4B (Right) Showing the in vivo off-target activity of the large-scale nuclease suicide system. The number of unique AAV integration sites in the host genome was identified and quantified in liver DNA samples obtained at 4 and 9 weeks after AAV9.hPCSK9 administration (left). From the obtained list, nine of the most common off-target sites were selected, and the percentage of insertions and deletions at these loci was calculated by NGS and plotted relative to the large-scale nuclease-only control (right).

[0017] Figures 5A to 5I This shows the results of administering the AAV suicide vector to rhesus macaques. The levels of PCSK9 in serum samples at different time points after vector administration are shown. Figure 5A ) and LDL protein levels ( Figure 5B For example, through Amplicon-Seq( Figure 5C ) or AMP-Seq method (Figure 5D ) detected indels in target sequences within the PCSK9 gene and the number of unique off-target sites at day 18 (d18) after vector administration. Figure 5E ) AAV genome copy number in NHP liver and meganuclease mRNA levels at d18. Figure 5F ) In situ hybridization in liver sections obtained from liver biopsies at d18 using probes specific for meganuclease DNA / mRNA. Figure 5G ) Editing in the PCSK9 target region induced PCSK9 reduction as previously observed (Nat Biotechnol. 2018 Sep; 36(8): 717-725), which also led to LDL levels reduction at different times after vector injection (respectively Figure 5H and Figure 5I ).

[0018] Figure 6 is a schematic of the AAV.TTR “suicide” vector. The AAV vector was constructed by insertion of a meganuclease in combination with a target sequence, a mutated target sequence containing 8 mismatches, and a PEST sequence.

[0019] Figures 7A to 7C Sequence analysis of AAV ITRs integrated into genomic DNA is provided. Figure 7A is a meta-analysis of on-target AMP-Seq data for all liver samples treated with AAV8-M1 PCSK9 and AAV8-M2 PCSK9 (SRR6343442). Our goal was to identify the most frequent ITR integration start sites within the vector ITRs. Figure 7B shows the secondary structure of the AAV2 5’ ITR (NC_001401.2). The position of the most frequently integrated start site is shown. The ITR-Seq primer GSP_ITR3.AAV2 binding site is highlighted in red. A-A’, B-B’ and C-C’, palindromic arms; RBE, rep binding element; TRS, terminal resolution site. Figure 7C is a schematic of the ITR-Seq protocol used for whole genome identification of ITR integration sites.

[0020] Figures 8A to 8C shows analysis of the targeting and off-target activity of AAV8-M1 PCSK9 and AAV8-M2 PCSK9 in vivo. Figure 8A is the ITR-Seq identified integration sites for liver samples treated with AAV8-M1 PCSK9 and AAV8-M2 PCSK9 collected at 17 and 128 days after vector administration. Figure 8BIt is a functional annotation of the ITR integration site for ITR identification, which shows the number of sites in exons, introns, transcription start sites (TSS) and transcription termination sites (TTS) between genes. Figure 8C The distribution of ITR integration sites in two animals treated with M1PCSK9 or M2PCSK9 (bars) or calculated random DNA sequences (dashed lines) at day 17 / 18, based on the number of nucleotides matching the expected target sequence, is shown as a percentage of the total number of identified sites.

[0021] Figures 9A to 9D This shows a comparison of off-target effects identified by GUIDE-Seq and ITR-Seq. On day 17 post-AAV administration, off-target effects were detected in AAV8-M1PCSK9 at a dose of 3 × 10⁻⁶. 13 GC / kg ( Figure 9A ), with 6×10 12 GC / kg ( Figure 9B Or from AAV8-M2PCSK9 at 6×10 12 GC / kg ( Figure 9C and Figure 9D The crossover points of the sample groups of identified target sites obtained by in vivo ITR-Seq (left side) and in vitro GUIDE-Seq (right side) for M1PCSK9 or M2PCSK9. Off-targets identified by ITR-Seq but not GUIDE-Seq (left side) are indicated as the percentage of the total number of off-targets identified by in vivo ITR-Seq. Off-targets identified by both ITR-Seq and GUIDE-Seq are shown as the white portion of the Venn diagram.

[0022] Figures 10A to 10B Examples are shown in which the target sequence or mutated target sequence is inserted into the enzyme coding sequence. Figure 10Ais an amino acid alignment showing a 10 amino acid nuclear localization signal (NLS) followed by a protein expressed from the target sequence (amino acids 10 to 20 followed by the active portion of the nuclease. The first amino acid sequence M2PCSK9 shows a fragment of the reference meganuclease with its NLS, reference to where other constructs will have space for insertion and the sequence of the nuclease beginning at position 18. M2PCSK9 [reverse target #1] shows the protein encoded when the target sequence is inserted on the antisense strand between the NLS and the enzyme. M2PCSK9 [target #1] shows the protein encoded when the target sequence is inserted on the sense strand between the NLS and the enzyme. M2PCSK9 [target #2] shows the protein encoded when a different target sequence is inserted on the sense strand between the NLS and the enzyme. M2PCSK9 [reverse target #2] shows the protein encoded when the (mutated) target sequence is inserted on the antisense strand between the NLS and the enzyme. [target] M2PCSK9 shows the protein encoded when a 22 bp target sequence replaces the coding sequence after the NLS, such that six amino acids of the enzyme are replaced. [reverse target] M2PCSK9 shows the protein encoded when a 22 bp target sequence replaces the coding sequence after the NLS on the opposite strand, such that seven amino acids of the enzyme are replaced. Figure 10B Design of PCSK9 meganuclease fusion proteins with a single protein degradation signal (degron) that is ubiquitin-independent (AR-6) and PCSK9 meganuclease fusion proteins with two protein degradation signals AR-6 and PEST are shown. DETAILED DESCRIPTION

[0023] The compositions and methods provided herein are designed to minimize off-target activity of the enzymes that persist after delivery of the expression cassettes and / or modulate the activity of the expressed enzymes. These compositions and methods are particularly desirable for use with non-secreted enzymes that can accumulate in cells and / or enzymes that accumulate at higher levels than desired prior to secretion. The compositions and methods of the present invention are particularly well suited for use with gene editing enzymes, particularly meganucleases. However, other applications will be apparent to those skilled in the art.

[0024] As used herein, suitable enzymes can be selected from meganucleases, zinc finger nucleases, transcription activator-like (TAL) effector nucleases (TALENs), clustered regularly interspaced short palindromic repeats (CRISPR) / endonucleases (Cas9, Cpfl, etc.). Examples of suitable meganucleases are described, for example, in U.S. Patent 8,445,251; US 9,340,777; US 9,434,931; US 9,683,257; and WO 2018 / 195449. Other suitable enzymes include the nuclease-inactivated S. pyogenes CRISPR / Cas9 that can be programmed with RNA (Nelles et al., Programmable RNA Tracking in Live Cells with CRISPR / Cas9, Cell, 165(2): P488-96 (April 2016)) and base editors (e.g., Levy et al., Cytosine and adenine base editing of the brain, liver, retina, heart and skeletal muscle of mice via adeno-associated viruses, Nature Biomedical Engineering, 4, 97-110 (January 2020)). In certain embodiments, the nuclease is not a zinc finger nuclease. In certain embodiments, the nuclease is not a CRISPR-associated nuclease. In certain embodiments, the nuclease is not a TALEN.

[0025] In certain embodiments, the nuclease is a member of the LAGLIDADG (SEQ ID NO: 1) family of homing endonucleases. In certain embodiments, the nuclease is a member of the I-CreI family of homing endonucleases that recognizes and cleaves a 22 base pair recognition sequence SEQ ID NO: 2 - CAAAACGTCGTGAGACAGTTTG. See, e.g., WO 2009 / 059195. Methods for rational design of single LAGLIDADG homing endonucleases are described that enable the de novo design of ICreI and other homing endonucleases to target a wide variety of DNA sites, including sites in mammalian, yeast, plant, bacterial, and viral genomes (WO 2007 / 047859).

[0026] Suitably, the coding sequence for a nuclease described herein is engineered into an expression cassette that further contains a self-inactivation component, a self-regulation component, or both, operably linked to a regulatory element that directs expression of the enzyme or enzyme-protein degradation signal (degron) in a cell containing a target site for the enzyme.

[0027] Self-inactivating nuclease fusion proteins

[0028] In one embodiment, a self-inactivating nuclease expression cassette contains a coding sequence for a nuclease fused in-frame to a protein degradation signal, such that a fusion protein comprising the nuclease and at least one protein degradation signal is produced. In one embodiment, a self-inactivating nuclease expression cassette contains a coding sequence for a meganuclease fused in-frame to a protein degradation signal, such that a fusion protein comprising the meganuclease and at least one protein degradation signal is produced. For convenience, the hyphenated phrase term "enzyme-protein degradation signal" is used to refer to such fusion proteins. This is not intended to limit the fusion protein to a protein degradation signal located at the carboxy terminus, or to only one such protein degradation signal present in the fusion protein. Similarly, it will be understood that a particular type of enzyme designation can be specified in the above phrase, e.g., "nuclease," "meganuclease," "endonuclease," etc., and is subject to the same interpretation as to the location and number of protein degradation signals without limitation, unless otherwise specified.

[0029] As described herein, a protein degradation signal is a portion of a protein that mediates degradation of the enzyme, which can also be referred to as a degron. Without wishing to be bound by theory, it is believed that a fusion protein containing an enzyme (e.g., nuclease) and at least one protein degradation signal reduces off-target activity without compromising on-target efficacy by facilitating removal of the enzyme from the cell when the presence of its natural substrate (target sequence) is reduced by enzymatic activity of the fusion protein. The protein degradation signal shortens the half-life of the protein and thus also reduces the level of accumulated protein. It is believed that this reduction in protein level contributes to the reduction in nuclease off-target activity. The peptide degradation signal acts independently of the activity of the enzyme (e.g., nuclease) in the target sequence.

[0030] Suitable protein degradation signals can include, for example, a PEST signal, a destruction box, or another destabilizing peptide.

[0031] A "PEST" sequence is a peptide sequence rich in proline (P), glutamic acid (E), serine (S), and threonine (T). This sequence has been shown to shorten the intracellular half-life of a protein, i.e., to act as a protein degradation signal. The examples provided herein utilize a PEST sequence having 42 amino acids (SEQ ID NO: 4 KLSHGPPEVEEQDDGTLPMSCAQESGMDRHPAACASARINV; coding sequence SEQ ID NO: 3). However, in certain embodiments, longer amino acid sequences containing this sequence or shorter amino acid fragments of this sequence can be selected, e.g., the sequence can be truncated at the amino terminus and / or carboxy terminus, e.g., to about 10, 15, 20, 25, 30, 35, 40, or 41 amino acids in length.

[0032] In certain embodiments, another protein degradation signal sequence can be selected. Such protein degradation signal sequences can be about 10 amino acids to about 50 amino acids in length, or values therebetween.

[0033] In certain embodiments, an ornithine decarboxylase (ODC) degradation determinant can be selected as a source of a suitable protein degradation signal sequence. [J Erales and P. Coffino, Biochimca et Biophsyica Acta, 1843 (2014) 216-221. In another embodiment, a ubiquitin degradation determinant can be selected as a protein degradation signal [KT Fortmann et al., J Mol Biol. 2015 Aug 28; 427(17): 2748-2756]. Yet other degradation determinants can be selected.

[0034] The engineered protein degradation signal can be at the amino (N-) terminus of the nuclease-encoding sequence. Optionally, multiple protein degradation signals can be present. Where two or more protein degradation signals are present, they can be the same or different. Further, where a fusion protein containing two or more protein degradation signals is provided, one or more of the signals can be at the N-terminus, one or more of the signals can be at the carboxy terminus, or a combination thereof (e.g., one signal can be at the N-terminus and one at the carboxy terminus of a single fusion protein; one signal can be at the N-terminus and two signals can be at the carboxy terminus of a single fusion protein, two protein degradation signals can be at the N-terminus and one signal can be at the carboxy terminus of a single fusion protein, etc.).

[0035] The examples herein illustrate the use of AAV vectors containing PEST sequences in the vector genome. In certain embodiments, the vector genome can be packaged into a different vector (e.g., a recombinant bocavirus). In certain embodiments, the expression cassette can be packaged into a different viral vector, packaged into a non-viral vector, and / or packaged into a different delivery system.

[0036] As used herein, an "expression cassette" refers to a nucleic acid molecule that includes a coding sequence, a promoter, and can include other regulatory sequences, which cassette can be engineered into genetic elements and / or packaged into the capsid of a viral vector (e.g., viral particle). Typically, such expression cassettes used to generate viral vectors contain the sequences described herein that flank the packaging signal of the viral genome and other expression control sequences, such as those described herein.

[0037] Expression cassettes typically contain a promoter sequence as part of the expression control sequence or regulatory sequence. In one embodiment, a tissue-specific promoter can be selected. For example, if a liver-specific promoter is desired, a liver-specific promoter can be selected from among the thyroid binding globulin (TBG), or other liver-specific promoters [see, e.g., The Liver Specific Gene Promoter Database, Cold Spring Harbor, rulai.schl.edu / LSPD], such as, e.g., alpha 1 antitrypsin (A1AT); human albumin (Miyatake et al., J. Virol., 71 :5124 32 (1997), humAlb); hepatitis B virus core promoter (Sandig et al., Gene Ther., 3:1002 9 (1996)); TTR minimal enhancer / promoter; alpha-antitrypsin promoter; T7 promoter; and LSP (845 nt)25 (for intronless scAAV). Other promoters can be used in the vectors described herein, such as viral promoters, constitutive promoters, regulated promoters [see, e.g., WO 2011 / 126808 and WO 2013 / 049493] or promoters that are responsive to physiological signals. In some embodiments, the promoter or promoter / enhancer is the promoter or promoter / enhancer set forth in any one of SEQ ID NOs: 11-16.

[0038] In addition to a promoter, the expression cassette and / or vector can contain other appropriate "regulatory elements" or "regulatory sequences" including, but not limited to, enhancers; transcription factors; transcription terminators; high efficiency RNA processing signals such as splicing and polyadenylation signals (poly A); sequences that stabilize cytoplasmic mRNA, for example, the Woodchuck Hepatitis Virus (WHP) post-transcriptional regulatory element (WPRE); sequences that enhance translation efficiency (i.e., the Kozak consensus sequence); sequences that enhance protein stability; and, if desired, sequences that enhance secretion of the encoded product. Examples of suitable poly A sequences include, for example, SV40, bovine growth hormone (bGH), and TK poly A. Examples of suitable enhancers include, for example, the alpha fetoprotein enhancer, the TTR minimal promoter / enhancer, the LSP (TH binding globulin promoter / alpha 1 -microglobulin / bikunin enhancer), and other enhancers. In one embodiment, the intron is the intron set forth in any one of SEQ ID NOs: 11-16. In one embodiment, the poly A is the poly A set forth in any one of SEQ ID NOs: 11-16.

[0039] These control sequences or regulatory sequences are operably linked to the protein and peptide coding sequences.

[0040] Self-regulating nuclease expression cassette

[0041] In certain embodiments, a self-regulating gene editing nuclease expression cassette is provided that contains a nuclease coding sequence operably linked to a regulatory sequence that directs nuclease expression upon delivery to a host cell having a sequence targeted by the nuclease; and at least one nuclease regulatory sequence selected from a target sequence of the nuclease or a mutated target sequence recognized by the nuclease upon its expression. Optionally, in these embodiments, the nuclease can be in the form of a fusion protein having a protein degradation signal as described in the foregoing sections by reference incorporated.

[0042] The term "recognition sequence" or "recognition site" refers to a nucleic acid sequence (e.g., DNA, including, for example, cDNA) that is bound and cleaved by an endonuclease. For certain meganucleases, the recognition sequence is 22 base pairs in length and includes a pair of inverted 9 base pair "half sites" separated by four base pairs. In the case of single-chain meganucleases, the N-terminal domain of the protein contacts the first half site and the C-terminal domain of the protein contacts the second half site. However, other nucleases and meganucleases can have shorter or longer recognition sites (e.g., about 10 base pairs to about 40 base pairs) as described herein.

[0043] As used herein, the term "target site" or "target sequence" refers to a region of a cell's DNA that includes the recognition sequence of a nuclease. In certain embodiments, the target site or target sequence is in the chromosomal DNA of a cell.

[0044] In certain embodiments, the nuclease modulating sequence has a sequence that is identical (100% identical) to the sequence of the nuclease recognition site in the cell over the full length of the recognition site in the cell. In certain embodiments, the nuclease modulating sequence is 100% identical to the target sequence, but is up to 5% to 10% shorter in length (e.g., for a 22 bp recognition site, about 18 to 20 bp in length).

[0045] In certain embodiments, the nuclease modulating sequence has a sequence that has a mutated sequence compared to the recognition site in the cell. Such sequences are designed to have mismatches in one or more base pairs.

[0046] In the examples below, the large range nucleases illustrated recognize 22 bp sequences, and thus, each enzyme modulating sequence selected in these expression cassettes is 22 bp. Mutant target sequences of these 22 bp sequences contain up to 35% to 37% mismatches, i.e., 8 bp that are different from the target sequence. However, other variations will be apparent to those skilled in the art.

[0047] A lower or higher percentage of mismatches can be selected. In certain embodiments, there is only 1 mismatch. In other embodiments, there are 2 to 12 mismatches. In certain embodiments, the mismatches are non-contiguous. In certain embodiments, two or three of the mismatches can be contiguous nucleotides. Optionally, the combination of a single mismatch with contiguous mismatches (e.g., 2 or 3) can be separated by unmutated sequence in the single mutant target. In certain embodiments, other mismatch sensitive nucleases can recognize shorter or longer sequences, e.g., about 12 base pairs to 40 base pairs in length. The mutant target sequence can have 0.5% to 45% mismatches (i.e., divergent nucleotide sequences) from the intended target sequence of the enzyme in the target cell. In certain embodiments, the mutant target sequence is engineered, e.g., by using off-target prediction / identification methods (such as GUIDE-Seq or ITR-Seq), and according to their ranking (% of indels in these sequences, or number of GUIDE-Seq reads, ITR-Seq reads), those mutant target sequences that can act at different levels or even better than the intended target sequence are selected.

[0048] In certain embodiments, the enzyme modulating sequence, e.g., a mutant target sequence or a target sequence, can be located downstream of a promoter sequence that directs expression of the enzyme.

[0049] In certain embodiments, the enzyme modulating sequence, e.g., a mutant target sequence or a target sequence, can be located downstream of the enzyme coding sequence.

[0050] In certain embodiments, for example, a mutant target sequence or an enzyme regulatory sequence of a target sequence can be located within the nuclease coding sequence.

[0051] In certain embodiments, the nuclease expression cassette (or vector genome containing the same) has multiple enzyme regulatory sequences that can be the same or different from each other. The enzyme regulatory sequences can be positioned in tandem, or can be separated from each other by one or more of the following: a non-coding spacer, between introns, or another regulatory element or enzyme coding sequence. For example, at least a first enzyme regulatory sequence can be located upstream of the enzyme coding sequence and a second enzyme regulatory sequence can be located downstream of the enzyme coding sequence (e.g., before the polyA). In the case where multiple different enzyme regulatory sequences are present, at least one enzyme regulatory sequence is a mutant target sequence. In such embodiments, two or more different mutant target sequences can be engineered in the nuclease expression cassette. In certain embodiments, the enzyme regulatory sequence is located within the nuclease coding sequence (e.g., at its 5’ or 3’ end). In one embodiment, the target sequence is SEQ ID NO: 5 - TGGACCTCTTTGCCCCAGGGGA. In one embodiment, the mutant target sequence is SEQ ID NO: 6 - TTGCCCTTTTTATTCCCAGGGA.

[0052] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising a PCSK9 meganuclease and a protein degradation signal (e.g., PEST) sequence. In certain embodiments, the meganuclease can be selected from those described in WO2018 / 195449A1.

[0053] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising a TTR meganuclease and a protein degradation signal (e.g., PEST) sequence. In certain embodiments, the TTR-target is SEQ ID NO: 7: GCTGGACTGGTATTTGTGTCTG. In certain embodiments, the TTR-mutant target has the sequence SEQ ID NO: 8: T C G GGACT TT T G TTTG CC TCT T .

[0054] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising a HOA meganuclease and a protein degradation signal (e.g., PEST) sequence.

[0055] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising a BCKDC meganuclease and a protein degradation signal (e.g., PEST) sequence.

[0056] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising an APOC3 meganuclease and a protein degradation signal (e.g., PEST) sequence.

[0057] In one embodiment, a nucleic acid molecule is provided that encodes a PCSK9 meganuclease and at least one target sequence or mutant target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0058] In one embodiment, a nucleic acid molecule is provided that encodes a fusion protein comprising a TTR meganuclease and at least one target sequence or mutant target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0059] Optionally, the expression cassette can comprise a miRNA target sequence in an untranslated region. The miRNA target sequence is designed to be specifically recognized by a miRNA present in a cell in which transgene expression is not desired and / or in which reduced transgene expression levels are desired. In certain embodiments, the expression cassette comprises a miRNA target sequence that specifically reduces expression of the nuclease in the dorsal root ganglion. In certain embodiments, the miRNA target sequence is in the 3’ UTR, the 5’ UTR, and / or both the 3’ UTR and the 5’ UTR, in some embodiments, the miRNA target sequence is operably linked to a regulatory sequence in the expression cassette. In certain embodiments, the expression cassette comprises at least two tandem repeats of a drg-specific miRNA target sequence, wherein the at least two tandem repeats comprise at least one first miRNA target sequence and at least one second miRNA target sequence, which can be the same or different. In certain embodiments, the tandem miRNA target sequences are contiguous, or are separated by a spacer having from 1 to 10 nucleic acids, wherein the spacer is not a miRNA target sequence. A discussion of miRNA DRG-specific sequences is provided in International Patent Application No. PCT / US19 / 67872, filed December 20, 2019, entitled “Compositions for DRG-Specific Reduction of Transgene Expression,” which is specifically incorporated herein by reference.

[0060] Viral and non-viral vectors

[0061] The expression cassettes described herein containing a nuclease-encoding sequence and at least one protein degradation signal and / or at least one target or mutant target site can be engineered into any suitable genetic element for delivery to a target cell.

[0062] A "vector" as used herein is a biological or chemical moiety that includes a nucleic acid sequence that can be introduced into an appropriate host cell for replication or expression of the nucleic acid sequence. Vectors include both non-viral vectors and viral vectors. As used herein, non-viral systems can be selected from the group consisting of nanoparticles, electroporation systems and novel biomaterials, naked DNA, bacteriophage, transposon, plasmid, cosmid (Phillip McClean, www.ndsu.edu / pubweb / ~mcclean / -plsc731 / cloning / cloning4.htm) and artificial chromosome (Gong, Shiaoching, et al. "A gene expression atlas of the central nervous system based on bacterial artificial chromosomes." Nature 425.6961 (2003): 917-925).

[0063] A "plasmid" or "plasmid vector" is designated herein generally by the lowercase p preceding and / or following the name of the vector. Plasmids, other cloning and expression vectors, their properties, and methods of construction / manipulation thereof that can be used in accordance with the present application will be apparent to those skilled in the art. In one embodiment, a nucleic acid sequence as described herein or an expression cassette as described herein is engineered into a suitable genetic element (vector) suitable for producing viral vectors and / or for delivery to a host cell, e.g., naked DNA, bacteriophage, transposon, cosmid, episome, etc., that transfers the nucleic acid sequence carried thereon. The selected vector can be delivered by any suitable method, including transfection, electroporation, liposome delivery, membrane fusion techniques, high velocity DNA-coated pellets, viral infection, and protoplast fusion. Methods for preparing such constructs are known to those skilled in nucleic acid manipulation and include genetic engineering, recombinant engineering, and synthetic techniques. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY.

[0064] In certain embodiments, the expression cassette is in a vector genome for packaging into a viral capsid. For example, for an AAV vector genome, the components of the expression cassette are flanked by AAV inverted terminal repeat sequences at the 5’ and 3’ ends. For example, 5’ AAV ITR, expression cassette, 3’ AAV ITR. In other embodiments, a self-complementary AAV can be selected. In other embodiments, a retroviral system, lentiviral vector system, or adenoviral system can be used. In one embodiment, the vector genome is the vector genome set forth in any one of SEQ ID NOs: 11-16.

[0065] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a PCSK9 meganuclease and a PEST sequence. In certain embodiments, the meganuclease can be selected from those described in WO 2018 / 195449 Al.

[0066] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a TTR meganuclease and a PEST sequence.

[0067] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a HOA1-2 or HOA3-4 meganuclease and a PEST sequence.

[0068] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a BCKDC meganuclease and a protein degradation signal (e.g., PEST) sequence.

[0069] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes an APOC3 meganuclease and a protein degradation signal (e.g., PEST) sequence.

[0070] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule that encodes a PCSK9 meganuclease and at least one target sequence or mutant target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0071] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a TTR meganuclease and at least one target sequence or mutant target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0072] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a HOA1-2 or HOA3-4 meganuclease and at least one target sequence or mutated target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0073] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a APOC3 meganuclease and at least one target sequence or mutated target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0074] In one embodiment, a viral or non-viral vector is provided that includes a nucleic acid molecule encoding a fusion protein that includes a BCKDC meganuclease and at least one target sequence or mutated target sequence. In certain embodiments, the meganuclease is expressed as a meganuclease-PEST fusion protein.

[0075] Any suitable non-viral vector (e.g., plasmid or other genetic element) or viral vector can be selected. The viral vector can be a recombinant bocavirus, a recombinant lentivirus, a recombinant adenovirus, or a recombinant adeno-associated virus.

[0076] AAV vectors

[0077] In certain embodiments, a recombinant AAV is provided. A "recombinant AAV" or "rAAV" is a nuclease-resistant viral particle containing two elements, an AAV capsid and a vector genome containing at least non- AAV coding sequences packaged within the AAV capsid. This term can be used interchangeably with the phrase "rAAV vector" unless otherwise specified. The rAAV is a "replication-defective virus" or "viral vector" because it lacks any functional AAV rep genes or functional AAV cap genes and is unable to produce progeny. In certain embodiments, the only AAV sequences are the AAV inverted terminal repeat sequences (ITRs), typically at the 5' and 3' most ends of the vector genome, in order to allow the gene and regulatory sequences located between the ITRs to be packaged within the AAV capsid.

[0078] The source of the AAV capsid can be one of several dozen naturally occurring and available adeno-associated viruses as well as engineered AAVs. An adeno-associated virus (AAV) viral vector is an AAV DNAse-resistant particle having an AAV protein capsid in which a nucleic acid sequence for delivery to a target cell is packaged. The AAV capsid is composed of 60 capsid (cap) protein subunits, VP1, VP2, and VP3, arranged in icosahedral symmetry in a ratio of approximately 1:1:10 to 1:1:20, depending on the selected AAV. Various AAVs can be selected as the source of the capsid for the AAV viral vectors as identified above. See, e.g., U.S. Published Patent Application No. 2007-0036760-A1; U.S. Published Patent Application No. 2009-0197338-A1; EP 1310571. See also WO 2003 / 042397 (AAV7 and other simian AAVs), U.S. Patent 7790449 and U.S. Patent 7282199 (AAV8), WO 2005 / 033321 and US 7,906,111 (AAV9) and WO 2006 / 110689, WO 2003 / 042397 (rh.10), and WO 2018 / 160582 (AAVhu68). These documents also describe other AAVs that can be selected for production of AAVs, and are incorporated by reference. Unless otherwise specified, the AAV capsids, ITRs, and other selected AAV components described herein can be readily selected from any AAV, including but not limited to variants of any of the AAVs commonly identified as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV8bp, AAV7M8, AAV Anc80, AAVrh10, and AAVPHP.B, as well as known or mentioned AAVs or variants thereof or AAVs not yet discovered or variants or mixtures thereof. See, e.g., WO 2005 / 033321, incorporated herein by reference. In one embodiment, the AAV capsid is an AAV1 capsid or variant thereof, an AAV8 capsid or variant thereof, an AAV9 capsid or variant thereof, an AAVrh.10 capsid or variant thereof, an AAVrh64R1 capsid or variant thereof, an AAVhu.37 capsid or variant thereof, or an AAV3B or variant thereof.See also PCT / US19 / 19804 and PCT / US19 / 19861, each entitled “NOVEL ADENO- ASSOCIATED VIRUS (AAV) VECTORS, AAV VECTORS HAVING REDUCED CAPSID DEAMIDATION AND USES THEREFOR” and filed on February 27, 2019, which are incorporated herein by reference in their entirety.

[0079] As used herein, “vector genome” refers to the nucleic acid sequence packaged inside the rAAV capsid that forms the viral particle. This nucleic acid sequence contains the AAV inverted terminal repeat sequences (ITRs). In the examples herein, the vector genome contains at least an AAV 5’ ITR, a coding sequence, and an AAV 3’ ITR, from 5’ to 3’. The ITRs can be selected from AAV2, a different source AAV, or full-length ITRs other than the ITRs from the capsid. In certain embodiments, the ITRs are from the same AAV source as the AAV that provides the rep function during production or back- complement AAV. In addition, other ITRs can be used. In addition, the vector genome contains regulatory sequences that direct expression of the gene product. Suitable components of the vector genome are discussed in more detail herein.

[0080] For production of AAV viral vectors (e.g., recombinant (r)AAV), the expression cassette can be carried on any suitable vector (e.g., plasmid) delivered to the packaging host cell. Plasmids suitable for use in the present application can be engineered so that they are suitable for replication and packaging in prokaryotic cells, insect cells, mammalian cells, and other cells in vitro. Suitable transfection techniques and packaging host cells are known and / or can be readily designed by one of skill in the art.

[0081] Methods for producing and isolating AAV suitable for use as vectors are known in the art. See generally, e.g., Grieger & Samulski, 2005, “Adeno-associated virus as a gene therapy vector: Vector development, production and clinical applications,” Adv. Biochem. Engin / Biotechnol, 99: 119-145; Buning et al., 2008, “Recent developments in adeno-associated virus vector technology,” J. Gene Med 10: 717-733; and references cited below, each of which is incorporated herein by reference in its entirety. To package a transgene into a virion, ITRs are the only AAV components required in cis in the same construct as the nucleic acid molecule containing the expression cassette. The cap and rep genes can be supplied in trans.

[0082] The term “AAV intermediate” or “AAV vector intermediate” refers to an assembled rAAV capsid that lacks the desired genomic sequence packaged therein. It can also be referred to as an “empty” capsid. This capsid can contain no detectable genomic sequence of an expression cassette, or only a partially packaged genomic sequence insufficient to effect expression of a gene product. These empty capsids are non-functional to transfer a gene of interest to a host cell.

[0083] Recombinant adeno-associated viruses (AAV) described herein can be produced using known techniques. See, e.g., WO 2003 / 042397; WO 2005 / 033321, WO 2006 / 110689; US 7588772 B2. This method involves culturing host cells containing nucleic acid sequences encoding AAV capsid proteins; a functional rep gene; an expression cassette consisting of at least an AAV inverted terminal repeat (ITR) and a transgene; and sufficient helper functions to permit packaging of the expression cassette into AAV capsid proteins. Methods of producing capsids, their encoding sequences, and methods for producing rAAV viral vectors have been described. See, e.g., Gao et al., Proc. Natl. Acad. Sci. U.S.A. 100(10), 6081-6086 (2003) and US 2013 / 0045186 Al.

[0084] In one embodiment, a production cell culture suitable for producing recombinant AAV is provided. This cell culture contains a nucleic acid that expresses an AAV capsid protein in a host cell; a nucleic acid molecule suitable for packaging into an AAV capsid, such as a vector genome containing AAV ITRs and a non- AAV nucleic acid sequence encoding a gene product operably linked to sequences that direct expression of the product in the host cell; and sufficient AAV rep function and adenovirus helper function to permit packaging of the nucleic acid molecule into a recombinant AAV capsid. In one embodiment, the cell culture is comprised of mammalian cells (e.g., human embryonic kidney 293 cells and other cells) or insect cells (e.g., baculovirus).

[0085] Optionally, the rep function is provided by an AAV other than the AAV providing the capsid. For example, the rep can be, but is not limited to, an AAV1 rep protein, an AAV2 rep protein, an AAV3 rep protein, an AAV4 rep protein, an AAV5 rep protein, an AAV6 rep protein, an AAV7 rep protein, an AAV8 rep protein; or rep 78, rep 68, rep 52, rep 40, rep 68 / 78, and rep 40 / 52; or fragments thereof; or another source. Optionally, the rep and cap sequences are on the same genetic element in the cell culture. There can be a spacer between the rep sequence and the cap gene. Any of these AAV or mutant AAV capsid sequences can be under the control of an exogenous regulatory control sequence that directs its expression in the host cell.

[0086] In one embodiment, cells are manufactured in a suitable cell culture (e.g., HEK 293) cell. Methods for manufacturing the gene therapy vectors described herein include methods well known in the art, such as production of plasmid DNA for gene therapy vector production, production of vectors, and purification of vectors. In some embodiments, the gene therapy vector is an AAV vector, and the plasmids produced are AAV cis-plasmids encoding the AAV genome and gene of interest, AAV trans-plasmids containing the AAV rep and cap genes, and adenoviral helper plasmids. The vector production process can include, for example, the following method steps: initiation of cell culture, cell passaging, cell seeding, transfection of cells with plasmid DNA, replacement of post-transfection media with serum-free media, and harvesting of vector-containing cells and media. The harvested vector-containing cells and media are referred to herein as a crude cell harvest. In yet another system, gene therapy vectors are introduced into insect cells by infection with baculovirus-based vectors. For a review of these production systems, see generally, e.g., Zhang et al., 2009, “Adenovirus-adeno-associated virus hybrid for large-scale recombinant adeno-associated virus production,” Human Gene Therapy 20:922-929, the contents of each of which are incorporated by reference herein in their entirety. Methods of preparing and using these and other AAV production systems are also described in the following U.S. Patents, the contents of each of which are incorporated by reference herein in their entirety: 5,139,941; 5,741,683; 6,057,152; 6,204,059; 6,268,213; 6,491,907; 6,660,514; 6,951,753; 7,094,604; 7,172,893; 7,201,898; 7,229,823; and 7,439,065.

[0087] Thereafter, the crude cell harvest can be a method step of the present subject matter, such as concentrating the vector harvest, diafiltering the vector harvest, microfluidizing the vector harvest, nuclease digesting the vector harvest, filtering the microfluidized intermediate, crude purification by chromatography, crude purification by ultracentrifugation, buffer exchange and / or formulation by tangential flow filtration, and filtration to prepare bulk vector.

[0088] Two-step affinity chromatography purification at high salt concentration followed by purification of the vector drug product and removal of empty capsids using anion exchange resin chromatography. These methods are described in more detail in International Patent Application No. PCT / US2016 / 065970, filed December 9, 2016, and its priority documents, U.S. Patent Application No. 62 / 322,071, filed April 13, 2016, and No. 62 / 226,357, filed December 11, 2015, and entitled "Scalable Purification Method for AAV9," which are incorporated herein by reference. For purification methods for AAV8, see International Patent Application No. PCT / US2016 / 065976, filed December 9, 2016, and its priority documents, U.S. Patent Application No. 62 / 322,098, filed April 13, 2016, and No. 62 / 266,341, filed December 11, 2015, and for purification methods for rhlO, see International Patent Application No. PCT / US16 / 66013, filed December 9, 2016, and its priority documents, U.S. Patent Application No. 62 / 322,055, filed April 13, 2016, and No. 62 / 266,347, also filed December 11, 2015, and entitled "Scalable Purification Method for AAVrh10," and for purification methods for AAVl, see International Patent Application No. PCT / US2016 / 065974, filed December 9, 2016, and its priority documents, U.S. Patent Application No. 62 / 322,083, filed April 13, 2016, and No. 62 / 26,351, filed December 11, 2015, and entitled "Scalable Purification Method for AAVl," all of which are incorporated herein by reference.

[0089] To calculate empty and full particle content, the VP3 band volume of the selected sample (e.g., in the examples herein, the iodixanol gradient purified formulation where GC# = particle #) was plotted against the loaded GC particles. The resulting linear equation (y = mx + c) was used to calculate the number of particles in the band volume of the test article peak. The number of particles per 20 μΐ, (pt) loaded was then multiplied by 50 to give particles (pt) / mL. Pt / mL divided by GC / mL gives the ratio of particles to genome copies (pt / GC). Pt / mL - GC / mL gives empty pt / mL. Empty pt / mL divided by pt / mL and x 100 gives the percentage of empty particles.

[0090] In general, methods for assaying empty capsids and AAV vector particles with packaged genomes are known in the art. See, e.g., Grimm et al., Gene Ther. (1999) 6:1322-1330; Sommer et al., Molec. Ther. (2003) 7:122-128. To test denatured capsids, the method comprises subjecting the treated AAV stock to SDS-polyacrylamide gel electrophoresis (consisting of any gel capable of separating the three capsid proteins, e.g., a gradient gel containing 3-8% Tris-acetate in the buffer), followed by running the gel until the sample material is separated, and blotting the gel onto a nylon or nitrocellulose membrane, preferably nylon. Subsequently, an anti-AAV capsid antibody is used as a primary antibody that binds to the denatured capsid proteins, preferably an anti-AAV capsid monoclonal antibody, most preferably the B1 anti-AAV-2 monoclonal antibody (Wobus et al., J. Virol. (2000) 74:9281-9293). Subsequently, a secondary antibody is used that binds to the primary antibody and contains a means for detecting the binding to the primary antibody, more preferably an anti-IgG antibody that has a detection molecule covalently bound thereto, most preferably a sheep anti-mouse IgG antibody covalently linked to horseradish peroxidase. A method for detecting the binding is used to determine the binding between the primary and secondary antibodies semi-quantitatively, preferably a detection method capable of detecting the emission of a radioisotope, electromagnetic radiation, or a colorimetric change, most preferably a chemiluminescent detection kit. For example, for SDS-PAGE, samples from column fractions can be taken and heated in SDS-PAGE loading buffer containing a reducing agent (e.g., DTT), and the capsid proteins are resolved on a precast gradient polyacrylamide gel (e.g., Novex). Silver staining can be performed using SilverXpress (Invitrogen, CA) or other suitable staining methods (i.e., SYPRO Ruby or Coomassie staining) according to the manufacturer's instructions. In one embodiment, the concentration of AAV vector genomes (vg) in the column fractions can be measured by quantitative real-time PCR (Q-PCR). The samples are diluted and digested with DNase I (or another suitable nuclease) to remove extraneous DNA. After inactivation of the nuclease, primers and TaqMan probes specific for the DNA sequence between the primers are used to measure the amount of AAV vector genomes in the sample. The amount of AAV vector genomes in the sample is determined by comparison to a standard curve of known amounts of AAV vector genomes. The amount of AAV vector genomes in the sample is determined by comparison to a standard curve of known amounts of AAV vector genomes. TMThe fluorescent probe further dilutes and amplifies the sample. The number of cycles required to reach a defined level of fluorescence (threshold cycle, Ct) is measured on an Applied Biosystems Prism 7700 Sequence Detection System for each sample. Plasmid DNA containing the same sequence as contained in the AAV vector is used to generate a standard curve in the Q-PCR reaction. The cycle threshold (Ct) values obtained from the samples are used to determine the vector genome titer by normalizing them to the Ct values for the plasmid standard curve. End-point assays based on digital PCR can also be used.

[0091] In one aspect, an optimized q-PCR method is used that utilizes a broad-spectrum serine protease, such as proteinase K (commercially available from Qiagen). More specifically, the optimized qPCR genome titer assay is similar to the standard assay except that after DNase I digestion, the sample is diluted with proteinase K buffer and treated with proteinase K, followed by heat inactivation. The sample is suitably diluted with proteinase K buffer in an amount equal to the sample size. The proteinase K buffer can be concentrated up to 2-fold or more. Typically, the proteinase K treatment is about 0.2 mg / mL, but can vary between 0.1 g / mL to about 1 mg / mL. The treatment step is generally performed at about 55°C for about 15 minutes, but can be performed at a lower temperature (e.g., about 37°C to about 50°C) for a longer period of time (e.g., about 20 minutes to about 30 minutes), or at a higher temperature (e.g., up to about 60°C) for a shorter period of time (e.g., about 5 to 10 minutes). Similarly, the heat inactivation is generally at about 95°C for about 15 minutes, but the temperature can be reduced (e.g., about 70°C to about 90°C) and the time can be lengthened (e.g., about 20 minutes to about 30 minutes). The sample is then diluted (e.g., 1000-fold) and TaqMan analysis is performed as described in the standard assay.

[0092] Additionally or alternatively, droplet digital PCR (ddPCR) can be used. For example, a method for determining single-stranded and self-complementary AAV vector genome titers by ddPCR has been described. See, e.g., M. Lock et al., Hum Gene Ther Methods. 2014 Apr;25(2): 115-25. Digital Object Identifier: 10.1089 / hgtb.2013.131. Epub February 14, 2014.

[0093] Briefly, a method for isolating rAAV particles having packaged genomic sequences from genomic-deficient AAV intermediates involves subjecting a suspension comprising recombinant AAV viral particles and AAV capsid intermediates to high performance liquid chromatography, wherein the AAV viral particles and AAV intermediates bind to a strong anion exchange resin equilibrated at high pH and subjected to a salt gradient while monitoring the eluate for ultraviolet absorbance at about 260 and about 280. The pH can be adjusted according to the AAV selected. See, e.g., WO 2017 / 160360 (AAV9), WO 2017 / 100704 (AAVrh10), WO 2017 / 100676 (e.g., AAV8), and WO 2017 / 100674 (AAV1), which are incorporated herein by reference. In this method, when the ratio of A260 / A280 reaches an inflection point, the AAV full capsids are collected from the fraction eluted. In one example, for the affinity chromatography step, the diafiltered product can be applied to a Capture Select TM Poros-AAV2 / 9 affinity resin (Life Technologies) effective to capture AAV2 serotypes. Under these ionic conditions, a significant percentage of residual cellular DNA and proteins flow through the column while AAV particles are effectively captured.

[0094] Pharmaceutical compositions

[0095] A pharmaceutical composition includes one or more of an expression cassette, a vector (viral or non-viral) containing an expression cassette, or another system containing an expression cassette, and one or more of a carrier, a suspending agent, and / or an excipient.

[0096] In certain embodiments, a composition contains at least one rAAV stock (e.g., rAAV stock) and optionally a carrier, an excipient, and / or a preservative. An rAAV stock refers to a plurality of rAAV vectors in an amount the same as described, e.g., in the discussion below regarding concentration and dosage units. The transgene that can be delivered by the rAAV vectors can be formulated for delivery or encapsulated in a lipid particle, liposome, vesicle, nanosphere, or nanoparticle, or the like.

[0097] In certain embodiments, the expression cassette is delivered via a lipid nanoparticle. The term "lipid nanoparticle" refers to a lipid composition having a typical spherical structure with an average diameter of 10 to 1000 nanometers, e.g., 75 nm to 750 nm, or 100 nm and 350 nm, or 250 nm to about 500 nm. In some formulations, the lipid nanoparticle can include at least one cationic lipid, at least one non-cationic lipid, and at least one conjugate lipid. Lipid nanoparticles suitable for encapsulating nucleic acids, such as mRNA, known in the art can be used. The "average diameter" is the average size of a population of nanoparticles comprising a lipophilic phase and a hydrophilic phase. The average size of these systems can be measured by standard methods known to one of skill in the art. Examples of suitable lipid nanoparticles for gene therapy are described in, e.g., L. Battaglia and E. Ugazio, Journal of Nanomaterials, Volume 2019, Article Number 283441, Pages 1-22; US2012 / 0183589A1; and WO2012 / 170930, which are incorporated herein by reference in their entirety.

[0098] In certain embodiments, a composition is provided that includes a nucleic acid molecule encoding a nuclease-degrading peptide signal fusion protein and a pharmaceutically acceptable diluent, carrier, and / or excipient.

[0099] rAAV stock refers to a plurality of rAAV vectors in an amount that is the same as that described, e.g., in the discussion below regarding concentration and dosage units.

[0100] As used herein, "carrier" includes any and all solvents, dispersion media, vehicles, coatings, diluents, antibacterial and antifungal agents, isotonic and absorption delaying agents, buffers, carrier solutions, suspensions, gel, and the like. The use of such media and agents for pharmaceutical active substances is well known in the art. Supplementary active ingredients can also be incorporated into the compositions. The phrase "pharmaceutically-acceptable" refers to molecular entities and compositions that do not produce an allergic or similar untoward reaction when administered to a host. Delivery vehicles, such as liposomes, nanocapsules, microparticles, microspheres, lipid particles, vesicles, and the like, can be used to introduce the compositions of the present application into suitable host cells. In particular, the vector genome of the rAAV vector can be formulated for delivery, or encapsulated in a lipid particle, liposome, vesicle, nanosphere, or nanoparticle, or the like.

[0101] In one embodiment, the composition comprises a final formulation suitable for delivery to an individual, e.g., an aqueous liquid suspension buffered to a physiologically compatible pH and salt concentration. Optionally, one or more surfactants are present in the formulation. In another embodiment, the composition can be delivered as a concentrate that is diluted for administration to an individual. In other embodiments, the composition can be lyophilized and reconstituted at the time of administration.

[0102] Methods and agents for preparing formulations well known in the art are described, e.g., in "Remington's Pharmaceutical Sciences," Mack Publishing Company, Easton, Pa. The formulations can, for example, contain excipients, carriers, stabilizers, or diluents, such as sterile water; physiological saline; polyalkylene glycols, such as polyvinyl glycols; oils of vegetable origin or hydrogenated naphthalenes; preservatives (e.g., octadecyldimethylbenzyl ammonium chloride, hexamethonium chloride, benzalkonium chloride, benzethonium chloride, phenol, butyl or benzyl alcohol, alkyl parabens such as methyl or propyl paraben, catechol, resorcinol, cyclohexanol, 3-pentanol, and m-cresol); low molecular weight polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; hydrophilic polymers, such as polyvinylpyrrolidone; amino acids such as glycine, glutamine, asparagine, histidine, arginine, and lysine; monosaccharides; disaccharides and other carbohydrates including glucose, mannose, and dextrins; chelating agents, such as EDTA; sugars such as sucrose, mannitol, trehalose or sorbitol; salt-forming counter-ions such as sodium; metal complexes (e.g., Zn-protein complexes); and / or non-ionic surfactants such as TWEEN TM , PLURONICS TM or polyethylene glycol (PEG).

[0103] The active ingredients can also be entrapped in microcapsules prepared, for example, by coacervation techniques or by interfacial polymerization, for example, hydroxymethylcellulose or gelatin-microcapsules and poly-(methylmethacrylate) microcapsules, respectively, in colloidal drug delivery systems (for example, liposomes, albumin microspheres, microemulsions, nano- particles, and nanocapsules) or in macroemulsions. Such techniques are disclosed in Remington's Pharmaceutical Sciences 16th edition, Osol, A. Ed. (1980).

[0104] Suitable surfactants or combinations of surfactants can be selected from among non-toxic non-ionic surfactants. In one embodiment, a difunctional block copolymer surfactant that terminates in a primary hydroxyl group is selected, e.g., PLURONICS®. F68 [BASF], also known as Poloxamer 188, has a neutral pH with an average molecular weight of 8400. Other surfactants and other poloxamers, i.e., nonionic triblock copolymers, consisting of a central hydrophobic chain of polyoxypropylene (poly(propylene oxide)) flanked by two hydrophilic chains of polyoxyethylene (poly(ethylene oxide)) side groups, SOLUTOL HS 15 (polyethylene glycol-15 hydroxystearate), LABRASOL (polyoxylglycerides), polyoxyl 10 oleyl ether, TWEEN (polyoxyethylene sorbitan fatty acid esters), ethanol, and polyethylene glycol, can be selected. In one embodiment, the formulation contains a poloxamer. These copolymers are generally named with the letter "P" (for poloxamer) followed by three digits: the first two digits x 100 give the approximate molecular mass of the polyoxypropylene core, and the last digit x 10 gives the percentage of polyoxyethylene content. In one embodiment, poloxamer 188 is selected. The surfactant can be present in an amount of up to about 0.0005% to about 0.001% of the suspension.

[0105] The carrier is administered in an amount sufficient to transfect the cells and provide a sufficient level of gene transfer and expression to provide a therapeutic benefit without undue side effects or with medically acceptable physiological effects, which can be determined by one skilled in the medical arts. Conventional and pharmaceutically acceptable routes of administration include, but are not limited to, direct delivery to the desired organ (e.g., liver (optionally via the hepatic artery), lung, heart, eye, kidney), oral, inhalation, intranasal, intrathecal, intratracheal, intraarterial, intraocular, intravenous, intramuscular, subcutaneous, intradermal, and other parental routes of administration. Routes of administration can be combined as desired.

[0106] The dosage of the viral vector depends primarily on factors such as the condition being treated, the age, weight, and health of the patient, and thus can vary between patients. For example, a therapeutically effective human dose of the viral vector is generally in the range of about 25 to about 1000 microliters to about 100 mL of a solution containing a concentration of about 1 x 10 9 to 1 x 10 16 The dosage is adjusted to balance the therapeutic benefit with any side effects, and such dosages can vary depending on the therapeutic application of the recombinant vector employed. Expression levels of the transgene product can be monitored to determine the frequency of dosage of the viral vector, preferably an AAV vector containing a minigene. Optionally, a dosage regimen similar to that described for therapeutic purposes can be used for immunization using the compositions of the present application.

[0107] The replication-defective viral composition can be formulated in dosage units to contain from about 1.0 x 10 9 GC to about 1.0 x 1016 replication-defective virus in the range of 1.0 x 10 12 GC to 1.0 x 10 14 GC. In one embodiment, the composition is formulated to contain at least 1 x 10 9 , 2 x 10 9 , 3 x 10 9 , 4 x 10 9 , 5 x 10 9 , 6 x 10 9 , 7 x 10 9 , 8 x 10 9 , or 9 x 10 9 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 10 , 2 x 10 10 , 3 x 10 10 , 4 x 10 10 , 5 x 10 10 , 6 x 10 10 , 7 x 10 10 , 8 x 10 10 , or 9 x 10 10 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 11 , 2 x 10 11 , 3 x 10 11 , 4 x 10 11 , 5 x 10 11 , 6 x 10 11 , 7 x 10 11 , 8 x 10 11 , or 9 x 10 11 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 12 , 2 x 10 12 , 3 x 10 12 , 4 x 10 12 , 5 x 10 12 , 6 x 10 12 , 7 x 10 12 , 8 x 10 12 , or 9 x 10 12 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 13 , 2 x 10 13 , 3 x 10 13, 5 x 10 13 , 6 x 10 13 , 7 x 10 13 , 8 x 10 13 , 9 x 10 13 , or 9 x 10 13 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 14 , 2 x 10 14 , 3 x 10 14 , 4 x 10 14 , 5 x 10 14 , 6 x 10 14 , 7 x 10 14 , 8 x 10 14 , or 9 x 10 14 GC, including all integers or fractional amounts within the range. In another embodiment, the composition is formulated to contain at least 1 x 10 15 , 2 x 10 15 , 3 x 10 15 , 4 x 10 15 , 5 x 10 15 , 6 x 10 15 , 7 x 10 15 , 8 x 10 15 , or 9 x 10 15 GC, including all integers or fractional amounts within the range. In one embodiment, for human applications, the dose can range from about 1 x 10 10 to about 1 x 10 12 GC, including all integers or fractional amounts within the range.

[0108] These above-mentioned doses can be administered in various volumes of carrier, excipient, or buffer formulation, ranging from about 25 to about 1000 microliters or more in volume, including all numbers within the range, depending on the size of the area to be treated, the viral titer used, the desired effect of the route and method of administration.

[0109] Any suitable route of administration can be chosen (e.g., oral, inhalation, intranasal, intratracheal, intraarterial, intraocular, intravenous, intramuscular, intraperitoneal, and other parental routes). Thus, the pharmaceutical composition can be formulated for any appropriate route of administration, e.g., in the form of a liquid solution or suspension (as, e.g., for intravenous administration, for oral administration, etc.). Alternatively, the pharmaceutical composition can be in solid form (e.g., in the form of a tablet or capsule, e.g., for oral administration). In some embodiments, the pharmaceutical composition can be in the form of a powder, drops, spray, etc.

[0110] Methods

[0111] The compositions provided herein are suitable for reducing off-target activity of an enzyme delivered in vivo. In certain embodiments, the compositions are suitable for reducing off-target activity of an expressed enzyme following delivery of a non-viral mediated expression cassette comprising an enzyme-encoding sequence, an enzyme-PEST-encoding sequence, and / or an enzyme-PEST-encoding sequence having one or more enzyme-modulating (target or mutant target) sequences. In certain embodiments, the compositions are suitable for reducing off-target activity of an expressed enzyme following delivery of an AAV-mediated vector genome.

[0112] In certain embodiments, the efficacy of a protein degradation signal (e.g., PEST, cassette, or other degron) can be assessed in vitro. For example, the half-life of a fusion protein comprising an enzyme (e.g., a nuclease) and a protein degradation signal can be assessed in vitro (in cultured cells) by treating cells to stop translation of the protein (e.g., using cycloheximide (CHX)) and then performing a western blot at different times following treatment. Other suitable methods for assessing nuclease degradation can readily be determined by one of skill in the art.

[0113] Reduction of off-target nuclease activity (or increase in nuclease specificity) can be determined using a variety of methods that have been described in the literature. Such methods for determining nuclease specificity include cell-free methods, such as Site-Seq [Cameron, P. et al. (2017) Mapping the genomic landscape of CRISPR-Cas9 cleavage. Nat Methods, 14, 600-606], Digenome-seq [Kim, D. et al., (2015) Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells. Nat Methods, 12, 237-243, page 231 followed by page 243] and Circle-Seq [Tsai, S.Q. et al., (2017) CIRCLE-seq: a highly sensitive in vitro screen for genome-wide CRISPR-Cas9 nuclease off-targets. Nat Methods, 14, 607-614] and in vitro based methods such as, for example, GUIDE-Seq [Tsai (2017) Nat Methods, 14, 607-614] and integrase-deficient lentiviral vector capture (IDLV) [Gabriel, R. et al. (2011) An unbiased genome-wide analysis of zinc-finger nuclease specificity. Nat Biotechnol, 29, 816-823; Wang, X. et al., (2015) Unbiased detection of off-target cleavage by CRISPR-Cas9 and TALENs using integrase-defective lentiviral vectors. Nat Biotechnol, 33, 175-178].

[0114] Provided herein and incorporated by reference are improved assays that more accurately predict the amount and ratio of off-target activity in vivo.

[0115] The ITR-seq method we developed provides unbiased whole genome identification of ITR integration sites. The following examples use AAV ITRs as a tag for identifying DSBs, which we can measure in vivo for off-target activity of genome editing nucleases. However, it will be readily appreciated that any terminal repeat sequence (trs) (e.g., lentiviral terminal repeat sequences) or other common integration site, such as repeats located 5' and 3' of an expression cassette, can be used. Similarly, non-viral expression cassettes can be engineered to contain such common integration sites in order to allow detection using this assay (e.g., trs can be engineered into DNA expression cassettes delivered by non-viral or non- vector delivery systems).

[0116] The ITR-Seq protocol is a modified version of an anchor PCR reaction, in which a single primer is designed to anneal to the ITR sequence and amplify outwards from the ITR sequence (e.g., Figure 7C ). Following ITR integration in DNA, the primer can be used to amplify the junction of the host genome with the inserted vector ITR sequence (e.g., Figure 7B 、 7C ). To fully denature the ordered secondary structure of the integrated ITR, a high annealing temperature (e.g., 69°C) and longer adapter-specific primers are used. In certain embodiments, a "high annealing temperature" can be any temperature at which a PCR polymerase functions and a primer anneals to a target sequence, such as at 60°C to 75°C or about 68°C to 72°C. In certain embodiments, the length of the primer sequence can be 18 to about 42 nucleotides, possibly longer if specificity is retained. In certain embodiments, the length of the primer is at least 20 nucleotides to 40 nucleotides, at least 30 nucleotides to 40 nucleotides, at least 35 nucleotides to 40 nucleotides, or about 37 nucleotides.

[0117] For sample analysis using the ITR-Seq assay, DNA is isolated from a sample (e.g., tissue from an animal treated with a nuclease-expressing AAV vector). The DNA is sheared and ligated to Y adapters as described in a previous report [Tsai, S. Q. et al. (2015) GUIDE-seq enables genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nature Biotechnology, 33, 187-197]. After two rounds of PCR using the ITR-specific primers and adapter-specific primers described above, an NGS-compatible library is generated. After sequencing, the resulting amplicons containing the amplified ITR sequence and adjacent genomic DNA sequence are computationally determined. The location and frequency of the whole-genome ITR integration sites are also determined. By requiring the ITR to integrate on both the forward and reverse strand orientations, we further reduce the number of false positives and identify high-confidence ITR integration sites. For each sample, a ranked list of nuclease target sites (ITR-Seq rank) is generated, and the sites are sorted in descending order by the total number of ITR integration events observed at each locus (ITR-Seq reads). The ITR-Seq report is generated at the end of the computational analysis using the most likely off-target sequence (based on homology to the intended target sequence), the genomic location, and the ITR-Seq rank (according to the number of NGS reads mapping to the corresponding locus). Example 3 provides additional details of the assay and illustrates the use of the assay.

[0118] In one aspect, a method for editing a targeted gene is provided, comprising delivering a self-modulating nuclease expression cassette as described herein.

[0119] In one aspect, a method for editing a targeted gene is provided, comprising delivering a composition as described herein.

[0120] In one aspect, a method for editing a targeted gene is provided, comprising delivering a viral or non-viral vector as described herein.

[0121] In one aspect, a method for editing a targeted gene is provided, comprising delivering a rAAV as described herein.

[0122] In one aspect, a method is provided for treating a patient having a cholesterol-related disorder, such as hypercholesterolemia, using a self-modulating nuclease expression cassette comprising a meganuclease that recognizes a site within the human PCSK9 gene as described herein. In certain embodiments, the expression cassette encodes a fusion protein comprising a PCSK9 meganuclease and a protein degradation signal, e.g., a PEST sequence. In one embodiment, the fusion protein has the sequence set forth in SEQ ID NO: 18. In certain embodiments, the expression cassette encodes a PCSK9 meganuclease and at least one target sequence. In certain embodiments, the expression cassette comprises a mutated target sequence. In certain embodiments, the expression cassette comprises a fusion protein comprising a PCSK9 meganuclease and a protein degradation signal and at least one target sequence. Such expression cassettes can be delivered via a viral or non-viral vector. In certain embodiments, the expression cassette can be delivered using LNP.

[0123] In one aspect, a method is provided for treating a patient having a disorder associated with alanine-glyoxylate aminotransferase gene deficiency, such as primary hyperoxaluria type 1, using a self-modulating nuclease expression cassette comprising a meganuclease that recognizes a site within the human HAO gene as described herein. In certain embodiments, the expression cassette encodes a fusion protein comprising a HOA meganuclease and a protein degradation signal, e.g., a PEST sequence. In certain embodiments, the expression cassette encodes a HOA meganuclease and at least one target sequence. In certain embodiments, the expression cassette comprises a mutated target sequence. In certain embodiments, the expression cassette comprises a fusion protein comprising a HOA meganuclease and a protein degradation signal and at least one target sequence. Such expression cassettes can be delivered via a viral or non-viral vector. In certain embodiments, the expression cassette can be delivered using LNP. In certain embodiments, the disorder is primary hyperoxaluria (PH1).

[0124] In one aspect, a method is provided for treating a patient having a disorder associated with a defect in the transthyretin (TTR) gene using a self-modulating nuclease expression cassette comprising a meganuclease that recognizes a site within the human TTR gene as described herein. In certain embodiments, the expression cassette encodes a fusion protein comprising a TTR meganuclease and a protein degradation signal, such as a PEST sequence. In certain embodiments, the expression cassette encodes a TTR meganuclease and at least one target sequence. In certain embodiments, the expression cassette comprises a mutant target sequence. In certain embodiments, the expression cassette comprises a fusion protein comprising a TTR meganuclease and a protein degradation signal and at least one target sequence. Such expression cassettes can be delivered via viral or non-viral vectors. In certain embodiments, the expression cassette can be delivered using LNP. In certain embodiments, the disorder is TTR-associated hereditary amyloidosis.

[0125] In another aspect, a method is provided for treating a patient having a disorder associated with a defect in the apolipoprotein C-II (APOC3) gene using a self-modulating nuclease expression cassette comprising a meganuclease that recognizes a site within the human APOC3 gene as described herein. In certain embodiments, the expression cassette encodes a fusion protein comprising an APOC3 meganuclease and a protein degradation signal, such as a PEST sequence. In certain embodiments, the expression cassette encodes an APOC3 meganuclease and at least one target sequence. In certain embodiments, the expression cassette comprises a mutant target sequence. In certain embodiments, the expression cassette comprises a fusion protein comprising an APOC3 meganuclease and a protein degradation signal and at least one target sequence. Such expression cassettes can be delivered via viral or non-viral vectors. In certain embodiments, the expression cassette can be delivered using LNP.

[0126] In one aspect, a method is provided for treating a patient having a disorder associated with a branched-chain a-keto acid dehydrogenase complex (BCKDC) El alpha gene defect using a self-modulating nuclease expression cassette comprising a meganuclease that recognizes a site within the human BCKDC El alpha gene. In certain embodiments, the expression cassette encodes a fusion protein comprising a BCKDC meganuclease and a protein degradation signal, such as a PEST sequence. See WO 2020 / 056155 A2. In certain embodiments, the expression cassette encodes a BCKDC meganuclease and at least one target sequence. In certain embodiments, the expression cassette comprises a mutated target sequence. In certain embodiments, the expression cassette comprises a fusion protein comprising a BCKDC meganuclease and a protein degradation signal and at least one target sequence. Such expression cassettes can be delivered via a viral or non-viral vector. In certain embodiments, the expression cassette can be delivered using an LNP. In certain embodiments, the disorder is maple syrup urine disease.

[0127] In certain embodiments, nucleases other than meganucleases targeting any of the above genes are contemplated.

[0128] In certain embodiments, the nuclease expression cassettes, non-viral vectors, viral vectors (e.g., rAAV) as described herein can be administered for gene editing in a patient. In certain embodiments, the methods are applicable to non-embryonic gene editing. In certain embodiments, the patient is an infant (e.g., from birth to about 9 months). In certain embodiments, the patient is older than an infant, e.g., 12 months or older.

[0129] In certain embodiments, the pharmaceutical compositions as described herein can be administered for gene editing in a patient. In certain embodiments, the methods are applicable to non-embryonic gene editing. In certain embodiments, the patient is an infant (e.g., from birth to about 9 months). In certain embodiments, the patient is older than an infant, e.g., 12 months or older.

[0130] As used herein, “a” or “an” or “the” can mean one (kind) or more than one. For example, “a” cell can mean a single cell or a plurality of cells.

[0131] In certain embodiments, the term "meganuclease" refers to an endonuclease that binds double stranded DNA at a recognition sequence of greater than 12 base pairs. Preferably, the recognition sequence of a meganuclease of the present application is 22 base pairs. The meganuclease can be an endonuclease derived from I-Crel, and can refer to engineered variants of I-Crel modified with respect to, for example, DNA binding specificity, DNA cleavage activity, DNA binding affinity, or dimerization properties relative to native I-Crel. Methods for generating such I-Crel modified variants are known in the art. See, e.g., WO 2007 / 047859). A meganuclease as used herein binds to double stranded DNA as a heterodimer. The meganuclease can also be a "single-chain meganuclease", in which a pair of DNA binding domains are joined into a single polypeptide using a peptide linker. The term "homing endonuclease" is synonymous with the term "meganuclease". See WO 2018 / 195449, incorporated herein in its entirety, which describes certain PCSK9 meganucleases.

[0132] As used herein, the term "specificity" means the ability of a meganuclease to recognize and cleave a double stranded DNA molecule only at a specific sequence of base pairs, referred to as a recognition sequence, or only at a specific set of recognition sequences. The set of recognition sequences will share certain conserved positions or sequence motifs, but can be degenerate at one or more positions. A highly specific meganuclease is capable of cleaving only one or a very few recognition sequences. Specificity can be determined by any method known in the art.

[0133] The abbreviation "Sc" refers to self-complementary. "Self-complementary AAV" refers to constructs in which the coding region carried by the recombinant AAV nucleic acid sequence has been designed to form an intramolecular double-stranded DNA template. Upon infection, the two complementary halves of a scAAV will associate to form a double-stranded DNA (dsDNA) unit, ready for immediate replication and transcription, rather than waiting for cell-mediated synthesis of the second strand. See, e.g., D M McCarty et al., "Self-complementary recombinant adeno-associated virus (scAAV) vectors promote efficient transduction independently of DNA synthesis", Gene Therapy, (August 2001), vol. 8, no. 16, pp. 1248-1254. Self-complementary AAVs are described in, e.g., U.S. Patent Nos. 6,596,535; 7,125,717; and 7,456,683, each of which is incorporated by reference herein in its entirety.

[0134] As used herein, the term "operably linked" refers to both expression control sequences contiguous to a gene of interest and expression control sequences acting in trans or at a distance to control the gene of interest.

[0135] The term "exogenous" as used to describe a nucleic acid sequence or protein means that the nucleic acid or protein is not naturally present in the location in the chromosome or host cell in which it is present. An exogenous nucleic acid sequence also refers to a sequence derived from the same expression cassette or host cell and inserted into the same expression cassette or host cell, but which is present in a non-native state (e.g., different copy number) or under the control of different regulatory elements.

[0136] The term "heterologous" when used with reference to a protein or nucleic acid indicates that the protein or nucleic acid comprises two or more sequences or subsequences that are not found in the same relationship to each other in nature. For example, the nucleic acid is typically recombinantly produced, having two or more sequences from unrelated genes arranged to produce a new functional nucleic acid. For example, in one embodiment, the nucleic acid has a promoter from one gene arranged to direct expression of a coding sequence from a different gene.

[0137] As used herein, the term "host cell" can refer to a packaging cell line in which a vector (e.g., a recombinant AAV) is produced from a production plasmid. In the alternative, the term "host cell" can refer to any target cell in which expression of a transgene is desired. Thus, "host cell" refers to a prokaryotic or eukaryotic cell containing an exogenous or heterologous nucleic acid sequence that has been introduced into the cell by any means, e.g., electroporation, calcium phosphate precipitation, microinjection, transformation, viral infection, transfection, liposome delivery, membrane fusion techniques, high velocity DNA-coated pellets, viral infection, and protoplast fusion. In certain embodiments herein, the term "host cell" refers to cultures of cells of various mammalian species used to assess compositions described herein in vitro. In other embodiments herein, the term "host cell" refers to cells used to produce and package viral vectors or recombinant viruses. In yet another embodiment, the term "host cell" is intended to refer to target cells of an individual being treated in vivo for a disease or condition as described herein. In certain embodiments, the term "host cell" is a liver cell or hepatocyte.

[0138] A "replication-defective virus" or "viral vector" refers to a synthetic or artificial viral particle in which an expression cassette containing a gene of interest is packaged in a viral capsid or envelope, wherein any viral genomic sequence also packaged within the viral capsid or envelope is replication-defective; i.e., it is unable to produce progeny virions, but retains the ability to infect target cells. In one embodiment, the genome of the viral vector does not contain genes encoding enzymes required for replication (the genome can be engineered to be "gutless" - containing only the gene of interest flanked by signals required for amplification and packaging of the artificial genome), but these genes can be supplied during production. Thus, it is considered safe for use in gene therapy because no replication and infection by progeny virions can occur except for the presence of viral enzymes required for replication.

[0139] In the context of nucleic acid sequences, the term "sequence identity," "percent sequence identity," or "percent identity" refers to residues in two sequences that are the same when aligned for maximum correspondence. The length of sequence identity comparison can be over the full length of a genome, the full length of a gene coding sequence, or a fragment of at least about 500 to 5000 nucleotides is desirable. However, identity in smaller fragments, such as at least about nine nucleotides, typically at least about 20 to 24 nucleotides, at least about 28 to 32 nucleotides, at least about 36 or more nucleotides is also desirable. Similarly, "percent sequence identity" of amino acid sequences, protein full lengths, or fragments thereof can be readily determined. Suitably, the fragment is at least about 8 amino acids in length, and can be up to about 700 amino acids in length. Examples of suitable fragments are described herein.

[0140] The term "substantial homology" or "substantial similarity" when referring to an amino acid or fragment thereof indicates that there is amino acid sequence identity in at least about 95% to 99% of the aligned sequences when optimally aligned with appropriate amino acid insertions or deletions with another amino acid (or its complementary strand). Preferably, the homology is over the full length sequence or its protein, such as a cap protein, rep protein, or a fragment thereof that is at least 8 amino acids in length, or more desirably at least 15 amino acids in length. Examples of suitable fragments are described herein.

[0141] The term "highly conserved" means at least 80% identity, preferably at least 90% identity, and more preferably over 97% identity. One skilled in the art can readily determine identity by means of algorithms and computer programs known to those skilled in the art.

[0142] In general, when referring to "identity," "homology," or "similarity" between two different adeno-associated viruses, the "identity," "homology," or "similarity" is determined with reference to "aligned" sequences. "Aligned" sequences or "alignment" refers to a plurality of nucleic acid sequences or protein (amino acid) sequences that, compared to a reference sequence, typically contain corrections for missing or extra bases or amino acids. In examples, AAV alignment is performed using the published AAV9 sequence as a reference point. Alignment is performed using any of a variety of published or commercially available multiple sequence alignment programs. Examples of such programs include "Clustal Omega," "Clustal W," "CAP sequence assembly," "MAP," and "MEME," which are accessible through a website server on the internet. Other sources of such programs are known to those of skill in the art. Alternatively, the Vector NTI utility is also used. There are also a number of algorithms known in the art that can be used to measure the identity of nucleotide sequences, including those contained in the programs described above. As another example, the program Fasta TM in GCG version 6.1 can be used to compare a plurality of nucleotide sequences. Fasta TM provides the best overlap between the query and search sequences and the percent sequence identity. For example, the percent sequence identity between nucleic acid sequences can be determined using Fasta TM as provided in GCG version 6.1 with its default parameters (a word length of 6 and the NOPAM factor of the scoring matrix). A number of sequence alignment programs can also be used for amino acid sequences, such as the "Clustal Omega," "Clustal X," "MAP," "PIMA," "MSA," "BLOCKMAKER," "MEME," and "Match-Box" programs. In general, any of these programs are used with default settings, but a person of skill in the art can vary these settings as desired. Alternatively, a person of skill in the art can utilize another algorithm or computer program that provides at least the same level of identity or alignment as the referenced algorithms and programs. See, e.g., J. D. Thomson et al., Nucl. Acids. Res., "A comprehensive comparison of multiple sequence alignments," 27(13):2682-2690 (1999).

[0143] As used herein, the term "about" refers to a variance of ±10% from a reference integer and values in between. For example, "about" 40 base pairs includes ±4 (i.e., 36 to 44, which includes the integers 36, 37, 38, 39, 40, 41, 42, 43, 44). For other values, especially when referring to percentages (e.g., 90% identity, about 10% difference, or about 36% mismatches), the term "about" includes all values within a range that includes both integers and fractions.

[0144] As used throughout this specification and claims, the terms "comprise", "contain", "include" and variations thereof mean other components, elements, integers, steps, etc. are not excluded. In contrast, the term "consist of and variations thereof, are not inclusive of other components, elements, integers, steps, etc.

[0145] Unless otherwise defined in this specification, technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary skill in the art and reference is made to the disclosure of the patent document for a general guide to the meanings of the terms used in this application.

[0146] Examples

[0147] We hypothesized that low intracellular levels of meganucleases are sufficient for on-target genome editing, and that beyond this threshold, the likelihood of off-target editing increases. To test this hypothesis, we developed a self-inactivating or "suicide" system to limit the expression of meganucleases.

[0148] In summary, limiting meganuclease expression by self-inactivation through insertion of target sequences and / or addition of protein degradation signals reduces off-target activity without compromising on-target potency. Inclusion of this suicide system in gene editing methods increases the safety profile of AAV-delivered genome editing nucleases.

[0149] Example 1-

[0150] In this example, engineered meganucleases described in WO 2018 / 195449 are used to illustrate the application. These meganucleases are engineered to recognize and cleave PCS 7-8 recognition sequences. The PCS 7-8 recognition sequences are located within the PCSK9 gene. These engineered meganucleases include a first subunit that includes a first hypervariable (HVR1) region and a second subunit that includes a second hypervariable (HVR2) region. In addition, the first subunit binds to a first recognition half-site (e.g., PCS7 half-site) in the recognition sequence and the second subunit binds to a second recognition half-site (e.g., PCS7 half-site) in the recognition sequence. In embodiments where the engineered meganuclease is a single-chain meganuclease, the first and second subunits can be oriented such that the first subunit including the HVR1 region and binding to the first half-site is positioned as the N-terminal subunit and the second subunit including the HVR2 region and binding to the second half-site is positioned as the C-terminal subunit. In alternative embodiments, the first and second subunits can be oriented such that the first subunit including the HVR1 region and binding to the first half-site is positioned as the C-terminal subunit and the second subunit including the HVR2 region and binding to the second half-site is positioned as the N-terminal subunit. See, e.g., Table 1 of WO 2018 / 195449.

[0151] We developed an AAV vector expressing M2 PCSK9 by inserting a 22 bp meganuclease target sequence after the promoter. With this design, the expressed M2 PCSK9 should edit the PCSK9 gene and cleave the AAV vector genome immediately after the promoter, preventing further transcription of the meganuclease transgene. We constructed alternative vectors by inserting additional target sequences before the polyA sequence or by inserting a mutated target sequence after the promoter. We also included a PEST sequence in frame with the M2 PCSK9 meganuclease, as this sequence should target the transgene protein for proteasomal degradation.

[0152] We intravenously administered AAV9 vectors expressing human PCSK9 to immunodeficient RAG knockout mice. Two weeks later, we re-administered the mice with the suicide system vectors. All versions of the suicide system reduced PCSK9 levels in serum, albeit to different extents and with different time courses following vector administration. As expected, M2 PCSK9 created indels in the target sequence in the PSCK9 gene and at the target sequence when present in the AAV genome. For some of the suicide system vectors, the on-target editing potency was comparable to that obtained with the parental AAV-M2 PCSK9 vector, with reduced protein expression determined by western blot and 20-fold reduction in off-target activity at 9 weeks post vector administration.

[0153] Plasmids

[0154] All constructs are based on AAV plasmids containing: AAV inverted terminal repeats (ITRs), human thyroid hormone binding globulin (TBG) promoter, Promega intron, PCS7-8L.197 (also known as ARCUS2 or M2 PCSK9) gene, woodchuck hepatitis virus (WHP) post-transcriptional regulatory element (WPRE), bovine growth hormone (bGH) polyA signal and a second AAV ITR sequence (this plasmid has been described in the aforementioned publication (Nat. Biotechnol. 2018 Sep; 36(8): 717-725).

[0155] Plasmids for AAV production:

[0156] • pAAV.TBG.PI.PCS7-8L.197.bGH (AAV.M2PCSK9): The WPRE sequence was removed from the previously described plasmid (Nat. Biotechnol. 2018 Sep; 36(8): 717-725). The final plasmid contains the TBG promoter, a synthetic intron, the coding sequence for M2PCSK9 (I-Cre-I engineered meganuclease) and a bovine growth hormone polyadenylation sequence. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 9. The amino acid sequence of M2PCSK9 is shown in SEQ ID NO: 10.

[0157] • pAAV.TBG.AP-T0.PI.PCS 7-8L.197.bGH (AAV.Target.M2PCSK9): The M2PCSK9 target sequence (SEQ ID NO: 5:

[0158] 5'-TGGACCTCTTTGCCCCAGGGGA-3') was cloned after the promoter sequence. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 11.

[0159] • pAAV.TBG.AP-T0.PI.PCS 7-8L.197-PEST.bGH (AAV.Target.M2PCSK9+PEST): The vector containing the target sequence as above and a proline-glutamate-serine-threonine (PEST)-rich sequence from mouse ornithine decarboxylase was cloned in frame with the M2PCSK9 coding sequence. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 12.

[0160] • pAAV.TBG.AP-T0.PI.PCS 7-8L.197-PEST.AP-T0.bGH (AAV.2x Target. M2 PCSK9 + PEST): Same as the previous vector (containing the target sequence after the promoter and the PEST sequence in frame with M2 PCSK9) plus an additional M2 PCSK9 target sequence cloned before the polyA signal. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 13.

[0161] • pAAV.TBG.AP-T8VL.PI.PCS 7-8L.197-PEST.bGH (AAV.Mutant Target. M2 PCSK9 + PEST): The mutant target sequence (5'-TTGCCCTTTTTATTCCCAGGGA-3') was cloned immediately after the promoter (similar to the AAV8.Target. M2 PCSK9 + PEST construct) replacing the parental target sequence (5'-TGGACCTCTTTGCCCCAGGGGA-3'). The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 14.

[0162] • pAAV.TBG.PI.PCS 7-8L.197-PEST.bGH (AAV.M2 PCSK9 + PEST): The PEST sequence was cloned in frame with the M2 PCSK9 coding sequence. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 15.

[0163] • pAAV.TBG.AP-T8VL.PI.PCS 7-8L.197.bGH (AAV.Mutant Target. M2 PCSK9): The mutant target sequence was cloned immediately after the promoter. The sequence of the expression cassette from this plasmid is shown in SEQ ID NO: 16.

[0164] To construct and produce the AAV suicide vectors, the WPRE element was removed. The PCS7-8L.197 target sequence (SEQ ID NO: 5 - TGGACCTCTTTGCCCCAGGGGA) was cloned after the TBG promoter and before the proUGI intron element. An additional target sequence was cloned after the M2 PCSK9 gene and before the polyA signal. The mutant target sequence (SEQ ID NO: 6: TTGCCCTTTTTATTCCCAGGGA) was identified in the GUIDE-Seq experiment in LLC-MK2 cells (Nat. Biotechnol. 2018 Sep; 36(8): 717-725) as a low-level off-target sequence for M2 PCSK9.

[0165] The ornithine decarboxylase (ODC) sequence rich in proline-glutamic acid-serine-threonine (PEST) is: SEQ ID NO:3: aagcttagcc atggcttccc gccggaggtg gaggagcagg atgatggcac gctgcccatgtcttgtgccc aggagagcgg gatggaccgt caccctgcag cctgtgcttc tgctaggatcaatgtgtagtaa) encodes the amino acid sequence: KLSHGFPPEVEEQDDGTLPMSCAQESGMDRHPAACASARINV (SEQ ID NO:4). This PEST sequence was obtained from mouse cells and cloned within the sequence frame of the M2PCSK9 gene in the vector.

[0166] AAV vectors were produced using the triple transfection technique as previously described.

[0167] Mouse experiment

[0168] Male 6- to 8-week-old Rag1 KO mice (The Jackson Laboratory) were administered 3.5 × 10⁻⁶ AAV serotype 9, encoding the human proprotein convertase subtilisin / kexin type 9 enzyme (AAV9.hPCSK9), via single intravenous tail injection. 10 One genome copy (GC). Two weeks later, 1×10⁻⁶ GC was injected via a single intravenous tail injection. 11 Or 1×10 12 GC encoding the M2PCSK9 nuclease or AAV serotype 8, the AAV suicide vector. Serum samples were collected weekly until the end of the study. One subgroup of mice was euthanized, and livers were collected at 4 or 9 weeks post-AAV9.hPCSK9 administration.

[0169] Non-human primate (NHP) experiments

[0170] Administer 6×10 intravenously to rhesus macaques 12 GC / Kg of AAV.M2PCSK9, AAV.target.M2PCSK9 and AAV.mutant target.M2PCSK9-PEST or 3×10 13 AAV mutant target M2PCSK9 with GC / Kg. Peripheral blood mononuclear cells (PBMCs) and serum samples were obtained at different time points before and after vector administration. Liver biopsies were collected on day 18 after vector administration. All blood tests including hPCSK9 measurements were performed as previously described (Nature Biotechnology, Sep 2018; 36(8):717-725).

[0171] Insertion and Missing Analysis

[0172] Indels in AAV9.hPCSK9 in the mouse experiment, in the AAV suicide vector, and in the target region in the PCSK9 gene, AAV.target.M2PCSK9, and AAV.mutant target.M2PCSK9-PEST vector were quantified as previously described (Nat. Biotechnol. 2018 Sep; 36(8): 717-725). Primers used for this assay are shown in Table 1.

[0173]

[0174] AMP-Seq analysis

[0175] Indels in the PCSK9 target region and ITR integration in the NHP experiment were determined by AMP-Seq analysis as previously described (Nat. Biotechnol. 2018 Sep; 36(8): 717-725).

[0176] Off-target identification and characterization

[0177] ITR-Seq was performed in the liver from mice and NHPs as described in Example 3 below.

[0178] Primers for the indicated off-target positions were designed for indel % calculation as previously described (Nat. Biotechnol. 2018 Sep; 36(8): 717-725).

[0179] Results

[0180] We evaluated whether the mutant target sequence, the PEST signal, or a combination of these elements in the M2PCSK9-expressing AAV vector caused a reduction in off-target activity of the nuclease without impairing the target activity in it. Figure 1 A schematic representation of the AAV vectors tested is shown. Figure 2A A timeline of the mouse and NHP studies described herein is shown. Figure 2B Low editing in the AAV genome occurs during AAV production of the suicide vector is shown.

[0181] To test the in vivo efficacy of these vectors, Rag1 KO mice were first injected with an AAV vector expressing hPCSK9 (as the mouse genome does not contain the M2PCSK9 target sequence), two weeks later, these mice were administered the AAV suicide vector at 1 x 1011vg / kg. The mice were euthanized two weeks (4 weeks from the first vector injection) or seven weeks (9 weeks total) later and the livers were collected. 11 GC / mouse. Two weeks (4 weeks from the first vector injection) or seven weeks (9 weeks total) later, the mice were euthanized and the livers were collected.

[0182] The region encompassing the M2 PCSK9 target sequence in the AAV.hPCSK9 vector was amplified by PCR, and the amplicon was analyzed by next-generation sequencing and bioinformatic analysis to determine the percentage of AAV.hPCSK9 derived amplicons containing an insertion or deletion (indel) in the target region Figure 3A ). All AAV suicide vectors tested induced indels in the AAV.hPCSK9 locus at both weeks 4 and 9. Additionally, we assessed whether the M2 PCSK9 nuclease also induced indels in the target or mutant target sequence present in the AAV vectors expressing the M2 PCSK9 nuclease. We found evidence of indels in the target region for all vectors containing the target sequence. Indels were present in vectors containing both the mutant target sequence and the PEST sequence, and the percentage was lower in vectors containing the mutant target sequence but no PEST sequence Figure 3B .

[0183] After assessing the on-target activity of the AAV suicide vectors, we determined the off-target activity of M2 PCSK9 when its expression was mediated by these vectors. We used a technique called ITR-Seq to determine the genomic double-strand breaks derived from the targeting and off-target activity of the M2 PCSK9 nuclease (BMC Genomics. 2020 Mar 17;21(l):239). ITR-Seq analysis of liver DNA Figure 4A identified approximately 160 off-target sites in mice treated with AAV.M2PCSK9, but the number of off-targets was reduced to 26 for the AAV suicide constructs, except for AAV.M2PCKS9+PEST (average of 52 off-targets) and AAV.mutant target.M2PCSK9 (128 off-targets).

[0184] We selected a subset of high-ranking off-targets from the identified off-targets for further analysis. We designed primers specific to amplify DNA regions encompassing these off-targets and calculated the % indel Figure 4B . The % indel for the off-target sites was reduced in mice treated with the AAV suicide vectors compared to mice treated with the AAV.M2PCSK9 vector. Similar to what was observed in the number of identified off-targets, the % indel for the selected off-targets was similar for mice treated with AAV.M2PCSK9 and AAV.mutant target.M2PCSK9 Figure 4B , indicating that the mutant target sequence itself was not sufficient to mediate a reduction in M2 PCSK9 off-target activity.

[0185] For further testing in non-human primates (NHP), in addition to the parental AAV M2PCSK9 vector, we selected two AAV suicide vectors with high target activity and reduced off-target activity: AAV.target.M2PCSK9 and AAV.mutant target.M2PCSK9+PEST. A schematic representation of these vectors is shown in [illustration / illustration]. Figure 1 middle.

[0186] Administer 6×10 intravenously to NHP 12 AAV at GC / kg dose and at higher doses (3×10) 13 AAV mutation target M2PCSK9+PEST (GC / kg). Liver biopsies were collected from treated NHP on days 18 and 128. Some NHP studies are ongoing and therefore day 128 biopsies have not yet been collected.

[0187] All tested AAV vectors induced insertions and deletions in the expected target regions of the PCSK9 gene, as detected by Amplicon-Seq or AMP-Seq methods (respectively). Figure 5C and Figure 5D Editing induction in the PCSK9 target region, as previously observed (Nature Biotechnology, September 2018; 36(8):717-725), also led to a decrease in LDL levels at different time points after vector injection (Figures 7 and 8, respectively).

[0188] Because editing of the target sequence present in the AAV suicide vector could potentially promote AAV vector degradation, leading to a subsequent decrease in transgenic RNA levels, we quantified the AAV genomic replica (GC) and M2CPSK9 RNA in liver biopsy samples at day 18 (Figure 9). As expected, AAV GC was quantified using 3 × 10⁻⁶ PCRs. 13 GC / kg dose of AAV. Mutant target + PEST treatment in NHP compared with 6×10 12 The remaining portion of the GC / kg treatment was higher; the amount of GC was similar to that of NHP treated with the same dosage. Even at 3×10 13 In GC / kg treated NHP, the amount of RNA was similar across all tested doses.

[0189] Although the editing of the target is similar ( Figure 6), but the number of off-target sites identified by ITR-Seq varied among the groups. For the AAV.M2PCSk9 group, the range was between 41 and 263 off-target sites (average 132). At the same dose, the AAV. Mutant Target.M2PCSK9+PEST group had 34 off-targets, and for the AAV. Target.M2PCSK9, the range was between 34 and 62 off-targets (average 48).

[0190] The basic results obtained in mouse and NHP experiments indicate that the AAV suicide system reduces M2PCSK9 off-target activity Figure 5E ), while preserving on-target activity.

[0191] Example 2 - TTR

[0192] Using the techniques described in Example 1, a TTR meganuclease or TTR meganuclease-PEST fusion protein recognizing the following site was utilized: SEQ ID NO: 7 - GCTGGACTGGTATTTGTGTCTG. See Figure 6 .

[0193] Example 3 - ITR-Seq: Next-generation sequencing assay to identify in vivo genome-wide DNA editing sites following genome editing

[0194] The publication Breton et al., ITR-Seq, a next-generation sequencing assay, identifies genome-wide DNA editing sites in vivo following adeno-associated viral vector-mediated genome editing, BMC Genomics, (2020): 21:239 is incorporated herein by reference in its entirety.

[0195] Materials and Methods

[0196] Animal Studies

[0197] All animal procedures were performed in accordance with protocols approved by the Institutional Animal Care and Use Committee of the University of Pennsylvania.

[0198] Rhesus macaque studies

[0199] DNA samples from previously published studies 15 AAV8 vectors driving expression of meganuclease M1 PCSK9 (AAV8.TBG.M1 PCSK9.WPRE) or M2 PCSK9 (AAV8.TBG.M2 PCSK9.WPRE) were administered via peripheral vein to rhesus macaques (n=4 for M1 PCSK9-treated animals and n=2 for M2 PCSK9-treated animals). Liver biopsies were performed at 17 and 129 days (for AAV8-M1 PCSK9) or 18 and 128 days (for AAV8-M2 PCSK9) after vector administration 15 . As untreated controls, we used DNA extracted from peripheral blood mononuclear cell (PBMC) samples collected prior to vector administration 15 .

[0200] To measure nuclease-independent AAV integration events, we analyzed liver DNA samples from previously published studies 17 . Briefly, one-week or one-month old male rhesus macaques were administered AAV8.TBG.EGFP at a dose of 3x1011 12 GC / kg. Animals were euthanized and livers were collected after vector administration.

[0201] Mouse studies

[0202] C57BL / 6J mice were co-administered via temporal vein injection with AAV expressing SaCas9 (AAV8.TBG.hSaCas9.bGH), LbCpf1 (AAV8.ABP2.TBG-S1.hLbCpf1.bGH), or AsCpf1 (AAV8.ABPS2.TBG-S1.hAsCpf1.PA75) at a dose of 3x1011 11 GC / mouse, and vectors expressing specific sgRNAs (AAV8.U6.sgRNA.mASS1.donor (mASS1)) or non-targeting sgRNAs (AAV8.U6.sgRNA-ctrl.mASS1.donor (mASS1)) as controls at a dose of 2x1011 12 GC / mouse. Mice were euthanized and livers were collected at 21 days after vector administration.

[0203] Additional neonatal mice (n=2 per group) were co-administered with vectors expressing SaCas9 or LbCpf1 as described above at a dose of 10 11 or 3x1011 11 GC / mouse, and vectors expressing specific sgRNAs (AAV8.U6.sgRNA.mASS1.donor (mASS1)) or non-targeting sgRNAs (AAV8.U6.sgRNA-ctrl.mASS1.donor (mASS1)) as controls at a dose of 1012 GC / mouse dose of a second vector expressing an ASS1 -specific sgRNA and a human coagulation factor IX (hFIX) transgene (AAV8.U6.sgRNA.mASS1.TBG.hFIX). Livers were collected 70 days after vector administration.

[0204] ITR-Seq

[0205] The developed ITR-Seq protocol is a modified version of the Anchored PCR reaction 18、19 , where a single primer is designed to anneal to the ITR sequence and amplify outwards from the ITR sequence Figure 7C . After ITR integration in the DNA, the primer can be used to amplify the junction of the host genome with the inserted vector ITR sequence Figure 7B , 7C . To allow for the full denaturation of the ordered secondary structure of the integrated ITR, a high annealing temperature of 69 °C and longer adaptor-specific primers were designed.

[0206] Amplifϊers were generated from purified genomic DNA isolated from liver tissue samples. DNA was sheared to an average size of 500 bp using a ME220 Focused-ultrasonicator (Covaris, Woburn, MA), purified at a 0.8x ratio using AMPure beads (Beckman Coulter, Indianapolis, IN), and eluted in 15 μΐ of elution buffer (Qiagen, Hilden, Germany). End-repair was then performed in a total volume of 22.5 μΐ containing: 1 μΐ of 5 mM dNTP mix (Thermo Fisher Scientific, Waltham, MA), 2.5 μΐ of 10x SLOW ligation buffer (Enzymatics, Beverly, MA), 2 μΐ of End-repair mix (low concentration; Enzymatics, Beverly, MA), 2 μΐ of 10x Taq polymerase buffer (New England BioLabs, Ipswich, MA), 0.5 μΐ of non-thermal start Taq polymerase (no MgCl2; Invitrogen, Carlsbad, CA), 0.5 μΐ of nuclease-free water (Life Technologies, Waltham, MA), and 14 μΐ of 400 ng of sheared genomic DNA. The mixture was incubated at 12°C for 15 min, 37°C for 15 min, 72°C for 15 min, and held at 4°C. Unique Y adapters, with molecular index tags annealed to MiSeq universal adapters (Illumina, San Diego, CA), were ligated to the end-repaired DNA in a mixture containing: 1 μΐ of 10 μΜ annealed A01-A16 Y adapters, 2 μΐ of T4 DNA ligase (Enzymatics, Beverly, MA), and 22.5 μΐ of pre-end-repaired DNA. The ligation program was 16°C for 30 min, 22°C for 30 min; the reaction was held at 4°C. The DNA was then purified by AMPure beads (Beckman Coulter, Indianapolis, IN) at a 0.7x ratio.End-repaired Y-adaptor ligated DNA fragments were amplified by PCR using ITR-specific primers and adaptor-specific primers (A01-A16_P5_FWD primers) in the following mix (amount per sample): 11.9 μΐ of nuclease-free water; 3 μΐ of 10X Taq Polymerase Buffer (without MgCl2, Invitrogen, Carlsbad, CA); 0.6 μΐ of 10 mM dNTP mix (Thermo Fisher Scientific, Waltham, MA); 1.2 μΐ of 50 mM MgCl2(Invitrogen, Carlsbad, CA); 0.3 μΐ of 5 U / μΐ Platinum Taq Polymerase (Invitrogen, Carlsbad, CA); 1 μΐ of 10 μΜ GSP_ITR3.AAV2 primer; 1.5 μΐ of 0.5 M TMAC (Sigma-Aldrich; St. Louis, MO); 0.5 μΐ of 10 μΜ A01-A16_P5_FWD primer, where the primer number matches the adaptor number (e.g., A01_P5_FWD primer is used with A01 Y-adaptor); and 10 μΐ of pre-purified DNA. The PCR program was 1 cycle of 95 °C for 5 min, 30 cycles of 95 °C for 30 s, 69 °C for 1 min, and 72 °C for 30 s, 1 cycle of 72 °C for 5 min; hold at 4 °C. The PCR products were purified using 0.7X AMPure beads (Beckman Coulter, Indianapolis, IN) and resuspended in 15 μΐ of elution buffer (Qiagen, Hilden, Germany).

[0207] NGS libraries were prepared by PCR in the following mix (amount per sample): 5.4 μΐ of nuclease-free water (Life Technologies, Waltham, MA); 3 μΐ of 10x Taq polymerase buffer (without MgCl2; Invitrogen, Carlsbad, CA); 0.6 μΐ of 10 mM dNTP mix (Thermo Fisher Scientific, Waltham, MA); 1.2 μΐ of 50 mM MgCl2(Invitrogen, Carlsbad, CA); 0.3 μΐ of 5 U / μΐ Platinum Taq polymerase (Invitrogen, Carlsbad, CA); 1 μΐ of 10 μΜ GSP_ITR3 primer; 1.5 μΐ of 0.5 M TMAC (Sigma-Aldrich, St. Louis, MO); 0.5 μΐ of 10 μΜ A01-A16_P5_FWD primer, where primer quantity was matched to adaptor quantity; 1.5 μΐ of 10 μΜ p701-16 primer; and 15 μΐ of pre-purified DNA (containing AMPure beads for the preceding PCR purification step). The PCR program was 1 cycle of 95 °C for 5 min; 10 cycles of 95 °C for 30 s, 75 °C for 2 min (-1 °C / cycle), and 72 °C for 30 s; 15 cycles of 95 °C for 30 s, 69 °C for 1 min, and 72 °C for 30 s; 1 cycle of 72 °C for 5 min; 4 °C hold. PCR products were purified using 0.7x AMPure beads (Beckman Coulter, Indianapolis, IN) and resuspended in 25 μΐ of elution buffer. Dual-indexed sequencing libraries were sequenced on an Illumina MiSeq cartridge (Illumina, San Diego, CA) v2 RGT kit 300 cyc PE-Bx 1 / 2, generating 2x150 bp paired-end reads. v2 RGT kit 300 cyc PE-Bx 1 / 2; Illumina, San Diego, CA) generating 2x150 bp paired-end reads.

[0208] Using Je 20 Sample demultiplexing and unique molecular identifier (UMI) tagging was performed on the raw fastq files, allowing for up to one mismatch on either index. FASTX Barcode Cutter (http: / / hannonlab.cshl.edu / fastx_toolkit / , allowing for up to five mismatches), fastq-pair (https: / / github.com / linsalrob / EdwardsLab / ), and FASTP 21 Read pairs were identified where read 2 started with the designed primer sequence, plus an additional 20 bp of flanking AAV2 ITR sequence Figures 7A to 7C). Selected read pairs were mapped to the reference genome for each sample (MM10 for mouse and RheMac8 for rhesus macaque samples) using NovoAlign (Novocraft, Selangor, Malaysia). Soft-clipped portions of reads were then mapped to the AAV2 reference genome using NovoAlign (Novocraft, Selangor, Malaysia). Only those original mapped reads found to contain soft-clipped read portions mapping to the AAV2 ITRs with mapping quality values greater than or equal to 30 were used to identify the genomic DNA insertion sites by using BEDtools 20 UMI read merging was performed to generate UMI-merged BAM files. Chimeric reads across ITR genomic DNA insertion sites were identified by determining split read junctions for each read using SE-MEI (https: / / github.com / dpryan79 / SE-MEI). Soft-clipped portions of reads were then mapped to the AAV2 reference genome using NovoAlign (Novocraft, Selangor, Malaysia). Only those original mapped reads found to contain soft-clipped read portions mapping to the AAV2 ITRs with mapping quality values greater than or equal to 30 were used to identify the genomic DNA insertion sites by using BEDtools 22、23 ITR integration sites found within 50 bp windows were merged into a single ITR integration site to identify ITR integration sites. Only those sites containing ITR integrations on both the forward and reverse strand orientations were considered as identified off-target sites. EMBOSS programs for semi-global alignment 24 were used to pair align the genomic DNA sequence under each identified ITR integration site to the on-target DNA sequence motifs from both the reverse and forward strand orientations; this allowed for the assessment of sequence homology and precise ITR integration site recognition by the nuclease of interest.

[0209] Results and discussion

[0210] Development of ITR-Seq assay to assess meganuclease activity in non-human primates

[0211] The development of NGS-based assays to identify and rank nuclease-induced DSBs following in vivo gene editing significantly advances our ability to assess the safety and efficacy of genome editing therapies translated into human clinical trials. Researchers have developed various methods to identify and quantify the targeting and off-target activity of genome editing nucleases to better understand the elements governing nuclease specificity and to improve the safety profile of these therapies 25 . One can study nuclease specificity and activity by first identifying DSBs in cultured cells or in created animal models as a result of their nuclease activity 26、27 . Some methods for determining this nuclease specificity include cell-free methods such as Site-Seq 28 , Digenome-seq 29 , and Circle-Seq 30 . Some in-vitro-based methods include GUIDE-Seq19 and integration-deficient lentiviral vector capture (IDLV) 31、32 However, these in vitro assays can not accurately predict the amount and ratio of off-target activity in vivo, as the conditions used in the in vitro assays do not represent the DNA accessibility and nuclease concentration present in the target organs of animal models.

[0212] Accordingly, we sought to develop a method for unbiased whole-genome identification of ITR integrations. By using AAV ITRs as a tag for identifying DSBs, we can measure off-target activity of genome-editing nucleases in vivo. Our approach is based on the aforementioned studies demonstrating integration of AAV ITR sequences into host genomic DNA after DSB occurrence 10-14、16、33-36 .

[0213] Recently, our group characterized the genome-editing efficiency of a large- nuclease targeting delivery to the PCSK9 gene in the liver of rhesus macaques using AAV8 vectors. This study showed a stable dose-dependent reduction of PCSK9, and 30% to 40% on-target indel percentage for higher and intermediate levels of the first-generation large-nuclease M1 PCSK9 or intermediate levels of the second-generation large-nuclease M2 PCSK9 15 We isolated edited PCSK9 alleles from liver biopsy samples of rhesus macaques pre-infused with both first- and second-generation large-nuclease constructs (AAV8-M1 PCSK9 and AAV8-M2 PCSK9, respectively) 15 When analyzing these alleles using AMP-Seq and amplicon sequencing assays, we noticed that genomic integration of ITR sequences occurred at a higher frequency 15 .

[0214] Here, we re-analyzed NGS reads generated when characterizing the target region 15 in the AAV8-M1 PCSK9 and AAV8-M2 PCSK9. Our goal was to identify the most common AAV ITR sequences that integrated into the large-nuclease on-target locus in the PCSK9 gene. Based on the absolute frequency peak at position 82 of the AAV2 reference genome Figure 7A ), we determined that the most common base position of ITR integration occurs 5' upstream of the Rep Binding Element (RBE). We used this information to design ITR-specific primers that hybridize 5' upstream of the observed ITR integration start site (shown in red in Figure 7B ). We used this primer in a novel NGS assay based on a modified version of anchor multiplex PCR to identify ITR-genome DNA junctions after insertional mutagenesis Figure 7CWe call this method ITR-Seq.

[0215] To perform sample analysis using ITR-Seq assays, we first isolated DNA from animal tissues treated with an AAV vector expressing nucleases. We cut the DNA and attached it to the Y-adaptor, as described in the previous report. 19 After two rounds of PCR using the ITR-specific and adaptor-specific primers described above, we generated an NGS-compatible library. Following sequencing, we computationally identified the resulting amplicon containing both the amplified ITR sequence and the adjacent genomic DNA sequence. We also determined the location and frequency of genome-wide ITR integration sites. By requiring ITR integration at any specific identified ITR integration site on both the forward and reverse strand orientations, we aimed to further reduce the number of false positives and identify high-confidence ITR integration sites. For each sample, we generated a sorted list of nuclease target sites (ITR-Seq rank) and sorted sites (in descending order) according to the total number of ITR integration events observed at each locus (ITR-Seq reads).

[0216] ITR-Seq identifies off-target sites of nucleases in rhesus macaques.

[0217] Following the development of ITR-Seq assays, we used this technique to further analyze the targeting and off-target effects of a wide range of nucleases in rhesus macaques previously administered AAV8-M1PCSK9 and AAV8-M2PCSK9. Figures 8A to 8C As previously described, rhesus macaques received three doses of AAV8-M1PCSK9 (3 × 10⁻⁶). 13 GC / kg, 6×10 12 GC / kg or 2×10 12 Any one or a single dose of AAV8-M2PCSK9 (6×10 GC / kg) 12 GC / kg) 15 We obtained liver biopsies from all rhesus monkeys on days 17 and 128 post-vector administration to assess targeted and off-target editing. The ITR-Seq report generated at the end of computational analysis of the NGS data included the most probable off-target sequences (based on homology with the intended target sequence); genomic location; and ITR-Seq rank (based on the number of NGS reads mapped to the corresponding locus). Among all samples from rhesus monkeys treated with a wide range of nucleases, the top ITR-Seq rank locus was the intermediate target locus (PCSK9 target sequence TGGACCTCTTTGCCCCAGGGGA, chr1:54708864-54708885). 15For both generations of large-scale nucleases (i.e., AAV8-M1PCSK9 and AAV8-M2PCSK9), the number of off-target sites identified by ITR-Seq depends on both the dose of the applied vector and the sampling time point. The number of off-target sites decreases in a dose-dependent manner over time (e.g., more off-target sites are observed on day 17 than on day 128; see [link to other documentation]). Figure 8A At the same dose, animals treated with the second-generation engineered large-scale nuclease M2PCSK9 had fewer identified off-target sites than animals treated with the first-generation large-scale nuclease M1PCSK9. Overall, at 6 × 10⁻⁶ doses... 12 Following administration of AAV8-M1PCSK9 at a dose of GC / kg, we observed 1,170 distinct off-target events. In contrast, at a dose of 6 × 10⁻⁶, we observed significantly more off-target events. 12 In two rhesus monkeys that received AAV8-M2PCSK9 at a dose of GC / kg, only 194 and 105 off-target events were observed, respectively. Figure 8A ).

[0218] From 3×10 13 and 6×10 12 Off-target effects were identified by ITR-Seq from a randomly selected subgroup of d18 liver samples from rhesus monkeys treated with AAV8-M1PCSK9 at a dose of GC / kg. These were subsequently analyzed by AMP-Seq as previously described. 15、18 The presence of ITR sequences at these selected loci was investigated using gene-specific primers with side-linked identified off-target sequences. AMP-Seq results (data not shown) were obtained by analyzing liver DNA samples from d18 (from rhesus monkeys treated with AAV8-M1PCSK9 at the same dose specified above). For the analyzed DNA, in 24 of the 27 studied loci (for 3 × 10⁻⁶), 13 GC / kg dose) or 21 of them (for 6×10) 12 Readings containing ITR sequences were found at GC / kg dose, with the highest ITR integration percentage (readings containing ITR) corresponding to those off-targets with high ITR-Seq grades. Importantly, since the highest levels of editing were observed at those loci with the highest ITR-Seq grades in both animals, there was also a significant correlation between ITR-Seq grade, ITR integration percentage, and insertion / deletion percentage. For some of these loci, we were unable to detect the integrated ITR sequence as determined by AMP-Seq, likely due to the lower sensitivity of AMP-Seq assays for detecting ITR integration when the ITR-Seq method uses the actual ITR sequence as the amplification starting point. Similarly, we were able to observe ITR sequences at several sites not shown in the ITR-Seq results (e.g., for those at 3 × 10⁻⁶ doses).13 Liver samples treated with AAV8-M1 PCSK9 at a dose of 6xlO11GC / kg, 20:359062-359285 and 7:165269225-165269449). However, these sites were found in animals treated with AAV8-M1 PCSK9 at a dose of 6xlO11 12 ITR-Seq results from animals treated with AAV8-M1 PCSK9 at a dose of 6xlO11GC / kg, indicating that our current protocol did not capture 100% of the ITR integration sites, and thus the sensitivity of the ITR-Seq method can be improved.

[0219] We then annotated the identified ITR integration sites based on the function of the DNA region ( Figure 7B ). Regardless of the number of nuclease target sites identified, the genomic distribution of the target sites (intergenic, intronic, or exonic regions) was reproducible among the macaques administered the meganuclease. In general, the majority of the target sites were found within introns, followed by intergenic regions of the genome ( Figure 8B ). The meganuclease we evaluated here has 22 nucleotide target sites within the PCSK9 gene. Thus, we evaluated the number of conserved nucleotides between the target DNA sequence and the DNA sequence of each off-target site ( Figure 8C ). The distribution of matches to the on-target sequence appears to follow a Gaussian distribution with an average of 15 to 16 nucleotides. This indicates that the majority of the ITR-Seq identified off-target sites have between 6 and 7 mismatches between the targeting DNA sequence motif and the genomic DNA sequence of each target site ( Figure 8C ). To assess whether this level of homology is caused by edits occurring in sequences similar to the meganuclease target sequence or is simply due to chance, we generated 10 million random DNA regions (40 bp in length) within the rhesus macaque genome. We then attempted to identify the sequence most similar to the meganuclease target site using the same algorithm used in the ITR-Seq protocol (dashed line, Figure 8C ). Unlike the ITR-Seq identified sequences, the random sequences share an average of 11 to 12 mismatches with the expected target sequence. This indicates that the ITRs are primarily integrating in targets with some degree of homology to the expected target sequence.

[0220] Comparison of GUIDE-Seq and ITR-Seq identification of nuclease on- and off-targets

[0221] Tools to identify nuclease on- and off-targets in vitro rely on the integration of exogenous DNA at the site of the DSB. This exogenous DNA can be a double-stranded oligodeoxynucleotide (dsODN; for GUIDE-Seq 19 ) or a lentiviral genome (for Integrase Defective Lentiviral Vector Capture or IDLV31、32 ). Amplicon libraries can be constructed by PCR or LAM-PCR (linear amplification mediated-PCR) using adaptors and primers specific for these exogenous sequences. The location of the DSB can be later identified by sequencing the constructed library using NGS, followed by bioinformatic analysis and mapping to the reference genome. One of the current preferred methods for characterizing the off-target activity of nucleases is GUIDE-Seq because 1) it requires a minimal number of components; 2) this software can be readily used to identify off-targets; and 3) it can detect low abundance off-targets. As mentioned previously, in vitro nuclease activity can not be predictive of in vivo nuclease activity, taking into account the dose, the length of the experiment and the cell type used in GUIDE-Seq are different from the characteristics of animal models. GUIDE-Seq allows us to rapidly compare a variety of sgRNA or guide RNA-independent nucleases, such as meganucleases. However, while it is possible to use amplicon sequencing to validate predicted off-targets, one cannot identify novel in vivo generated off-targets if they are not identified in vitro. In addition, these methods require an in vivo validation step, where PCR amplicons generated from primers covering the in vitro predicted off-target region are later sequenced by NGS to calculate the editing rate by indel identification.

[0222] For each meganuclease (M1 PCSK9 and M2 PCSK9), we have previously performed a GUIDE-Seq analysis on LLC-MK2 cells transfected with a plasmid to express the nuclease in order to identify in vitro the off-target sites 15 . We selected M1 PCSK9 and M2 PCSK9 off-target positions from the list of off-targets identified by in vitro GUIDE-Seq. Using Rhesus macaques, we then performed in vivo validation of these predicted off-targets using amplicon sequencing on the off-target loci 15 . Amplicon sequencing 15 and ITR-Seq( Figure 8A ) both exhibit a dose and time-dependent decrease in off-target editing efficiency. We used ITR-Seq to compare these previous results with our assessment of in vivo off-target sites Figures 9A to 9DBy identifying off-target sites that are not homologous to the expected target sequence, we were able to find approximately the same number of off-target sites in two independent experiments using M1PCSK9 (1093 and 1499, respectively, for GUIDE-Seq experiments 1 and 2) or M2PCSK9 (568 and 651, respectively, for GUIDE-Seq experiments 1 and 2). We compared sites identified by the two Guide-Seq in vitro experiments with off-target sites identified by ITR-Seq. For this study in rhesus macaques, we performed ITR-Seq on DNA samples from liver biopsies obtained 17 days after nuclease administration. We compared the results of 3 × 10⁻⁶ off-target sites with those from liver biopsies obtained 17 days after nuclease administration. 13 GC / kg of AAV8-M1PCSK9 ( Figure 9A ), 6×10 12 GC / kg of AAV8-M1PCSK9 ( Figure 9B The dosage of 6×10⁻⁶ was used on rhesus monkeys. 12 GC / kg of AAV8-M2PCSK9 ( Figure 9C and Figure 9D Off-target events identified by GUIDE-Seq and ITR-Seq in two animals. The majority (71.9% to 82.9%) of off-target sites were identified only by ITR-Seq, not by Guide-Seq (see [link to relevant documentation]). Figures 9A to 9D (The colored portion). Interestingly, in animals receiving the second-generation nuclease (M2PCSK9), both ITR-Seq and Guide-Seq identified fewer off-target sites (see [link to documentation]). Figures 9A to 9D (The white part).

[0223] We have previously validated a set of GUIDE-Seq analyses that classify high- and low-level off-target errors based on the number of readings. 15 We performed this by quantifying the percentage of insertions and deletions, which we determined using amplicon sequencing of off-target loci in samples obtained from rhesus monkeys administered a wide range of nucleases at days 17 / 18 and 128 / 129. 15 Off-target sites exhibiting a significantly higher percentage of insertions or deletions than the untreated PBMC DNA control sample will be counted as positive.

[0224] Here, we evaluate whether ITR-Seq can identify positive GUIDE-Seq off-targets. By reanalyzing the same DNA used to validate GUIDE-Seq off-targets, we were able to assess the sensitivity of our method in detecting positive off-target loci. (3 × 10⁻⁶) 13 GC / kg and 6×10 12Amongst the macaques administered AAV8-M1 PCSK9, ITR-Seq failed to identify only two high and two low rank positive GUIDE-Seq off-targets. Amongst the macaques administered AAV8-M2 PCSK9, ITR-Seq correctly identified most positive off-targets and only missed three low rank positive off-targets in one animal and three high rank positive off-targets in the other macaques. Taken together, these results clearly indicate that the vast majority of off-targets in vivo are identified by ITR-Seq but not by GUIDE-Seq. This demonstrates that ITR-Seq provides a more accurate assessment of the in vivo activity of AAV delivered nucleases than amplicon sequencing of GUIDE-Seq predicted off-targets.

[0225] ITR-Seq analysis of DNA samples from macaques administered AAV8-M1 PCSK9 and AAV8-M2 PCSK9 revealed the potential to characterize off-target sites of guide RNA independent nucleases in vivo. ITR-Seq has the potential to provide detailed characterization of genome editing nucleases. Furthermore, in vitro off-target data do not fully understand genome editing nuclease off-target activity in vivo. In fact, ITR-Seq not only identified most of the highest rank off-target sites identified by GUIDE-Seq that we characterized in our previous study 15 but also identified other off-targets Figures 9A to 9D that were not captured by in vitro GUIDE-Seq assays. By using AAV ITRs as a tag for identifying DSBs, we believe that the ITR-Seq method captures true off-targets as several sites were previously identified by Guide-Seq and amplicon sequencing. Furthermore, the sequences identified by ITR-Seq are similar to the expected target sequence. Thus, ITR-Seq can identify novel off-targets that were not previously detected by the combined method of Guide-Seq and subsequent amplicon sequencing.

[0226] In contrast to Guide-Seq, ITR-Seq is not a tool for predicting off-targets. In fact, ITR-Seq identifies novel sites in the genome where targeted and off-target nuclease activity occurs. In fact, this method directly identifies AAV ITR integration sites from DNA samples of animals treated with AAV expressing the nuclease. The identified off-targets can be further analyzed using amplicon sequencing to 1) accurately determine the percentage of editing and ITR integration; and 2) obtain a detailed landscape of nuclease activity in clinical relevant doses and animal models.

[0227] Evaluation of the ITR-Seq assay as a tool for identifying guide RNA dependent nuclease off-target sites in mice

[0228] We tested whether ITR-Seq could detect the targeting and off-target activity of a variety of guide RNA-dependent nucleases commonly used in preclinical studies, namely SaCas9, LbCpfl, and AsCpfl. We also wanted to assess whether the ITR-Seq analysis was compatible with the distinct types of DSB ends formed by these nucleases (blunt ends for SaCas9 and 5' overhangs for Cpf1).

[0229] C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq. 11 C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq. 12 C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq. 11 C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq. 11 C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq. 12 C57BL6 / J mice were co-administered with 3 x 1011vg of AAV8-SaCas9, AAV8-LbCpfl, or AAV8-AsCpfl and 2 x 1011vg of AAV8-ITR-Seq.

[0230] The on-target locus (mASSl) had the highest number of reads in the ITR-Seq analysis, except for one animal treated with AsCpfl-sgRNA2. In addition, all sgRNA-guided nucleases evaluated exhibited high specificity towards the on-target locus with a maximum of 33% on-target indel percentage (Table 2). Importantly, mice treated with AAV8-SaCas9 exhibited the highest frequency of on-target ITR integration events in all treated samples (Table 2 below). Although exhibiting comparable on-target editing, we observed lower on-target ITR integration events in livers treated with AAV8-LbCpfl compared to AAV8-SaCas9. AAV8-AsCpfl had very low editing efficiency at the on-target locus with a maximum indel percentage of 1.22% as assessed by targeted amplicon sequencing (Table 2 below).

[0231]

[0232] All ITR-Seq identified off-target sites for CRISPR nucleases reside within annotated mouse genes, with the most common sites being the on-target locus (data not shown). We identified some low frequency off-targets for the sgRNAs tested, most of which have low homology to the target sequence. The off-target site identified for AAV8-SaCas9 sgRNA1 is within the locus of the known oncogene NOTCH2 (data not shown). Importantly, this off-target was found in mice administered 3 x 1010 11 GC / mouse of AAV8-SaCas9 and 2 x 1010 12 GC / mouse of AAV-sgRNA1. However, this off-target was not present after administration of 10 12 GC / mouse of AAV-sgRNA (twice lower dose). We needed a higher dose of AAV8-sgRNA1 to observe editing at this low abundance off-target (Table 2). The rate of AAV ITR integration can be affected by 1) the homology between the target sequences; or 2) the blunt or overhanging nature of the DNA ends due to nuclease cleavage. We need to conduct detailed studies in the future to fully understand the kinetics of ITR integration.

[0233] Analysis of nuclease-independent events in mice and non-human primates

[0234] ITR integration occurred in the locus targeted by sgRNA among mice injected with AAV8 expressing CRISPR-associated nuclease and sgRNA. Interestingly, ITR integration also occurred in a seemingly sgRNA-independent manner, as we observed AAV ITR integration in control samples in the absence of functional sgRNA. In mice, the most common nuclease-independent ITR integration events occurred in Gm10800 and the albumin gene. ITR integration in the albumin gene is consistent with previous reports showing that the albumin gene is very susceptible to AAV integration 37 . AAV integration has been reported in genes that are transcriptionally active in the liver 38 .

[0235] To assess the frequency and extent of nuclease-independent ITR integration in non-human primates, we assessed the liver DNA of rhesus macaques administered 3 x 1010 12 GC / kg dose of AAV-eGFP. We detected nuclease-independent integration events (but not in the genes identified) as positive ITR insertions. Importantly, our method appears to only identify integrated ITR sequences that are due to in vivo events. We have reached this conclusion based on the following finding: analysis of DNA isolated from PBMCs of naive rhesus macaques in addition to AAV.eGFP DNA did not yield identification of off-targets (data not shown).

[0236] While the observed ITR integration is the result of nuclease-induced DSBs, other researchers have shown that DSBs induced by restriction enzymes, drugs, or gamma radiation can also produce AA V-ITR insertions at these breakpoints 16 Not only for genome editing, but also for gene therapy research, the identification and characterization of nuclease-independent AAV integration sites is of particular importance to form a more complete AAV safety profile.

[0237] As BLISS 6 , BLESS 39、40 , and End-Seq 7 alternative approaches can identify the site of DSBs by capturing the ends of DSBs formed by the activity of nucleases. These methods can accurately identify nuclease-induced DSBs in vivo. However, these methods have the limitation that they can only capture DSBs present at a single time point. This limitation can be overcome in part by analyzing DSBs at multiple time points. Non- limiting LAM-PCR coupled with NGS 41 (analogous method to ITR-Seq) can detect AAV-ITR integration sites. In this technique, linear amplification is first performed using ITR-specific biotinylated primers. This single-stranded DNA is isolated by streptavidin beads and ligated to a known adaptor. Using secondary LTR-specific and adaptor-specific primers, one can amplify, sequence, and identify regions of LTR integration 41 . This method was used to identify integration sites of AAV1-LPLS447X vectors developed for the treatment of lipoprotein lipase deficiency (LPLD) in mouse and human DNA samples 34 While in theory, nrLAM-PCR can detect off-target activity of nucleases, researchers have not performed a direct comparison of ITR-Seq and nrLAM-PCR to understand the strengths and limitations of both techniques.

[0238] Conclusion

[0239] In summary, we have developed an NGS assay to assess the in vivo specificity of any AAV-expressing nuclease. The ITR-Seq assay can be used to identify in vivo AAV insertion sites in gene editing studies. If locus-specific analysis is required, such as calculating the percentage of indels in the identified off-target region or characterizing the integrated ITR sequence, a deep sequencing analysis is recommended. However, it is important to keep in mind that the detection of ITR integration events by ITR-Seq is more sensitive than other NGS-based methods such as AMP-Seq.

[0240] Given the need to use only AAV as a delivery vehicle, ITR-Seq can be used to measure nuclease specificity in nearly any organism with an annotated reference genome. ITR-Seq can be used as a companion diagnostic in preclinical and clinical studies to assess nuclease target sites in longitudinal animal studies with different doses and / or routes of administration. This technology can provide valuable insights into the safety and efficacy of gene editing therapies and ultimately better understand the design of gene editing therapies.

[0241] This same approach can be used to assess AAV integration events in traditional AAV gene therapy studies. While the risk of insertional mutagenesis from AAV gene therapy is considered low 8、42 , it is reasonable to use it to treat rare, debilitating, and lethal diseases, and these potential risks can be more relevant as the field evolves to treat less serious acquired diseases.

[0242] The following list of references corresponds to the references in Example 3.

[0243] Example 4 - AAV.PCSK9 self-inactivating vector

[0244] Figures 10A to 10B The embodiments shown where a target sequence or a mutant target sequence is inserted into the enzyme coding sequence. Figure 10A is an amino acid alignment showing a 10 amino acid nuclear localization signal (NLS) followed by a protein expressed from the target sequence (amino acids 10 to 20 followed by the active portion of the nuclease. The first amino acid sequence M2PCSK9 shows a fragment of the reference meganuclease with its NLS, the space where other constructs will have an insertion and the sequence of the nuclease starting at position 18. M2PCSK9 [reverse target #1] shows the protein encoded when the target sequence is inserted on the antisense strand between the NLS and the enzyme. M2PCSK9 [target #1] shows the protein encoded when the target sequence is inserted on the sense strand between the NLS and the enzyme. M2PCSK9 [target #2] shows the protein encoded when a different target sequence is inserted on the sense strand between the NLS and the enzyme. M2PCSK9 [reverse target #2] shows the protein encoded when the (mutant) target sequence is inserted on the antisense strand between the NLS and the enzyme. [target] M2PCSK9 shows the protein encoded when a 22 bp target sequence replaces the coding sequence after the NLS such that six amino acids of the enzyme are replaced. [reverse target] M2PCSK9 shows the protein encoded when a 22 bp target sequence replaces the coding sequence after the NLS on the opposite strand such that seven amino acids of the enzyme are replaced. Figure 10BDesign of PCSK9 meganuclease fusion proteins with a single protein degradation signal (degron) that is ubiquitin-independent (AR-6) and PCSK9 meganuclease fusion proteins with two protein degradation signals AR-6 and PEST.

[0245] All documents cited in this specification are hereby incorporated by reference as if each had been individually incorporated by reference. While the application has been described with reference to particular embodiments, it will be appreciated that variations and modifications can be made within the spirit and scope of the application. Such variations and modifications are intended to fall within the scope of the appended claims.

[0246] Features of the Sequence Listing:

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253] References

[0254] 1. Kosicki, M., Tomberg, K. & Bradley, A. Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nat Biotechnol 36, 765-771 (2018).

[0255] 2. Sander, J. D. & Joung, J. K. CRISPR-Cas systems for editing, regulating and targeting genomes. Nat Biotechnol 32, 347-355 (2014).

[0256] 3. Waryah, C. B., Moses, C., Arooj, M. & Blancafort, P. Zinc Fingers, TALEs, and CRISPR Systems: A Comparison of Tools for Epigenome Editing. Methods Mol Biol 1767, 19-63 (2018).

[0257] 4. Arnould, S. et al. The I-CreI meganuclease and its engineered derivatives: applications from cell modification to gene therapy. Protein Eng Des Sel 24, 27-31 (2011).

[0258] 5. Silva, G. et al. Meganucleases and other tools for targeted genome engineering: perspectives and challenges for gene therapy. Curr Gene Ther 11, 11-27 (2011).

[0259] 6. Yan, W. X. et al. BLISS is a versatile and quantitative method for genome-wide profiling of DNA double-strand breaks. Nat Commun 8, 15058 (2017).

[0260] 7. Canela, A. et al. DNABreaks and End Resection Measured Genome-wide by End Sequencing. Mol Cell 63, 898-911 (2016).

[0261] 8. Chandler, R.J., Sands, M.S. & Venditti, J.P. Recombinant Adeno-Associated Viral Integration and Genotoxicity: Insights from Animal Models. Hum Gene Ther 28, 314-322 (2017).

[0262] 9. Janovitz, T., Oliveira, T., Sadelain, M. & Falck-Pedersen, E. Highly divergent integration profile of adeno-associated virus serotype 5 revealed by high-throughput sequencing. J Virol 88, 2481-2488 (2014).

[0263] 10. Kotin, R.M. et al. Site-specific integration by adeno-associated virus. Proc Natl Acad Sci U S A 87, 2211-2215 (1990).

[0264] 11. Miller, D.G. et al. Large-scale analysis of adeno-associated virus vector integration sites in normal human cells. J Virol 79, 11434-11442 (2005).

[0265] 12. Nakai, H. et al. Large-scale molecular characterization of adeno-associated virus vector integration in mouse liver. J Virol 79, 3606-3614 (2005).

[0266] 13. Nault, J.C. et al. Recurrent AAV2-related insertional mutagenesis in human hepatocellular carcinomas. Nat Genet 47, 1187-1193 (2015).

[0267] 14. Samulski, R.J. et al. Targeted integration of adeno-associated virus (AAV) into human chromosome 19. EMBO J 10, 3941-3950 (1991).

[0268] 15. Wang, L. et al. Meganuclease targeting of PCSK9 in macaque liver leads to stable reduction in serum cholesterol. Nat Biotechnol 36, 717-725 (2018).

[0269] 16. Miller, D.G., Petek, L.M. & Russell, D.W. Adeno-associated virus vectors integrate at chromosome breakage sites. Nat Genet 36, 767-773 (2004).

[0270] 17. Wang, L. et al. AAV8-mediated hepatic gene transfer in infant rhesus monkeys (Macaca mulatta). Mol Ther 19, 2012-2020 (2011).

[0271] 18. Zheng, Z. et al. Anchored multiplex PCR for targeted next-generation sequencing. Nat Med 20, 1479-1484 (2014).

[0272] 19. Tsai, S.Q. et al. Genome-wide profiling of off-target cleavage by CRISPR-Cas nucleases. Nat Biotechnol 33, 187-197 (2015).

[0273] 20. Girardot, C., Scholtalbers, J., Sauer, S., Su, S.Y. & Furlong, E.E.Je, a versatile suite to handle multiplexed NGS libraries with unique molecular identifiers. BMC Bioinformatics 17, 419 (2016).

[0274] 21. Chen, S., Zhou, Y., Chen, Y. & Gu, J. fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884-i890 (2018).

[0275] 22. Quinlan, A. R. BEDTools: The Swiss-Army Tool for Genome Feature Analysis. Current protocols in bioinformatics 47, 11.12.11-34 (2014).

[0276] 23. Quinlan, A. R. BEDTools: The Swiss-Army Tool for Genome Feature Analysis. Current protocols in bioinformatics 47, 11 12 11-34 (2014).

[0277] 24. Rice, P., Longden, I. & Bleasby, A. EMBOSS: the European Molecular Biology Open Software Suite. Trends Genet 16, 276-277 (2000).

[0278] 25. Martin, F., Sanchez-Hernandez, S., Gutierrez-Guerrero, A., Pinedo-Gomez, J. & Benabdellah, K. Biased and Unbiased Methods for the Detection of Off-Target Cleavage by CRISPR / Cas9: An Overview. Int J Mol Sci 17 (2016).

[0279] 26. Lazzarotto, C. R. et al. Defining CRISPR-Cas9 genome-wide nuclease activities with CIRCLE-seq. Nat Protoc 13, 2615-2642 (2018).

[0280] 27. Yee, J. K. Off-target effects of engineered nucleases. FEBS J 283, 3239-3248 (2016).

[0281] 28. Cameron, P. et al. Mapping the genomic landscape of CRISPR-Cas9 cleavage. Nat Methods 14, 600-606 (2017).

[0282] 29. Kim, D. et al. Digenome-seq: genome-wide profiling of CRISPR-Cas9 off-target effects in human cells. Nat Methods 12, 237-243, page 231 followed by page 243 (2015).

[0283] 30. Tsai, S. Q. et al. CIRCLE-seq: a highly sensitive in vitro screen for genome- wide CRISPR-Cas9 nuclease off-targets. Nat Methods 14, 607-614 (2017).

[0284] 31. Gabriel, R. et al. An unbiased genome-wide analysis of zinc-finger nuclease specificity. Nat Biotechnol 29, 816-823 (2011).

[0285] 32. Wang, X. et al. Unbiased detection of off-target cleavage by CRISPR-Cas9 and TALENs using integrase-defective lentiviral vectors. Nat Biotechnol 33, 175-178 (2015).

[0286] 33. Donsante, A. et al. AAV vector integration sites in mouse hepatocellular carcinoma. Science 317, 477 (2007).

[0287] 34. Kaeppel, C. et al. A largely random AAV integration profile after LPLD gene therapy. Nat Med 19, 889-891 (2013).

[0288] 35. Kotin, R. M., Menninger, J. C., Ward, D. C. & Berns, K. I. Mapping and direct visualization of a region-specific viral DNA integration site on chromosome 19q13-qter. Genomics 10, 831-834 (1991).

[0289] 36. Rosas, L.E., et al. Patterns of scAAV vector insertion associated with oncogenic events in a mouse model for genotoxicity. Mol Ther 20, 2098-2110 (2012).

[0290] 37. Chandler, R.J., et al. Vector design influences hepatic genotoxicity after adeno-associated virus gene therapy. J Clin Invest 125, 870-880 (2015).

[0291] 38. Nakai, H., et al. AAV serotype 2 vectors preferentially integrate into active genes in mice. Nat Genet 34, 297-302 (2003).

[0292] 39. Crosetto, N., et al. Nucleotide-resolution DNA double-strand break mapping by next-generation sequencing. Nat Methods 10, 361-365 (2013).

[0293] 40. Ran, F.A., et al. In vivo genome editing using Staphylococcus aureus Cas9. Nature 520, 186-191 (2015).

[0294] 41. Paruzynski, A. et al. Genome-wide high-throughput integrome analyses by nrLAM-PCR and next-generation sequencing. Nat Protoc 5, 1379-1395 (2010).

[0295] 42. Mingozzi, F. & High, K. A. Therapeutic in vivo gene transfer for genetic disease using AAV: progress and challenges. Nat Rev Genet 12, 341-355 (2011).

Claims

1. A self-regulating gene editing nuclease expression cassette, comprising: (a) A nucleic acid sequence comprising a PCSK9 wide-range nuclease coding sequence operatively linked to a regulatory sequence, the regulatory sequence guiding the expression of the nuclease upon delivery to a host cell having a sequence targeted by the nuclease; (b) At least one nuclease regulatory sequence selected from the target sequence of said nuclease or a mutant target sequence recognized by said nuclease after its expression, wherein said nuclease regulatory sequence is selected from SEQ ID NO:5 and SEQ ID NO:6; and (c) The coding sequence of the peptide degradation signal, wherein the peptide degradation signal is framed with the coding sequence of the nuclease so that the nuclease-peptide degradation signal is expressed as a fusion protein, wherein the protein degradation signal sequence is the PEST sequence of SEQ ID NO:

4.

2. The self-regulating nuclease expression cassette according to claim 1, wherein the first nuclease regulatory sequence is located upstream of the nuclease coding sequence.

3. The self-regulating nuclease expression cassette according to claim 1 or 2, wherein the nuclease regulatory sequence is located downstream of the nuclease coding sequence.

4. The self-regulating nuclease expression cassette according to claim 1 or 2, wherein the nuclease regulatory sequence is located within the nuclease coding sequence.

5. The self-regulating nuclease expression cassette according to claim 1 or 2, wherein the expression cassette comprises at least two nuclease regulatory sequences.

6. The self-regulating nuclease expression cassette according to claim 1 or 2, wherein the expression cassette comprises at least two different nuclease regulatory sequences.

7. A pharmaceutical composition comprising a self-regulating nuclease expression cassette according to any one of claims 1 to 6, and one or more of a carrier, a suspending agent, and / or an excipient.

8. The pharmaceutical composition of claim 7, wherein the expression is in a non-viral delivery system.

9. The pharmaceutical composition of claim 8, wherein the non-viral delivery system is lipid nanoparticles.

10. A viral vector comprising a self-regulating nuclease expression cassette according to any one of claims 1 to 6.

11. A recombinant AAV suitable for gene editing, comprising an AAV capsid and a vector genome packaged within the AAV capsid, the vector genome comprising: (a) The expression box according to any one of claims 1 to 6, and (b) The required AAV inverted terminal repeat sequence is packaged into the capsid.

12. A composition comprising a viral vector according to claim 10 or a recombinant AAV according to claim 11, and one or more of a carrier, a diluent and / or an excipient.

13. Use of the self-regulating nuclease expression cassette according to any one of claims 1 to 6, the composition according to any one of claims 7 to 9, the viral vector according to claim 10, or the rAAV according to claim 11 in the preparation of a medicament for editing a target gene.