Methods of nuclease-initiated homology directed repair and replacement and compositions and uses thereof

Nuclease-initiated HDR methods with modified DNA repair templates and flanking homology arms provide a solution for accurately correcting genetic diseases by replacing genomic sequences with modified nucleic acids, overcoming the limitations of traditional HDR techniques.

WO2026083329A1PCT designated stage Publication Date: 2026-04-23PRECISION BIOSCIENCES INC
View PDF 58 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
PRECISION BIOSCIENCES INC
Filing Date
2025-10-16
Publication Date
2026-04-23

Smart Images

  • Figure IMGF000066_0001
    Figure IMGF000066_0001
  • Figure IMGF000068_0001
    Figure IMGF000068_0001
  • Figure IMGF000070_0001
    Figure IMGF000070_0001
Patent Text Reader

Abstract

The present invention encompasses compositions and methods for performing nuclease-mediated homology-directed repair for the repair and replacement of genomic DNA sequences. The methods and compositions allow for the insertion of exogenous nucleic acid sequences into the genome of a cell.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS OF NUCLEASE-INITIATED HOMOLOGY DIRECTED REPAIR AND REPLACEMENT AND COMPOSITIONS AND USES THEREOF

[0002] FIELD OF THE INVENTION

[0003] The present disclosure pertains to the field of molecular biology and recombinant nucleic acid technology. In particular, the present disclosure pertains to methods of genomic DNA editing using nuclease-initiated homology directed repair.

[0004] REFERENCE TO A SEQUENCE LISTING SUBMITTED ELECTRONICALLY AS AN XML FILE

[0005] The instant application contains a Sequence Listing which has been submitted in XML format via USPTO Patent Center and is hereby incorporated by reference in its entirety. Said XML copy, created on October 18, 2024, is named “P893392050USPl.xml”, and is 26,229 bytes in size.

[0006] BACKGROUND OF THE INVENTION

[0007] With the advent of precision medicine, there is an increased need for methods that can accurately edit the genome. Modification of genomic DNA at precise locations in genes comprising pathogenic mutations can permanently alleviate genetic diseases across many therapeutic areas. One mechanism that allows for modification of genomic DNA is homology directed repair (HDR), which promotes the insertion of DNA templates into the genome at a double strand break (DSB). Typically, such DSBs are generated by a site-specific engineered nuclease, and homologous recombination of the DNA template at the site of the DSB is enabled, in part, by the inclusion of upstream and downstream homology arms that are homologous to sequences flanking the chromosomal location of the DSB. However, insertion of a DNA template at a specified site in the genome is not always sufficient for correction of a genetic disease, or sufficient to achieve a desired gene modifications in vitro or in vivo. Thus, a need remains for novel methods that utilize HDR- based editing to modify genomic sequences.

[0008] SUMMARY OF THE INVENTION

[0009] The present disclosure provides for methods of nuclease-initiated HDR, referred to herein as HDR repair and replacement, that are capable of replacing a specified endogenous genomic sequence with a modified replacement nucleic acid sequence. Such methods enable, for example, the correction of one or more pathogenic mutations in a gene, including nucleotide changes, insertions, and / or deletions. The methods utilize a nuclease that generates a double strand break

[0010] 1

[0011] WBD (US) 4882-8896-8689vl

[0012] P893392050WO (01287) (DSB) and a DNA repair template comprising a replacement nucleic acid sequence flanked by homology arms. By altering the length of the homology arms, and the position of the genomic sequences that the homology arms align with, the present methods enable the replacement of either small (fewer than lObp) or large (>10kb) stretches of genomic DNA that are downstream of a nuclease -induced DSB. Furthermore, the replacement nucleic acid of the methods and compositions herein is a variant of the endogenous genomic sequence it is replacing, comprising a plurality of modified nucleotides and / or modified codons that differ from corresponding endogenous nucleotides or codons in the endogenous genomic sequence being replaced, and encoding the same amino acid sequence as the replaced endogenous genomic sequence with at least one replacement modification (e.g., a nucleotide change, insertion, or deletion).

[0013] Thus, in one aspect, disclosed herein is a cell comprising an exogenous polynucleotide, wherein the cell comprises in its genome a double-strand break (DSB), wherein the exogenous polynucleotide comprises, from 5’ to 3’ a first homology arm, a replacement nucleic acid sequence, and a second homology arm. The first homology arm has homology to a first genomic sequence and is adjacent to the DSB and the second homology arm has homology to a second genomic sequence and is adjacent to the DSB, with the second genomic sequence and the first genomic sequence being on opposite sides of the double strand break and with the DSB and the second genomic sequence separated by an intervening genomic sequence.

[0014] In one aspect, the disclosure provides a cell comprising an exogenous polynucleotide, wherein the cell comprises in its genome a double-strand break, wherein the exogenous polynucleotide comprises, from 5' to 3': (a) a first homology arm; (b) a replacement nucleic acid sequence; and (c) a second homology arm; wherein the first homology arm has homology to a first genomic sequence that is adjacent to the double-strand break, wherein the second homology arm has homology to a second genomic sequence that is not adjacent to the double-strand break, and wherein the double-strand break and the second genomic sequence are separated by an intervening genomic sequence.

[0015] In some embodiments, the first genomic sequence is 5' upstream of the double-strand break. In some embodiments, the second genomic sequence is 3' downstream of the double-strand break. In some embodiments, the first genomic sequence is 5' upstream of the double-strand break and the second genomic sequence is 3' downstream of the double-strand break.

[0016] In some embodiments, the replacement nucleic acid sequence comprises a plurality of modified nucleotides and / or modified codons that differ from corresponding endogenous nucleotides or codons in the intervening genomic sequence, and wherein the replacement nucleic

[0017] 2

[0018] WBD (US) 4882-8896-8689vl

[0019] P893392050WO (01287) acid sequence encodes the same amino acid sequence as the intervening genomic sequence except for at least one replacement modification, wherein the replacement modification comprises: a) one or more nucleotide modifications that result in an altered codon, wherein the altered codon encodes a different amino acid than a corresponding endogenous codon in the intervening genomic sequence, or wherein the altered codon is a stop codon; b) deletion of one or more nucleotides or codons that are present in the intervening genomic sequence; or c) insertion of one or more nucleotides or codons that are not present in the intervening genomic sequence.

[0020] In some embodiments, the replacement nucleic acid sequence has 1 replacement modification. In some embodiments, the replacement nucleic acid sequence has 2 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 3 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 4 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 5 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 6 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 7 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 8 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 9 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 10 replacement modifications.

[0021] In some embodiments, the replacement nucleic acid sequence has up to 10 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 20 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 30 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 40 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 50 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 100 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 200 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 500 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 1000 replacement modifications.

[0022] In some embodiments, the replacement nucleic acid sequence is a variant of the intervening genomic sequence.

[0023] In some embodiments, at least 20% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 30% of the nucleotides in the replacement nucleic acid sequence are

[0024] 3

[0025] WBD (US) 4882-8896-8689vl

[0026] P893392050WO (01287) modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, between about 30% to 40% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 33% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 50% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

[0027] In some embodiments, at least 75% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at least 80% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at least 90% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at least 95% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, all codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

[0028] In some embodiments, each of the modified codons comprises one nucleotide change, or two nucleotide changes, relative to a corresponding codon in the intervening genomic sequence.

[0029] In some embodiments, the replacement nucleic acid sequence has less than 80% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 70% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has between about 60% to 70% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 60% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 50% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 40% homology to the intervening genomic sequence.

[0030] In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 50 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 25 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 10 consecutive nucleotides in common. In

[0031] 4

[0032] WBD (US) 4882-8896-8689vl

[0033] P893392050WO (01287) some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 5 consecutive nucleotides in common.

[0034] In some embodiments, the replacement modification comprises one or more nucleotide modifications that result in an altered codon, wherein the altered codon encodes a different amino acid than a corresponding endogenous codon in the intervening genomic sequence, or wherein the altered codon is a stop codon. In some such embodiments, the altered codon comprises one nucleotide change, or two nucleotide changes, relative to the corresponding endogenous codon in the intervening genomic sequence. In some such embodiments, the double-strand break is present within a mutant gene encoding a mutant protein, wherein the endogenous codon encodes a mutant amino acid, and wherein the altered codon encodes a wild-type amino acid or a stop codon.

[0035] In some embodiments, the replacement modification comprises a deletion of one or more nucleotides or codons that are present in the intervening genomic sequence. In some such embodiments, the double-strand break is within a gene encoding a protein, and wherein the deletion inhibits expression of a full-length and / or active version of the protein. In some such embodiments, the double-strand break is within a mutant gene encoding a mutant protein, wherein the mutant gene comprises one or more mutant nucleotides and / or mutant codons not normally present in a wild-type gene, and wherein the replacement modification does not comprise the mutant nucleotides and / or mutant codons.

[0036] In some embodiments, the replacement modification comprises an insertion of one or more nucleotides or codons that are not present in the intervening genomic sequence. In some such embodiments, the double-strand break is within a mutant gene encoding a mutant protein, wherein the mutant gene lacks one or more deleted nucleotides or deleted codons normally present in a wild-type gene, and wherein the replacement modification comprises the deleted nucleotides or deleted codons. In some such embodiments, the replacement modification comprises an insertion of a nucleic acid sequence encoding a transgene, or a portion thereof. In some such embodiments, the replacement modification comprises an insertion of an intron. In some such embodiments, the replacement modification comprises an insertion of a nucleic acid sequence encoding a regulatory element, or a portion thereof, or an inhibitory element, or a portion thereof.

[0037] In some embodiments, the first homology arm and the second homology arm are approximately the same length. In some embodiments, the first homology arm and the second homology arm are between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 400

[0038] WBD (US) 4882-8896-8689vl

[0039] P893392050WO (01287) to 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 400 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 500 base pairs in length.

[0040] In some embodiments, the first homology arm and the second homology arm are different lengths. In some embodiments, the second homology arm is longer than the first homology arm. In some embodiments, the first homology arm is between about 100 to 1000 base pairs in length, and the second homology arm is between about 1000 to 3000 base pairs in length. In some embodiments, the first homology arm is between about 200 to 800 base pairs in length, and the second homology arm is between about 1200 to 1800 base pairs in length. In some embodiments, the first homology arm is between about 400 to 600 base pairs in length, and the second homology arm is between about 1400 to 1600 base pairs in length. In some embodiments, the first homology arm is about 400 or about 500 base pairs in length, and the second homology arm is about 1500 or about 1600 base pairs in length.

[0041] In some embodiments, the first homology arm has at least 98%, at least 99%, or 100% sequence homology to the first genomic sequence.

[0042] In some embodiments, the second homology arm has at least 98%, at least 99%, or 100% sequence homology to the second genomic sequence.

[0043] In some embodiments, the intervening genomic sequence is between about 10 to about 25000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 100 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1000 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 2500 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 500 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 250 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 100 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 50 base pairs in length.

[0044] In some embodiments, the replacement nucleic acid sequence is between about 1 to about 5000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 4500 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 4000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 3000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 2000 base pairs in

[0045] WBD (US) 4882-8896-8689vl

[0046] P893392050WO (01287) length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 1000 base pairs in length.

[0047] In some embodiments, the double-strand break is positioned in a cleaved exon of a gene. In some such embodiments, the second genomic sequence begins within the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the cleaved exon. In some such embodiments, the second genomic sequence begins within a second exon that is 3' downstream of the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the second exon. In some such embodiments, the second genomic sequence spans the second exon and one or more 3' downstream introns or exons. In some such embodiments, the second genomic sequence begins within an intron that is 3' downstream of the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the intron. In some such embodiments, the second genomic sequence spans the intron and one or more 3' downstream exons or introns. In some such embodiments, the replacement genomic sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 5' start of the intron and the 5' start of the second genomic sequence. In some such embodiments, the replacement intron sequence comprises a splice donor site that differs from the endogenous splice donor site but is capable of pairing with the endogenous splice acceptor site. In some such embodiments, the replacement intron sequence is a synthetic intron sequence. In some such embodiments, the replacement intron sequence is from a different location within the same gene or from a different gene.

[0048] In some embodiments, the double-strand break is positioned in a cleaved intron of a gene. In some such embodiments, the second genomic sequence is entirely comprised within the cleaved intron. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the double-strand break and the 5' end of the second genomic sequence. In some such embodiments, the second genomic sequence begins within an exon that is 3' downstream of the cleaved intron. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the double-strand break and the 5' end of the second genomic sequence. In some such embodiments, the second genomic sequence is entirely comprised within the exon. In some such embodiments, the second genomic

[0049] WBD (US) 4882-8896-8689vl

[0050] P893392050WO (01287) sequence spans the exon and one or more 3' downstream introns or exons. In some such embodiments, the second genomic sequence begins within an intron that is 3' downstream of the cleaved exon. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the doublestrand break and the 5' end of the second genomic sequence. In some such embodiments, the second genomic sequence is entirely comprised within the intron. In some such embodiments, the second genomic sequence spans the intron and one or more 3' downstream exons or introns. In some such embodiments, the replacement intron sequence comprises: (a) a splice donor site that differs from an endogenous splice donor of the gene site but is capable of pairing with an endogenous splice acceptor site of the gene; (b) a splice acceptor site that differs from an endogenous splice acceptor site of the gene but is capable of pairing with an endogenous splice donor site; (c) an intron enhancer; or (d) any combination of (a)-(c). In some such embodiments, the replacement intron sequence is a synthetic intron sequence. In some such embodiments, the replacement intron sequence is a heterologous intron sequence from a different location within the same gene or from a different gene.

[0051] In some embodiments, each end of the double-strand break has a 3' overhang. In some embodiments, the 3' overhang is a 4 base pair 3' overhang.

[0052] In some embodiments, each end of the double-strand break is a blunt end.

[0053] In some embodiments, each end of the double-strand break has a 5' overhang.

[0054] In some embodiments, the exogenous polynucleotide comprises a nucleic acid sequence encoding a nuclease. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 5' upstream of the first homology arm. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 3' downstream of the second homology arm. In some embodiments, the nuclease is capable of binding and cleaving the genome of the cell to generate the double-strand break.

[0055] In some embodiments, the exogenous polynucleotide comprises a promoter that is operably linked to the nucleic acid sequence encoding the nuclease.

[0056] In some embodiments, the nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, a compact TALEN, or a megaTAL.

[0057] In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo.

[0058] 8

[0059] WBD (US) 4882-8896-8689vl

[0060] P893392050WO (01287) In some embodiments, the exogenous polynucleotide is comprised by a delivery vehicle. In some embodiments, the delivery vehicle is a recombinant virus and the exogenous polynucleotide is comprised by a viral genome. In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).

[0061] In some embodiments, the delivery is a lipid nanoparticle. In some embodiments, the exogenous polynucleotide is a messenger RNA (mRNA), a single-stranded DNA, or a doublestranded DNA.

[0062] In another aspect, the disclosure provides a method for genetically modifying a cell, the method comprising introducing into the cell an exogenous polynucleotide and a nuclease or a gene encoding a nuclease, wherein the nuclease is expressed in the cell and generates a double-strand break in the genome of the cell, wherein the exogenous polynucleotide comprises, from 5' to 3': (a) a first homology arm; (b) a replacement nucleic acid sequence; and (c) a second homology arm; wherein the first homology arm has homology to a first genomic sequence that is adjacent to the double-strand break, wherein the second homology arm has homology to a second genomic sequence that is not adjacent to the double-strand break, wherein the double-strand break and the second genomic sequence are separated by an intervening genomic sequence, and wherein the intervening genomic sequence is replaced in the genome by the replacement nucleic acid sequence.

[0063] In some embodiments, the first genomic sequence is 5' upstream of the double-strand break. In some embodiments, the second genomic sequence is 3' downstream of the double-strand break. In some embodiments, the first genomic sequence is 5' upstream of the double-strand break and the second genomic sequence is 3' downstream of the double-strand break.

[0064] In some embodiments, the replacement nucleic acid sequence comprises a plurality of modified nucleotides and / or modified codons that differ from corresponding endogenous nucleotides or codons in the intervening genomic sequence, and wherein the replacement nucleic acid sequence encodes the same amino acid sequence as the intervening genomic sequence except for at least one replacement modification, wherein the replacement modification comprises: a) one or more nucleotide modifications that result in an altered codon, wherein the altered codon encodes a different amino acid than a corresponding endogenous codon in the intervening genomic sequence, or wherein the altered codon is a stop codon; b) deletion of one or more nucleotides or codons that are present in the intervening genomic sequence; or c) insertion of one or more nucleotides or codons that are not present in the intervening genomic sequence.

[0065] In some embodiments, the replacement nucleic acid sequence is a variant of the intervening genomic sequence.

[0066] 9

[0067] WBD (US) 4882-8896-8689vl

[0068] P893392050WO (01287) In some embodiments, the replacement nucleic acid sequence has 1 replacement modification. In some embodiments, the replacement nucleic acid sequence has 2 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 3 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 4 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 5 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 6 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 7 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 8 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 9 replacement modifications. In some embodiments, the replacement nucleic acid sequence has 10 replacement modifications.

[0069] In some embodiments, the replacement nucleic acid sequence has up to 10 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 20 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 30 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 40 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 50 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 100 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 200 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 500 replacement modifications. In some embodiments, the replacement nucleic acid sequence has up to 1000 replacement modifications.

[0070] In some embodiments, at least 20% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 30% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, between about 30% to 40% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 33% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide. In some embodiments, at least 50% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

[0071] In some embodiments, at least 75% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at

[0072] 10

[0073] WBD (US) 4882-8896-8689vl

[0074] P893392050WO (01287) least 80% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at least 90% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, at least 95% of the codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon. In some embodiments, all codons in the replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

[0075] In some embodiments, each of the modified codons comprises one nucleotide change, or two nucleotide changes, relative to a corresponding codon in the intervening genomic sequence.

[0076] In some embodiments, the replacement nucleic acid sequence has less than 80% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 70% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has between about 60% to 70% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 60% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 50% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 40% homology to the intervening genomic sequence.

[0077] In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 50 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 25 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 10 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening nucleic acid sequence do not share a sequence having more than 5 consecutive nucleotides in common.

[0078] In some embodiments, the replacement modification comprises one or more nucleotide modifications that result in an altered codon, wherein the altered codon encodes a different amino acid than a corresponding endogenous codon in the intervening genomic sequence, or wherein the altered codon is a stop codon. In some such embodiments, the altered codon comprises one nucleotide change, or two nucleotide changes, relative to the corresponding endogenous codon in the intervening genomic sequence. In some such embodiments, the double-strand break is generated within a mutant gene encoding a mutant protein, wherein the endogenous codon encodes

[0079] 11

[0080] WBD (US) 4882-8896-8689vl

[0081] P893392050WO (01287) a mutant amino acid, and wherein (a) the altered codon encodes a wild-type amino acid and the mutant gene is corrected to encode a wild-type protein; or (b) the altered codon encodes a stop codon that disrupts expression of a full-length and / or active version of the mutant protein.

[0082] In some embodiments, the replacement modification comprises a deletion of one or more nucleotides or codons that are present in the intervening genomic sequence. In some such embodiments, the double-strand break is within a gene encoding a protein, and wherein the deletion disrupts expression of a full-length and / or active version of the protein. In some such embodiments, the double-strand break is generated within a mutant gene encoding a mutant protein, wherein the mutant gene comprises one or more mutant nucleotides and / or mutant codons not normally present in a wild-type gene, wherein the replacement modification does not comprise the mutant nucleotides and / or mutant codons, and wherein the mutant gene is corrected to encode a wild-type protein.

[0083] In some embodiments, the replacement modification comprises an insertion of one or more nucleotides or codons that are not present in the intervening genomic sequence. In some embodiments, the double-strand break is within a mutant gene encoding a mutant protein, wherein the mutant gene lacks one or more deleted nucleotides or deleted codons normally present in a wild-type gene, wherein the replacement modification comprises the deleted nucleotides or deleted codons, and wherein the mutant gene is corrected to encode a wild-type protein. In some such embodiments, the replacement modification comprises an insertion of a nucleic acid sequence encoding a transgene, or a portion thereof. In some such embodiments, the replacement modification comprises an insertion of an intron. In some such embodiments, the replacement modification comprises an insertion of a nucleic acid sequence encoding a regulatory element, or a portion thereof, or an inhibitory element, or a portion thereof.

[0084] In some embodiments, the first homology arm and the second homology arm are approximately the same length. In some embodiments, the first homology arm and the second homology arm are between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 400 to 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 400 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 500 base pairs in length.

[0085] In some embodiments, the first homology arm and the second homology arm are different lengths. In some embodiments, the second homology arm is longer than the first homology arm. In

[0086] 12

[0087] WBD (US) 4882-8896-8689vl

[0088] P893392050WO (01287) some embodiments, the first homology arm is between about 100 to 1000 base pairs in length, and the second homology arm is between about 1000 to 3000 base pairs in length. In some embodiments, the first homology arm is between about 200 to 800 base pairs in length, and the second homology arm is between about 1200 to 1800 base pairs in length. In some embodiments, the first homology arm is between about 400 to 600 base pairs in length, and the second homology arm is between about 1400 to 1600 base pairs in length. In some embodiments, the first homology arm is about 400 or about 500 base pairs in length, and the second homology arm is about 1500 or about 1600 base pairs in length.

[0089] In some embodiments, the first homology arm has at least 98%, at least 99%, or 100% sequence homology to the first genomic sequence.

[0090] In some embodiments, the second homology arm has at least 98%, at least 99%, or 100% sequence homology to the second genomic sequence.

[0091] In some embodiments, the intervening genomic sequence is between about 10 to about 25000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 100 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1000 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 2500 to about 10000 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 500 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 250 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 100 base pairs in length. In some embodiments, the intervening genomic sequence is between about 1 to about 50 base pairs in length.

[0092] In some embodiments, the replacement nucleic acid sequence is between about 1 to about 5000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 4500 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 4000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 3000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 2000 base pairs in length. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 1000 base pairs in length.

[0093] In some embodiments, the double-strand break is positioned in a cleaved exon of a gene. In some such embodiments, the second genomic sequence begins within the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the cleaved exon. In

[0094] 13

[0095] WBD (US) 4882-8896-8689vl

[0096] P893392050WO (01287) some such embodiments, the second genomic sequence begins within a second exon that is 3' downstream of the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the second exon. In some such embodiments, the second genomic sequence spans the second exon and one or more 3' downstream introns or exons. In some such embodiments, the second genomic sequence begins within an intron that is 3' downstream of the cleaved exon. In some such embodiments, the second genomic sequence is entirely comprised within the intron. In some such embodiments, the second genomic sequence spans the intron and one or more 3' downstream exons or introns. In some such embodiments, the replacement genomic sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 5' start of the intron and the 5' start of the second genomic sequence. In some such embodiments, the replacement intron sequence comprises a splice donor site that differs from the endogenous splice donor site but is capable of pairing with the endogenous splice acceptor site. In some such embodiments, the replacement intron sequence is a synthetic intron sequence. In some such embodiments, the replacement intron sequence is from a different location within the same gene or from a different gene.

[0097] In some embodiments, the double-strand break is positioned in a cleaved intron of a gene. In some such embodiments, the second genomic sequence is entirely comprised within the cleaved intron. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the double-strand break and the 5' end of the second genomic sequence. In some such embodiments, the second genomic sequence begins within an exon that is 3' downstream of the cleaved intron. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the double-strand break and the 5' end of the second genomic sequence. In some such embodiments, the second genomic sequence is entirely comprised within the exon. In some such embodiments, the second genomic sequence spans the exon and one or more 3' downstream introns or exons. In some such embodiments, the second genomic sequence begins within an intron that is 3' downstream of the cleaved exon. In some such embodiments, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by the intervening genomic sequence that is positioned between the 3' end of the double-strand break and the 5' end of the

[0098] 14

[0099] WBD (US) 4882-8896-8689vl

[0100] P893392050WO (01287) second genomic sequence. In some such embodiments, the second genomic sequence is entirely comprised within the intron. In some such embodiments, the second genomic sequence spans the intron and one or more 3' downstream exons or introns. In some such embodiments, the replacement intron sequence comprises: (a) a splice donor site that differs from an endogenous splice donor of the gene site but is capable of pairing with an endogenous splice acceptor site of the gene; (b) a splice acceptor site that differs from an endogenous splice acceptor site of the gene but is capable of pairing with an endogenous splice donor site; (c) an intron enhancer; or (d) any combination of (a)-(c). In some such embodiments, the replacement intron sequence is a synthetic intron sequence. In some such embodiments, the replacement intron sequence is a heterologous intron sequence from a different location within the same gene or from a different gene.

[0101] In some embodiments, each end of the double-strand break generated by the nuclease has a 3' overhang. In some embodiments, the 3' overhang is a 4 base pair 3' overhang. In some embodiments, the nuclease is an engineered meganuclease.

[0102] In some embodiments, each end of the double-strand break generated by the nuclease is a blunt end or has a 5' overhang.

[0103] In some embodiments, the nuclease is a CRISPR-system nuclease, a zinc finger nuclease, a TALEN, a compact TALEN, or a megaTAL.

[0104] In some embodiments, the exogenous polynucleotide comprises the gene encoding the nuclease. In some embodiments, the gene encoding the nuclease is positioned 5' upstream of the first homology arm. In some embodiments, the gene encoding the nuclease is positioned 3' downstream of the second homology arm. In some embodiments, the exogenous polynucleotide comprises a promoter that is operably linked to the gene encoding the nuclease.

[0105] In some embodiments, the exogenous polynucleotide is introduced into the cell by a recombinant virus. In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).

[0106] In some embodiments, the exogenous polynucleotide is introduced into the cell by non-viral delivery.

[0107] In some embodiments, the exogenous polynucleotide is introduced into the cell by a lipid nanoparticle.

[0108] In some embodiments, the gene encoding the nuclease is introduced into the cell by a recombinant virus. In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).

[0109] 15

[0110] WBD (US) 4882-8896-8689vl

[0111] P893392050WO (01287) In some embodiments, the nuclease or the gene encoding the nuclease is introduced into the cell by non-viral delivery.

[0112] In some embodiments, the nuclease or the gene encoding the nuclease is introduced into the cell by a lipid nanoparticle.

[0113] In some embodiments, the exogenous polynucleotide is introduced into the cell by a recombinant virus, and the nuclease or the gene encoding the nuclease is introduced into the cell by a lipid nanoparticle. In some embodiments, the recombinant virus is a recombinant AAV.

[0114] In some embodiments, the exogenous polynucleotide is introduced into the cell by a first recombinant virus, and the gene encoding the nuclease is introduced into the cell by a second recombinant virus. In some embodiments, the first recombinant virus and the second recombinant virus are each a recombinant AAV.

[0115] In some embodiments, the exogenous polynucleotide is introduced into the cell by a first lipid nanoparticle, and the nuclease or the gene encoding the nuclease is introduced into the cell by a second lipid nanoparticle.

[0116] In some embodiments, the exogenous polynucleotide and the nuclease or the gene encoding the nuclease are introduced into the cell by a lipid nanoparticle.

[0117] In some embodiments, the method is performed in vitro.

[0118] In some embodiments, the method is performed in vivo.

[0119] In some embodiments, the method is a method for modifying a gene in a target cell in a subject, the method comprising delivering to the target cell the exogenous polynucleotide described herein and the nuclease or the gene encoding the nuclease described herein.

[0120] In some embodiments, the method is a method for treating a disease in a subject, the method comprising administering to the subject a therapeutically-effective amount of the exogenous polynucleotide described herein and a therapeutically-effective amount of the nuclease or the gene encoding the nuclease described herein.

[0121] BRIEF DESCRIPTION OF THE FIGURES

[0122] Figure 1 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, and the first and second genomic sequences are within the same cleaved exon.

[0123] Figure 2 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of

[0124] 16

[0125] WBD (US) 4882-8896-8689vl

[0126] P893392050WO (01287) the DSB in the same cleaved exon, and the second genomic sequence is within a downstream exon. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. In the particular configuration shown in the figure, the second genomic sequence does not begin at the start of the downstream exon; therefore, the 5’ region of the downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence.

[0127] Figure 3 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence begins within a downstream exon. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins in the downstream exon and further spans into an adjacent downstream intron.

[0128] Figure 4 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence is within a downstream intron. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene.

[0129] Figure 5 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence begins within a downstream intron. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene.

[0130] Figure 6 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, and the first and second genomic sequences are within the same cleaved intron.

[0131] Figure 7 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same intron, and the second genomic sequence is within a downstream exon. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. In the particular configuration shown in the figure, the second genomic sequence does not begin at the start of the downstream exon; therefore, the 5’ region of the

[0132] 17

[0133] WBD (US) 4882-8896-8689vl

[0134] P893392050WO (01287) downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence.

[0135] Figure 8 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence begins within a downstream exon. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins in the downstream exon and further spans into an adjacent downstream intron.

[0136] Figure 9 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence is within a downstream intron. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins at the start of the downstream intron, so no intron replacement sequence is necessary in the replacement nucleic acid sequence.

[0137] Figure 10 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence begins within a downstream intron. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins at the start of the downstream intron, so no intron replacement sequence is necessary in the replacement nucleic acid sequence.

[0138] Figure 11 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises a mutant nucleotide or codon, leading to knockout of a gene or production of a mutant protein. As shown, by introducing a modified nucleotide or codon into the replacement nucleic acid sequence, the method can be utilized to change an endogenous mutant nucleotide or codon to a wild-type nucleotide or codon that allows the gene to encode a wild-type protein.

[0139] Figure 12 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises a deleted nucleotide or codon, leading to knockout of the gene or production of a mutant protein. As shown, by introducing the deleted nucleotide or codon into the replacement nucleic acid sequence, the method can be utilized to re-introduce the deleted nucleotide or codon into the gene, allowing it to encode a wild-type protein.

[0140] 18

[0141] WBD (US) 4882-8896-8689vl

[0142] P893392050WO (01287) Figure 13 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises one or more extra nucleotides, such as a nucleotide expansion or nucleotide repeat expansion, leading to knockout of the gene or production of a mutant protein. As shown, by introducing the correct number of nucleotides into the replacement nucleic acid sequence, the method can be utilized to remove the extra / expanded nucleotides from the mutant gene, allowing it to encode a wild-type protein.

[0143] Figure 14A provides a graphic depiction of the genomic region targeted for nuclease- initiated repair and replace as described in Example 1.

[0144] Figure 14B provides a graphic depiction of a replacement nucleic acid sequence used in Example 1.

[0145] Figure 14C provides a graphic depiction of the predicted genome following repair and replacement with the replacement nucleic acid sequence of Example 1.

[0146] Figure 15A provides an overview of the experimental design for the flow cytometry analysis of Example 1.

[0147] Figure 15B provides an exemplary flow cytometry plot of cells treated with AAV only (repair template) from Example 1.

[0148] Figure 15C provides an exemplary flow cytometry plot of cells treated with nuclease and repair template from Example 1.

[0149] Figure 15D provides exemplary read length as measured by long-range Nanopore sequencing for cell populations from Example 1. The amplified region of interest is between the black and gray lines.

[0150] Figure 16 provides a graphic depiction of a replacement nucleic acid sequence used in Example 2.

[0151] Figure 17A provides an exemplary flow cytometry plot of cells treated with nuclease only from Example 2.

[0152] Figure 17B provides an exemplary flow cytometry plot of cells treated with nuclease and repair template from Example 2.

[0153] Figure 17C provides exemplary read length and mechanism of insertion as measured by long-range Nanopore sequencing for cell populations from Example 2. The amplified region of interest is between the black and gray lines.

[0154] Figure 18 provides a graphic depiction of the replacement nucleic acid sequence used in Example 3.

[0155] 19

[0156] WBD (US) 4882-8896-8689vl

[0157] P893392050WO (01287) Figure 19A provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a 2.8kb downstream skip.

[0158] Figure 19B provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a 2.3kb upstream skip.

[0159] Figure 20A provides a graphic depiction of the replacement nucleic acid sequence with a 5kb downstream skip according to Example 4.

[0160] Figure 20B provides a graphic depiction of the replacement nucleic acid sequence with a 7.5kb downstream skip according to Example 4.

[0161] Figure 20C provides a graphic depiction of the replacement nucleic acid sequence with a lOkb downstream skip according to Example 4.

[0162] Figure 21A provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a 2.8kb downstream skip.

[0163] Figure 21B provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a 5kb downstream skip.

[0164] Figure 21C provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a 7.5kb downstream skip.

[0165] Figure 21D provides an exemplary flow cytometry plot of cells treated with nuclease and a repair construct with a lOkb downstream skip.

[0166] Figure 22A provides a comparison of the repair constructs tested across Examples 1-4 as measured by the percent of B2M-CD3+ cells.

[0167] Figure 22B provides a comparison of the repair constructs tested across Examples 1-4 as measured by the percent of B2M-CD3+ cells as a percentage of the total population of edited cells.

[0168] BRIEF DESCRIPTION OF THE SEQUENCES

[0169] SEQ ID NO: 1 sets forth the nucleic acid sequence of a repair construct of Example 1 corresponding to the graphic provided in Figure 14C.

[0170] SEQ ID NO: 2 sets forth the nucleic acid sequence of a repair construct of Example 2 corresponding to the graphic provided in Figure 16.

[0171] SEQ ID NO: 3 sets forth the nucleic acid sequence of a repair construct of Example 3 corresponding to the graphic provided in Figure 18.

[0172] SEQ ID NO: 4 sets forth the nucleic acid sequence of a repair construct of Example 4 corresponding to the graphic provided in Figure 20A.

[0173] SEQ ID NO: 5 sets forth the nucleic acid sequence of a repair construct of Example 4 corresponding to the graphic provided in Figure 20B.

[0174] 20

[0175] WBD (US) 4882-8896-8689vl

[0176] P893392050WO (01287) SEQ ID NO: 6 sets forth the nucleic acid sequence of a repair construct of Example 4 corresponding to the graphic provided in Figure 20C.

[0177] DETAILED DESCRIPTION OF THE INVENTION

[0178] 1, References and Definitions

[0179] The patent and scientific literature referred to herein establishes knowledge that is available to those of skill in the art. The issued US patents, allowed applications, published foreign applications, and references, including GenBank database sequences, which are cited herein are hereby incorporated by reference to the same extent as if each was specifically and individually indicated to be incorporated by reference.

[0180] The present disclosure can be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. For example, features illustrated with respect to one embodiment can be incorporated into other embodiments, and features illustrated with respect to a particular embodiment can be deleted from that embodiment. In addition, numerous variations and additions to the embodiments suggested herein will be apparent to those skilled in the art in light of the instant disclosure, which do not depart from the instant disclosure.

[0181] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in the description of the disclosure herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure.

[0182] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference herein in their entirety.

[0183] As used herein, “a,” “an,” or “the” can mean one or more than one. For example, “a” cell can mean a single cell or a multiplicity of cells.

[0184] As used herein, unless specifically indicated otherwise, the word “or” is used in the inclusive sense of “and / or” and not the exclusive sense of “either / or.”

[0185] As used herein, the term “recombinant” or “engineered” with respect to a protein means having an altered amino acid sequence as a result of the application of genetic engineering techniques to nucleic acids that encode the protein and cells or organisms that express the protein. With respect to a nucleic acid, the term “recombinant” or “engineered” means having an altered

[0186] 21

[0187] WBD (US) 4882-8896-8689vl

[0188] P893392050WO (01287) nucleic acid sequence as a result of the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning technologies; transfection, transformation, and other gene transfer technologies; homologous recombination; site- directed mutagenesis; and gene fusion. In accordance with this definition, a protein having an amino acid sequence identical to a naturally-occurring protein but produced by cloning and expression in a heterologous host, is not considered recombinant or engineered.

[0189] As used herein, the term “modified nucleotides” refers to one or more nucleotides present in a replacement nucleic acid sequence that differ from one or more corresponding endogenous nucleotides in a genomic nucleic acid sequence. One or more modified nucleotides in a replacement nucleic acid sequence may result in one or more modified codons.

[0190] As used herein, the term “modified codons” or “altered codons,” used interchangeably, refers to one or more codons present in a replacement nucleic acid sequence that differ from one or more corresponding endogenous codons in a genomic nucleic acid sequence. In some cases, the one or more modified / altered codons in the replacement nucleic acid sequence encode the same amino acid as the one or more corresponding endogenous codons. In other cases, where the one or more modified / altered codons are intended to be a replacement modification described herein, the modified / altered codons encode a different amino acid than the one or more corresponding endogenous codons.

[0191] As used herein, the term “wild-type” refers to the most common naturally occurring allele (i.e., polynucleotide sequence) in the allele population of the same type of gene, wherein a polypeptide encoded by the wild-type allele has its original functions. The term “wild-type” also refers to a polypeptide encoded by a wild-type allele. Wild-type alleles (i.e., polynucleotides) and polypeptides are distinguishable from mutant or variant alleles and polypeptides, which comprise one or more mutations and / or substitutions relative to the wild-type sequence(s). Whereas a wildtype allele or polypeptide can confer a normal phenotype in an organism, a mutant or variant allele or polypeptide can, in some instances, confer an altered phenotype. The term “wild-type” can also refer to a cell, an organism, and / or a subject which possesses a wild-type allele of a particular gene, or a cell, an organism, and / or a subject used for comparative purposes.

[0192] As used herein, the term “mutant” refers to a locus having one or more mutations and / or substitutions relative to the wild-type locus. It would be understood that a locus may be associated with a disease and / or disorder. As such, the term mutant can refer to a locus that is associated with a disease and / or a disorder in a subject (i.e., a diseased state). The term mutant can also refer to a locus that differs from the wild-type locus and is not associated with a disease and / or a disorder in a

[0193] 22

[0194] WBD (US) 4882-8896-8689vl

[0195] P893392050WO (01287) subject (i.e., a diseased state). In some embodiments, a mutant locus is related and / or associated with a diseased state, and a wild-type locus is related and / or associated with a non-diseased state. In some embodiments, a mutant locus is not related and / or associated with a diseased state, and a wild-type locus is related and / or associated with a non-diseased state. A mutant polynucleotide may have a methylation state that differs from a wild-type polynucleotide. In some embodiments, a mutant locus is replaced with a wild-type locus to correct a disease and / or disorder in a subject. The terms “wild-type” and “mutant” can refer to any locus (polynucleotide) or polypeptide disclosed herein, including for instance a nucleotide, codon, exon sequence, intron sequence, such as splice donor sequence, splice acceptor sequence, splice enhancer sequence, non-gene coding sequence, and amino acid.

[0196] As used herein, the term “mutant nucleotide” refers to a nucleotide present in a mutant gene that differs from the corresponding wild-type nucleotide found in a wild-type gene. For example, an allele or gene having a mutant nucleotide is considered a variant of the wild-type allele or gene. In some cases, a mutant nucleotide can disrupt expression of the mutant gene and / or lead to expression of a mutant protein (e.g., not full-length, and / or not functional).

[0197] As used herein, the term “mutant codon” refers to a codon present in a mutant gene that differs in sequence from the corresponding wild-type codon found in a wild-type gene. In some cases, the mutant codon can encode a wild-type amino acid leading to expression of a wild-type protein. In other cases, the mutant codon can encode a mutant amino acid that differs from the corresponding wild-type amino acid. A mutant gene can comprise, for example, one or more mutant codons. A mutant gene comprising a mutant codon can, in some cases, not express a protein, express a mutant protein that is not full length, and / or express a mutant protein is not functional relative to the wild-type protein.

[0198] As used herein, the term “mutant gene” refers to a gene that differs in sequence from the corresponding wild-type gene sequence. A mutant gene can comprise, for example, one or more mutant nucleotides, one or more mutant codons, one or more nucleotide insertions relative to the wild-type sequence, one or more nucleotide deletions relative to the wild-type sequence, or any combination thereof. A mutant gene can, in some cases, not express a protein, express a mutant protein that is not full length, and / or express a mutant protein is not functional relative to the wildtype protein.

[0199] As used herein, the term “mutant protein” refers to a protein having one or more amino acids present in the polypeptide that differ from the corresponding, wild-type amino acid. That is, the mutant protein may comprise one or more “mutant amino acids.” As used herein, the term

[0200] 23

[0201] WBD (US) 4882-8896-8689vl

[0202] P893392050WO (01287) “mutant amino acid” refers to an amino acid present in a polypeptide that differs from the corresponding wild-type polypeptide. A polypeptide sequence (i.e., a protein) having a mutant amino acid is considered a variant of the wild-type polypeptide sequence.

[0203] As used herein, the term “replacement nucleic acid sequence” refers to a nucleic acid sequence that is designed to be inserted into the genome of a cell, an organism, and / or a subject, preferably via homologous recombination. Following homologous recombination, the replacement nucleic acid sequence is introduced into the genome of a cell, an organism, and / or a subject, thereby replacing an endogenous intervening nucleic acid sequence described herein (i.e., the sequence between the double strand break and the beginning of the second genomic region recognized by the second homology arm). The replacement nucleic acid sequence comprises one or more modified nucleotides and / or modified codons as compared to the corresponding endogenous nucleotides or codons in the intervening genomic sequence. The replacement nucleic acid sequence encodes the same amino acid sequence as the intervening genomic sequence with at least one replacement modification. In some cases, the replacement nucleic acid sequence comprises a replacement intron sequence that differs from the endogenous intron sequence found in the intervening genomic sequence. In some cases, the replacement nucleic acid sequence comprises altered splice donor sequences and splice acceptor sequences on either an exon or an intron such that splicing of an intron and exon is altered.

[0204] As used herein, the term “replacement intron sequence” refers to an intronic sequence in a replacement nucleic acid sequence described herein that differs from an endogenous intronic sequence that is part of an intervening genomic sequence described herein. A replacement intron sequence can include, for example, intronic elements necessary for proper splicing and / or expression of a gene following introduction of the replacement intron sequence into the genome. Such elements can include, for example, splice donor sequences, splice acceptor sequences, and intronic enhancer sequences. A replacement intron sequence can, for example, originate from a different location of the same gene being modified, a different gene than that being modified, or can be a synthetic intron sequence.

[0205] As used herein, the phase “encodes the same amino acid sequence” refers to the ability of a replacement nucleic acid sequence described herein to produce the same amino acid sequence following transcription, translation, and intron splicing, as an intervening genomic sequence described herein. By way of example, in some cases, a double strand break is generated in an exon, and the second genomic sequence (i.e., the sequence having homology to the second homology arm) is within a downstream exon. In such an example, the replacement nucleic acid sequence that

[0206] 24

[0207] WBD (US) 4882-8896-8689vl

[0208] P893392050WO (01287) is introduced into the genome between the double strand break and the second genomic sequence may lack one or more introns that may have been present in the genome; however, due to intron splicing following transcription and translation, that replacement nucleic acid sequence still encodes the same amino acid sequence as the endogenous intervening genomic sequence that was replaced (with one or more replacement modifications, as described herein).

[0209] A replacement nucleic acid sequence described herein comprises one or more replacement modifications. As used herein, the term a “replacement modification” refers to a change in the replacement nucleotide sequence relative to the endogenous intervening genomic sequence that, when incorporated into the gene following HDR repair and replacement, results in a modified gene in the genome, and in some examples can result in expression of a modified protein from the gene.

[0210] The methods described herein utilize replacement nucleic acid sequences to introduce variant nucleic acid sequences into the genome of a cell, an organism, and / or a subject. As such, the replacement nucleic acid sequence may comprise one or more nucleotides not normally present in an endogenous gene; the replacement nucleic acid sequence may comprise one or more codons not normally present in an endogenous gene; the replacement nucleic acid sequence may comprise one or more nucleotide deletions that are not normally present in an endogenous gene; and / or the replacement nucleic acid sequence may comprise one or more nucleotide insertions that are not normally present in an endogenous gene.

[0211] As used herein with respect to both amino acid sequences and nucleic acid sequences, the terms “percent identity,” “sequence identity,” “percentage similarity,” “sequence similarity” and the like refer to a measure of the degree of similarity of two sequences based upon an alignment of the sequences that maximizes similarity between aligned amino acid residues or nucleotides, and which is a function of the number of identical or similar residues or nucleotides, the number of total residues or nucleotides, and the presence and length of gaps in the sequence alignment. A variety of algorithms and computer programs are available for determining sequence similarity using standard parameters. As used herein, sequence similarity is measured using the BLASTp program for amino acid sequences and the BLASTn program for nucleic acid sequences, both of which are available through the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov / ), and are described in, for example, Altschul et al. (1990), J. Mol. Biol. 215:403-410; Gish and States (1993), Nature Genet. 3:266-272; Madden et al. (1996), Meth. Enzymol.266: 131-141; Altschul et al. (1997), Nucleic Acids Res. 25:33 89-3402); Zhang et al. (2000), J. Comput. Biol. 7( l-2):203- 14. As used herein, percent similarity of two amino acid sequences is the score based upon the following parameters for the BLASTp algorithm: word size=3; gap opening penalty=-l l; gap

[0212] 25

[0213] WBD (US) 4882-8896-8689vl

[0214] P893392050WO (01287) extension penalty=-l; and scoring matrix=BLOSUM62. As used herein, percent similarity of two nucleic acid sequences is the score based upon the following parameters for the BLASTn algorithm: word size=l l; gap opening penalty=-5; gap extension penalty=-2; match reward=l; and mismatch penalty=-3.

[0215] As used herein with respect to modifications of two proteins or amino acid sequences, the term “corresponding to” is used to indicate that a specified modification in the first protein is a substitution of the same amino acid residue as in the modification in the second protein, and that the amino acid position of the modification in the first protein corresponds to or aligns with the amino acid position of the modification in the second protein when the two proteins are subjected to standard sequence alignments (e.g., using the BLASTp program) and aligned for maximum sequence identity across the entire subunit or protein. Thus, the modification of residue “X” to amino acid “A” in the first protein will correspond to the modification of residue “Y” to amino acid “A” in the second protein if residues X and Y correspond to each other in a sequence alignment, and despite the fact that X and Y may be at different positions relative to the N-terminus or the C- terminus.

[0216] As used herein, the term “homologous recombination” (HR) or “homology-directed repair” (HDR), used interchangeably, refers to the natural, cellular process in which a double -stranded DNA-break is repaired using a homologous DNA sequence as the repair template (see, e.g., Cahill et al. (2006), Front. Biosci. 11: 1958-1976). As used herein, the term “non-homologous endjoining” (NHEJ) refers to the natural, cellular process in which a double-stranded DNA-break is repaired by the direct joining of two non-homologous DNA-break is repaired by the direct joining of two non-homologous DNA segments (see, e.g. Cahill et al. (2006), Front. Biosci. 11: 1958- 1976). DNA repair by NHEJ is error prone and frequently results in the untemplated addition or deletion of DNA sequences at the site of repair.

[0217] As used herein, the term “exogenous polynucleotide” refers to a polynucleotide that originates from outside of a cell, an organism, and / or a subject. An exogeneous polynucleotide may be integrated, in whole or in part, into the genome of a cell, and organism, and / or a subject. In some embodiments, the exogenous polynucleotide is integrated into the genome via homologous recombination.

[0218] To perform HDR, an exogenous polynucleotide is introduced into a cell, an organism, and / or a subject, wherein the exogenous polynucleotide comprises a first homology arm, a replacement nucleic acid sequence, and a second homology arm.

[0219] 26

[0220] WBD (US) 4882-8896-8689vl

[0221] P893392050WO (01287) As used herein, a “homology arm” refers to a sequence flanking the 5 ’ and 3 ’ ends of a nucleic acid molecule, i.e., a replacement nucleic acid sequence, which promotes integration of a replacement nucleic acid sequence into the genome of a cell, an organism, and / or a subject.

[0222] As used herein, the term “first homology arm” refers to a homology arm that flanks the 5 ’ end of a replacement nucleic acid sequence. The first homology arm has homology to a first genomic sequence. As used herein, the term “first genomic sequence” refers to a region of the genome to which the first homology arm shares homology. The first genomic sequence is upstream of a double strand break, preferably a double strand break generated by an engineered nuclease. The first genomic sequence (and therefore the first homology arm) may be located upstream, and directly adjacent to, the double strand break.

[0223] As used herein, the term “second homology arm” refers to a homology arm that flanks the 3’ end of a replacement nucleic acid sequence. The second homology arm has homology to a second genomic sequence. As used herein, the term “second genomic sequence” refers a region of the genome to which the second homology arm shares homology. The second genomic sequence is located downstream of a double strand break, preferably a double strand break generated by an engineered nuclease. The second genomic sequence (and therefore the second homology arm) is located downstream of the double strand break but is not adjacent to the double strand break.

[0224] As used herein, the term “intervening genomic sequence” refers to genomic sequence between the double strand break and the 5 ’ end of the second genomic sequence, which is downstream of the double strand break. In some embodiments, the intervening genomic sequence is between the double strand break and the 3’ end of the first homology arm.

[0225] The first and second homology arms may be approximately the same length. As used herein, the term “approximately the same length” refers to a first homology arm and a second homology arm that are within 10% of length of each other as measured by the number of base pairs (nucleotides). Alternatively, the first and second homology arms may be different lengths. As used herein, the term “different lengths” refers to a first homology arm and a second homology arm that have a greater than 10% difference in length of each other as measured by the number of base pairs (nucleotides). For example, but not by way of limitation, a first homology arm having a length of 500 base pairs and a second homology arm having a length of 600 base pairs would be classified as having approximately the same length. A first homology arm having a length of 500 base pairs and a second homology arm having a length of 1000 base pairs would be classified as having different lengths.

[0226] 27

[0227] WBD (US) 4882-8896-8689vl

[0228] P893392050WO (01287) As used herein, the term “flanked” or “flanking” in relation to a particular nucleic acid sequence refers to a sequence (i.e., a flanking sequence) being on each side of a particular nucleic acid sequence. For example, a first homology arm and a second homology arm are flanking sequences to the replacement nucleic acid sequence.

[0229] As used herein, the terms “transfected” or “transformed” or “transduced” or “nucleofected” refer to a process by which exogenous nucleic acid is transferred or introduced into the host cell. A “transfected” or “transformed” or “transduced” cell is one which has been transfected, transformed, or transduced with exogenous nucleic acid. The cell includes the primary subject cell and its progeny.

[0230] The methods described herein use a site-directed nuclease that cleaves a site in the genome to generate a double strand break. As used herein, the terms “cleave” or “cleavage” refer to the hydrolysis of phosphodiester bonds within the backbone of a recognition sequence within a target sequence that results in a double-stranded break within the target sequence, referred to herein as a “cleavage site.”

[0231] As used herein, the term “nuclease” refers to a naturally-occurring or engineered enzyme, which cleaves a phosphodiester bond within a polynucleotide chain, thereby generating a double strand break. Examples of nucleases include, but are not limited to, engineered meganucleases, CRISPR system nucleases, TALE nucleases (TALENs), zinc finger nucleases, compact TALENs, and megaTALs.

[0232] As used herein, the terms “recognition sequence” or “recognition site” refers to a DNA sequence that is bound and cleaved by a nuclease. In the case of a meganuclease, a recognition sequence comprises a pair of inverted, 9 base pair “half sites” which are separated by four base pairs. In the case of a single-chain meganuclease, the N-terminal domain of the protein contacts a first half-site and the C-terminal domain of the protein contacts a second half-site. Cleavage by a meganuclease produces four base pair 3' overhangs. “Overhangs,” or “sticky ends” are short, single-stranded DNA segments that can be produced by endonuclease cleavage of a doublestranded DNA sequence. In the case of meganucleases and single-chain meganucleases derived from I-Crel, the overhang comprises bases 10-13 of the 22 base pair recognition sequence. In the case of a compact TALEN, the recognition sequence comprises a first CNNNGN sequence that is recognized by the I-TevI domain, followed by a non-specific spacer 4-16 base pairs in length, followed by a second sequence 16-22 bp in length that is recognized by the TAL-effector domain (this sequence typically has a 5' T base). Cleavage by a compact TALEN produces two base pair 3' overhangs. In the case of a CRISPR nuclease, the recognition sequence is the sequence, typically

[0233] 28

[0234] WBD (US) 4882-8896-8689vl

[0235] P893392050WO (01287) 16-24 base pairs, to which the guide RNA binds to direct cleavage. Full complementarity between the guide sequence and the recognition sequence is not necessarily required to effect cleavage. Cleavage by a CRISPR nuclease can produce blunt ends (such as by a class 2, type II CRISPR nuclease) or overhanging ends (such as by a class 2, type V CRISPR nuclease), depending on the CRISPR nuclease. In those embodiments wherein a Cpfl CRISPR nuclease is utilized, cleavage by the CRISPR complex comprising the same will result in 5' overhangs and in certain embodiments, 5 nucleotide 5' overhangs. Each CRISPR nuclease enzyme also requires the recognition of a PAM (protospacer adjacent motif) sequence that is near the recognition sequence complementary to the guide RNA. The precise sequence, length requirements for the PAM, and distance from the target sequence differ depending on the CRISPR nuclease enzyme, but PAMs are typically 2-5 base pair sequences adjacent to the target / recognition sequence. PAM sequences for particular CRISPR nuclease enzymes are known in the art (see, for example, U.S. Patent No. 8,697,359 and U.S. Publication No. 20160208243, each of which is incorporated by reference in its entirety) and PAM sequences for novel or engineered CRISPR nuclease enzymes can be identified using methods known in the art, such as a PAM depletion assay (see, for example, Karvelis et al. (2017) Methods 121-122:3-8, which is incorporated herein in its entirety). In the case of a zinc finger, the DNA binding domains typically recognize an 18-bp recognition sequence comprising a pair of nine base pair “half-sites” separated by 2-10 base pairs and cleavage by the nuclease creates a blunt end or a 5' overhang of variable length (frequently four base pairs).

[0236] As used herein, the term “meganuclease” refers to an endonuclease that binds doublestranded DNA at a recognition sequence that is greater than 12 base pairs. A meganuclease can be an endonuclease that is derived from I-Crel and can refer to an engineered variant of I-Crel that has been modified relative to natural I-Crel with respect to, for example, DNA-binding specificity, DNA cleavage activity, DNA-binding affinity, or dimerization properties. Methods for producing such modified variants of I-Crel are known in the art (e.g., WO 2007 / 047859, incorporated by reference in its entirety). A meganuclease as used herein binds to double-stranded DNA as a heterodimer. A meganuclease may also be a “single-chain meganuclease” in which a pair of meganuclease DNA-binding domains is joined into a single polypeptide using a peptide linker. The term “homing endonuclease” is synonymous with the term “meganuclease.” Meganucleases used in the presently disclosed methods and compositions are substantially non-toxic when expressed in the targeted cells as described herein such that cells can be transfected and maintained at 37°C without observing deleterious effects on cell viability or significant reductions in meganuclease cleavage activity.

[0237] 29

[0238] WBD (US) 4882-8896-8689vl

[0239] P893392050WO (01287) As used herein, the term “TALE nuclease” or “TALEN” refers to an endonuclease comprising a DNA-binding domain comprising a plurality of TAL domain repeats fused to a nuclease domain or an active portion thereof from an endonuclease or exonuclease, including but not limited to a restriction endonuclease, homing endonuclease, S 1 nuclease, mung bean nuclease, pancreatic DNAse I, micrococcal nuclease, and yeast HO endonuclease. See, for example, Christian et al. (2010) Genetics 186:757-761, which is incorporated by reference in its entirety. Nuclease domains useful for the design of TALENs include those from a Type Ils restriction endonuclease, including but not limited to FokI, FoM, StsI, Hhal, Hindlll, Nod, BbvCI, EcoRI, Bgll, and AlwI. Additional Type Ils restriction endonucleases are described in International Publication No. WO 2007 / 014275, which is incorporated by reference in its entirety. In some embodiments, the nuclease domain of the TALEN is a FokI nuclease domain or an active portion thereof. TAL domain repeats can be derived from the TALE (transcription activator-like effector) family of proteins used in the infection process by plant pathogens of the Xanthomonas genus. TAL domain repeats are 33-34 amino acid sequences with divergent 12th and 13th amino acids. These two positions, referred to as the repeat variable dipeptide (RVD), are highly variable and show a strong correlation with specific nucleotide recognition. Each base pair in the DNA target sequence is contacted by a single TAL repeat with the specificity resulting from the RVD. In some embodiments, the TALEN comprises 16-22 TAL domain repeats. DNA cleavage by a TALEN requires two DNA recognition regions (i.e., “half-sites”) flanking a nonspecific central region (i.e., the “spacer”). The term “spacer” in reference to a TALEN refers to the nucleic acid sequence that separates the two nucleic acid sequences recognized and bound by each monomer constituting a TALEN. The TAL domain repeats can be native sequences from a naturally-occurring TALE protein or can be redesigned through rational or experimental means to produce a protein that binds to a pre-determined DNA sequence (see, for example, Boch et al. (2009) Science 326(5959): 1509- 1512 and Moscou and Bogdanove (2009) Science 326(5959): 1501, each of which is incorporated by reference in its entirety). See also, U.S. Publication No. 20110145940 and International Publication No. WO 2010 / 079430 for methods for engineering a TALEN to recognize and bind a specific sequence and examples of RVDs and their corresponding target nucleotides. In some embodiments, each nuclease (e.g., FokI) monomer can be fused to a TAL effector sequence that recognizes and binds a different DNA sequence, and only when the two recognition sites are in close proximity do the inactive monomers come together to create a functional enzyme. It is understood that the term “TALEN” can refer to a single TALEN protein or, alternatively, a pair of TALEN proteins (i.e., a left TALEN protein and a right TALEN protein) which bind to the

[0240] 30 WBD (US) 4882-8896-8689vl

[0241] P893392050WO (01287) upstream and downstream half-sites adjacent to the TALEN spacer sequence and work in concert to generate a cleavage site within the spacer sequence. Given a predetermined DNA locus or spacer sequence, upstream and downstream half-sites can be identified using a number of programs known in the art (Komel Labun; Tessa G. Montague; James A. Gagnon; Summer B. Thyme; Eivind Valen. (2016). CHOPCHOP v2: a web tool forthe next generation of CRISPR genome engineering. Nucleic Acids Research; doi: 10.1093 / nar / gkw398; Tessa G. Montague; Jose M. Cruz; James A. Gagnon; George M. Church; Eivind Valen. (2014). CHOPCHOP: a CRISPR / Cas9 and TALEN web tool for genome editing. Nucleic Acids Res. 42. W401-W407). It is also understood that a TALEN recognition sequence can be defined as the DNA binding sequence (i.e., half-site) of a single TALEN protein or, alternatively, a DNA sequence comprising the upstream half-site, the spacer sequence, and the downstream half-site.

[0242] As used herein, the term “compact TALEN” refers to an endonuclease comprising a DNA- binding domain with one or more TAL domain repeats fused in any orientation to any portion of the I-TevI homing endonuclease or any of the endonucleases listed in Table 2 in U.S. Application No. 20130117869 (which is incorporated by reference in its entirety), including but not limited to Mmel, EndA, Endl, I-BasI, I-TevII, I-TevIII, I-Twol, MspI, Mval, NucA, and NucM. Compact TALENs do not require dimerization for DNA processing activity, alleviating the need for dual target sites with intervening DNA spacers. In some embodiments, the compact TALEN comprises 16-22 TAL domain repeats.

[0243] As used herein, the term “megaTAL” refers to a single-chain endonuclease comprising a transcription activator-like effector (TALE) DNA binding domain with an engineered, sequencespecific homing endonuclease.

[0244] As used herein, the terms “CRISPR nuclease” or “CRISPR system nuclease” refers to a CRISPR (clustered regularly interspaced short palindromic repeats)-associated (Cas) endonuclease or a variant thereof, such as Cas9, that associates with a guide RNA that directs nucleic acid cleavage by the associated endonuclease by hybridizing to a recognition site in a polynucleotide. In certain embodiments, the CRISPR nuclease is a class 2 CRISPR enzyme. In some of these embodiments, the CRISPR nuclease is a class 2, type II enzyme, such as Cas9. In other embodiments, the CRISPR nuclease is a class 2, typeV enzyme, such as Cpfl . The guide RNA comprises a direct repeat and a guide sequence (often referred to as a spacer in the context of an endogenous CRISPR system), which is complementary to the target recognition site. In certain embodiments, the CRISPR system further comprises a tracrRNA (trans-activating CRISPR RNA) that is complementary (fully or partially) to the direct repeat sequence (sometimes referred to as a

[0245] 31

[0246] WBD (US) 4882-8896-8689vl

[0247] P893392050WO (01287) tracr-mate sequence) present on the guide RNA. In particular embodiments, the CRISPR nuclease can be mutated with respect to a corresponding wild-type enzyme such that the enzyme lacks the ability to cleave one strand of a target polynucleotide, functioning as a nickase, cleaving only a single strand of the target DNA. Non-limiting examples of CRISPR enzymes that function as a nickase include Cas9 enzymes with a D10A mutation within the RuvC I catalytic domain, or with a H840A, N854A, or N863A mutation. Given a predetermined DNA locus, recognition sequences can be identified using a number of programs known in the art (Komel Labun; Tessa G. Montague; James A. Gagnon; Summer B. Thyme; Eivind Valen. (2016). CHOPCHOP v2: a web tool for the next generation of CRISPR genome engineering. Nucleic Acids Research; doi: 10.1093 / nar / gkw398; Tessa G. Montague; Jose M. Cruz; James A. Gagnon; George M. Church; Eivind Valen. (2014). CHOPCHOP: a CRISPR / Cas9 and TALEN web tool for genome editing. Nucleic Acids Res. 42. W401-W407).

[0248] As used herein, the terms “zinc finger nuclease” or “ZFN” refers to a chimeric protein comprising a zinc finger DNA-binding domain fused to a nuclease domain from an endonuclease or exonuclease, including but not limited to a restriction endonuclease, homing endonuclease, S 1 nuclease, mung bean nuclease, pancreatic DNAse I, micrococcal nuclease, and yeast HO endonuclease. Nuclease domains useful for the design of zinc finger nucleases include those from a Type Ils restriction endonuclease, including but not limited to FokI, FoM, and StsI restriction enzyme. Additional Type Ils restriction endonucleases are described in International Publication No. WO 2007 / 014275, which is incorporated by reference in its entirety. The structure of a zinc finger domain is stabilized through coordination of a zinc ion. DNA binding proteins comprising one or more zinc finger domains bind DNA in a sequence -specific manner. The zinc finger domain can be a native sequence or can be redesigned through rational or experimental means to produce a protein which binds to a pre-determined DNA sequence ~I8 base pairs in length, comprising a pair of nine base pair half-sites separated by 2-10 base pairs. See, for example, U.S. Pat. Nos. 5,789,538, 5,925,523, 6,007,988, 6,013,453, 11,311,574, and International Publication Nos. WO 95 / 19431, WO 96 / 06166, WO 98 / 53057, WO 98 / 54311, WO 00 / 27878, WO 01 / 60970, WO 01 / 88197, and WO 02 / 099084, each of which is incorporated by reference in its entirety. By fusing this engineered protein domain to a nuclease domain, such as FokI nuclease, it is possible to target DNA breaks with genome-level specificity. The selection of target sites, zinc finger proteins and methods for design and construction of zinc finger nucleases are known to those of skill in the art and are described in detail in U.S. Publications Nos. 20030232410, 20050208489, 2005064474, 20050026157, 20060188987 and International Publication No. WO 07 / 014275, each of which is

[0249] 32

[0250] WBD (US) 4882-8896-8689vl

[0251] P893392050WO (01287) incorporated by reference in its entirety. In the case of a zinc finger, the DNA binding domains typically recognize an 18-bp recognition sequence comprising a pair of nine base pair “half-sites” separated by a 2-10 base pair “spacer sequence”, and cleavage by the nuclease creates a blunt end or a 5' overhang of variable length (frequently four base pairs). It is understood that the term “zinc finger nuclease” can refer to a single zinc finger protein or, alternatively, a pair of zinc finger proteins (i.e., a left ZFN protein and a right ZFN protein) that bind to the upstream and downstream half-sites adjacent to the zinc finger nuclease spacer sequence and work in concert to generate a cleavage site within the spacer sequence. Given a predetermined DNA locus or spacer sequence, upstream and downstream half-sites can be identified using a number of programs known in the art (Mandell JG, Barbas CF 3rd. Zinc Finger Tools: custom DNA-binding domains fortranscription factors and nucleases. Nucleic Acids Res. 2006 Jul 1 ;34 (Web Server issue):W516-23). It is also understood that a zinc finger nuclease recognition sequence can be defined as the DNA binding sequence (i.e., half-site) of a single zinc finger nuclease protein or, alternatively, a DNA sequence comprising the upstream half-site, the spacer sequence, and the downstream half-site.

[0252] As used herein, the term “specificity” means the ability of a nuclease to bind and cleave double-stranded DNA molecules only at a particular sequence of base pairs referred to as the recognition sequence, or only at a particular set of recognition sequences. The set of recognition sequences will share certain conserved positions or sequence motifs but may be degenerate at one or more positions. A highly-specific nuclease is capable of cleaving only one or a very few recognition sequences. Specificity can be determined by any method known in the art.

[0253] As used herein, the terms “target site” or “target sequence” refers to a region of the chromosomal DNA of a cell comprising a recognition sequence for a nuclease.

[0254] As used herein, the term “recognition half-site,” “recognition sequence half-site,” or simply “half-site” means a nucleic acid sequence in a double-stranded DNA molecule that is recognized and bound by a monomer of a homodimeric or heterodimeric meganuclease or by one subunit of a single-chain meganuclease or by one subunit of a single-chain meganuclease, or by a monomer of a TALEN or zinc finger nuclease.

[0255] As used herein, the term “5' portion” when referring to a nuclease recognition sequence is intended to mean the nucleotides of the recognition sequence that are 5' upstream of a cleavage site generated by an engineered nuclease. Similarly, the term “3' portion” when referring to a nuclease recognition sequence is intended to mean the nucleotides of the recognition sequence that are 3' downstream of a cleavage site generated by an engineered nuclease. By way of example, where an engineered nuclease, such as a CRISPR / Cas9 nuclease system, generates a blunt end cleavage site

[0256] 33

[0257] WBD (US) 4882-8896-8689vl

[0258] P893392050WO (01287) within a recognition sequence, the 5' portion of the recognition sequence comprises the nucleotides of the sequence that are 5' upstream of the cleavage site, while the 3' portion of the recognition sequence comprises the nucleotides of the sequence that are 3' downstream of the cleavage site. For nucleases that generate 5' or 3' overhangs (e.g., engineered meganucleases), each 5' portion or 3' portion can, in some examples, include only the nucleotides that are 5' upstream or 3' downstream of the cleavage site, respectively.

[0259] In other examples, each 5' portion or 3' portion can include the nucleotides that are 5' upstream or 3' downstream of the cleavage site, respectively, as well as the nucleotides of the 5' or 3' overhang. By way of example, an I-Crel-derived engineered meganuclease generates a cleavage site between nucleotide positions 13 and 14 of its 22 base pair recognition sequence, wherein positions 9-13 represent the 4 base pair 3' overhang. Thus, in some examples, the 5' portion of the sequence can comprise the nucleotides of positions 1-13, and the 3' portion of the sequence can comprise the nucleotides of positions 14-22. In other examples, the 5' portion of the sequence can comprise the nucleotides of positions 1-13, and the 3' portion of the sequence can comprise the nucleotides of positions 10-22, which includes the 4 base pair overhang and the 9 base pairs of the 3' half site. In some cases, the inclusion or exclusion of the overhang nucleotides as components of a homology arm may be used to affect homology directed repair. It is understood that in examples wherein a 5' portion of a recognition sequence and a 3' portion of a recognition sequence are positioned adjacent to one another for the purpose of generating a complete recognition sequence, only the portion that normally includes the overhang base pairs will comprise the overhang nucleotides.

[0260] As used herein, the term “operably linked” is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a nucleic acid sequence encoding a nuclease and a regulatory sequence (e.g., a promoter) is a functional link that allows for expression of the nucleic acid sequence encoding the nuclease. Operably linked elements may be contiguous or non-contiguous. When used to refer to the joining of two protein coding regions, by operably linked is intended that the coding regions are in the same reading frame.

[0261] As used herein, the term “promoter” or “regulatory sequence” refers to a nucleic acid sequence which is required for expression of a gene product operably linked to the promoter or regulatory sequence. In some instances, this sequence may be the core promoter sequence and in other instances, this sequence may also include an enhancer sequence and other regulatory elements which are required for expression of the gene product. The promoter / regulatory sequence may, for example, be one which expresses the gene product in a tissue specific manner.

[0262] 34

[0263] WBD (US) 4882-8896-8689vl

[0264] P893392050WO (01287) As used herein, the term “delivery vehicle” refers to a delivery mechanism for contacting a cell, an organism, and / or a subject with one or more exogenous polynucleotides. Delivery vehicles may be, for example, recombinant viruses, such as an adeno-associated virus (AAV), or lipid nanoparticles.

[0265] As used herein, the term “adeno-associated virus,” “adeno-associated viral particles,” “adeno-associated viral vectors,” or “AAV” or “AAV particles” or “AAV vectors” are used interchangeably and refer to viruses and vectors derived from the Parvoviridcie family. Adeno- associated viruses can infect both dividing and quiescent cells. Adeno-associated viruses of any serotype can be used in the presently disclosed methods and compositions, as well as variants thereof. As used herein, the term “serotype” refers to a distinct variant within a species of virus that is determined based on the viral cell surface antigens. Known serotypes of AAV include, for example, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11 (Weitzman and Linden (2011) In Snyder and Moullier Adeno-associated virus methods and protocols. Totowa, NJ: Humana Press). Other serotypes of AAV are known, and a person of skill in the art could determine their usefulness in the present disclosure, including, for example, AAV 13, AAV14, AAV15, AAV16, AAV.rh8, AAV.rhlO, AAV.rh20, AAV.rh39, AAV.rh74, AAV.rh79, AAV.RHM4-1, AAV.hu37, AAV.hu68, AAV.Anc80, AAV.Anc90L65, AAV.7m8, AAV. PHP. B, AAV2.5, AAV2tYF, AAV3B, AAV.LK03, AAV.HsCl, AAV.HSC2, AAV.HSC3, AAV.HSC4, AAV.HSC5, AAV.HSC6, AAV.HSC7, AAV.HSC9, AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAVMYO, MYOAAV and AAV.HSC16 serotypes, such as those described in PCT Publication No. WO 2022 / 133051, incorporated herein by reference. An AAV particle may be of a single serotype or a combination of serotypes, or derivatives thereof. Generally, adeno-associated viruses comprise a transgene or portion thereof that is flanked by parvoviral or AAV inverted terminal repeat sequences (ITRs). Such AAV vectors can be replicated and packaged into infectious viral particles when present in a host cell that is expressing AAV replication (rep) and capsid (cap) gene products. Methods of producing recombinant AAV particles are described in the art, including, for example, PCT Publication No. WO 2022 / 133051.

[0266] As used herein, the term “inverted terminal repeat” or “ITR” refers to regions found at the 5' and 3' termini of an AAV genome. The ITRs are about 145 nt each that flank a transgene or portion thereof within an AAV genome. The ITRs are self-complementary and organized so that an energetically stable intramolecular duplex forming a T-shaped hairpin may be formed. These hairpin structures function as an origin for viral DNA replication, serving as primers for the cellular

[0267] 35

[0268] WBD (US) 4882-8896-8689vl

[0269] P893392050WO (01287) DNA polymerase complex. The ITRs also aid in concatamer formation in the nucleus and integration into the genome. Sequences of AAV-associated ITRs are known in the art, for example those disclosed by Yan et al., J. Virol. 79(l):364-379 (2005). ITR sequences that find use herein may be full length, wild-type AAV ITRs or fragments thereof that retain functional capability, or may be sequence variants of full-length, wild-type AAV ITRs that are capable of functioning in cis as origins of replication. AAV ITRs useful in the presently disclosed methods and compositions may derive from any known AAV serotype.

[0270] As used herein, the term “D sequence” refers to D-sequence, a stretch of nucleotides, which in some cases can be 20 nucleotides in length, that are associated with an AAV inverted terminal repeat but do not play a role in hairpin formation. The D sequence has been shown to play a role in the life cycle of AAV in a number of ways. For example, the D sequence acts as the packaging signal for AAV. Further, the first 10 nucleotides of a D sequence have been shown to be necessary for AAV DNA replication (see, Kwon et al., Human Gene Therapy (2020), Vol. 31 (9-10): 565- 574).

[0271] As used herein, the term “lipid nanoparticle” refers to a lipid composition having atypically spherical structure with an average diameter between 10 and 1000 nanometers. In some formulations, lipid nanoparticles can comprise at least one cationic lipid, at least one non-cationic lipid, and at least one conjugated lipid. Lipid nanoparticles known in the art that are suitable for encapsulating nucleic acids, such as mRNA, are contemplated for use in the presently disclosed methods. A nucleic acid, such as an mRNA, may be encapsulated in the lipid portion of the lipid nanoparticle or aqueous space enveloped by some or all of the lipid portion of the lipid nanoparticle. This affords protection from enzymatic degradation or other undesired effects induced by a cell, an organism, and / or a subject contacted with the lipid nanoparticle. Lipid nanoparticles may further be conjugated with a targeting moiety, such as an antibody or a ligand, to direct the lipid nanoparticle to a target cell or target tissue.

[0272] As used herein, the term “therapeutically effective amount” refers to an amount sufficient to effect beneficial or desirable biological and / or clinical results, thereby treating a subject.

[0273] As used herein, the term “treat,” “treating,” or “treatment” refers to the administration of a pharmaceutical composition disclosed herein, comprising, for example an exogenous polynucleotide encoding a replacement nucleic acid sequence and an exogenous polynucleotide encoding a nuclease, to a subject having a disease, disorder, or condition. For example, the subject may have a disease characterized by a pathogenic allele, wherein the treatment corrects the pathogenic allele to a wild-type allele.

[0274] 36

[0275] WBD (US) 4882-8896-8689vl

[0276] P893392050WO (01287) As used herein, the recitation of a numerical range for a variable is intended to convey that the present disclosure may be practiced with the variable equal to any of the values within that range. Thus, for a variable which is inherently discrete, the variable can be equal to any integer value within the numerical range, including the end-points of the range. Similarly, for a variable which is inherently continuous, the variable can be equal to any real value within the numerical range, including the end-points of the range. As an example, and without limitation, a variable which is described as having values between 0 and 2 can take the values 0, 1 or 2 if the variable is inherently discrete, and can take the values 0.0, 0.1, 0.01, 0.001, or any other real values =0 and =2 if the variable is inherently continuous.

[0277] 2, Principle of the Invention

[0278] The present disclosure is based, at least in part, on the discovery that a novel nuclease- initiated homology directed repair (HDR) method, referred to herein as HDR repair and replacement, can be used to predictably replace an endogenous genomic sequence downstream of a double strand break with a specified replacement nucleic acid sequence.

[0279] Unlike other approaches, HDR repair and replacement is not designed to replace one genomic sequence with an entirely different genomic sequence, such as replacing one gene with a different gene. Rather, the donor template utilized for HDR repair and replacement is designed to replace a specific genomic region, downstream of a double strand break, with a replacement nucleic acid sequence that: (i) is a nucleotide and / or codon-modified version of that same genomic region, and (ii) encodes the same amino acid sequence as the endogenous genomic sequence with at least one replacement modification (e.g., a nucleotide and / or codon change, insertion, or deletion). A replacement modification can be, for example, a modified nucleotide or a modified codon which, when incorporated into the gene, encodes a different amino acid than the endogenous nucleotide or codon. In other examples, a replacement modification can be one or more nucleotides or codons present in the replacement nucleic acid sequence that are not present in the intervening genomic sequence. When incorporated, such added nucleotides or codons can, for example, result in an insertion that alters and / or corrects a mutant gene sequence, or can introduce a new sequence. Conversely, such added nucleotides or codons can, for example, result in an insertion that alters a wild type gene sequence to encode a non-wild type gene sequence. Thus, the resulting modified gene will express the same protein as encoded by the endogenous genomic sequence including the desired modification(s).

[0280] 37

[0281] WBD (US) 4882-8896-8689vl

[0282] P893392050WO (01287) The donor template utilized in HDR repair and replacement comprises the replacement nucleic acid sequence flanked by upstream and downstream homology arms, which have homology to sequences in the genome upstream and downstream of the double strand break and help enable homologous recombination. The sequences of the homology arms are designed to specify the endogenous genomic region that will be replaced. Particularly, the downstream homology arm has homology to a sequence, referred to herein as the second genomic sequence, that is downstream, and not adjacent to, the double strand break. The intervening genomic sequence between the double strand break and the second genomic sequence is ultimately the region of the genome that is substituted with the replacement nucleic acid sequence provided by the donor template. The homology arms can be designed to allow this intervening genomic sequence to be as small as few as several base pairs or as large as 10,000 base pairs or more.

[0283] As previously discussed, the replacement nucleic acid is a nucleotide and / or codon- modified version of the genomic region being replaced (with at least one replacement modification). This modification of the replacement sequence relative to the endogenous sequence prevents the occurrence of premature cross-over events before the second homology arm. Such premature cross-over events would result in a partial replacement of the endogenous sequence, rather than introduction of the entire replacement nucleic acid sequence. Reducing the overall homology between the intervening genomic sequence and the replacement nucleic acid sequence reduces or eliminate microhomologies between the two where cross-over events could occur.

[0284] A variety of predictable outcomes can be achieved by HDR repair and replacement including, for example, modification of a mutant gene (e.g., a gene comprising one or more mutant nucleotides / codons, one or more nucleotide deletions, and / or one or more nucleotide insertions) such that it expresses a wild-type protein.

[0285] 3 , Methods of Nuclease-Initiated Homology Directed Repair and Replacement

[0286] Site-specific nucleases generate a double strand break in the genome of a cell, which can result in permanent modification of the genome via homologous recombination with a donor DNA sequence. The use of nucleases to induce a double strand break in a target locus stimulates homologous recombination of DNA sequences that are flanked by sequences having homology to the genomic target; a process known as homology directed repair (HDR). The natural mechanism of HDR can be utilized to integrate exogenous DNA, such as the replacement nucleic acid sequences and replacement modifications described herein, at a specific site in the genome. As discussed, the HDR repair and replacement approach described herein utilizes a nuclease to induce

[0287] 38

[0288] WBD (US) 4882-8896-8689vl

[0289] P893392050WO (01287) a double strand break and a DNA repair template comprising a replacement nucleic acid sequence flanked by homology arms.

[0290] 3 , 1 Nucleases

[0291] The presently disclosed methods and compositions utilize nucleases to cleave nuclease recognition sequences, generating a double strand break, thereby allowing the integration of an exogenous nucleic acid sequence, i.e., a replacement nucleic acid sequence described herein, into a genomic locus via HDR. Non-limiting examples of nucleases useful in the present disclosure include engineered meganucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), compact TALENs, megaTALs, and CRISPR system nucleases.

[0292] Engineered Meganucleases

[0293] In one aspect, the methods and compositions described herein use an engineered meganuclease to generate a DSB in the genome. Meganucleases are homing endonucleases that recognize between 12 to 40 base pairs, thus conferring stringent site specificity. Meganucleases can be engineered to alter DNA-binding specificity, DNA cleavage activity, DNA-binding affinity, or dimerization properties. A non-limiting example of a meganuclease is the I-Crel meganuclease, and engineered meganucleases derived from the I-Crel meganuclease that are modified (e.g., as described above) to alter DNA-binding specificity, DNA cleavage activity, and / or DNA-binding affinity. Methods for producing modified I-Crel meganucleases are known in the art, see, e.g., PCT Publication Nos. WO 2007 / 047859, WO 2017 / 062439, WO 2017 / 062451, and WO 2019 / 200122, incorporated herein by reference. I-Crel-derived engineered meganucleases cleave a 22 base pair recognition sequence in DNA that comprises two half-sites, each comprising nine base pairs, separated by a four base pair center sequence. Cleavage by I-Crel-derived engineered meganucleases occurs between position 13 and 14 of the recognition sequence and results in the formation of a cleavage site having a four base pair, 3’ overhang on each strand. More specifically, on each strand, a cleavage site is formed that comprises positions 1-13 on the 5’ end and positions 14-22 on the 3’ end. The presence of a 3’ overhang is believed to enhance the efficiency of homologous recombination, and thereby contribute to the efficiency of HDR repair and replacement. Thus, in some particular embodiments, the double strand break utilized in HDR repair and replacement comprises a 3’ overhang, and particularly a four base pair 3’ overhang.

[0294] 39

[0295] WBD (US) 4882-8896-8689vl

[0296] P893392050WO (01287) Nucleases referred to as megaTALs are single-chain endonucleases comprising a transcription activator-like effector (TALE) DNA binding domain with an engineered, sequencespecific homing endonuclease.

[0297] Zinc Finger Nucleases (ZFNs)

[0298] In some embodiments, HDR repair and replacement utilizes a zinc finger nuclease (ZFN) to generate a DSB in the genome. ZFNs are chimeric proteins comprising a zinc finger DNA-binding domain fused to a nuclease domain derived from an endonuclease or exonuclease, such as the Type II restriction endonuclease Fokl. The zinc finger DNA-binding domain can be a native sequence or can be redesigned through rational or experimental means to produce a protein which binds to a pre -determined DNA sequence of approximately 18 base pairs in length. By fusing the zinc finger DNA-binding domain to the nuclease domain, the resultant chimeric ZFN can be used to cleave the genome at the target site, i.e., the pre-determined DNA sequence. The engineering of ZFNs, particularly the zinc finger DNA-binding domain is further described in U.S. Pat. Nos. 6,453,242; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,030,215; 6,794,136; 7,067,317; 7,262,054; 7,070,934; 7,361,635; and 7,253,273, and U.S. Patent Application No. US 2023 / 0295563, all incorporated herein by reference.

[0299] Transcription Activator-Like Effector Nucleases (TALENs)

[0300] In one aspect, the methods and compositions described herein use a transcription activatorlike effector nuclease (TAEEN) to generate a DSB in the genome. TAL-effector nucleases (TALENs) are chimeric proteins comprising a site-specific DNA-binding domain, a TAL effector, fused to an endonuclease or exonuclease, such as the Type II restriction endonuclease Fokl. The DNA-binding domain of TALENs comprises TAL domain repeats, which are 33-34 amino acid sequences with divergent 12th and 13th amino acids. These two positions, referred to as the repeat variable dipeptide (RVD), are highly variable and show a strong correlation with specific nucleotide recognition. Each base pair in the DNA target sequence is contacted by a single TAL repeat with the specificity resulting from the RVD. As the tandem array of TAL-effector domains each recognize a single DNA base pair, the TALEN can be engineered to cleave the genome at the target site conferred by the DNA-binding domain.

[0301] Clustered Regularly Interspaced Short Palindromic Repeats (CRISP R)-Cas Nucleases

[0302] 40

[0303] WBD (US) 4882-8896-8689vl

[0304] P893392050WO (01287) In one aspect, the methods and compositions described herein use a CRISPR / Cas systemnuclease to generate a DSB in the genome. CRISPR / Cas systems comprise two components: (1) a CRISPR nuclease; and (2) a guide RNA (gRNA) comprising a ~20 nucleotide targeting sequence that confers site -specificity to the nuclease. The CRISPR system may further comprise a trans-activating RNA (tracrRNA) which complexes with the guide RNA. The gRNA and the tracrRNA, if present, may be the same or separate molecules, referred to as a single guide RNA (sgRNA). Methods for engineering gRNAs to confer target site-specificity are known in the art, and further described in, for example, PCT Publication No. WO 2016 / 182959, incorporated herein by reference.

[0305] There are a number of CRISPR nucleases known in the art that can be used to generate DSB in the genome of a cell, a subject, and / or an organism (see, e.g., Xu and Li (2020) Comput Struct Biotechnol J. 18:2401-2415). In some embodiments, the nuclease is a Cas9 nuclease. Cas9 nucleases comprise a recognition domain (REC) and a nuclease domain. The REC domain is important for recognition of the gRNA, thereby conferring site-specificity. The nuclease domain comprises a RuvC domain, a HNH domain, and a PAM-interacting domain. The RuvC domain is structurally similar to retroviral integrase superfamily members and cleaves a single strand, e.g., the non-complementary strand targeted by the gRNA. The HNH domain shares structural similarity with HNH endonucleases, and cleaves a single strand, e.g., the complementary strand targeted by the gRNA. The PAM-interacting domain interacts with the PAM targeted by the gRNA. Cas9 nucleases are further described in PCT Publication No. WO 2016 / 182959. Non-limiting examples of Cas nucleases Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, CaslO, Cpfl, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, and Cmr6.

[0306] 3,2 Exogenous Polynucleotide Repair Construct and Replacement Nucleic Acid Sequences

[0307] The HDR repair and replacement approach described herein utilizes a donor template (or repair construct, used interchangeably) comprising a replacement nucleic acid sequence in tandem with a nuclease to replace a specified region of the genome of a cell, an organism, and / or a subject.

[0308] The repair constructs generally comprise the following:

[0309] 5’ [First homology arm] -[replacement nucleic acid sequence]-[second homology arm] 3’. The replacement nucleic acid sequences may comprise one or more replacement modifications, such as modified nucleotides. In some embodiments, the replacement nucleic acid sequence comprises one or more modified nucleotides. In some embodiments, at least 20% of the

[0310] 41

[0311] WBD (US) 4882-8896-8689vl

[0312] P893392050WO (01287) nucleotides in the replacement nucleic acid sequence are modified nucleotides. In some embodiments, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least

[0313] 35%, at least 40%, at least 45%, or at least 50% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides. In some embodiments, between 1-50% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides. In some embodiments, between 1-2%,

[0314] 2-3%, 3-4%, 4-5%, 5-6%, 6-7%, 7-8%, 8-9%, 9-10%, 10-15%, 15-20%, 20-25%, 25-30%, 30-35%,

[0315] 35-40%, 40-45%, or 45-50% of the nucleotides in the replacement nucleic acid sequence are modified nucleotides.

[0316] In yet other embodiments, the replacement nucleic acid sequence comprises one or more modified codons. In some embodiments, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% of the codons in the replacement nucleic acid sequence are modified codons. In some embodiments, at least 75% of the codons in the replacement nucleic acid sequence are modified codons. In some embodiments, between 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40- 45%, 45-50%, 50-55%, 55-60%, 60-65%, 65-70%, 70-75%, 75-80%, 80-85%, 85-90%, 90-95%, or 95-99% of the codons in the replacement nucleic acid sequence are modified codons. In some embodiments, all of the codons in the replacement nucleic acid sequence are modified codons.

[0317] The first and second homology arms, which have homology to a first and second genomic sequence, respectively, allow for recombination of the replacement sequence into the chromosome via homologous directed repair (HDR).

[0318] In some embodiments, the first homology arm comprises 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 nucleotides, or any number in between the values.

[0319] In some embodiments, the second homology arm comprises 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2100, 2200, 2300, 2400, 2500, 2600, 2700, 2800, 2900, or 3000 nucleotides, or any number in between the values.

[0320] In one aspect, the first and second homology arms are approximately the same length. In some embodiments, the first homology arm and the second homology arm are between about 100 to 2000 nucleotides in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 nucleotides in length. In some embodiments, the first

[0321] 42

[0322] WBD (US) 4882-8896-8689vl

[0323] P893392050WO (01287) homology arm and the second homology arm are between about 400 to 600 nucleotides in length. In some embodiments, the first homology arm and the second homology arm are about 400 nucleotides in length. In some embodiments, the first homology arm and the second homology arm are about 500 nucleotides in length.

[0324] In one aspect, the first and second homology arms are different lengths. In some embodiments, the first homology arm is longer than the second homology arm. In some embodiments, the second homology arm is longer than the first homology arm. In some embodiments, the first homology arm is between about 100 to 1000 base pairs in length and the second homology arm is between about 1000 to 3000 base pairs in length. In some embodiments, the first homology arm is between about 200 to 800 base pairs in length and the second homology arm is between about 1200 to 1800 base pairs in length. In some embodiments, the first homology arm is between about 400 to 600 base pairs in length and the second homology arm is between about 1400 to 1600 base pairs in length. In some embodiments, the first homology arm is about 400 or about 500 base pairs in length and the second homology arm is about 1500 or about 1600 base pairs in length. In some embodiments, the first homology arm is between about 400 to 600 base pairs in length and the second homology arm is between about 1800 to 2500 base pairs in length. In some embodiments, the first homology arm is between about 400 to 600 base pairs in length and the second homology arm is between about 2000 to 2500 base pairs in length. In some embodiments, the second homology arm is between about 100 to 1000 base pairs in length and the first homology arm is between about 1000 to 3000 base pairs in length. In some embodiments, the second homology arm is between about 200 to 800 base pairs in length and the first homology arm is between about 1200 to 1800 base pairs in length. In some embodiments, the second homology arm is between about 400 to 600 base pairs in length and the first homology arm is between about 1400 to 1600 base pairs in length. In some embodiments, the second homology arm is about 400 or about 500 base pairs in length and the first homology arm is about 1500 or about 1600 base pairs in length. In some embodiments, the second homology arm is between about 400 to 600 base pairs in length and the first homology arm is between about 1800 to 2500 base pairs in length. In some embodiments, the second homology arm is between about 400 to 600 base pairs in length and the first homology arm is between about 2000 to 2500 base pairs in length.

[0325] In some embodiments, the first homology arm is configured such that it is immediately upstream and adjacent to the DSB (i.e., the target site of the nuclease).

[0326] In one aspect, the DSB and a homology arm are separated by an intervening genomic sequence. That is, the 5’ end of the second homology arm and the DSB may be separated by an

[0327] 43

[0328] WBD (US) 4882-8896-8689vl

[0329] P893392050WO (01287) intervening genomic sequence. Additionally or alternatively, the 3’ end of the first homology arm and the DSB may be separated by an intervening genomic sequence.

[0330] Without wishing to be bound by any particular theory, it is believed that increasing the length of the homology arm allows for increased length of the intervening genomic sequence while maintain efficient homology directed repair and replacement with the replacement nucleic acid sequence.

[0331] Accordingly, the intervening genomic sequence may vary in length. In some embodiments, the intervening genomic sequence is between about 500 to about 25,000 base pairs. In some embodiments, the intervening genomic sequence is between about 500 to 12,500 base pairs. In some embodiments, the intervening genomic sequence is between about 1000 to 10,000 base pairs. In some embodiments, the intervening genomic sequence is between about 2500 to 10,000 base pairs. In some embodiments, the intervening genomic sequence is about 500, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 9500, 10,000, 10,500, 11,000, 11,500, 12,000, or 12,500 base pairs, or any number in between the values. In some embodiments, the intervening genomic sequence is between 2000-3000 base pairs. In some embodiments, the intervening genomic sequence is 4,500 to 5,500 base pairs.

[0332] Similarly, the replacement nucleic acid sequence may vary in length as well. In some embodiments, the replacement nucleic acid sequence is between about 1 to about 100 base pairs, about 1 to about 250 base pairs, about 1 to about 500 base pairs, about 1 to about 750 base pairs, about 1 to about 1000 base pairs, about 1 to about 1250 base pairs, about 1 to about 1500 base pairs, about 1 to about 1750 base pairs, about 1 to about 2000 base pairs, about 1 to about 2250 base pairs, about 1 to about 2500 base pairs, about 1 to about 2750 base pairs, about 1 to about 3000 base pairs, about 1 to about 3250 base pairs, about 1 to about 3500 base pairs, about 1 to about 3750 base pairs, about 1 to about 4000 base pairs, about 1 to about 4250 base pairs, about 1 to about 4500 base pairs, about 1 to about 4750 base pairs, or about 1 to about 5000 base pairs or any number in between the values. In some embodiments, the replacement nucleic acid sequence is between about 10 to about 100 base pairs, about 25 to about 250 base pairs, about 50 to about 500 base pairs, about 75 to about 750 base pairs, about 100 to about 1000 base pairs, about 250 to about 1250 base pairs, about 500 to about 1500 base pairs, about 750 to about 1750 base pairs, about 1000 to about 2000 base pairs, about 1250 to about 2250 base pairs, about 1500 to about 2500 base pairs, about 1750 to about 2750 base pairs, about 2000 to about 3000 base pairs, about 2250 to about 3250 base pairs, about 2500 to about 3500 base pairs, about 2750 to about 3750 base pairs, about 3000 to about 4000 base pairs, about 3250 to about 4250 base pairs, about 3500 to about

[0333] 44

[0334] WBD (US) 4882-8896-8689vl

[0335] P893392050WO (01287) 4500 base pairs, about 3750 to about 4750 base pairs, or about 4000 to about 5000 base pairs or any number in between the values.

[0336] In one aspect, the replacement nucleic acid sequence and the intervening genomic sequence differ in homology. In some embodiments, the replacement nucleic acid sequence has less than 70% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has less than 60%, less than 50%, or less than 40% homology to the intervening genomic sequence. In some embodiments, the replacement nucleic acid sequence has between about 1-5%, 5-10%, 10-20%, 20-30%, 30-40%, or 40-50% homology to the intervening genomic sequence.

[0337] The replacement nucleic acid sequence and the intervening genomic sequence also differ in the number of shared consecutive nucleotides. This difference is important to avoid the repair of the double-strand break via microhomology-mediated end joining (MMEJ). MMEJ is an alternative pathway for repairing DSB, relying on microhomologous sequences to align broken strands. Unlike HDR, MMEJ is highly error prone and can result in undesired insertions and deletions. MMEJ is further described in Sfeir et al., Annu Rev Cell Dev Biol (2024) “Microhomology-Mediated End Joining Chronicles Tracing the Evolutionary Footprints of Genome Protection”. To minimize repair via MMEJ mechanisms, the replacement nucleic acid sequence and the intervening genomic sequence do not share more than 50 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening genomic sequence do not share more than 25 consecutive nucleotides in common, more than 10 consecutive nucleotides in common, or more than 5 consecutive nucleotides in common. In some embodiments, the replacement nucleic acid sequence and the intervening genomic sequence do not share more than 1-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, or 45-50 consecutive nucleotides in common.

[0338] Figures 1-13 provide exemplary, non-limiting depictions of replacement nucleic acid sequence constructs and the resultant homology-directed repair and replacement (referred to herein as “HDR and Replacement” according to methods of the present disclosure. For the purposes of Figures 1-13, reference is made to hypothetical exons, introns, and the like.

[0339] Figure 1 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, and the first and second genomic sequences are within the same cleaved exon. As shown, the repair construct comprises a first homology arm to a first genomic sequence in exon 1, a replacement nucleic acid sequence comprising a replacement modification, and a second homology arm to a second genomic sequence in exon 1. In this example, the HDR repair and replacement approach leads to the replacement of

[0340] 45

[0341] WBD (US) 4882-8896-8689vl

[0342] P893392050WO (01287) an intervening genomic sequence within the cleaved exon with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0343] Figure 2 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence is within a downstream exon. The repair construct comprises a first homology arm to a first genomic sequence in exon 1, a replacement nucleic acid sequence comprising a replacement modification and a 5 ’ region of a downstream exon 2, and a second homology arm to a second genomic sequence in a portion of exon 2. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. In the particular configuration shown in the figure, the second genomic sequence does not begin at the start of the downstream exon; therefore, the 5 ’ region of the downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence. In other configurations, the second genomic sequence could begin at the start of the downstream exon, in which case a modified variant of a portion of the downstream exon would not be necessary in the replacement nucleic acid sequence. In either case, in some examples, any intron(s) between the cleaved exon and the downstream exon can be omitted from the replacement nucleic acid sequence, resulting in their removal in the final genomic sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved exon and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0344] Figure 3 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence begins within a downstream exon. The repair construct comprises a first homology arm to a first genomic sequence in exon 1, a replacement nucleic acid sequence comprising a replacement modification and a 5 ’ region of a downstream exon 2, and a second homology arm to a second genomic sequence in a portion of exon 2 and a downstream intron. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins in the downstream exon and further spans into an adjacent downstream intron. In this particular configuration, the second genomic sequence does not

[0345] 46

[0346] WBD (US) 4882-8896-8689vl

[0347] P893392050WO (01287) begin at the start of the downstream exon; therefore, the 5 ’ region of the downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence. In other configurations, the second genomic sequence could begin at the start of the downstream exon, in which case a modified variant of a portion of the downstream exon would not be necessary in the replacement nucleic acid sequence. In either case, in some examples, any intron(s) between the cleaved exon and the downstream exon can be omitted from the replacement nucleic acid sequence, resulting in their removal in the final genomic sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved exon and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0348] Figure 4 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of the DSB in the same cleaved exon, and the second genomic sequence is within a downstream intron. The repair construct comprises a first homology arm to a first genomic sequence in exon 1, a replacement nucleic acid sequence comprising a replacement modification and a portion of an intron (e.g., replacement intron), and a second homology arm to a second genomic sequence in a portion of an intron. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence does not begin at the start of the downstream intron; therefore, the 5’ region of the downstream exon is replaced by replacement intron sequence that is included in the replacement nucleic acid sequence. The replacement intron sequence should comprise any necessary intronic elements that ensure proper gene transcription (e.g., a splice donor sequence, enhancers, etc.) but differ from the endogenous intron sequence being replaced. In other configurations, the second genomic sequence could begin at the start of the downstream intron, in which case a replacement intron sequence would not be necessary in the replacement nucleic acid sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved exon and one or more downstream introns (and potentially exons), with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0349] Figure 5 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an exon, the first genomic sequence is upstream of

[0350] 47

[0351] WBD (US) 4882-8896-8689vl

[0352] P893392050WO (01287) the DSB in the same cleaved exon, and the second genomic sequence begins within a downstream intron. The repair construct comprises a first homology arm to a first genomic sequence in exon 1, a replacement nucleic acid sequence comprising a replacement modification and a portion of an intron (e.g., replacement intron), and a second homology arm to a second genomic sequence in a portion of an intron and downstream exon 2. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins in the downstream intron and further spans into an adjacent downstream exon. Similar to Figure 4, this figure shows a particular example in which the second genomic sequence does not begin at the start of the downstream intron; therefore, the 5 ’ region of the downstream exon is replaced by replacement intron sequence that is included in the replacement nucleic acid sequence. The replacement intron sequence should comprise any necessary intronic elements that ensure proper gene transcription (e.g., a splice donor sequence, enhancers, etc.) but differ from the endogenous intron sequence being replaced. In other configurations, the second genomic sequence could begin at the start of the downstream intron, in which case a replacement intron sequence would not be necessary in the replacement nucleic acid sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved exon and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0353] Figure 6 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, and the first and second genomic sequences are within the same cleaved intron. The repair construct comprises a first homology arm to a first genomic sequence in intron 1, a replacement nucleic acid sequence comprising a replacement modification, and a second homology arm to a second genomic sequence in intron 1. In this example, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence within the cleaved intron with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0354] Figure 7 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same intron, and the second genomic sequence is within a downstream exon. The repair construct comprises a first homology arm to a first genomic sequence in intron 1, a

[0355] 48

[0356] WBD (US) 4882-8896-8689vl

[0357] P893392050WO (01287) replacement nucleic acid sequence comprising a replacement modification and a 5 ’ region of a downstream exon, and a second homology arm to a second genomic sequence in the downstream exon. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. In the particular configuration shown in the figure, the second genomic sequence does not begin at the start of the downstream exon; therefore, the 5 ’ region of the downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence. In other configurations, the second genomic sequence could begin at the start of the downstream exon, in which case a modified variant of a portion of the downstream exon would not be necessary in the replacement nucleic acid sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved intron and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0358] Figure 8 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence begins within a downstream exon. The repair construct comprises a first homology arm to a first genomic sequence in intron 1, a replacement nucleic acid sequence comprising a replacement modification and a 5’ region of a downstream exon, and a second homology arm to a second genomic sequence in the downstream exon and intron 2. The downstream exon can be the next exon in the gene (as shown) or can be an exon that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins in the downstream exon and further spans into an adjacent downstream intron. In this particular configuration, the second genomic sequence does not begin at the start of the downstream exon; therefore, the 5 ’ region of the downstream exon is replaced by a modified variant that is included in the replacement nucleic acid sequence. In other configurations, the second genomic sequence could begin at the start of the downstream exon, in which case a modified variant of a portion of the downstream exon would not be necessary in the replacement nucleic acid sequence. In either case, in some examples, any intron(s) between the cleaved intron and the downstream exon can be omitted from the replacement nucleic acid sequence, resulting in their removal in the final genomic sequence. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved intron and one or more downstream introns and exons, with the replacement nucleic

[0359] 49

[0360] WBD (US) 4882-8896-8689vl

[0361] P893392050WO (01287) acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0362] Figure 9 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence is within a downstream intron. The repair construct comprises a first homology arm to a first genomic sequence in intron 1, a replacement nucleic acid sequence comprising a replacement modification and a modified downstream exon, and a second homology arm to a second genomic sequence in downstream intron 2. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins at the start of the downstream intron, so no intron replacement sequence is necessary in the replacement nucleic acid sequence. In other examples, the second genomic sequence does not begin at the start of the downstream intron; therefore, the 5 ’ region of the downstream exon is replaced by replacement intron sequence that is included in the replacement nucleic acid sequence. The replacement intron sequence should comprise any necessary intronic elements that ensure proper gene transcription (e.g., a splice donor sequence, enhancers, etc.) but differ from the endogenous intron sequence being replaced. Also as shown, any intervening exons between the cleaved and downstream introns should be modified variants that encode the same amino acid sequence, with the exception that the replacement modification can either be introduced into the cleaved intron or into the modified exon. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved intron and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0363] Figure 10 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene, wherein the DSB is generated within an intron, the first genomic sequence is upstream of the DSB in the same cleaved intron, and the second genomic sequence begins within a downstream intron. The repair construct comprises a first homology arm to a first genomic sequence in intron 1, a replacement nucleic acid sequence comprising a replacement modification and a modified downstream exon, and a second homology arm to a second genomic sequence in downstream intron 2 and a portion of a downstream exon. The downstream intron can be the next intron in the gene (as shown) or can be an intron that is further downstream in the gene. The figure shows a particular example in which the second genomic sequence begins at the start of the downstream

[0364] 50

[0365] WBD (US) 4882-8896-8689vl

[0366] P893392050WO (01287) intron, so no intron replacement sequence is necessary in the replacement nucleic acid sequence. In other examples, the second genomic sequence does not begin at the start of the downstream intron; therefore, the 5 ’ region of the downstream exon is replaced by replacement intron sequence that is included in the replacement nucleic acid sequence. The replacement intron sequence should comprise any necessary intronic elements that ensure proper gene transcription (e.g., a splice donor sequence, enhancers, etc.) but differ from the endogenous intron sequence being replaced. Also as shown, any intervening exons between the cleaved and downstream introns should be modified variants that encode the same amino acid sequence, with the exception that the replacement modification can either be introduced into the cleaved intron or into the modified exon. As further shown, the second genomic sequence can span into the exon immediately downstream of the downstream intron. In other examples, the second genomic sequence can further span into additional introns and exons further downstream. In the example shown, the HDR repair and replacement approach leads to the replacement of an intervening genomic sequence, which spans the cleaved intron and one or more downstream introns and exons, with the replacement nucleic acid sequence, which comprises at least one replacement modification. The result is that the gene encodes the same protein with the replacement modification(s).

[0367] Figure 11 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises a mutant nucleotide or codon, leading to knockout of a gene or production of a mutant protein. The repair construct comprises a first homology arm to a first genomic sequence in a portion of exon 1, a replacement nucleic acid sequence comprising a replacement modification, and a second homology arm to a second genomic sequence in a portion of exon 1. As shown, by introducing a modified nucleotide or codon into the replacement nucleic acid sequence, the method can be utilized to change an endogenous mutant nucleotide or codon to a wild-type nucleotide or codon that allows the gene to encode a wild-type protein. This figure uses the configuration of Figure 1 as an example but can be applied to any configurations illustrated and / or described herein.

[0368] Figure 12 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises a deleted nucleotide or codon, leading to knockout of the gene or production of a mutant protein. The repair construct comprises a first homology arm to a first genomic sequence in a portion of exon 1, a replacement nucleic acid sequence comprising a replacement modification, and a second homology arm to a second genomic sequence in a portion of exon 1. As shown, by introducing the deleted nucleotide or codon into the replacement nucleic acid sequence, the method can be utilized to re-introduce the deleted nucleotide or codon into the

[0369] 51

[0370] WBD (US) 4882-8896-8689vl

[0371] P893392050WO (01287) gene, allowing it to encode a wild-type protein. This figure uses the configuration of Figure 1 as an example but can be applied to any configurations illustrated and / or described herein.

[0372] Figure 13 provides a graphic depiction of a nuclease-initiated HDR repair and replacement in a gene that comprises one or more extra nucleotides, such as a nucleotide expansion or nucleotide repeat expansion, leading to knockout of the gene or production of a mutant protein. The repair construct comprises a first homology arm to a first genomic sequence in a portion of exon 1, a replacement nucleic acid sequence comprising a replacement modification, and a second homology arm to a second genomic sequence in a portion of exon 1. As shown, by introducing the correct number of nucleotides into the replacement nucleic acid sequence, the method can be utilized to remove the extra / expanded nucleotides from the mutant gene, allowing it to encode a wild-type protein. This figure uses the configuration of Figure 1 as an example but can be applied to any configurations illustrated and / or described herein.

[0373] 4, Methods for Delivering Nucleases and Repair Constructs

[0374] In one aspect, an exogenous polynucleotide comprising the nuclease and the repair constructs are administered to a cell, a subject, and / or an organism. In some embodiments, the exogenous polynucleotide comprises a gene encoding a nuclease. In some embodiments, the exogenous polynucleotide further comprises a repair construct, wherein the repair construct comprises a first homology arm, a replacement nucleic acid sequence, and a second homology arm. That is, the exogenous polynucleotide comprises both a gene encoding a nuclease and the repair construct.

[0375] In some embodiments, the gene encoding the nuclease is positioned 5 ’ upstream of the first homology arm of the repair construct. In some embodiments, the gene encoding the nuclease is positioned 3’ downstream of the second homology arm. In some embodiments, the exogenous polynucleotide comprises a promoter operably linked to the gene encoding the nuclease.

[0376] Alternatively, the nuclease and the repair construct may be on separate exogenous polynucleotides. That is, a first exogenous polynucleotide comprises a gene encoding a nuclease, and a second exogenous polynucleotide comprises a repair construct. In some embodiments, the first exogenous polynucleotide comprises a promoter operably linked to the gene encoding the nuclease.

[0377] The nucleases and repair constructs described herein (i.e., the polynucleotides comprising the nucleases and / or repair constructs) may be administered to a cell, a subject, and / or an organism using a number of delivery methods known in the art, including recombinant viruses such as an

[0378] 52

[0379] WBD (US) 4882-8896-8689vl

[0380] P893392050WO (01287) adeno-associated virus (AAV) or non-viral delivery particles such as lipid nanoparticles (i.e., non- viral). The administration may be performed in vitro or in vivo.

[0381] In some embodiments, an mRNA encodes a nuclease, such as an engineered meganuclease, is administered to a cell to reduce the likelihood that the gene encoding the nuclease will integrate into the genome. Such mRNA can be produced using methods known in the art such as in vitro transcription. In some embodiments, the mRNA is 5' capped using 7-methyl-guanosine, antireverse cap analogs (ARCA) (US 7,074,596), CleanCap® analogs such as Cap 1 analogs (Trilink, San Diego, CA), or enzymatically capped using vaccinia capping enzyme or similar. In some embodiments, the mRNA may be polyadenylated. The mRNA may contain various 5' and 3' untranslated sequence elements to enhance expression the encoded engineered meganuclease and / or stability of the mRNA itself. Such elements can include, for example, posttranslational regulatory elements such as a woodchuck hepatitis virus posttranslational regulatory element (WPRE). The mRNA may contain nucleoside analogs or naturally-occurring nucleosides, such as pseudouridine, 5 -methylcytidine, N6-methyladenosine, 5 -methyluridine, or 2-thiouridine. Additional nucleoside analogs include, for example, those described in US 8,278,036.

[0382] 4, 1 Recombinant Viruses

[0383] In one aspect, the exogenous polynucleotides comprising the repair construct and / or gene encoding the nuclease are administered to a cell, a subject, and / or an organism using a recombinant virus (i.e., a recombinant viral vector). Recombinant viruses are known in the art and include recombinant AAVs. Recombinant AAVs useful in the presently disclosed methods and compositions can have any serotype that allows for transduction of the virus into a target cell type and expression of the nuclease gene in the target cell. For example, in some embodiments, recombinant AAVs have a serotype (i.e., a capsid) of AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11. In other embodiments, the recombinant AAVs have a serotype of AAV 13, AAV14, AAV15, AAV16, AAV.rh8, AAV.rhlO, AAV.rh20, AAV.rh39, AAV.rh74, AAV.rh79, AAV.RHM4-1, AAV.hu37, AAV.hu68, AAV.Anc80, AAV.Anc90U65, AAV.7m8, AAV.PHP.B, AAV2.5, AAV2tYF, AAV3B, AAV.UK03, AAV.HsCl, AAV.HSC2, AAV.HSC3, AAV.HSC4, AAV.HSC5, AAV.HSC6, AAV.HSC7, AAV.HSC9, AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAVMYO, MYOAAV, and AAV.HSC16 serotypes, such as those described in PCT Publication No. WO 2022 / 133051, incorporated herein by reference. An AAV particle may be of a single serotype or a combination of serotypes, or derivatives thereof. Generally, adeno-associated

[0384] 53

[0385] WBD (US) 4882-8896-8689vl

[0386] P893392050WO (01287) viruses comprise a transgene or portion thereof that is flanked by parvoviral or AAV inverted terminal repeat sequences (ITRs). Such AAV vectors can be replicated and packaged into infectious viral particles when present in a host cell that is expressing AAV replication (rep) and capsid (cap) gene products. It is known in the art that different AAVs tend to localize to different tissues. Polynucleotides delivered by recombinant AAVs can include left (5') and right (3') ITRs as part of the viral genome. In some embodiments, the recombinant viruses are injected directly into target tissues. In alternative embodiments, the recombinant viruses are delivered systemically via the circulatory system.

[0387] The recombinant AAVs of the disclosure referred to as AAV herein may exhibit transduction and / or activity in a multitude of tissue types, including but not limited to liver tissue, spleen tissue, adrenal tissue, lung tissue and heart tissue.

[0388] If a polynucleotide comprising a gene encoding a nuclease is delivered to a cell by a recombinant virus (e.g., an AAV), the nucleic acid sequence (the gene) encoding the nuclease may be operably linked to a promoter. In some embodiments, the promoter may be a viral promoter that is endogenous to the recombinant virus, or any other promoter that is expressed in the desired cell type.

[0389] In some embodiments, the AAV may comprise a polynucleotide that comprises both a gene encoding a nuclease and a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm according to the disclosure.

[0390] Alternatively, a first AAV may comprise a polynucleotide that comprises a gene encoding a nuclease and a second AAV may comprise a polynucleotide comprising a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm according to the disclosure.

[0391] Alternatively, a first AAV may comprise an mRNA encoding a nuclease and a second AAV may comprise a polynucleotide comprising a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm according to the disclosure.

[0392] Alternatively, the AAV may comprise an mRNA encoding a nuclease and a polynucleotide comprising a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm.

[0393] 4,2 Lipid Nanoparticles (LNPs)

[0394] In one aspect, the exogenous polynucleotides comprising the repair construct and / or gene encoding the nuclease are administered to a cell, a subject, and / or an organism using a lipid

[0395] 54

[0396] WBD (US) 4882-8896-8689vl

[0397] P893392050WO (01287) nanoparticle. Examples of lipids and lipid nanoparticles compositions comprising the same are well known in the art, and are described in, for example, PCT Publication Nos. WO 2016164609, WO 2021 / 236930, WO 2022 / 115645, WO 2022 / 120388, WO 2023 / 183616, WO 2023 / 240156, WO 2018 / 089540, and WO 2020 / 257611, the entireties of which are incorporated herein by reference as they pertain to lipids and lipid nanoparticle compositions comprising the same.

[0398] If a polynucleotide comprising a gene encoding a nuclease is delivered to a cell by an LNP, the nucleic acid sequence (the gene) encoding the nuclease may be operably linked to a promoter. In some embodiments, the promoter may be any promoter that is expressed in the desired cell type.

[0399] In some embodiments, the LNP may comprise a polynucleotide that comprises both a gene encoding a nuclease and a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm of the disclosure.

[0400] Alternatively, a first LNP may comprise a polynucleotide that comprises a gene encoding a nuclease and a second LNP may comprise a polynucleotide comprising a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm of the disclosure.

[0401] Alternatively, a first LNP may comprise an mRNA encoding a nuclease and a second LNP may comprise a polynucleotide comprising a repair construct comprising a first homology arm, a replacement nucleic acid sequence, and a second homology arm of the disclosure.

[0402] Some lipid nanoparticles contemplated for use comprise at least one cationic lipid, at least one non-cationic lipid, and at least one conjugated lipid. In more particular examples, lipid nanoparticles can comprise from about 50 mol % to about 85 mol % of a cationic lipid, from about 13 mol % to about 49.5 mol % of a non-cationic lipid, and from about 0.5 mol % to about 10 mol % of a lipid conjugate, and are produced in such a manner as to have a non-lamellar (i.e., non- bilayer) morphology. In other particular examples, lipid nanoparticles can comprise from about 40 mol % to about 85 mol % of a cationic lipid, from about 13 mol % to about 49.5 mol % of a noncationic lipid, and from about 0.5 mol % to about 10 mol % of a lipid conjugate, and are produced in such a manner as to have a non-lamellar (i.e., non-bilayer) morphology.

[0403] Cationic lipids can include, for example, one or more of the following: palmitoyi-oleoyl- nor-arginine (PONA), MPDACA, GUADACA, ((6Z,9Z,28Z,31Z)-heptatriaconta-6,9,28,31- tetraen-19-yl 4-(dimethylamino)butanoate) (MC3), LenMC3, CP-LenMC3, y-LenMC3, CP-y- LenMC3, MC3MC, MC2MC, MC3 Ether, MC4 Ether, MC3 Amide, Pan-MC3, Pan-MC4 and Pan MC5, l,2-dilinoleyloxy-N,N-dimethylaminopropane (DLinDMA), l,2-dilinolenyloxy-N,N- dimethylaminopropane (DLenDMA), 2,2-dilinoleyl-4-(2-dimethylaminoethyl)-[l,3]-dioxolane

[0404] 55

[0405] WBD (US) 4882-8896-8689vl

[0406] P893392050WO (01287) (DLin-K-C2-DMA; “XTC2”), 2,2-dilinoleyl-4-(3-dimethylaminopropyl)-[l,3]-dioxolane (DLin-K- C3-DMA), 2,2-dilinoleyl-4-(4-dimethylaminobutyl)-[l,3]-dioxolane (DLin-K-C4-DMA), 2,2- dilinoleyl-5-dimethylaminomethyl-[ l,3]-dioxane (DLin-K6-DMA), 2,2-dilinoleyl-4-N- methylpepiazino- [ 1 ,3] -dioxolane (DLin-K-MPZ), 2,2-dilinoleyl-4-dimethylaminomethyl- [1,3]- dioxolane (DLin-K-DMA), l,2-dilinoleylcarbamoyloxy-3-dimethylaminopropane (DLin-C-DAP), l,2-dilinoleyoxy-3-(dimethylamino)acetoxypropane (DLin-DAC), l,2-dilinoleyoxy-3- morpholinopropane (DLin-MA), l,2-dilinoleoyl-3 -dimethylaminopropane (DLinDAP), 1,2- dilinoleylthio-3-dimethylaminopropane (DLin-S-DMA), l-linoleoyl-2-linoleyloxy-3- dimethylaminopropane (DLin-2-DMAP), l,2-dilinoleyloxy-3 -trimethylaminopropane chloride salt (DLin-TMA.Cl), l,2-dilinoleoyl-3-trimethylaminopropane chloride salt (DLin-TAP.Cl), 1,2- dilinoleyloxy-3-(N-methylpiperazino)propane (DLin-MPZ), 3-(N,N-dilinoleylamino)-l,2- propanediol (DLinAP), 3-(N,N-dioleylamino)-l,2-propanedio (DOAP), l,2-dilinoleyloxo-3-(2- N,N-dimethylamino)ethoxypropane (DLin-EG-DMA), N,N-dioleyl-N,N-dimethylammonium chloride (DODAC), l,2-dioleyloxy-N,N-dimethylaminopropane (DODMA), l,2-distearyloxy-N,N- dimethylaminopropane (DSDMA), N-(l-(2,3-dioleyloxy)propyl)-N,N,N -trimethylammonium chloride (DOTMA), N,N-distearyl-N,N-dimethylammonium bromide (DDAB), N-(l-(2,3- dioleoyloxy)propyl)-N,N,N-trimethylammonium chloride (DOTAP), 3-(N-(N',N'- dimethylaminoethane)-carbamoyl)cholesterol (DC-Chol), N-(l,2-dimyristyloxyprop-3-yl)-N,N- dimethyl-N-hydroxyethyl ammonium bromide (DMRIE), 2,3-dioleyloxy-N-[2(spermine- carboxamido)ethyl] -N,N-dimethyl- 1 -propanaminiumtrifluoroacetate (DOSPA), dioctadecylamidoglycyl spermine (DOGS), 3-dimethylamino-2-(cholest-5-en-3-beta-oxybutan-4- oxy)-l-(cis,cis-9, 12-octadecadienoxy)propane (CLinDMA), 2-[5'-(cholest-5-en-3-beta-oxy)-3'- oxapentoxy)-3-dimethy-l-(cis,cis-9',l-2'-octadecadienoxy)propane (CpLinDMA), N,N-dimethyl- 3,4-dioleyloxybenzylamine (DMOBA), l,2-N,N'-dioleylcarbamyl-3 -dimethylaminopropane (DOcarbDAP), l,2-N,N'-dilinoleylcarbamyl-3-dimethylaminopropane (DLincarbDAP), or mixtures thereof. The cationic lipid can also be DLinDMA, DLin-K-C2-DMA (“XTC2”), MC3, LenMC3, CP-LenMC3, y-LenMC3, CP-y-LenMC3, MC3MC, MC2MC, MC3 Ether, MC4 Ether, MC3 Amide, Pan-MC3, Pan-MC4, Pan MC5, or mixtures thereof.

[0407] In various embodiments, the cationic lipid may comprise from about 50 mol % to about 90 mol %, from about 50 mol % to about 85 mol %, from about 50 mol % to about 80 mol %, from about 50 mol % to about 75 mol %, from about 50 mol % to about 70 mol %, from about 50 mol % to about 65 mol %, or from about 50 mol % to about 60 mol % of the total lipid present in the particle.

[0408] 56

[0409] WBD (US) 4882-8896-8689vl

[0410] P893392050WO (01287) In other embodiments, the cationic lipid may comprise from about 40 mol % to about 90 mol %, from about 40 mol % to about 85 mol %, from about 40 mol % to about 80 mol %, from about 40 mol % to about 75 mol %, from about 40 mol % to about 70 mol %, from about 40 mol % to about 65 mol %, or from about 40 mol % to about 60 mol % of the total lipid present in the particle.

[0411] The non-cationic lipid may comprise, e.g., one or more anionic lipids and / or neutral lipids. In particular embodiments, the non-cationic lipid comprises one of the following neutral lipid components: (1) cholesterol or a derivative thereof; (2) a phospholipid; or (3) a mixture of a phospholipid and cholesterol or a derivative thereof. Examples of cholesterol derivatives include, but are not limited to, cholestanol, cholestanone, cholestenone, coprostanol, cholesteryl-2'- hydroxyethyl ether, cholesteryl-4'-hydroxybutyl ether, and mixtures thereof. The phospholipid may be a neutral lipid including, but not limited to, dipalmitoylphosphatidylcholine (DPPC), distearoylphosphatidylcholine (DSPC), dioleoylphosphatidylethanolamine (DOPE), palmitoyloleoyl-phosphatidylcholine (POPC), pahnitoyloleoyl-phosphatidylethanolamine (POPE), palmitoyloleyol-phosphatidylglycerol (POPG), dipalmitoyl-phosphatidylethanolamine (DPPE), dimyristoyl-phosphatidylethanolamine (DMPE), distearoyl-phosphatidylethanolamine (DSPE), monomethyl-phosphatidylethanolamine, dimethyl-phosphatidylethanolamine, dielaidoylphosphatidylethanolamine (DEPE), stearoyloleoyl-phosphatidylethanolamine (SOPE), egg phosphatidylcholine (EPC), and mixtures thereof. In certain embodiments, the phospholipid is DPPC, DSPC, or mixtures thereof.

[0412] In some embodiments, the non-cationic lipid (e.g., one or more phospholipids and / or cholesterol) may comprise from about 10 mol % to about 60 mol %, from about 15 mol % to about 60 mol %, from about 20 mol % to about 60 mol %, from about 25 mol % to about 60 mol %, from about 30 mol % to about 60 mol %, from about 10 mol % to about 55 mol %, from about 15 mol % to about 55 mol %, from about 20 mol % to about 55 mol %, from about 25 mol % to about 55 mol %, from about 30 mol % to about 55 mol %, from about 13 mol % to about 50 mol %, from about 15 mol % to about 50 mol % or from about 20 mol % to about 50 mol % of the total lipid present in the particle. When the non-cationic lipid is a mixture of a phospholipid and cholesterol or a cholesterol derivative, the mixture may comprise up to about 40, 50, or 60 mol % of the total lipid present in the particle.

[0413] The conjugated lipid that inhibits aggregation of particles may comprise, e.g., one or more of the following: a polyethyleneglycol (PEG)-lipid conjugate, a polyamide (ATTA)-lipid conjugate, a cationic-polymer-lipid conjugates (CPLs), or mixtures thereof. In one particular embodiment, the

[0414] 57

[0415] WBD (US) 4882-8896-8689vl

[0416] P893392050WO (01287) nucleic acid-lipid particles comprise either a PEG-lipid conjugate or an ATTA-lipid conjugate. In certain embodiments, the PEG-lipid conjugate or ATTA-lipid conjugate is used together with a CPL. The conjugated lipid that inhibits aggregation of particles may comprise a PEG-lipid including, e.g., a PEG-diacylglycerol (DAG), a PEG dialkyloxypropyl (DAA), a PEG- phospholipid, a PEG-ceramide (Cer), or mixtures thereof. The PEG-DAA conjugate may be PEG- di lauryloxypropyl (C12), a PEG-dimyristyloxypropyl (C14), a PEG-dipalmityloxypropyl (C16), a PEG-distearyloxypropyl (Cl 8), or mixtures thereof.

[0417] Additional PEG-lipid conjugates suitable for use include, but are not limited to, mPEG2000-l,2-di-0-alkyl-sn3-carbomoylglyceride (PEG-C-DOMG). The synthesis of PEG-C- DOMG is described in PCT Application No. PCT / US08 / 88676. Yet additional PEG-lipid conjugates suitable for use include, without limitation, l-[8'-(l,2-dimyristoyl-3-propanoxy)- carboxamido-3',6'-dioxaoctanyl]carbamoyl-co-methyl-poly(ethylene glycol) (2KPEG-DMG). The synthesis of 2KPEG-DMG is described in U.S. Pat. No. 7,404,969.

[0418] In some cases, the conjugated lipid that inhibits aggregation of particles (e.g., PEG-lipid conjugate) may comprise from about 0. 1 mol % to about 2 mol %, from about 0.5 mol % to about 2 mol %, from about 1 mol % to about 2 mol %, from about 0.6 mol % to about 1.9 mol %, from about 0.7 mol % to about 1.8 mol %, from about 0.8 mol % to about 1.7 mol %, from about 1 mol % to about 1.8 mol %, from about 1.2 mol % to about 1.8 mol %, from about 1.2 mol % to about 1.7 mol %, from about 1.3 mol % to about 1.6 mol %, from about 1.4 mol % to about 1.5 mol %, or about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mol % (or any fraction thereof or range therein) of the total lipid present in the particle. Typically, in such instances, the PEG moiety has an average molecular weight of about 2,000 Daltons. In other cases, the conjugated lipid that inhibits aggregation of particles (e.g., PEG-lipid conjugate) may comprise from about 5.0 mol % to about 10 mol %, from about 5 mol % to about 9 mol %, from about 5 mol % to about 8 mol %, from about 6 mol % to about 9 mol %, from about 6 mol % to about 8 mol %, or about 5 mol %, 6 mol %, 7 mol %, 8 mol %, 9 mol %, or 10 mol % (or any fraction thereof or range therein) of the total lipid present in the particle. Typically, in such instances, the PEG moiety has an average molecular weight of about 750 Daltons.

[0419] In other embodiments, the composition may comprise amphoteric liposomes, which contain at least one positive and at least one negative charge carrier, which differs from the positive one, the isoelectric point of the liposomes being between 4 and 8. This objective is accomplished owing to the fact that liposomes are prepared with a pH-dependent, changing charge.

[0420] 58

[0421] WBD (US) 4882-8896-8689vl

[0422] P893392050WO (01287) Liposomal structures with the desired properties are formed, for example, when the amount of membrane -forming or membrane-based cationic charge carriers exceeds that of the anionic charge carriers at a low pH and the ratio is reversed at a higher pH. This is always the case when the ionizable components have a pKa value between 4 and 9. As the pH of the medium drops, all cationic charge carriers are charged more and all anionic charge carriers lose their charge.

[0423] Cationic compounds useful for amphoteric liposomes include those cationic compounds previously described herein above. Without limitation, strongly cationic compounds can include, for example: DC-Chol 3-p-[N-(N',N'-dimethylmethane) carbamoyl] cholesterol, TC-Chol 3-P-[N- (N', N', N'-trimethylaminoethane) carbamoyl cholesterol, BGSC bisguanidinium-spermidine- cholesterol, BGTC bis-guadinium-tren-cholesterol, DOTAP (l,2-dioleoyloxypropyl)-N,N,N- trimethylammonium chloride, DOSPER (l,3-dioleoyloxy-2-(6-carboxy-spermyl)-propylamide, DOTMA (l,2-dioleoyloxypropyl)-N,N,N-trimethylamronium chloride) (Lipofectin®), DORIE l,2-dioleoyloxypropyl)-3 -dimethylhydroxyethylammonium bromide, DOSC (l,2-dioleoyl-3- succinyl-sn-glyceryl choline ester), DOGSDSO (l,2-dioleoyl-sn-glycero-3 -succinyl -2- hydroxyethyl disulfide ornithine), DDAB dimethyldioctadecylammonium bromide, DOGS ((C18)2GlySper3+) N,N-dioctadecylamido-glycol-spermin (Transfectam®) (C18)2Gly+ N,N- dioctadecylamido-glycine, CTAB cetyltrimethylammonium bromide, CpyC cetylpyridinium chloride, DOEPC l,2-dioleoly-sn-glycero-3-ethylphosphocholine or other O-alkyl- phosphatidylcholine or ethanolamines, amides from lysine, arginine or ornithine and phosphatidyl ethanolamine.

[0424] Examples of weakly cationic compounds include, without limitation: His-Chol (histaminyl- cholesterol hemisuccinate), Mo-Chol (morpholine-N-ethylamino-cholesterol hemisuccinate), or histidinyl-PE.

[0425] Examples of neutral compounds include, without limitation: cholesterol, ceramides, phosphatidyl cholines, phosphatidyl ethanolamines, tetraether lipids, or diacyl glycerols.

[0426] Anionic compounds useful for amphoteric liposomes include those non-cationic compounds previously described herein. Without limitation, examples of weakly anionic compounds can include: CHEMS (cholesterol hemisuccinate), alkyl carboxylic acids with 8 to 25 carbon atoms, or diacyl glycerol hemisuccinate. Additional weakly anionic compounds can include the amides of aspartic acid, or glutamic acid and PE as well as PS and its amides with glycine, alanine, glutamine, asparagine, serine, cysteine, threonine, tyrosine, glutamic acid, aspartic acid or other amino acids or aminodicarboxylic acids. According to the same principle, the esters of hydroxycarboxylic acids or hydroxydicarboxylic acids and PS are also weakly anionic compounds.

[0427] 59

[0428] WBD (US) 4882-8896-8689vl

[0429] P893392050WO (01287) In some embodiments, amphoteric liposomes may contain a conjugated lipid, such as those described herein above. Particular examples of useful conjugated lipids include, without limitation, PEG-modified phosphatidylethanolamine and phosphatidic acid, PEG-ceramide conjugates (e.g., PEG-CerC14 or PEG-CerC20), PEG-modified dialkylamines and PEG-modified 1,2- diacyloxypropan-3 -amines. Particular are PEG-modified diacylglycerols and dialkylglycerols.

[0430] In some embodiments, the neutral lipids may comprise from about 10 mol % to about 60 mol %, from about 15 mol % to about 60 mol %, from about 20 mol % to about 60 mol %, from about 25 mol % to about 60 mol %, from about 30 mol % to about 60 mol %, from about 10 mol % to about 55 mol %, from about 15 mol % to about 55 mol %, from about 20 mol % to about 55 mol %, from about 25 mol % to about 55 mol %, from about 30 mol % to about 55 mol %, from about 13 mol % to about 50 mol %, from about 15 mol % to about 50 mol % or from about 20 mol % to about 50 mol % of the total lipid present in the particle.

[0431] In some cases, the conjugated lipid that inhibits aggregation of particles (e.g., PEG-lipid conjugate) may comprise from about 0. 1 mol % to about 2 mol %, from about 0.5 mol % to about 2 mol %, from about 1 mol % to about 2 mol %, from about 0.6 mol % to about 1.9 mol %, from about 0.7 mol % to about 1.8 mol %, from about 0.8 mol % to about 1.7 mol %, from about 1 mol % to about 1.8 mol %, from about 1.2 mol % to about 1.8 mol %, from about 1.2 mol % to about 1.7 mol %, from about 1.3 mol % to about 1.6 mol %, from about 1.4 mol % to about 1.5 mol %, or about 1, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 mol % (or any fraction thereof or range therein) of the total lipid present in the particle. Typically, in such instances, the PEG moiety has an average molecular weight of about 2,000 Daltons. In other cases, the conjugated lipid that inhibits aggregation of particles (e.g., PEG-lipid conjugate) may comprise from about 5.0 mol % to about 10 mol %, from about 5 mol % to about 9 mol %, from about 5 mol % to about 8 mol %, from about 6 mol % to about 9 mol %, from about 6 mol % to about 8 mol %, or about 5 mol %, 6 mol %, 7 mol %, 8 mol %, 9 mol %, or 10 mol % (or any fraction thereof or range therein) of the total lipid present in the particle. Typically, in such instances, the PEG moiety has an average molecular weight of about 750 Daltons.

[0432] Considering the total amount of neutral and conjugated lipids, the remaining balance of the amphoteric liposome can comprise a mixture of cationic compounds and anionic compounds formulated at various ratios. The ratio of cationic to anionic lipid may selected in order to achieve the desired properties of nucleic acid encapsulation, zeta potential, pKa, or other physicochemical property that is at least in part dependent on the presence of charged lipid components.

[0433] 60

[0434] WBD (US) 4882-8896-8689vl

[0435] P893392050WO (01287) 5 , Pharmaceutical Compositions

[0436] Compositions

[0437] In some embodiments, the disclosure provides a pharmaceutical composition comprising a pharmaceutically acceptable carrier, a nuclease, and a polynucleotide that comprises a repair construct of the disclosure. In other embodiments, the disclosure provides a pharmaceutical composition comprising a pharmaceutically acceptable carrier, a polynucleotide that comprises a nucleic acid sequence encoding a nuclease, and a polynucleotide that comprises a repair construct of the disclosure.

[0438] The polynucleotides of the pharmaceutical compositions may be comprised by a lipid nanoparticle or by a recombinant virus (e.g., a recombinant AAV). The pharmaceutical compositions can be useful for treating a subject, wherein the treatment comprises replacing a region of the genome with a repair construct of the disclosure.

[0439] Pharmaceutical compositions can be prepared in accordance with known techniques, see e.g., Remington, The Science and Practice of Pharmacy (21stEdition, Philadelphia, Lippincott, Williams & Wilkins, 2005). In the manufacture of a pharmaceutical formulation according to the disclosure, polynucleotides encoding a nuclease and / or polynucleotides comprising a repair construct of the disclosure, or cells expressing the same, are typically admixed with a pharmaceutically acceptable carrier and the resulting composition is administered to a subject.

[0440] In some embodiments, a recombinant AAV comprising a polynucleotide that comprises a nucleic acid sequence encoding a nuclease, and a recombinant AAV comprising a polynucleotide that comprises repair construct of the disclosure are administered to a cell, an organism, and / or a subject. In some embodiments, the recombinant AAV comprises a nuclease.

[0441] In some embodiments, a lipid nanoparticle comprising a polynucleotide that comprises a nucleic acid sequence encoding a nuclease, and a lipid nanoparticle comprising a polynucleotide that comprises a repair construct of the disclosure are administered to a cell, an organism, and / or a subject. In some embodiments, the lipid nanoparticles comprise a nuclease.

[0442] The pharmaceutical compositions described herein, when administered to a cell, an organism, and / or a subject, can result in the integration of a repair construct of the disclosure into the genome of the cell, the organism, and / or the subject. In preferred embodiments, the integration occurs via homology-directed repair (i.e., homologous recombination).

[0443] Administration

[0444] 61

[0445] WBD (US) 4882-8896-8689vl

[0446] P893392050WO (01287) In some embodiments, the pharmaceutical compositions described herein are delivered to a cell in vitro. In some embodiments, the pharmaceutical compositions described herein are delivered to a target cell in a subject in vivo. In some embodiments, the pharmaceutical compositions are supplied to target cells via injection directly to the target tissue, or alternatively systemically via the circulatory system.

[0447] In various embodiments of the methods, the compositions described herein can be administered via any suitable route of administration known in the art. Such routes of administration can include, for example, intravenous, intramuscular, intraperitoneal, subcutaneous, intrahepatic, transmucosal, transdermal, intraarterial, and sublingual. Other suitable routes of administration can be readily determined by the treating physician, as necessary.

[0448] In one aspect, described herein are methods of modifying a gene in a target cell in a subject, wherein an exogenous polynucleotide comprising a repair construct and a gene encoding a nuclease is delivered to the target cell.

[0449] In one aspect, described herein are methods of treating a disease in a subject, wherein a therapeutically effective amount of an exogenous polynucleotide comprising a repair construct and a therapeutically effective amount of a gene encoding a nuclease are administered to the subject.

[0450] EXAMPLES

[0451] This invention is further illustrated by the following examples, which should not be construed as limiting. Those skilled in the art will recognize, or be able to ascertain, using no more than routine experimentation, numerous equivalents to the specific substances and procedures described herein. Such equivalents are intended to be encompassed in the scope of the claims that follow the examples below.

[0452] Example 1 Replacement of Genomic DNA using an Engineered Meganuclease and a Replacement Nucleic Acid Sequence having 500 bp of Homology

[0453] 1. Methods

[0454] The engineered meganucleases of the disclosure are small homing endonucleases with programmable sequence specificity that support high rates of targeted insertion via HDR. The cleavage activity of these meganucleases leaves a 4 base pair overhang at the region of the double-

[0455] 62

[0456] WBD (US) 4882-8896-8689vl

[0457] P893392050WO (01287) strand break (DSB); DNA repair at this site requires an AAV repair template to share homology with the regions flanking the DSB site in the genome.

[0458] Typically, repair constructs are designed to have homology arms directly flanking the meganuclease cut site. Here, the function of a repair construct having a homology arm at a set distance (such as 2.8kb away from the cut site) was determined. An engineered meganuclease with sequence specificity to T cell receptor alpha (TRAC) locus was used. The TRAC locus has 4 exons with a stop codon in exon 3 and a putative polyA in exon 4. The engineered meganuclease used, referred to herein as TRC1-2L.2307 as described in PCT International Patent Application Publication No: WO2024148167 has a recognition site in exon 1 of TRAC. A repair construct was designed to contain 500 bp homology arms that sit at distinct regions of the gene (see Figure 14A). The left homology arm (i.e., a first homology arm) flanks the meganuclease cut site and is composed of some intergenic sequence and some exon 1 coding sequence. The right homology arm (i.e., a second homology arm) targets a region that is 2.8kb downstream of the meganuclease cut site and is composed of some of exon 3 and some of the intron between exons 3 and 4. The repair template (i.e., the replacement nucleic acid sequence) used in this example includes the following: (i) a rebuilt 3 ’ end of exon 1 that has been codon-switched to make it non-homologous, (ii) a synthetic intron containing a P-2 microglobulin (B2M) targeting miR as a marker for AAV insertion (cells containing the insert will have B2M knocked down), and (iii) a codon-switched exon 2 in-frame with exon 3 (see Figure 14B). An allele carrying the correct insert will have a genomic sequence as detailed in Figure 14C. Specifically, the desired allele will feature a transgenic exon 1 containing some native (blue) and some wobbled DNA sequence (green), an artificial intron with the B2M knock-down reporter, and a wobbled transgenic exon 2 fused to native exon 3 and parts of intron 3.

[0459] The function of the construct was tested in human T cells. Human peripheral blood mononuclear cells (PBMCs) were thawed and incubated overnight in media containing IL-2 growth factor. The following day, cells were stimulated with a human CD3 / CD28 / CD2 T cell activator for an additional three days, after which they were electroporated with meganuclease mRNA using the Lonza nucleofector. DNA repair template was delivered at a multiplicity of infection (MOI) of 20,000 vg / cell. One week after electroporation, the frequency of edited and inserted cells was measured by flow cytometry, using fluorescent antibodies against CD3 and B2M. To characterize the DNA repair mechanism used in the cells, the cells were sorted and gDNA purified. gDNA was amplified across the TRC1-2 site and sequenced with long-range Nanopore sequencing (available from Oxford Nanopore Technologies).

[0460] 63

[0461] WBD (US) 4882-8896-8689vl

[0462] P893392050WO (01287) . Results

[0463] As described above, the edited cells were analyzed by flow cytometry, sorted, and sequenced. A correctly edited allele has an indel in exon 1 causing a disruption in the T cell receptor (TCR), phenotypically observed by a loss in CD3 expression. Cells that have also captured an AAV transgene should regain TCR / CD3 expression from that allele and will have reduced B2M levels due to the synthetic miR intron. That is, correctly edited cells will be CD3+B2M- as detected by flow cytometry (Figure 15 A; bottom right quadrant). Edited cells that did not capture the insert will be CD3-B2M+ (Figure 15 A; top left quadrant). Cells that did not receive any edit will be CD3+B2M+ (Figure 15 A; top right quadrant).

[0464] As shown in Figure 15, the CD3+B2M- population represents a fraction of correctly edited and inserted cells; however, the TRAC gene includes both an active and silent TCR allele. The active TCR allele allows for productive VDJ T cell arrangement, producing a surviving signal to T cells and expression of the desired insert (in this case, displayed as B2M-). In contrast, in the silent TCR allele, VDJ recombination has either failed or not been attempted, resulting in potential epigenetic silencing and no expression. In this case, TCR / CD3 expression may be disrupted by insertion, resulting in a loss of CD3 but expression of the synthetic B2M miR intron, thereby displaying the cells as B2M-CD3- (Figure 15 A; bottom left quadrant). That is, it is possible that both CD3-B2M- and CD3+B2M- populations may contain the correct insert.

[0465] The frequency of cells that utilized the insert described above was determined. As a control, T cells treated with only the repair template and no nuclease was used. This control population displayed a majority of the cells as unedited, presenting as CD3+B2M+ (Figure 15B). Cells treated with the nuclease and repaired with template AAV resulted in cells displaying the phenotypes as described above and shown in Figure 15A. Approximately 50% of cells were edited but displayed no insert, while 7% of cells were edited and repaired correctly with the insert and replace construct (Figure 15C). An additional 22% of cells were displayed in the lower left quadrant, indicating possible insertion of the silent allele.

[0466] To determine the mechanism of DNA repair, cells were sorted into CD3+B2M-, CD3-B2M- , and CD3-B2M+ populations. gDNA was amplified across the TRC1-2 site resulting in a 9.6kb amplicon, which was subjected to long-range Nanopore sequencing (available from Oxford Nanopore Technologies). The majority of CD3+B2M- cells were full length sequences containing the insert (Figure 15D; gray line), while both the CD3-B2M- and CD3-B2M+ populations predominantly encompassed full length, unedited genomic DNA (Figure 14D; black line). Of the

[0467] 64

[0468] WBD (US) 4882-8896-8689vl

[0469] P893392050WO (01287) sequenced populations, 49% of the sorted CD3+B2M- cells contained some insert, of which 95% of the integrants were repaired by HDR as shown in Table 1. In contrast, only 14% of the CD3-B2M- population contained some inserts which were predominantly repaired by HDR. A summary of the sequencing results is provided in Table 1.

[0470] Table 1: Mechanism of Repair Observed in Edited Cells

[0471] 3 , Conclusion

[0472] The above study demonstrates that the tested construct was capable of editing of T cells having the correct insert using the desired mechanism of HDR with an efficiency of ~7%.

[0473] Example 2 Replacement of Genomic DNA using an Engineered Meganuclease with a Replacement Nucleic Acid Sequence having an Extended Right Homology Arm

[0474] 1. Methods

[0475] A new repair template construct (i.e., replacement nucleic acid sequence) was designed targeting the TRAC locus. The engineered meganuclease targeting the TRAC locus described in example 1 was utilized. The construct has a similar design as detailed in Example 1, with the 500 bp left homology arm flanking the meganuclease cut site and comprised of some intergenic sequence and some exon 1 coding sequence. The right homology arm was extended to 2.6kb in length but targeting the same genomic region 2.8kb downstream of the meganuclease cut site. The right homology arm comprises a rebuilt 3’ end of exon 1 that has been codon-switched to make it non-homologous, a synthetic intron containing a P-2 microglobulin (B2M) targeting miR as a

[0476] WBD (US) 4882-8896-8689vl

[0477] P893392050WO (01287) marker for AAV insertion, a codon-switched exon 2, exon 3, and the following intronic sequence between exons 3 and 4 (see Figure 16).

[0478] The construct was tested for its ability to edit human T cells. Thawed human PBMCs were incubated overnight in media containing IL-2 growth factor. The following day, cells were stimulated with a human CD3 / CD28 / CD2 T cell activator for an additional three days, after which they were electroporated with meganuclease mRNA using the Lonza nucleofector. DNA repair template was delivered at an MOI of 20,000 vg / cell. One week after electroporation, the frequency of edited and inserted cells were measured by flow cytometry using fluorescent antibodies against CD3 and B2M as described in Example 1. Cells were subsequently sorted and sequenced as described in Example 1.

[0479] 2, Results

[0480] The frequency of cells that were edited was determined as described in Example 1. In the meganuclease only control group there was as expected no detected repair lower right quadrant of Figure 17A. Of the cells that were treated with the nuclease and repair template, -61% were edited and repaired correctly with the insert and replace construct lower right quadrant of Figure 17B; an approximate 9X increase as compared to the 7% using the repair template of Example 1. The CD3- B2M+ population was sequenced using long-range Nanopore sequencing, which showed a 92% rate of HDR-mediated repair (Figure 17C).

[0481] 3 , Conclusion

[0482] These data demonstrate that an extension of the right homology arm from 500 bp (the construct of Example 1) to 2.6kb resulted in an -9X increase in the frequency of edited cells with repair driven by HDR.

[0483] Example 3 Replacement of Genomic DNA Upstream of the Meganuclease Cleavage Site with Replacement Nucleic Acid Sequence having an Extended Left Homology Arm

[0484] 1. Methods

[0485] It was next determined whether the methods described herein could be used to edit upstream of the meganuclease cut site. The engineered meganuclease targeting the TRAC locus described in example 1 was utilized. A repair construct (i.e., a replacement nucleic acid sequence) was designed

[0486] 66

[0487] WBD (US) 4882-8896-8689vl

[0488] P893392050WO (01287) targeting the TRAC locus. The repair construct was designed to have a 1.6kb left homology arm, with the length of the homology arm limited by upstream T cell recombination sequences. The construct further comprises a synthetic intron containing a B-2 microglobulin (B2M) targeted miR as a marker for AAV insertion, the 5 ’ end of exon 1 that has been codon switched to make it non- homologous, the TRC1-2 binding site, and a 423bp right homology arm (Figure 18).

[0489] The construct was tested for its ability to edit human T cells. Thawed human PBMCs were incubated overnight in media containing IL-2 growth factor. The following day, cells were stimulated with a human CD3 / CD28 / CD2 T cell activator for an additional three days, after which they were electroporated with meganuclease mRNA using the Lonza nucleofector. DNA repair template was delivered at an MOI of 20,000 vg / cell. One week after electroporation, the frequency of edited and inserted cells were measured by flow cytometry using fluorescent antibodies against CD3 and B2M as described in Example 1. Cells were subsequently sorted and sequenced as described in Example 1.

[0490] 2, Results

[0491] The frequency of cells that were edited was determined as described in Example 1. The upstream construct resulted in approximately 30% of cells containing the proper insert (Figure 19; right). The CD3-B2M+ population was sequenced using long-range Nanopore sequencing (available from Oxford Nanopore Technologies), which showed an 86% rate of HDR-mediated repair as shown in Table 2.

[0492] Table 2: Mechanism of Repair of Edited Cells

[0493] 3 , Conclusion

[0494] 67

[0495] WBD (US) 4882-8896-8689vl

[0496] P893392050WO (01287) These data demonstrate that the methods described herein may be used to edit genomic DNA upstream of the engineered meganuclease cut site.

[0497] Example 4 Evaluation of Insertion and Replacement Efficiency Based on the Location of the Homology Sequence Downstream of the Meganuclease Cleavage Site

[0498] 1. Methods

[0499] It was next determined how many bases downstream of the nuclease cut site can be skipped without losing efficiency. Three constructs (i.e., replacement nucleic acid sequences) were designed: a 5kb genomic skip (Figure 20A), a 7.5kb genomic skip (Figure 20B), and a lOkb genomic skip (Figure 20C). Constructs were designed targeting the TRAC locus using the engineered meganuclease described in example 1. The repair construct has a 500 bp left homology arm that flanks the meganuclease cut site and comprises some intergenic sequence and some exon 1 coding sequence. The repair construct also includes a rebuilt 3’ end of exon 1 that has been codon- switched to make it non-homologous, a P-2 microglobulin (B2M) targeting miR as a marker for AAV insertion, and adjacent codon-switched exons 2, 3, and 4 with no intergenic sequence. The right homology arm was kept constant at 2.8kb; however, the sequences were adjusted to account for the genomic distance skipped.

[0500] The constructs were tested for the ability to edit human T cells. Thawed human PBMCs were incubated overnight in media containing IL-2 growth factor. The following day, cells were stimulated with a human CD3 / CD28 / CD2 T cell activator for an additional three days, after which they were electroporated with meganuclease mRNA using the Lonza nucleofector. DNA repair template was delivered at an MOI of 20,000 vg / cell. One week after electroporation, the frequency of edited and inserted cells were measured by flow cytometry using fluorescent antibodies against CD3 and B2M as described in Example 1. Cells were subsequently sorted and sequenced as described in Example 1.

[0501] 2, Results

[0502] The frequency of cells that were edited was determined as described in Example 1. As a control, cells were also treated with the construct of Example 2 (a 2.8kb skip region). In the present experiment, -43% of cells treated with the 2.8kb construct contained the correct edit and insert (B2M-CD3+, Figure 21A). For the 5kb skip construct, -33% of cells had the correct edit and insert

[0503] 68

[0504] WBD (US) 4882-8896-8689vl

[0505] P893392050WO (01287) (Figure 21B). For the 7.5kb skip construct, -19% of cells had the correct edit and insert (Figure 21C). For the lOkb skip construct, -7% of cells had the correct edit and insert (Figure 2 ID). That is, while cells could be edited at up to lOkb distances, the editing efficiency decreased with increased distance.

[0506] The CD3+B2M- and CD3-B2M- populations were sorted and sequenced using long-range Nanopore sequencing. A majority of DNA repair occurred by the HDR mechanism (Table 3); however, the efficiency of HDR-mediated repair decreased slightly with the longer skipped distances. Higher percentages of miscellaneous reads (which mostly represent partial inserts) also increase with the longest skipped distance.

[0507] Table 3: Mechanism of Repair Observed in Edited Cells

[0508] 3. Conclusion

[0509] These data demonstrate that the methods described herein can be used to skip large sections of genomic DNA as guided by the location of the homology arm downstream of the meganuclease cut site. Up to lOkb of genomic DNA can be skipped and replaced with functional sequence using HDR-mediated repair. There was a linear correlation between the amount of DNA skipped and the efficiency of sequence insertion.

[0510] Figure 22A-22B provides a summary of the constructs tested and the resultant efficiency across Examples 1-4. Looking at B2M-CD3+ cells (i.e., correctly edited and inserted), there is an increase in the percentage of editing as the length of the homology arm increases. There is also a linear correlation between insert efficiency and the length of skipped DNA (Figure 22A). When visualized as the fraction of total edited cells that received the correct insert, the same trends are observed (Figure 22B).

[0511] WBD (US) 4882-8896-8689vl

[0512] P893392050WO (01287)

[0513] WBD (US) 4882-8896-8689vl

[0514] P893392050WO (01287) Sequence Listing

[0515] SEQ ID NO: 1

[0516] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA

[0517] GACCGTGACTTGCCAGCCCCACAGAGCCCCGCCCTTGTCCATCACTGGCATCTGGACTC CAGCCTGGGTTGGGGCAAAGAGGGAAATGAGATCATGTCCTAACCCTGATCCTCTTGT CCCACAGATATCCAGAACCCTGACCCTGCCGTGTACCAGCTGAGAGACTCTAAATCCA

[0518] GTGACAAGTCTGTCTGCCTATTCACCGATTTTGATTCTCAAACAAATGTGTCACAAAGT

[0519] AAGGATTCTGATGTGTATATCACAGACAAAACTGTGCTAGACATGAGGTCTATGGACT TCAAGAGCAACAGTGCTGTGGCCTGGAGCAATAAGAGCGATTTCGCCTGCGCCAATGC ATTTAATAATTCTATCATCCCCGAGGATACATTTTTTCCGTCACCCGGTAAGTATTAAT

[0520] GTTACAAGACAGGTTTAAGGACACCAATAGAAACTGGGCTTGTCGAGACAGAGAGTA CCTGTTTGAATGAGGCTTCAGTACTTTACAGAATCGTTGCCTGCACATCTTGGAAACAC

[0521] TTGCTGGGATTACTTCGACTTCTTAACCCAACAGAAGGCTCGAGAAGGTATATTGCTGT TGACAGTGAGCGCGGACATGATCTTCTTTATAATTAGTGAAGCCACAGATGTAATTAT AAAGAAGATCATGTCCATGCCTACTGCCTCGGACTTCAAGGGGCTAGAATTCGAGCAA

[0522] TTATCTTGTTTACTAAAACTGAATACCTTGCTATCTCTTTGATACATTTTTACAAAGCTG AATTAAAATGGTATAAATTAAATCACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCA CCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAGAATCCAGTTGCGACGT

[0523] GAAACTCGTTGAAAAGTCATTCGAGACCGATACGAACCTAAACTTTCAAAACCTGTCA GTGATTGGGTTCCGAATCCTCCTCCTGAAAGTGGCCGGGTTTAATCTGCTCATGACGCT GCGGCTGTGGTCCAGCTGAGGTGAGGGGCCTTGAAGCTGGGAGTGGGGTTTAGGGAC

[0524] GCGGGTCTCTGGGTGCATCCTAAGCTCTGAGAGCAAACCTCCCTGCAGGGTCTTGCTTT TAAGTCCAAAGCCTGAGCCCACCAAACTCTCCTACTTCTTCCTGTTACAAATTCCTCTT GTGCAATAATAATGGCCTGAAACGCTGTAAAATATCCTCATTTCAGCCGCCTCAGTTGC

[0525] ACTTCTCCCCTATGAGGTAGGAAGAACAGTTGTTTAGAAACGAAGAAACTGAGGCCCC ACAGCTAATGAGTGGAGGAAGAGAGACACTTGTGTACACCACATGCCTTGTGTTGTAC TTCTCTCACCGTGTAACCTCCTCATGTCCTCTCTCCCCAGTACGGCTCTCTTAGCTCAGT

[0526] AG

[0527] SEQ ID NO: 2

[0528] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA

[0529] GACCGTGACTTGCCAGCCCCACAGAGCCCCGCCCTTGTCCATCACTGGCATCTGGACTC CAGCCTGGGTTGGGGCAAAGAGGGAAATGAGATCATGTCCTAACCCTGATCCTCTTGT CCCACAGATATCCAGAACCCTGACCCTGCCGTGTACCAGCTGAGAGACTCTAAATCCA

[0530] GTGACAAGTCTGTCTGCCTATTCACCGATTTTGATTCTCAAACAAATGTGTCACAAAGT AAGGATTCTGATGTGTATATCACAGACAAAACTGTGCTAGACATGAGGTCTATGGACT TCAAGAGCAACAGTGCTGTGGCCTGGAGCAATAAGAGCGATTTCGCCTGCGCCAATGC

[0531] ATTTAATAATTCTATCATCCCCGAGGATACATTTTTTCCGTCACCCGGTAAGTATTAAT

[0532] GTTACAAGACAGGTTTAAGGACACCAATAGAAACTGGGCTTGTCGAGACAGAGAGTA CCTGTTTGAATGAGGCTTCAGTACTTTACAGAATCGTTGCCTGCACATCTTGGAAACAC

[0533] TTGCTGGGATTACTTCGACTTCTTAACCCAACAGAAGGCTCGAGAAGGTATATTGCTGT TGACAGTGAGCGCGGACATGATCTTCTTTATAATTAGTGAAGCCACAGATGTAATTAT AAAGAAGATCATGTCCATGCCTACTGCCTCGGACTTCAAGGGGCTAGAATTCGAGCAA

[0534] 71

[0535] WBD (US) 4882-8896-8689vl

[0536] P893392050WO (01287) TTATCTTGTTTACTAAAACTGAATACCTTGCTATCTCTTTGATACATTTTTACAAAGCTG AATTAAAATGGTATAAATTAAATCACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCA CCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAGAATCCAGTTGCGACGT

[0537] GAAACTCGTTGAAAAGTCATTCGAGACCGATACGAACCTAAACTTTCAAAACCTGTCA GTGATTGGGTTCCGAATCCTCCTCCTGAAAGTGGCCGGGTTTAATCTGCTCATGACGCT GCGGCTGTGGTCCAGCTGAGGTGAGGGGCCTTGAAGCTGGGAGTGGGGTTTAGGGAC

[0538] GCGGGTCTCTGGGTGCATCCTAAGCTCTGAGAGCAAACCTCCCTGCAGGGTCTTGCTTT TAAGTCCAAAGCCTGAGCCCACCAAACTCTCCTACTTCTTCCTGTTACAAATTCCTCTT GTGCAATAATAATGGCCTGAAACGCTGTAAAATATCCTCATTTCAGCCGCCTCAGTTGC

[0539] ACTTCTCCCCTATGAGGTAGGAAGAACAGTTGTTTAGAAACGAAGAAACTGAGGCCCC ACAGCTAATGAGTGGAGGAAGAGAGACACTTGTGTACACCACATGCCTTGTGTTGTAC TTCTCTCACCGTGTAACCTCCTCATGTCCTCTCTCCCCAGTACGGCTCTCTTAGCTCAGT

[0540] AGAAAGAAGACATTACACTCATATTACACCCCAATCCTGGCTAGAGTCTCCGCACCCT CCTCCCCCAGGGTCCCCAGTCGTCTTGCTGACAACTGCATCCTGTTCCATCACCATCAA AAAAAAACTCCAGGCTGGGTGCGGGGGCTCACACCTGTAATCCCAGCACTTTGGGAGG

[0541] CAGAGGCAGGAGGAGCACAGGAGCTGGAGACCAGCCTGGGCAACACAGGGAGACCC CGCCTCTACAAAAAGTGAAAAAATTAACCAGGTGTGGTGCTGCACACCTGTAGTCCCA GCTACTTAAGAGGCTGAGATGGGAGGATCGCTTGAGCCCTGGAATGTTGAGGCTACAA

[0542] TGAGCTGTGATTGCGTCACTGCACTCCAGCCTGGAAGACAAAGCAAGATCCTGTCTCA AATAATAAAAAAAATAAGAACTCCAGGGTACATTTGCTCCTAGAACTCTACCACATAG CCCCAAACAGAGCCATCACCATCACATCCCTAACAGTCCTGGGTCTTCCTCAGTGTCCA

[0543] GCCTGACTTCTGTTCTTCCTCATTCCAGATCTGCAAGATTGTAAGACAGCCTGTGCTCC CTCGCTCCTTCCTCTGCATTGCCCCTCTTCTCCCTCTCCAAACAGAGGGAACTCTCCTAC CCCCAAGGAGGTGAAAGCTGCTACCACCTCTGTGCCCCCCCGGCAATGCCACCAACTG

[0544] GATCCTACCCGAATTTATGATTAAGATTGCTGAAGAGCTGCCAAACACTGCTGCCACC CCCTCTGTTCCCTTATTGCTGCTTGTCACTGCCTGACATTCACGGCAGAGGCAAGGCTG CTGCAGCCTCCCCTGGCTGTGCACATTCCCTCCTGCTCCCCAGAGACTGCCTCCGCCAT

[0545] CCCACAGATGATGGATCTTCAGTGGGTTCTCTTGGGCTCTAGGTCCTGCAGAATGTTGT

[0546] GAGGGGTTTATTTTTTTTTAATAGTGTTCATAAAGAAATACATAGTATTCTTCTTCTCAA GACGTGGGGGGAAATTATCTCATTATCGAGGCCCTGCTATGCTGTGTATCTGGGCGTGT TGTATGTCCTGCTGCCGATGCCTTCATTAAAATGATTTGGAAGAGCAGAGACTGTGCCT

[0547] CTGTTTGACTGGGTTTGGTAGGAGTCATTTTCTGCTTGCTGGTGATCACTAGCTGGGCA GAGAAAAACCAAGGCATTTGTCTATGATGCTGTCCAGGAAGCCTCATTCAACAAGCTG CCTAAGTCAACCTCTTCTTGGAATAACCTCTAAAAGCTTCCGCTTAGCAGGCTATGCTG

[0548] AGGGCCAGGAAAACCCACCTACCAGCTTGGACCCCTCCTCTCCCACTCTCATGCCACG

[0549] CCACGGGACCACCCATAACAGGAGCCCACACACATGGGTGGCAGTGACCTGCGGCAG ACAGGGACCACACAGCAAGTGTCCCCAAAATGCCACCCACTGTCTCCTGCCCTCCAGG AGCATTTCCTTTGCCTCTCCTCTCAGACTGGGTTTCCACTGAAACTGTGCATTGTCTCAC

[0550] AAATTCGTGGCTGGGGACCACCCACCACTCTGCTGCCTGATCCAGCCCCACGCCAGCC CTTTGAGGTGCCCAAGCTGACACCAGGAGCAAGGTTGAGAGGAAGCTGTGACCCCAG CAGGACTTTATGTTCCCACCATCCCGGATGTGAGAATGAGGAAAAAAGGAGATGAGCT

[0551] GTCTCCCCACAAGCCCAGAGATTTGACCGAGGAGAGTAGAGGCCTCGAGCTCTCACCT AAGAGAAAAGACATGGGGCTTCCTGGGGTCCACAGCTCACTGCGCTCTCCCTCCTGAG ACTCCTGCTGCCAGAGCACCTTTCCCCAGGGTCATGGATGCTGAGGGAAACACAACTT

[0552] AGAGACCACTCCACCATCCACCCAGCAAGCACAGCTACCAAGACCCAAAGCTGAGGC TTACCAATGCCCAGGGTGGGAGGGGGTTCCATCCCTGAATAACTCCATGGTTCCCCTAT GCGTCTGACCATCCCAGCCAGAAATACATAAATCATCTCAGCTACAATTCAGGCCTGC

[0553] TTCTTTTCATAGGGATGAAGCTACAGGTTGAGTATCCCTTATCTGAAATGCTTGGTACT A

[0554] 72

[0555] WBD (US) 4882-8896-8689vl

[0556] P893392050WO (01287) SEQ ID NO: 3

[0557] CAACAGCCACATGTAGCTAGAGACTATTATACCAGACAGAGCAGCCTAGATCTTCTCC

[0558] AGTCTGACACCCACCAGCCCCAGGACTTGAGTGAGTGTTTAACCAGGACTCAAAGTTG

[0559] GGTTTCTGCCCCACAAGGCCACCCCCTTTCCTCTTTAAAGCCAACCTGCATCTGGTGGC

[0560] CCCTGATCCCCTGCCTTGAGGATCGGCACTTCCAGACTCCTCTCCCCCTCTGCAGTGCT

[0561] GTCCAGTACCCCCACTGATGACTAACAATCAGGGGGATGTGTTGGTAGAGCTAATGGC

[0562] TTTCTGTCTGTCCCTTCCCAGCAAAGGAACTATGCCTTAGGGCCTTCACCCAGAGTGAT

[0563] GTCAGGCTGCCCAAGCATGAGGAGGGAAGTAGGCAGAATCCTCTGGAGCCAAAGCTC

[0564] TGGATGTCTCTCCCCTCTGACCATGGAGCCCACCCCTGCTCCACTGCTCCAGGGACAGC

[0565] CCTATGCTGCAGGCAGCTCTGCCCCCACTCAGCATCCCAGGGGCTGATTTCTTTGGTTT

[0566] TGGATCCAGCTGGATGTCTGCATTGCCGAGGCCACCAGGGCTGGCTCAGCAACTGTCG

[0567] GGGAATCACCAGGGTCTGAGAAATCTTGTGCGCATGTGAGGGGCTGTGGGAGCAGAG

[0568] AACCACTGGGTGGGAAATTCTAATCCCCACCCTGCTGGAAACTCTCTGGGTGGCCCCA

[0569] ACATGCTAATCCTCCGGCAAACCTCTGTTTCCTCCTCAAAAGGCAGGAGGTCGGAAAG

[0570] AATAAACAATGAGAGTCACATTAAAAACACAAAATCCTACGGAAATACTGAAGAATG

[0571] AGTCTCAGCACTAAGGAAAAGCCTCCAGCAGCTCCTGCTTTCTGAGGGTGAAGGATAG

[0572] ACGCTGTGGCTCTGCATGACTCACTAGCACTCTATCACGGCCATATTCTGGCAGGGTCA

[0573] GTGGCTCCAACTAACATTTGTTTGGTACTTTACAGTTTATTAAATAGATGTTTATATGG

[0574] AGAAGCTCTCATTTCTTTCTCAGAAGAGCCTGGCTAGGAAGGTGGATGAGGCACCATA

[0575] TTCATTTTGCAGGTGAAATTCCTGAGATGTAAGGAGCTGCTGTGACTTGCTCAAGGCCT

[0576] TATATCGAGTAAACGGTAGTGCTGGGGCTTAGACGCAGGTGTTCTGATTTATAGTTCA

[0577] AAACCTCTATCAATGAGAGAGCAATCTCCTGGTAATGTGATAGATTTCCCAACTTAAT

[0578] GCCAACATACCATAAACCTCCCATTCTGCTAATGCCCAGCCTAAGTTGGGGAGACCAC

[0579] TCCAGATTCCAAGATGTACAGTTTGCTTTGCTGGGCCTTTTTCCCATGCCTGCCTTTACT

[0580] CTGCCAGAGTTATATTGCTGGGGTTTTGAAGAAGATCCTATTAAATAAAAGAATAAGC

[0581] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG

[0582] TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG

[0583] AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA

[0584] GACCGTGACTTGCCAGCCCCACGTAAGTATTAATGTTACAAGACAGGTTTAAGGACAC

[0585] CAATAGAAACTGGGCTTGTCGAGACAGAGAGTACCTGTTTGAATGAGGCTTCAGTACT

[0586] TTACAGAATCGTTGCCTGCACATCTTGGAAACACTTGCTGGGATTACTTCGACTTCTTA

[0587] ACCCAACAGAAGGCTCGAGAAGGTATATTGCTGTTGACAGTGAGCGCGGACATGATCT

[0588] TCTTTATAATTAGTGAAGCCACAGATGTAATTATAAAGAAGATCATGTCCATGCCTACT

[0589] GCCTCGGACTTCAAGGGGCTAGAATTCGAGCAATTATCTTGTTTACTAAAACTGAATA

[0590] CCTTGCTATCTCTTTGATACATTTTTACAAAGCTGAATTAAAATGGTATAAATTAAATC

[0591] ACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCACCTATTGGTCTTACTGACATCCAC

[0592] TTTGCCTTTCTCTCCACAGATATTCAAAATCCAGATCCAGCAGTCTATCAACTCAGGGA

[0593] TAGTAAGAGTTCCGATAAATCCGTTTGTCTCTTTACTGACTTCGACTCACAGACTAACG

[0594] TCAGTCAGTCAAAAGACAGCGACGTCTACATTACCGATAAGACCGTTCTTGATATGAG

[0595] ATCCATGGATTTTAAAAGTAATTCTGCAGTTGCATGGAGCAACAAATCTGACTTTGCAT

[0596] GTGCAAACGCCTTCAACAACAGCATTATTCCAGAAGACACCTTCTTCCCCAGCCCAGG

[0597] TAAGGGCAGCTTTGGTGCCTTCGCAGGCTGTTTCCTTGCTTCAGGAATGGCCAGGTTCT

[0598] GCCCAGAGCTCTGGTCAATGATGTCTAAAACTCCTCTGATTGGTGGTCTCGGCCTTATC

[0599] CATTGCCACCAAAACCCTCTTTTTACTAAGAAACAGTGAGCCTTGTTCTGGCAGTCCAG

[0600] AGAATGACACGGGAAAAAAGCAGATGAAGAGAAGGTGGCAGGAGAGGGCACGTGGC

[0601] CCAGCCTCAGTCTCTCCAACTGAGTTCCTGCCTGCCTGCCTTTGCTCAGACTGTTTGCCC

[0602] CTTACTGCTCTTCTAGGCCTCATTCTAAGCCCCTTCTCCAAGTTGCCTCTCCTTATTTCT

[0603] CCCTGTCTGCCAAAAAATCTTTCCCAGCTCACTAAGTCAGTCTCACGCAGTCACTCATT

[0604] AACCCACCAATCACT

[0605] 73

[0606] WBD (US) 4882-8896-8689vl

[0607] P893392050WO (01287) SEQ ID NO: 4

[0608] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG

[0609] TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG

[0610] AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA

[0611] GACCGTGACTTGCCAGCCCCACAGAGCCCCGCCCTTGTCCATCACTGGCATCTGGACTC

[0612] CAGCCTGGGTTGGGGCAAAGAGGGAAATGAGATCATGTCCTAACCCTGATCCTCTTGT

[0613] CCCACAGATATCCAGAACCCTGACCCTGCCGTGTACCAGCTGAGAGACTCTAAATCCA

[0614] GTGACAAGTCTGTCTGCCTATTCACCGATTTTGATTCTCAAACAAATGTGTCACAAAGT

[0615] AAGGATTCTGATGTGTATATCACAGACAAAACTGTGCTAGACATGAGGTCTATGGACT

[0616] TCAAGAGCAACAGTGCTGTGGCCTGGAGCAATAAGAGCGATTTCGCCTGCGCCAATGC

[0617] ATTTAATAATTCTATCATCCCCGAGGATACATTTTTTCCGTCACCCGGTAAGTATTAAT

[0618] GTTACAAGACAGGTTTAAGGACACCAATAGAAACTGGGCTTGTCGAGACAGAGAGTA

[0619] CCTGTTTGAATGAGGCTTCAGTACTTTACAGAATCGTTGCCTGCACATCTTGGAAACAC

[0620] TTGCTGGGATTACTTCGACTTCTTAACCCAACAGAAGGCTCGAGAAGGTATATTGCTGT

[0621] TGACAGTGAGCGCGGACATGATCTTCTTTATAATTAGTGAAGCCACAGATGTAATTAT

[0622] AAAGAAGATCATGTCCATGCCTACTGCCTCGGACTTCAAGGGGCTAGAATTCGAGCAA

[0623] TTATCTTGTTTACTAAAACTGAATACCTTGCTATCTCTTTGATACATTTTTACAAAGCTG

[0624] AATTAAAATGGTATAAATTAAATCACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCA

[0625] CCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAGAATCCAGTTGCGACGT

[0626] GAAACTCGTTGAAAAGTCATTCGAGACCGACACTAATCTTAATTTCCAGAATCTCTCTG

[0627] TCATCGGATTTCGCATTCTCCTGCTTAAGGTTGCTGGCTTCAACCTTCTGATGACCTTGC

[0628] GCTTGTGGTCTAGTTGAGACTTGCAGGACTGCAAAACCGCTTGCGCACCATCTCTTCTC

[0629] CCACTTCACTGTCCATCCTCCCCATCACCCAATCGCGGTAATTCACCCACTCCTAAAGA

[0630] AGTTAAGGCCGCAACTACATCCGTTCCTCCACGCCAGTGTCATCAGCTCGACCCCACA

[0631] AGAATCTACGACTAAGACTGTTGAAGGGCCGCTAAGCATTGTTGTCATCCTCTGTGCA

[0632] GTCTCATCGCCGCCTGCCATTGTTTGACCTTTACTGCCGAAGCTCGGCTTCTCCAACCC

[0633] CCATTGGCCGTTCATATCCCTAGCTGTTCTCCTGAAACAGCTTCTGCTATTCCCCAAAT

[0634] GATGGACCTGCAATGGGTCCTTCTTGGAAGTCGGAGTTGTAGGATGCTGTGAGGCGTC

[0635] TACTTCTTCTTGATCGTTTTTATCAAAAAGTATATCGTCTTTTTTTTTAGCAGGCGCGGA

[0636] GGTAAGCTGAGTCACTACCGCGGACCCGCCATGTTGTGCATTTGGGCTTGCTGCATGA

[0637] GTTGTTGTAGATGTCTCCACTAAAACGACCTGGAGGAACAAACCGGATGTGAGAATGA

[0638] GGAAAAAAGGAGATGAGCTGTCTCCCCACAAGCCCAGAGATTTGACCGAGGAGAGTA

[0639] GAGGCCTCGAGCTCTCACCTAAGAGAAAAGACATGGGGCTTCCTGGGGTCCACAGCTC

[0640] ACTGCGCTCTCCCTCCTGAGACTCCTGCTGCCAGAGCACCTTTCCCCAGGGTCATGGAT

[0641] GCTGAGGGAAACACAACTTAGAGACCACTCCACCATCCACCCAGCAAGCACAGCTACC

[0642] AAGACCCAAAGCTGAGGCTTACCAATGCCCAGGGTGGGAGGGGGTTCCATCCCTGAAT

[0643] AACTCCATGGTTCCCCTATGCGTCTGACCATCCCAGCCAGAAATACATAAATCATCTCA

[0644] GCTACAATTCAGGCCTGCTTCTTTTCATAGGGATGAAGCTACAGGTTGAGTATCCCTTA

[0645] TCTGAAATGCTTGGTACTAGAAGTGTTCAGGTTTTGGATTTTTTTTTTTTTTTTTTGAGG

[0646] CTGGGGGTGGAATATTAGCATTATACTTACCAATTCGGCATCCCTAATCTAAAAATCTG

[0647] AAATCCAAGATGCCCCAGTGAGTGAGCATTTCCTTTGAGCATCATGTTGGCATTTAAA

[0648] AAGTTTCAGATTTTGGAGCATTTCCAATTTCAGATTTTTGAATTAGGGATGCTTAACCT

[0649] GTACCAGCTTTAATAGGTGACCTAGAGCACATCCCTCCCCTCTACAGGCTTATGTGTTG

[0650] CCACTTACGAAATGTCGGGTAGGACTGGAGGCTACCTCCCACTCCCTGCACCTCTGATT

[0651] CTGTGACCCCTGCAGAATTAAGAACCAGTGTCCCTTCTCCTGGTCAGCCCTACTGACGG

[0652] GAGTCACAGAATCCCCAGTTCTTTCTAAGCTGCCCCATCTCTCACCTAAATACAATCCC

[0653] CTTAAATAACACCAAAGGGAAAGGGCTCAGACCTCCATCACAGCAGGGTCACTCTCGC

[0654] ACGTGGTAGGACCATACCACCTTCACAAAGGCGGGGTTTGCCACCTTTGGGGAGCTCT

[0655] GGGGGGCCTCTACCTCCTCACCAGTCTGACACAATGCCAGAGATTCCACCACTGGGAA

[0656] TTTCTTATAATTAAAGCATTCTCTTCCCTGTATTAAATGAAAATGCCCCTGGGAGAGTT

[0657] 74

[0658] WBD (US) 4882-8896-8689vl

[0659] P893392050WO (01287) AATAGAGCAAGCCTTTATCAACCATTAAAAACTGAGGGCCAGTTAGTTTCTCTTTCTTT

[0660] TCCCCCTGAAGTGGTACTTCATTTTGTTTTATAGAAAAAAGATTCAGGCAAGGGAAGT

[0661] GTGGGTGGCTGGGGAGGCAGGTCTGCTCCTTTGAGTTGGCTGCAGTGACATGGAAGTC

[0662] ACAGGGCTGAGGGAAGGAGACAAGAGCCTGGACAGCAGTGAAGGGGTCAAAGACAG

[0663] ACCCCTCCAAGAACCTCAGAGGAGACCCGGACTGCAGGAGACCTGCAGGAGGCCCGT

[0664] GGGAGCCTGTGGAGGCCTGTGGAGGCCCGCGGGAGCCTGTGGAGGCCTGTGGAGGCC

[0665] CGCGGGAGCCTGTGGAGGCCTGTGGAGGCCTGTGGAGGTCTGCGGGAGCCTGTGGAG

[0666] GCCTGTGGAGGCCTGCGGGAGCCTGTGGAGGCCTGTGGAGGCCTGTGGAGGTCTGCGG

[0667] GAGCCTGTGGAGGTCTGCGGGAGCCTGTGGAGGCCTGTGGAGGCCTGTGGAGGCCTGC

[0668] GGGAGCCTGTGGAGGTCTGCGGGAGCCTGTGGAGGCCTGTGGAGGCCTGTGGAGGTCT

[0669] GCGGGAGCCTGTGGAGGCCTGTGGAGGCCTGTGGAGGTCTGCGGGAGCCTGTGGAGG

[0670] CCTGTGGAGGCCTGCCAGCCCAGTGCCCTCAGGCAGCAAAGCCCAGCAGTCACTGGTC

[0671] CCCCAAGGCCGCCCACAGATCTATGCAAGGTGATGCAGCAGGGGCTATGGACCCACCC

[0672] ACACCCAAGCAGGGAGGCACGATGACAAGGCCCGGGCAGGTAGGGGGACATCCGGAA

[0673] AGCAGGCTCAGCTCCACCCCTGAGAAATTTCCGTCTAAATCCAGGAGATCCTTGATTC

[0674] AGCAAGAACCTTCCCTCATGTCCGAGGCTGTTAAACCTAAAGCCAGTTAGGGCAGAGG

[0675] TCAGAGGGGTGGACCAGGAGCAAGTGGGTCGAGGGTGAGCAAGTGTTGGGAGTGGGG

[0676] ATAGAAACTCCATCCATCTCCGACTGCTCTGCTCGCATATCTGGCTCCAGGGCCCACCT

[0677] GGTATGTCAGAGGGAATTAGAACGGCCTTGTGAGGAGGCTCAGGAGAAGGGCTCCTG

[0678] TGCCCACCGGCCAAGTCAGCACTGGGCCTAACGCCACCACAGCAAAGCCCCTCAGTGC

[0679] AGAGGGCCCAGCTATACG

[0680] SEQ ID NO: 5

[0681] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG

[0682] TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG

[0683] AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA

[0684] GACCGTGACTTGCCAGCCCCACAGAGCCCCGCCCTTGTCCATCACTGGCATCTGGACTC

[0685] CAGCCTGGGTTGGGGCAAAGAGGGAAATGAGATCATGTCCTAACCCTGATCCTCTTGT

[0686] CCCACAGATATCCAGAACCCTGACCCTGCCGTGTACCAGCTGAGAGACTCTAAATCCA

[0687] GTGACAAGTCTGTCTGCCTATTCACCGATTTTGATTCTCAAACAAATGTGTCACAAAGT

[0688] AAGGATTCTGATGTGTATATCACAGACAAAACTGTGCTAGACATGAGGTCTATGGACT

[0689] TCAAGAGCAACAGTGCTGTGGCCTGGAGCAATAAGAGCGATTTCGCCTGCGCCAATGC

[0690] ATTTAATAATTCTATCATCCCCGAGGATACATTTTTTCCGTCACCCGGTAAGTATTAAT

[0691] GTTACAAGACAGGTTTAAGGACACCAATAGAAACTGGGCTTGTCGAGACAGAGAGTA

[0692] CCTGTTTGAATGAGGCTTCAGTACTTTACAGAATCGTTGCCTGCACATCTTGGAAACAC

[0693] TTGCTGGGATTACTTCGACTTCTTAACCCAACAGAAGGCTCGAGAAGGTATATTGCTGT

[0694] TGACAGTGAGCGCGGACATGATCTTCTTTATAATTAGTGAAGCCACAGATGTAATTAT

[0695] AAAGAAGATCATGTCCATGCCTACTGCCTCGGACTTCAAGGGGCTAGAATTCGAGCAA

[0696] TTATCTTGTTTACTAAAACTGAATACCTTGCTATCTCTTTGATACATTTTTACAAAGCTG

[0697] AATTAAAATGGTATAAATTAAATCACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCA

[0698] CCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAGAATCCAGTTGCGACGT

[0699] GAAACTCGTTGAAAAGTCATTCGAGACCGACACTAATCTTAATTTCCAGAATCTCTCTG

[0700] TCATCGGATTTCGCATTCTCCTGCTTAAGGTTGCTGGCTTCAACCTTCTGATGACCTTGC

[0701] GCTTGTGGTCTAGTTGAGACTTGCAGGACTGCAAAACCGCTTGCGCACCATCTCTTCTC

[0702] CCACTTCACTGTCCATCCTCCCCATCACCCAATCGCGGTAATTCACCCACTCCTAAAGA

[0703] AGTTAAGGCCGCAACTACATCCGTTCCTCCACGCCAGTGTCATCAGCTCGACCCCACA

[0704] AGAATCTACGACTAAGACTGTTGAAGGGCCGCTAAGCATTGTTGTCATCCTCTGTGCA

[0705] GTCTCATCGCCGCCTGCCATTGTTTGACCTTTACTGCCGAAGCTCGGCTTCTCCAACCC

[0706] CCATTGGCCGTTCATATCCCTAGCTGTTCTCCTGAAACAGCTTCTGCTATTCCCCAAAT

[0707] GATGGACCTGCAATGGGTCCTTCTTGGAAGTCGGAGTTGTAGGATGCTGTGAGGCGTC

[0708] 75

[0709] WBD (US) 4882-8896-8689vl

[0710] P893392050WO (01287) TACTTCTTCTTGATCGTTTTTATCAAAAAGTATATCGTCTTTTTTTTTAGCAGGCGCGGA

[0711] GGTAAGCTGAGTCACTACCGCGGACCCGCCATGTTGTGCATTTGGGCTTGCTGCATGA

[0712] GTTGTTGTAGATGTCTCCACTAAAACGACCTGGAGGAACAAAGGACTCTCTGAAACTT

[0713] CACTTGAAGACTTGACCTCTGAGTGCAACTGGAGTCACAGCCACTCAGGAGAAAGGGC

[0714] AGGCTATGCAAGCACGTGCTTCAGGAGCATCCTGGTGAGGTCTTCAGCTCTGACAGAG

[0715] TC TA TA TC TC TC AC GT GC TT CG CC AT AG AC CT AA TC AC AC ATC AACT ^CG AG AG CG TCG CC CA CG TC GC GA GT AC TT CG AA CG GT TC AT TT AA CA AG GG AA CG AT GTG GC CT CT TA A

[0716] TAGCACCTTATTATCAGGTACCATCCTAGTGTGGAAGTCCCCTGGCCTGTGTGCTCTGT

[0717] GGGATGGGAATGTCCCTTGTCATCAATGTGTCCCCAGGAACTAGAACAGTGCCTGGCA

[0718] CACAACAGATGCCGTTAAATGAATAAGTGAATGACTCCTGCAGTTAGGCTCTGGGTGG

[0719] GGGGAGCTGAGTGTGAGGAGTGGGGAGTACAGGAGAAAGGAAAGTGGGGATTAGTA

[0720] GAGGAAGTGGTCAGGTGAGTCAATGCTTCACAGTTTTGCCAAGAGCTAAAGTAACTTG

[0721] TCAAGAACTTCAGAACAAGGAAAGCCTGAGTGGTGGCCCCTTGCTCAGGGCAGCCCCC

[0722] AGGCAGGCACACATCTATCCAACAAAGCCTGGAGGACAGGGACTGTGGTAAATCACG

[0723] CAGGGTACCTGGGAGTCTGACCTCACTGTGTGACCACACTGAGGTCCAGAGCCAACAG

[0724] TCACTTTATGTGGGTTCGCCTCAGTTTCTCCATGAGTCCTACTCTAAAACCCATCTACCT

[0725] CCGTGAGCCATGTGTGGGAAGACGAGCTTCTGACTGGGATGAGTCTGGGGGAGCAAG

[0726] CACATGGAGACCCAGTGCTGTCATTCTGGCTGCCGGAAGGCTGCAGGACCACCAGCGC

[0727] CCTGCGTCTGTGTCAAGGCCACCTGGGCAACGCGGCGAGGGGTGGGGAGTGCAGGCCT

[0728] CCTTTCACACAGAGACACACAGGCCTGAGGCCTCTCTGGGTAGCACACATAGAGCTAT

[0729] ATTTTGCAATTAATAAACAGGTCAGACATTGAGTCAGTAGCCACAAGTAGAACAGGAA

[0730] GTGGAAAAAGTCTCCCACTTCCCTCCAGGTGTTTGGGGGCTGGGGCGGTCCCCTCCCAT

[0731] TTCCATGACGTCATGGTTACCAAGAGGGGCAAGTAGGGCACCCTTTGAAGCTCTCCCG

[0732] CAGAAGCCACATCCTCTGGAAAGAAGAGTTTAAAATACTGAGTTAGAGATAGCATCGC

[0733] CCCAGGCCACGTGCCGAGGGGAGCAGGCTGGGCCGTTACACCACCCCCCAACCGCAG

[0734] GTGCAGCAAGGCCAACATGCCAGGCTGGGAGGGGCTGCCGGCCCCTCGGTGACTTGCA

[0735] CCGGGCCCGCAAGCAGGCTGGGGAGGTGCTGCTGGCCCCACCGCTGAGCCGCCTCCCT

[0736] CACCGCACCTCTGCTGCCCCTCCAAACCAGCCAAGATCAGACCTTGATCTCACCTGATT

[0737] AAGCGCTTGCTTAGTACAGGGCAGACCCTCGAGACTTATGCTCAGGCCAGTGTCTCCT

[0738] GGGAGGGAAGGAGGTGGGGAGGCCGTATTTGCCCCCACACAGAGCTAGGGCTGATTC

[0739] TCTGAGGGACCCCCCCCCAATAGCACCTCCCCCCTTACCCTGATGCTTGTCCATTAGGT

[0740] GATGGCCCTAATGGACGCCTTGATGGACACGTCCGGGACCCAGGGACCCGGAGACCA

[0741] GGAATCTGCTGGTCTCAGGGAGCCGCAGAACCTGAAACCCAAAAATACCAGGCTTGAT

[0742] CGCCTGCTCTTCGGTAGTGTCGCCATCCACCCGCATGGGGGCAGCAGCGCTCCACCAA

[0743] AAAGACCTCTTCTCCCTTTGTCTGTTAATTCTGGAGATAGAAAGCCCCCAAAGTTTAGG

[0744] AAAAGCTCTTTTCCGAGGAAAATAAAAGGGCAGGTGACTTTTATCCCACAAGTCCCTG

[0745] CCCTGTGCCCGTCACTGCGCTGGGCGCAAGGCCTGTGTCATCCTGTCACATCCCCCCCC

[0746] AAAGGAGAGGAGGGCCTGGGGCCTGGGCCGGCAGCAGGGGGCGCTGTTCGTCCGGGA

[0747] TCGTCAGACCCGCTCCACCCGGAGCCCTGGAGGAAGTGCCCAGGGAGGGCTGGGGGC

[0748] CCGGAGGAGGCCTTGGGTTCGGAATTAGGCTGGCTGGAGTCCGGGATCCACCACATCA

[0749] TACACTCGTGCCTTTGGCAGGTCACTAACCGCCCACCGCCTCAGTTTCCCGGTGGGACG

[0750] AACTCTCTCTCACAGGTAGTTTAAGAACCAGGCTTTCTTTGGCATTGTATACCTGCCAT GTGCCAGGCTGTGTCGTGC

[0751] SEQ ID NO: 6

[0752] AGTATTATTAAGTAGCCCTGCATTTCAGGTTTCCTTGAGTGGCAGGCCAGGCCTGGCCG

[0753] TGAACGTTCACTGAAATCATGGCCTCTTGGCCAAGATTGATAGCTTGTGCCTGTCCCTG

[0754] AGTCCCAGTCCATCACGAGCAGCTGGTTTCTAAGATGCTATTTCCCGTATAAAGCATGA GACCGTGACTTGCCAGCCCCACAGAGCCCCGCCCTTGTCCATCACTGGCATCTGGACTC CAGCCTGGGTTGGGGCAAAGAGGGAAATGAGATCATGTCCTAACCCTGATCCTCTTGT

[0755] 76

[0756] WBD (US) 4882-8896-8689vl

[0757] P893392050WO (01287) CCCACAGATATCCAGAACCCTGACCCTGCCGTGTACCAGCTGAGAGACTCTAAATCCA

[0758] GTGACAAGTCTGTCTGCCTATTCACCGATTTTGATTCTCAAACAAATGTGTCACAAAGT

[0759] AAGGATTCTGATGTGTATATCACAGACAAAACTGTGCTAGACATGAGGTCTATGGACT

[0760] TCAAGAGCAACAGTGCTGTGGCCTGGAGCAATAAGAGCGATTTCGCCTGCGCCAATGC

[0761] ATTTAATAATTCTATCATCCCCGAGGATACATTTTTTCCGTCACCCGGTAAGTATTAAT

[0762] GTTACAAGACAGGTTTAAGGACACCAATAGAAACTGGGCTTGTCGAGACAGAGAGTA

[0763] CCTGTTTGAATGAGGCTTCAGTACTTTACAGAATCGTTGCCTGCACATCTTGGAAACAC

[0764] TTGCTGGGATTACTTCGACTTCTTAACCCAACAGAAGGCTCGAGAAGGTATATTGCTGT

[0765] TGACAGTGAGCGCGGACATGATCTTCTTTATAATTAGTGAAGCCACAGATGTAATTAT

[0766] AAAGAAGATCATGTCCATGCCTACTGCCTCGGACTTCAAGGGGCTAGAATTCGAGCAA

[0767] TTATCTTGTTTACTAAAACTGAATACCTTGCTATCTCTTTGATACATTTTTACAAAGCTG

[0768] AATTAAAATGGTATAAATTAAATCACTTTGCGGCCAGACTCTTGCGTTTCTGATAGGCA

[0769] CCTATTGGTCTTACTGACATCCACTTTGCCTTTCTCTCCACAGAATCCAGTTGCGACGT

[0770] GAAACTCGTTGAAAAGTCATTCGAGACCGACACTAATCTTAATTTCCAGAATCTCTCTG

[0771] TCATCGGATTTCGCATTCTCCTGCTTAAGGTTGCTGGCTTCAACCTTCTGATGACCTTGC

[0772] GCTTGTGGTCTAGTTGAGACTTGCAGGACTGCAAAACCGCTTGCGCACCATCTCTTCTC

[0773] CCACTTCACTGTCCATCCTCCCCATCACCCAATCGCGGTAATTCACCCACTCCTAAAGA

[0774] AGTTAAGGCCGCAACTACATCCGTTCCTCCACGCCAGTGTCATCAGCTCGACCCCACA

[0775] AGAATCTACGACTAAGACTGTTGAAGGGCCGCTAAGCATTGTTGTCATCCTCTGTGCA

[0776] GTCTCATCGCCGCCTGCCATTGTTTGACCTTTACTGCCGAAGCTCGGCTTCTCCAACCC

[0777] CCATTGGCCGTTCATATCCCTAGCTGTTCTCCTGAAACAGCTTCTGCTATTCCCCAAAT

[0778] GATGGACCTGCAATGGGTCCTTCTTGGAAGTCGGAGTTGTAGGATGCTGTGAGGCGTC

[0779] TACTTCTTCTTGATCGTTTTTATCAAAAAGTATATCGTCTTTTTTTTTAGCAGGCGCGGA

[0780] GGTAAGCTGAGTCACTACCGCGGACCCGCCATGTTGTGCATTTGGGCTTGCTGCATGA

[0781] GTTGTTGTAGATGTCTCCACTAAAACGACCTGGAGGAACAAATTAGGGCACTGTT GC

[0782] CTCTGGGGTAGAAGGATGGATTGAAAAGGGGCACGAGGAAACCTTCAGGGGTGGTGG

[0783] AAATGTTCCATATCTTTATCTTGGGTGATGGCTATATATATATCAAATGCATCATTAGT

[0784] GTATTTTCATGTAATTATGCCTCAAGAAAGTACTGCTCCCATTAAATACACACACACAC

[0785] ACACACACACACACACACAGAGTTTCTGTCTTCAAGGAGTTAACAGTCCAGTGAGTAG

[0786] TGGGTATGAAAGAGATGTGTTAATGGCTGCAATTGCAAGACAGAGGCTTTTGTTACTA

[0787] TATTCTCCTCAGGCAGGAGTCCTCAGGCAGGAGGCAGGAGAACCAGTTGTTTGTCCCT

[0788] AAGCGTCACTGCCCCTCACTAGCCTTCTCTCCTGGGGCCTTTAATTTTTCTGAGTCCCCA

[0789] TGTGTTATTATTCACATAGAGATAAAAATACCTGCCCACCTATTCTCAGTGTTGTTGAA

[0790] GGGATTGAAAGCAGTAATGGGTGTGGAATGCTTAGTAAACTGAAATCCCAGTGCAAAT

[0791] ATGAAGTGGGAATGCATCAGAATGCCCAGGAAGACGAACCATGACCCGAACAAGATG

[0792] GGCTGGGCCAGGTCAGATGGCACTGCCATACTGAGAAAAAAGTGGGCTGGCTGCGGC

[0793] TTCACACAGCCCGGGAGAGCTCAGAACTCCCCAGGCCCTGCCGGTGCGCTGGTGTTTC

[0794] TCATAGTATGTGGCCACCTCTCCAGTCACTTCGGCCCTAGGAGCTTCTTGTTACACCAT

[0795] GATTTTGTCACTACCTCCGCAGCTTCCCACCTCCTGCTTCCGGCTTCCAGAAGGAAGCG

[0796] CGTGATGGAGGGGAAATGACAGGGCCAGGAGGCAGCGGTCAGGCTGGGCCGAGTTTC

[0797] TGACAGAAGCCAACACTTCTGGCCCCACAGTGCTCACGGCCTCCCTGACTCCTCACCA

[0798] CATCACAGCTGAAAAGGCTCACGTCTTATCTCCCCTCTCAACCTCTGGGGGCTTCTCTG

[0799] CTTTCAGAAGGAGCCCCTGGAGCCCCTGGCCGAACACATCTGAAAGCACGGAGCCAG

[0800] ACCTCTGCAGCGCCCTCCAGCGGGCAGGTTCTGCGGGCAGCCATCTCCTCGCAGGCCA

[0801] GAACACCCCACGTCACTCCAGGACGGTTCGCCACCTTTGAGAAGAGAGGACTCCAGTG

[0802] TCTGTTCGCTCTGGAAATGTGGGCATGTTCCCATTCCAAAGGTGATAGGATCTGAGGG

[0803] GAGGATCCTTCCCATTTTCAAAGCCTTCCTAAAATCCCACTGATTAGTGGGGAGGAGT

[0804] GCAAGTCTACCCTTTCCCCGAACCACCCACATGAAATTTAGCAACAACAACAACAAAA

[0805] AAGGAAAATATTAGAGGAAGATCATATAAGTGTAAGAAGACAAGGCTGGCAGGACCT

[0806] AAGAGAGTCTAGGTAGATGAAGGAATATGGTTTACATAATGGGGAGCACCTGGTCTCC

[0807] 77

[0808] WBD (US) 4882-8896-8689vl

[0809] P893392050WO (01287) ATTTCATGCTTCCAGCCTTGATGGGAGGAAAGTAGGTAAACCTTAAAAGGGACTTCTT

[0810] TCCAGTGGAGGCTGCAAGATTCTATTTTTACCAACCTCTAAGGTGTCCACCTTATCCTT

[0811] GCTGACCCAAGTTATTCCAGTGATTTACCAAGGGAAATGGGTGGAACGACATCCTTAT

[0812] CTGGACTCCTTATTAACAATATCCTACATCTGTACAACACCAAAGCAAACATCAACTCT TTTGATCAAAATGATAAACCCCAGTGGTAGACAGAGTATTCTCTCCTCTTTTACAAATC

[0813] AGAAGGCACATGGAGCGGCTTGCACAGGATCACGGTGAACTAGGCCCTGATCCCCACC

[0814] AGACTCGGATCCCAGGCTTTCCACCAAGCTCTGTGGACTCAGAAATACCAAGTGGTTC

[0815] CAGATTCAGCCATCAGAGGAAGGAGCATGAATTGACTGGCATTCCCAGGCCCCCTAAG

[0816] ACTAAGTGTCTTGTTTCCTGCCTCCGTGGCCCTTCCCTCTTCCTGACTCTGCCTGCATGG ATTTGATCCCCCTTCAGGATGAGATCAGAATGGCAAGGCAGAGCTCTGTACCTTTAGG

[0817] AGGCACTCCCTCAATAAATATTTGATGTTGCCGGGCACGGTGGCTCACGCCTGTAATCC

[0818] CAGCACTTTGGGAAGCCAAGGCGGGCAGATCACTTGAGGTCAGGAGTTCAAGACTAG

[0819] CCTGGCAAACATGGTGAAACCCCATCTCTAGTAAAAATACAAAAAATTAGCTGGGCAT

[0820] GGTGGCGGGCGCCTGTAATCCCAGCTACTAGGGAGGCTGAGGCAGGAGAATTGCTTGA ACCCAGGAAGTGGAG

[0821] 78

[0822] WBD (US) 4882-8896-8689vl

[0823] P893392050WO (01287)

Claims

CLAIMSWhat is claimed is:

1. A cell comprising an exogenous polynucleotide, wherein said cell comprises in its genome a double-strand break, wherein said exogenous polynucleotide comprises, from 5' to 3':(a) a first homology arm;(b) a replacement nucleic acid sequence; and(c) a second homology arm; wherein said first homology arm has homology to a first genomic sequence that is adjacent to said double-strand break, wherein said second homology arm has homology to a second genomic sequence that is not adjacent to said double-strand break, and wherein said double-strand break and said second genomic sequence are separated by an intervening genomic sequence.

2. The cell of claim 1, wherein said first genomic sequence is 5' upstream of said double strand break.

3. The cell of claim 1 or claim 2, wherein said second genomic sequence is 3' downstream of said double strand break.

4. The cell of any one of claims 1-3, wherein said replacement nucleic acid sequence comprises a plurality of modified nucleotides and / or modified codons that differ from corresponding endogenous nucleotides or codons in said intervening genomic sequence, and wherein said replacement nucleic acid sequence encodes the same amino acid sequence as said intervening genomic sequence except for at least one replacement modification, wherein said replacement modification comprises: a) one or more nucleotide modifications that result in an altered codon, wherein said altered codon encodes a different amino acid than a corresponding endogenous codon in said intervening genomic sequence, or wherein said altered codon is a stop codon;79 WBD (US) 4882-8896-8689vlP893392050WO (01287)b) deletion of one or more nucleotides or codons that are present in said intervening genomic sequence; or c) insertion of one or more nucleotides or codons that are not present in said intervening genomic sequence.

5. The cell of any one of claims 1-4, wherein at least 20% of the nucleotides in said replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

6. The cell of any one of claims 1-5, wherein at least 33% of the nucleotides in said replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

7. The cell of any one of claims 1-6, wherein at least 75% of the codons in said replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

8. The cell of any one of claims 1-7, wherein all codons in said replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

9. The cell of claim 7 or claim 8, wherein each of said modified codons comprises one nucleotide change, or two nucleotide changes, relative to a corresponding codon in said intervening genomic sequence.

10. The cell of any one of claims 1-9, wherein said replacement nucleic acid sequence has less than 80% homology to said intervening genomic sequence.

11. The cell of any one of claims 1-10, wherein said replacement nucleic acid sequence has less than 70% homology to said intervening genomic sequence.

12. The cell of any one of claims 1-11, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 50 consecutive nucleotides in common.80WBD (US) 4882-8896-8689vlP893392050WO (01287)13. The cell of any one of claims 1-12, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 25 consecutive nucleotides in common.

14. The cell of any one of claims 1-13, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 10 consecutive nucleotides in common.

15. The cell of any one of claims 1-14, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 5 consecutive nucleotides in common.

16. The cell of any one of claims 1-15, wherein said replacement modification comprises one or more nucleotide modifications that result in an altered codon, wherein said altered codon encodes a different amino acid than a corresponding endogenous codon in said intervening genomic sequence, or wherein said altered codon is a stop codon.

17. The cell of claim 16, wherein said altered codon comprises one nucleotide change, or two nucleotide changes, relative to said corresponding endogenous codon in said intervening genomic sequence.

18. The cell of claim 16 or claim 17, wherein said double-strand break is present within mutant gene encoding a mutant protein, wherein said endogenous codon encodes a mutant aminocid, and wherein said altered codon encodes a wild-type amino acid or a stop codon.

19. The cell of any one of claims 1-15, wherein said replacement modification comprises a deletion of one or more nucleotides or codons that are present in said intervening genomic sequence.

20. The cell of claim 19, wherein said double-strand break is within a gene encoding a protein, and wherein said deletion inhibits expression of a full-length and / or active version of said protein.81WBD (US) 4882-8896-8689vlP893392050WO (01287)21. The cell of claim 19 or claim 20, wherein said double-strand break is within a mutant gene encoding a mutant protein, wherein said mutant gene comprises one or more mutant nucleotides and / or mutant codons not normally present in a wild-type gene, and wherein said replacement modification does not comprise said mutant nucleotides and / or mutant codons.

22. The cell of any one of claims 1-15, wherein said replacement modification comprises an insertion of one or more nucleotides or codons that are not present in said intervening genomic sequence.

23. The cell of claim 22, wherein said double-strand break is within a mutant gene encoding a mutant protein, wherein said mutant gene lacks one or more deleted nucleotides or deleted codons normally present in a wild-type gene, and wherein said replacement modification comprises said deleted nucleotides or deleted codons.

24. The cell of claim 22 or claim 23, wherein said replacement modification comprises an insertion of a nucleic acid sequence encoding a transgene, or a portion thereof.

25. The cell of claim 22 or claim 23, wherein said replacement modification comprises n insertion of an intron.

26. The cell of claim 22 or claim 23, wherein said replacement modification comprises an insertion of a nucleic acid sequence encoding a regulatory element, or a portion thereof, or an inhibitory element, or a portion thereof.

27. The cell of any one of claims 1-26, wherein said first homology arm and said second homology arm are approximately the same length.

28. The cell of claim 27, wherein said first homology arm and said second homology arm are between about 100 to 2000 base pairs in length.

29. The cell of claim 27 or claim 28, wherein said first homology arm and said second homology arm are between about 200 to 800 base pairs in length.82WBD (US) 4882-8896-8689vlP893392050WO (01287)30. The cell of any one of claims 27-29, wherein said first homology arm and said second homology arm are between about 400 to 600 base pairs in length.

31. The cell of any one of claims 27-30, wherein said first homology arm and said second homology arm are about 400 base pairs in length.

32. The cell of any one of claims 27-31, wherein said first homology arm and said second homology arm are about 500 base pairs in length.

33. The cell of any one of claims 1-26, wherein said first homology arm and said second homology arm are different lengths.

34. The cell of claim 33, wherein said second homology arm is longer than said firstomology arm.

35. The cell of claim 33 or claim 34, wherein said first homology arm is between about 100 to 1000 base pairs in length and said second homology arm is between about 1000 to 3000 base pairs in length.

36. The cell of any one of claims 33-35, wherein said first homology arm is between about 200 to 800 base pairs in length and said second homology arm is between about 1200 to 1800 base pairs in length.

37. The cell of any one of claims 33-36, wherein said first homology arm is between about 400 to 600 base pairs in length and said second homology arm is between about 1400 to 1600 base pairs in length.

38. The cell of any one of claims 33-37, wherein said first homology arm is about 400 or about 500 base pairs in length and said second homology arm is about 1500 or about 1600 base pairs in length.83WBD (US) 4882-8896-8689vlP893392050WO (01287)39. The cell of any one of claims 1-38, wherein said first homology arm has at least 98%, at least 99%, or 100% sequence homology to said first genomic sequence.

40. The cell of any one of claims 1-39, wherein said second homology arm has at least%, at least 99%, or 100% sequence homology to said second genomic sequence.

41. The cell of any one of claims 1-40, wherein said intervening genomic sequence is between about 10 to about 25000 base pairs in length.

42. The cell of any one of claims 1-41, wherein said intervening genomic sequence is between about 100 to about 10000 base pairs in length.

43. The cell of any one of claims 1-42, wherein said intervening genomic sequence is between about 1000 to about 10000 base pairs in length.

44. The cell of any one of claims 1-43, wherein said intervening genomic sequence is between about 2500 to about 10000 base pairs in length.

45. The cell of any one of claims 1-44, wherein said intervening genomic sequence is between about 1 to about 500 base pairs in length.

46. The cell of any one of claims 1-45, wherein said intervening genomic sequence is between about 1 to about 250 base pairs in length.

47. The cell of any one of claims 1-46, wherein said intervening genomic sequence is between about 1 to about 100 base pairs in length.

48. The cell of any one of claims 1-47, wherein said intervening genomic sequence is between about 1 to about 50 base pairs in length.

49. The cell of any one of claims 1-48, wherein said replacement nucleic acid sequence is between about 1 to about 5000 base pairs in length.WBD (US) 4882-8896-8689vlP893392050WO (01287)50. The cell of any one of claims 1-49, wherein said replacement nucleic acid sequence is between about 1 to about 4500 base pairs in length.

51. The cell of any one of claims 1-50, wherein said replacement nucleic acid sequence is between about 1 to about 4000 base pairs in length.

52. The cell of any one of claims 1-51, wherein said replacement nucleic acid sequence is between about 1 to about 3000 base pairs in length.

53. The cell of any one of claims 1-52, wherein said replacement nucleic acid sequence is between about 1 to about 2000 base pairs in length.

54. The cell of any one of claims 1-53, wherein said replacement nucleic acid sequence is between about 1 to about 1000 base pairs in length.

55. The cell of any one of claims 1-54, wherein said double-strand break is positioned in a cleaved exon of a gene.

56. The cell of claim 55, wherein said second genomic sequence begins within saidleaved exon.

57. The cell of claim 55 or claim 56, wherein said second genomic sequence is entirely comprised within said cleaved exon.

58. The cell of claim 55, wherein said second genomic sequence begins within a second exon that is 3' downstream of said cleaved exon.

59. The cell of claim 58, wherein said second genomic sequence is entirely comprised within said second exon.

60. The cell of claim 58, wherein said second genomic sequence spans said second exon and one or more 3' downstream introns or exons.85WBD (US) 4882-8896-8689vlP893392050WO (01287)61. The cell of claim 55, wherein said second genomic sequence begins within an intron that is 3' downstream of said cleaved exon.

62. The cell of claim 61, wherein said second genomic sequence is entirely comprised within said intron.

63. The cell of claim 61, wherein said second genomic sequence spans said intron and one or more 3' downstream exons or introns.

64. The cell of any one of claims 61-63, wherein said replacement genomic sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 5' start of said intron and the 5' start of said second genomic sequence.

65. The cell of claim 64, wherein said replacement intron sequence comprises a splice donor site that differs from the endogenous splice donor site but is capable of pairing with the endogenous splice acceptor site.

66. The cell of claim 64 or claim 65, wherein said replacement intron sequence is a synthetic intron sequence.

67. The cell of claim 64 or claim 65, wherein said replacement intron sequence is from a different location within the same gene or from a different gene.

68. The cell of any one of claims 1-54, wherein said double-strand break is positioned in a cleaved intron of a gene.

69. The cell of claim 68, wherein said second genomic sequence is entirely comprised within said cleaved intron.

70. The cell of claim 68 or claim 69, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron86WBD (US) 4882-8896-8689vlP893392050WO (01287)sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.

71. The cell of claim 68, wherein said second genomic sequence begins within an exon that is 3' downstream of said cleaved intron.

72. The cell of claim 71, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.

73. The cell of claim 71 or claim 72, wherein said second genomic sequence is entirely comprised within said exon.

74. The cell of claim 71 or claim 72, wherein said second genomic sequence spans said exon and one or more 3' downstream introns or exons.

75. The cell of claim 68, wherein said second genomic sequence begins within an intron that is 3' downstream of said cleaved exon.

76. The cell of claim 75, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.

77. The cell of claim 75 or claim 76, wherein said second genomic sequence is entirely comprised within said intron.

78. The cell of claim 75 or claim 76, wherein said second genomic sequence spans saidntron and one or more 3' downstream exons or introns.

79. The cell of any one of claims 70, 72-74, or 76-78, wherein said replacement intron sequence comprises:87WBD (US) 4882-8896-8689vlP893392050WO (01287)(a) a splice donor site that differs from an endogenous splice donor of said gene site but is capable of pairing with an endogenous splice acceptor site of said gene;(b) a splice acceptor site that differs from an endogenous splice acceptor site of said gene but is capable of pairing with an endogenous splice donor site;(c) an intron enhancer; or(d) any combination of (a) -(c).

80. The cell of any one of claims 70, 72-74, or 76-79, wherein said replacement intron sequence is a synthetic intron sequence.

81. The cell of any one of claims 70, 72-74, or 76-79, wherein said replacement intron sequence is a heterologous intron sequence from a different location within the same gene or from a different gene.

82. The cell of any one of claims 1-81, wherein each end of said double-strand break has a 3' overhang.

83. The cell of claim 82, wherein said 3' overhang is a 4 base pair 3' overhang.

84. The cell of any one of claims 1-81, wherein each end of said double-strand break is a blunt end.

85. The cell of any one of claims 1-81, wherein each end of said double-strand break has a 5' overhang.

86. The cell of any one of claims 1-85, wherein said exogenous polynucleotide comprises a nucleic acid sequence encoding a nuclease.

87. The cell of claim 86, wherein said nucleic acid sequence encoding said nuclease isositioned 5' upstream of said first homology arm.

88. The cell of claim 86, wherein said nucleic acid sequence encoding said nuclease is positioned 3' downstream of said second homology arm.88WBD (US) 4882-8896-8689vlP893392050WO (01287)89. The cell of any one of claims 86-88, wherein said nuclease is capable of binding and cleaving the genome of said cell to generate said double-strand break.

90. The cell of any one of claims 86-89, wherein said exogenous polynucleotide comprises a promoter that is operably linked to said nucleic acid sequence encoding said nuclease.

91. The cell of any one of claims 86-90, wherein said nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, a ompact TALEN, or a megaTAL.

92. The cell of any one of claims 1-91, wherein said cell is in vitro.

93. The cell of any one of claims 1-91, wherein said cell is in vivo.

94. The cell of any one of claims 1-93, wherein said exogenous polynucleotide is comprised by a delivery vehicle.

95. The cell of claim 94, wherein said delivery vehicle is a recombinant virus and said exogenous polynucleotide is comprised by a viral genome.

96. The cell of claim 95, wherein said recombinant virus is a recombinant adeno- associated virus (AAV).

97. The cell of claim 94, wherein said delivery is a lipid nanoparticle.

98. The cell of claim 97, wherein said exogenous polynucleotide is a messenger RNA (mRNA), a single -stranded DNA, or a double-stranded DNA.

99. A method for genetically modifying a cell, said method comprising introducing into said cell an exogenous polynucleotide and a nuclease or a gene encoding a nuclease, wherein said nuclease is expressed in said cell and generates a double-strand break in the genome of said cell,89WBD (US) 4882-8896-8689vlP893392050WO (01287)wherein said exogenous polynucleotide comprises, from 5' to 3':(a) a first homology arm;(b) a replacement nucleic acid sequence; and(c) a second homology arm; wherein said first homology arm has homology to a first genomic sequence that is adjacent to said double-strand break, wherein said second homology arm has homology to a second genomic sequence that is not adjacent to said double-strand break, wherein said double-strand break and said second genomic sequence are separated by an intervening genomic sequence, and wherein said intervening genomic sequence is replaced in the genome by said replacement nucleic acid sequence.

100. The method of claim 99, wherein said first genomic sequence is 5' upstream of said double strand break.

101. The method of claim 99 or claim 100, wherein said second genomic sequence is 3' downstream of said double strand break.

102. The method of any one of claims 99-101, wherein said replacement nucleic acid sequence comprises a plurality of modified nucleotides and / or modified codons that differ from corresponding endogenous nucleotides or codons in said intervening genomic sequence, and wherein said replacement nucleic acid sequence encodes the same amino acid sequence as said intervening genomic sequence except for at least one replacement modification, wherein said replacement modification comprises: a) one or more nucleotide modifications that result in an altered codon, wherein said altered codon encodes a different amino acid than a corresponding endogenous codon in said intervening genomic sequence, or wherein said altered codon is a stop codon; b) deletion of one or more nucleotides or codons that are present in said intervening genomic sequence; or c) insertion of one or more nucleotides or codons that are not present in said intervening genomic sequence.90WBD (US) 4882-8896-8689vlP893392050WO (01287)103. The method of any one of claims 99-102, wherein at least 20% of the nucleotides in said replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

104. The method of any one of claims 99-103, wherein at least 33% of the nucleotides in said replacement nucleic acid sequence are modified nucleotides that differ from a corresponding endogenous nucleotide.

105. The method of any one of claims 99-104, wherein at least 75% of the codons in said replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

106. The method of any one of claims 99-105, wherein all codons in said replacement nucleic acid sequence are modified codons that differ from a corresponding endogenous codon.

107. The method of claim 105 or claim 106, wherein each of said modified codons comprises one nucleotide change, or two nucleotide changes, relative to a corresponding codon in said intervening genomic sequence.

108. The method of any one of claims 99-107, wherein said replacement nucleic acid sequence has less than 80% homology to said intervening genomic sequence.

109. The method of any one of claims 99-108, wherein said replacement nucleic acid sequence has less than 70% homology to said intervening genomic sequence.

110. The method of any one of claims 99-109, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 50 consecutive nucleotides in common.

111. The method of any one of claims 99- 110, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 25 consecutive nucleotides in common.WBD (US) 4882-8896-8689vlP893392050WO (01287)112. The method of any one of claims 99-111, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 10 consecutive nucleotides in common.

113. The method of any one of claims 99-112, wherein said replacement nucleic acid sequence and said intervening nucleic acid sequence do not share a sequence having more than 5 consecutive nucleotides in common.

114. The method of any one of claims 99-113, wherein said replacement modification comprises one or more nucleotide modifications that result in an altered codon, wherein said altered codon encodes a different amino acid than a corresponding endogenous codon in said intervening genomic sequence, or wherein said altered codon is a stop codon.

115. The method of claim 114, wherein said altered codon comprises one nucleotide change, or two nucleotide changes, relative to said corresponding endogenous codon in said intervening genomic sequence.

116. The method of claim 114 or claim 115, wherein said double-strand break is generated within a mutant gene encoding a mutant protein, wherein said endogenous codon encodes a mutant amino acid, and wherein (a) said altered codon encodes a wild-type amino acid and said mutant gene is corrected to encode a wild-type protein; or (b) said altered codon encodes a stop codon that disrupts expression of a full-length and / or active version of said mutant protein.

117. The method of any one of claims 99-113, wherein said replacement modification comprises a deletion of one or more nucleotides or codons that are present in said intervening genomic sequence.

118. The method of claim 117, wherein said double-strand break is within a gene encoding a protein, and wherein said deletion disrupts expression of a full-length and / or active version of said protein.

119. The method of claim 117 or claim 118, wherein said double-strand break is generated within a mutant gene encoding a mutant protein, wherein said mutant gene comprises92WBD (US) 4882-8896-8689vlP893392050WO (01287)one or more mutant nucleotides and / or mutant codons not normally present in a wild-type gene, wherein said replacement modification does not comprise said mutant nucleotides and / or mutant codons, and wherein said mutant gene is corrected to encode a wild-type protein.

120. The method of any one of claims 99-113, wherein said replacement modification comprises an insertion of one or more nucleotides or codons that are not present in said intervening genomic sequence.

121. The method of claim 120, wherein said double-strand break is within a mutant gene encoding a mutant protein, wherein said mutant gene lacks one or more deleted nucleotides or deleted codons normally present in a wild-type gene, wherein said replacement modification comprises said deleted nucleotides or deleted codons, and wherein said mutant gene is corrected to encode a wild-type protein.

122. The method of claim 120 or claim 121, wherein said replacement modification comprises an insertion of a nucleic acid sequence encoding a transgene, or a portion thereof.

123. The method of claim 120 or claim 121, wherein said replacement modification comprises an insertion of an intron.

124. The method of claim 120 or claim 121, wherein said replacement modification comprises an insertion of a nucleic acid sequence encoding a regulatory element, or a portion thereof, or an inhibitory element, or a portion thereof.

125. The method of any one of claims 99-124, wherein said first homology arm and said second homology arm are approximately the same length.

126. The method of claim 125, wherein said first homology arm and said second homology arm are between about 100 to 2000 base pairs in length.

127. The method of claim 125 or claim 126, wherein said first homology arm and said second homology arm are between about 200 to 800 base pairs in length.93WBD (US) 4882-8896-8689vlP893392050WO (01287)128. The method of any one of claims 125-127, wherein said first homology arm and said second homology arm are between about 400 to 600 base pairs in length.

129. The method of any one of claims 125-128, wherein said first homology arm and said second homology arm are about 400 base pairs in length.

130. The method of any one of claims 125-129, wherein said first homology arm and said second homology arm are about 500 base pairs in length.

131. The method of any one of claims 99- 124, wherein said first homology arm and said second homology arm are different lengths.

132. The method of claim 131, wherein said second homology arm is longer than said first homology arm.

133. The method of claim 131 or claim 132, wherein said first homology arm is between about 100 to 1000 base pairs in length and said second homology arm is between about 1000 to 3000 base pairs in length.

134. The method of any one of claims 131-133, wherein said first homology arm is between about 200 to 800 base pairs in length and said second homology arm is between about 1200 to 1800 base pairs in length.

135. The method of any one of claims 131-134, wherein said first homology arm is between about 400 to 600 base pairs in length and said second homology arm is between about 1400 to 1600 base pairs in length.

136. The method of any one of claims 131-135, wherein said first homology arm is about 400 or about 500 base pairs in length and said second homology arm is about 1500 or about 1600 base pairs in length.

137. The method of any one of claims 99-136, wherein said first homology arm has at least 98%, at least 99%, or 100% sequence homology to said first genomic sequence.94WBD (US) 4882-8896-8689vlP893392050WO (01287)138. The method of any one of claims 99-137, wherein said second homology arm has at least 98%, at least 99%, or 100% sequence homology to said second genomic sequence.

139. The method of any one of claims 99-138, wherein said intervening genomic sequence is between about 10 to about 25000 base pairs in length.

140. The method of any one of claims 99-139, wherein said intervening genomic sequence is between about 100 to about 10000 base pairs in length.

141. The method of any one of claims 99-140, wherein said intervening genomic sequence is between about 1000 to about 10000 base pairs in length.

142. The method of any one of claims 99-141, wherein said intervening genomic sequence is between about 2500 to about 10000 base pairs in length.

143. The method of any one of claims 99-142, wherein said intervening genomic sequence is between about 1 to about 500 base pairs in length.

144. The method of any one of claims 99-143, wherein said intervening genomic sequence is between about 1 to about 250 base pairs in length.

145. The method of any one of claims 99-144, wherein said intervening genomic sequence is between about 1 to about 100 base pairs in length.

146. The method of any one of claims 99-145, wherein said intervening genomic sequence is between about 1 to about 50 base pairs in length.

147. The method of any one of claims 99-146, wherein said replacement nucleic acid sequence is between about 1 to about 5000 base pairs in length.

148. The method of any one of claims 99-147, wherein said replacement nucleic acid sequence is between about 1 to about 4500 base pairs in length.95WBD (US) 4882-8896-8689vlP893392050WO (01287)149. The method of any one of claims 99-148, wherein said replacement nucleic acid sequence is between about 1 to about 4000 base pairs in length.

150. The method of any one of claims 99-149, wherein said replacement nucleic acid sequence is between about 1 to about 3000 base pairs in length.

151. The method of any one of claims 99-150, wherein said replacement nucleic acid sequence is between about 1 to about 2000 base pairs in length.

152. The method of any one of claims 99-151, wherein said replacement nucleic acid sequence is between about 1 to about 1000 base pairs in length.

153. The method of any one of claims 99-152, wherein said double-strand break isositioned in a cleaved exon of a gene.

154. The method of claim 153, wherein said second genomic sequence begins within said cleaved exon.

155. The method of claim 153 or claim 154, wherein said second genomic sequence is entirely comprised within said cleaved exon.

156. The method of claim 153, wherein said second genomic sequence begins within a second exon that is 3' downstream of said cleaved exon.

157. The method of claim 156, wherein said second genomic sequence is entirely comprised within said second exon.

158. The method of claim 156, wherein said second genomic sequence spans said secondxon and one or more 3' downstream introns or exons.

159. The method of claim 153, wherein said second genomic sequence begins within an intron that is 3' downstream of said cleaved exon.96WBD (US) 4882-8896-8689vlP893392050WO (01287)160. The method of claim 159, wherein said second genomic sequence is entirely comprised within said intron.

161. The method of claim 159, wherein said second genomic sequence spans said intron and one or more 3' downstream exons or introns.

162. The method of any one of claims 159-161, wherein said replacement genomic sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 5' start of said intron and the 5' start of said second genomic sequence.

163. The method of claim 162, wherein said replacement intron sequence comprises a splice donor site that differs from the endogenous splice donor site but is capable of pairing withhe endogenous splice acceptor site.

164. The method of claim 162 or claim 163, wherein said replacement intron sequence is a synthetic intron sequence.

165. The method of claim 162 or claim 163, wherein said replacement intron sequence is from a different location within the same gene or from a different gene.

166. The method of any one of claims 99-152, wherein said double-strand break is positioned in a cleaved intron of a gene.

167. The method of claim 166, wherein said second genomic sequence is entirely comprised within said cleaved intron.

168. The method of claim 166 or claim 167, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.97WBD (US) 4882-8896-8689vlP893392050WO (01287)169. The method of claim 166, wherein said second genomic sequence begins within an exon that is 3' downstream of said cleaved intron.

170. The method of claim 169, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.

171. The method of claim 169 or claim 170, wherein said second genomic sequence isntirely comprised within said exon.

172. The method of claim 169 or claim 170, wherein said second genomic sequence spans said exon and one or more 3' downstream introns or exons.

173. The method of claim 166, wherein said second genomic sequence begins within an intron that is 3' downstream of said cleaved exon.

174. The method of claim 173, wherein said replacement nucleic acid sequence comprises a replacement intron sequence that differs from a corresponding endogenous intron sequence comprised by said intervening genomic sequence that is positioned between the 3' end of said double-strand break and the 5' end of said second genomic sequence.

175. The method of claim 173 or claim 174, wherein said second genomic sequence is entirely comprised within said intron.

176. The method of claim 173 or claim 174, wherein said second genomic sequence spans said intron and one or more 3' downstream exons or introns.

177. The method of any one of claims 168, 170-172, or 174-176, wherein said replacement intron sequence comprises:(a) a splice donor site that differs from an endogenous splice donor of said gene site but is capable of pairing with an endogenous splice acceptor site of said gene;98WBD (US) 4882-8896-8689vlP893392050WO (01287)(b) a splice acceptor site that differs from an endogenous splice acceptor site of said gene but is capable of pairing with an endogenous splice donor site;(c) an intron enhancer; or(d) any combination of (a)-(c).

178. The method of any one of claims 168, 170-172, or 174-177, wherein said replacement intron sequence is a synthetic intron sequence.

179. The method of any one of claims 168, 170-172, or 174-177, wherein said replacement intron sequence is a heterologous intron sequence from a different location within the same gene or from a different gene.

180. The method of any one of claims 99-179, wherein each end of said double-strand break generated by said nuclease has a 3' overhang.

181. The method of claim 180, wherein said 3' overhang is a 4 base pair 3' overhang.

182. The method of claim 180 or claim 181, wherein said nuclease is an engineered meganuclease.

183. The method of any one of claims 99-179, wherein each end of said double-strand break generated by said nuclease is a blunt end or has a 5' overhang.

184. The method of claim 183, wherein said nuclease is a CRISPR-system nuclease, ainc finger nuclease, a TALEN, a compact TALEN, or a megaTAL.

185. The method of any one of claims 99-184, wherein said exogenous polynucleotide comprises said gene encoding said nuclease.

186. The method of claim 185, wherein said gene encoding said nuclease is positioned 5' upstream of said first homology arm.WBD (US) 4882-8896-8689vlP893392050WO (01287)187. The method of claim 185, wherein said gene encoding said nuclease is positioned 3' downstream of said second homology arm.

188. The method of any one of claims 185-187, wherein said exogenous polynucleotide comprises a promoter that is operably linked to said gene encoding said nuclease.

189. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by a recombinant virus.

190. The method of claim 189, wherein said recombinant virus is a recombinant adeno- associated virus (AAV).

191. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by non-viral delivery.

192. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by a lipid nanoparticle.

193. The method of any one of claims 99-192, wherein said gene encoding said nuclease introduced into said cell by a recombinant virus.

194. The method of claim 193, wherein said recombinant virus is a recombinant adeno- associated virus (AAV).

195. The method of any one of claims 99-192, wherein said gene encoding said nuclease is introduced into said cell by non-viral delivery.

196. The method of any one of claims 99-192, wherein said gene encoding said nuclease is introduced into said cell by a lipid nanoparticle.

197. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by a recombinant virus, and wherein said gene encoding said nuclease is introduced into said cell by a lipid nanoparticle.100WBD (US) 4882-8896-8689vlP893392050WO (01287)198. The method of claim 197, wherein said recombinant virus is a recombinant AAV.

199. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by a first recombinant virus, and wherein said gene encoding said nuclease is introduced into said cell by a second recombinant virus.

200. The method of claim 199, wherein said first recombinant virus and said second recombinant virus are each a recombinant AAV.

201. The method of any one of claims 99-188, wherein said exogenous polynucleotide is introduced into said cell by a first lipid nanoparticle, and wherein said gene encoding said nuclease is introduced into said cell by a second lipid nanoparticle.

202. The method of claim 201, wherein said exogenous polynucleotide and said gene encoding said nuclease are introduced into said cell by a lipid nanoparticle.

203. The method of any one of claims 99-202, wherein said method is performed in vitro.

204. The method of any one of claims 99-202, wherein said method is performed in vivo.

205. The method of any one of claims 99-204, wherein said method is a method for modifying a gene in a target cell in a subject, said method comprising delivering to said target cell said exogenous polynucleotide and said gene encoding said nuclease of any one of claims 99-204.

206. The method of any one of claims 99-205, wherein said method is a method for treating a disease in a subject, said method comprising administering to said subject a therapeutically-effective amount of said exogenous polynucleotide and a therapeutically-effective amount of said gene encoding said nuclease of any one of claims 99-204.101WBD (US) 4882-8896-8689vlP893392050WO (01287)

Citation Information

Patent Citations

  • Methods and compositions for targeted cleavage and recombination

    US11311574B2

  • Methods and compositions for using zinc finger endonucleases to enhance homologous recombination

    US20030232410A1

  • Use of chimeric nucleases to stimulate gene targeting

    US20050026157A1

  • Methods and compositions for targeted cleavage and recombination

    US20050064474A1

  • Targeted chromosomal mutagenasis using zinc finger nucleases

    US20050208489A1