Base editing nucleotide sequences using homology directed repair
Nuclease-initiated HDR with modified homology arms allows for precise genomic editing by introducing nucleotide changes, insertions, and deletions, overcoming the limitations of traditional HDR methods in correcting genetic diseases.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PRECISION BIOSCIENCES INC
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for genomic DNA editing using homology directed repair (HDR) are insufficient for accurately correcting genetic diseases or achieving desired gene modifications, as they often fail to insert DNA templates at specified sites in the genome.
A method of nuclease-initiated HDR that utilizes a DNA repair template with homology arms containing nucleotide modifications, including mismatches, deletions, or insertions, to facilitate precise nucleotide changes, insertions, and deletions at genomic locations.
Enables the correction of pathogenic mutations by introducing single or multiple base changes, insertions, and deletions, thereby addressing the limitations of traditional HDR methods in genomic editing.
Smart Images

Figure IB2025060570_23042026_PF_FP_ABST
Abstract
Description
[0001] BASE EDITING NUCLEOTIDE SEQUENCES USING HOMOLOGY DIRECTED REPAIR
[0002] FIELD OF THE INVENTION
[0003] The present disclosure pertains to the field of molecular biology and recombinant nucleic acid technology. In particular, the present disclosure pertains to methods of genomic DNA editing using nuclease-initiated homology directed repair.
[0004] REFERENCE TO A SEQUENCE LISTING SUBMITTED ELECTRONICALLY AS AN XML FILE
[0005] The instant application contains a Sequence Listing which has been submitted in XML format via USPTO Patent Center and is hereby incorporated by reference in its entirety. Said XML copy, created on October 18, 2024, is named “P893392070USPl.xml”, and is 38,599 bytes in size.
[0006] BACKGROUND OF THE INVENTION
[0007] With the advent of precision medicine, there is an increased need for methods that can accurately edit the genome. Modification of genomic DNA at precise locations in genes comprising pathogenic mutations can permanently alleviate genetic diseases across many therapeutic areas. One mechanism that allows for modification of genomic DNA is homology directed repair (HDR), which can promote the insertion of DNA templates into the genome at a double strand break (DSB). Typically, such DSBs are generated by a site-specific engineered nuclease, and homologous recombination of the DNA template at the site of the DSB is enabled, in part, by the inclusion of upstream and downstream homology arms that are homologous to sequences flanking the chromosomal location of the DSB. However, insertion of a DNA template at a specified site in the genome is not always sufficient for correction of a genetic disease, or sufficient to achieve a desired gene modifications in vitro or in vivo. Thus, a need remains for novel methods that utilize HDR- based editing to modify genomic sequences.
[0008] SUMMARY OF THE INVENTION
[0009] The present disclosure provides for methods of nuclease-initiated HDR, referred to herein as HDR base editing, that can result in nucleotide modifications including single or multiple base changes, nucleotide insertions, and nucleotide deletions. Such methods enable, for example, the correction of one or more pathogenic mutations in a gene, including nucleotide changes, insertions, and / or deletions. The methods utilize a nuclease that generates a double strand break (DSB) and a
[0010] 1
[0011] P89339 2070WO (01288) DNA repair template comprising homology arms that have one or more base pair mismatches, nucleotide deletions or nucleotide insertions. Optionally the repair template can contain a heterologous nucleic acid sequence for insertion described further herein.
[0012] Thus, in one aspect of the disclosure is a cell comprising an exogenous polynucleotide, wherein the cell comprises in its genome a double-strand break, wherein the exogenous polynucleotide comprises, from 5’ to 3’, a first homology arm and a second homology arm, wherein the first homology arm has homology to a first genomic sequence that is 5 ’ upstream of the double-strand break; wherein the second homology arm has homology to a second genomic sequence that is 3’ downstream of the double-strand break, and wherein: a) the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence; b) the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence; or c) the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence, and the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence.
[0013] In some embodiments, the first genomic sequence is adjacent to the double-strand break. In some embodiments, the first genomic sequence is not adjacent to the double-strand break. In some embodiments, the second genomic sequence is adjacent to the double-strand break. In some embodiments, the second genomic sequence is not adjacent to the double-strand break. In some embodiments, the double-strand break comprises a 3’ overhang. In some embodiments, the doublestrand break comprises a four base pair 3’ overhang. In some embodiments, the double-strand break comprises a blunt end. In some embodiments, the double-strand break comprises a 5' overhang.
[0014] In some embodiments, the exogenous polynucleotide comprises, from 5’ to 3’, the first homology arm, a heterologous nucleic acid sequence, and the second homology arm.
[0015] In some embodiments, the double-strand break is positioned within an exon of a gene, an intron of a gene, or a non-protein coding region of a gene.
[0016] In some embodiments, the nucleotide modification comprises a mismatched nucleotide that differs from a corresponding endogenous nucleotide of the first genomic sequence or the second genomic sequence.
[0017] In some embodiments, the corresponding endogenous nucleotide is comprised with a gene coding sequence, a non-gene coding sequence, and / or intron sequence of a gene.
[0018] 2
[0019] P89339 2070WO (01288) In some embodiments, the corresponding endogenous nucleotide is an A, and the mismatched nucleotide is a T, C, or G. In some embodiments, the corresponding endogenous nucleotide is a T, and the mismatched nucleotide is an A, C, or G. In some embodiments, the corresponding endogenous nucleotide is a C, and the mismatched nucleotide is an A, T, or G. In some embodiments, the corresponding endogenous nucleotide is a G, and the mismatched nucleotide is an A, T, or C.
[0020] In some embodiments, the first homology arm comprises between 1-20 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the first homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the first homology arm comprises between 11-13 mismatched nucleotides. In some embodiments, the first homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the first homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the first homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the first homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the first homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the first homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the first homology arm comprises between 18-20 mismatched nucleotides.
[0021] In some embodiments, the first homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched
[0022] 3
[0023] P89339 2070WO (01288) nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides.
[0024] In some embodiments, the first homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-5 mismatched nucleotides.
[0025] In some embodiments, the first homology arm comprises 1 mismatched nucleotide. In some embodiments, the first homology arm comprises 2 mismatched nucleotides. In some embodiments, the first homology arm comprises 3 mismatched nucleotides. In some embodiments, the first homology arm comprises 4 mismatched nucleotides. In some embodiments, the first homology arm comprises 5 mismatched nucleotides.
[0026] In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the second homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the second homology arm comprises between 11- 13 mismatched nucleotides. In some embodiments, the second homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the second homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the second homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the second homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the second homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the second homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the second homology arm comprises between 18-20 mismatched nucleotides.
[0027] 4
[0028] P89339 2070WO (01288) In some embodiments, the second homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides.
[0029] In some embodiments, the second homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-5 mismatched nucleotides.
[0030] In some embodiments, the second homology arm comprises 1 mismatched nucleotide. In some embodiments, the second homology arm comprises 2 mismatched nucleotides. In some embodiments, the second homology arm comprises 3 mismatched nucleotides. In some embodiments, the second homology arm comprises 4 mismatched nucleotides. In some embodiments, the second homology arm comprises 5 mismatched nucleotides.
[0031] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide. In some embodiments, the mismatched nucleotide is a wild-type nucleotide. In some embodiments, the corresponding endogenous nucleotide is comprised by an exon.
[0032] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0033] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the mutant codon.
[0034] In some embodiments, the mismatched nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0035] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a wild-type amino acid.
[0036] In some embodiments, the mismatched nucleotide is comprised by a stop codon.
[0037] 5
[0038] P89339 2070WO (01288) In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0039] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the wild-type codon.
[0040] In some embodiments, the mismatched nucleotide is a mutant nucleotide comprised by a mutant codon.
[0041] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes the wild-type amino acid.
[0042] In some embodiments, the mismatched nucleotide is comprised by a stop codon.
[0043] In some embodiments, the corresponding endogenous nucleotide is comprised by an intron.
[0044] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0045] In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0046] In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0047] In some embodiments, the mismatched nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0048] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide.
[0049] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0050] In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0051] In some embodiments, the mismatched nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0052] In some embodiments, the corresponding endogenous nucleotide is comprised by a nongene protein coding sequence. In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element. In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0053] In some embodiments, the corresponding endogenous nucleotide is comprised by a mutant non-gene protein coding sequence.
[0054] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0055] In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0056] In some embodiments, the mismatched nucleotide differs from a wild-type nucleotide.
[0057] In some embodiments, the corresponding endogenous nucleotide is comprised by a wildtype non-gene protein coding sequence.
[0058] 6
[0059] P89339 2070WO (01288) In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0060] In some embodiments, the nucleotide modification comprises a nucleotide deletion, wherein the first homology arm and / or the second homology arm lacks at least one corresponding endogenous nucleotide that is present in the first genomic sequence or the second genomic sequence.
[0061] In some embodiments, the first homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 17-29 corresponding
[0062] 7
[0063] P89339 2070WO (01288) endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0064] In some embodiments, the first homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0065] In some embodiments, the first homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0066] In some embodiments, the first homology arm lacks 1 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0067] In some embodiments, the first homology arm lacks 2 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0068] In some embodiments, the first homology arm lacks 3 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0069] In some embodiments, the first homology arm lacks 4 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0070] 8
[0071] P89339 2070WO (01288) In some embodiments, the first homology arm lacks 5 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0072] In some embodiments, the second homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 17-29 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments,
[0073] 9
[0074] P89339 2070WO (01288) the second homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0075] In some embodiments, the second homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between
[0076] 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0077] In some embodiments, the second homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between
[0078] 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0079] In some embodiments, the second homology arm lacks 1 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0080] In some embodiments, the second homology arm lacks 2 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0081] In some embodiments, the second homology arm lacks 3 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0082] In some embodiments, the second homology arm lacks 4 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0083] 10
[0084] P89339 2070WO (01288) In some embodiments, the second homology arm lacks 5 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0085] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0086] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0087] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an exon sequence.
[0088] In some embodiments, the exon sequence is a mutant exon sequence.
[0089] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0090] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a codon that encodes a different amino acid than the mutant codon.
[0091] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a wild-type codon.
[0092] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a codon that encodes a wild-type amino acid.
[0093] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a stop codon.
[0094] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0095] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a mutant codon that encodes a different amino acid than the wild-type codon.
[0096] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a codon that encodes the same amino acid as the wild-type codon.
[0097] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a stop codon.
[0098] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an intron sequence.
[0099] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0100] 11
[0101] P89339 2070WO (01288) In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0102] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0103] In some embodiments, the first homology arm and / or the second homology arm comprises a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0104] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide.
[0105] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0106] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide.
[0107] In some embodiments, the first homology arm and / or the second homology arm comprises a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0108] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a non-gene coding sequence.
[0109] In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0110] In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0111] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0112] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0113] In some embodiments, the non-gene coding sequence is a wild-type non-gene coding sequence.
[0114] In some embodiments, the first homology arm and / or the second homology arm does not comprise at least one corresponding wild-type nucleotide.
[0115] In some embodiments, the nucleotide modification comprises a nucleotide insertion, wherein the first homology arm and / or the second homology arm comprises at least one inserted nucleotide that is not present in the first genomic sequence or the second genomic sequence.
[0116] In some embodiments, the first homology arm comprises between 1-20 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm
[0117] 12
[0118] P89339 2070WO (01288) comprises between 1-3 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-4 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-6 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 5-7 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 6-8 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 7-9 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 8-10 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 9-11 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 10-12 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 11-13 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 12-14 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 13-15 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 14-16 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 15-17 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 16-18 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 17-29 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 18-20 inserted nucleotides that are not present in the first genomic sequence.
[0119] In some embodiments, the first homology arm comprises between 1-10 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 1-3 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-4 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some
[0120] 13
[0121] P89339 2070WO (01288) embodiments, the first homology arm comprises between 4-6 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 5-7 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 6-8 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 7-9 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 8-10 inserted nucleotides that are not present in the first genomic sequence.
[0122] In some embodiments, the first homology arm comprises between 1-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-5 inserted nucleotides that are not present in the first genomic sequence.
[0123] In some embodiments, the first homology arm comprises 1 inserted nucleotide that is not present in the first genomic sequence.
[0124] In some embodiments, the first homology arm comprises 2 inserted nucleotides that is not present in the first genomic sequence.
[0125] In some embodiments, the first homology arm comprises 3 inserted nucleotides that is not present in the first genomic sequence.
[0126] In some embodiments, the first homology arm comprises 4 inserted nucleotides that is not present in the first genomic sequence.
[0127] In some embodiments, the first homology arm comprises 5 inserted nucleotides that is not present in the first genomic sequence.
[0128] In some embodiments, the second homology arm comprises between 1-20 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 1-3 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-4 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 4- 6 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 5-7 inserted nucleotides that are not present in the
[0129] 14
[0130] P89339 2070WO (01288) second genomic sequence. In some embodiments, the second homology arm comprises between 6- 8 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 7-9 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 8- 10 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 9-11 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 10-12 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 11-13 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 12-14 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 13-15 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 14-16 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 15-17 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 16-18 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 17-29 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 18-20 inserted nucleotides that are not present in the second genomic sequence.
[0131] In some embodiments, the second homology arm comprises between 1-10 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 1-3 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-4 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 4- 6 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 5-7 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 6- 8 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 7-9 inserted nucleotides that are not present in the
[0132] 15
[0133] P89339 2070WO (01288) second genomic sequence. In some embodiments, the second homology arm comprises between 8- 10 inserted nucleotides that are not present in the second genomic sequence.
[0134] In some embodiments, the second homology arm comprises between 1-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 4-5 inserted nucleotides that are not present in the second genomic sequence.
[0135] In some embodiments, the second homology arm comprises 1 inserted nucleotide that is not present in the second genomic sequence.
[0136] In some embodiments, the second homology arm comprises 2 inserted nucleotides that is not present in the second genomic sequence.
[0137] In some embodiments, the second homology arm comprises 3 inserted nucleotides that is not present in the second genomic sequence.
[0138] In some embodiments, the second homology arm comprises 4 inserted nucleotides that is not present in the second genomic sequence.
[0139] In some embodiments, the second homology arm comprises 5 inserted nucleotides that is not present in the second genomic sequence.
[0140] In some embodiments, the first genomic sequence and / or the second genomic sequence lacks at least one wild-type nucleotide.
[0141] In some embodiments, the inserted nucleotide is the wild-type nucleotide.
[0142] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an exon sequence that lacks at least one wild-type nucleotide.
[0143] In some embodiments, the exon sequence comprises a mutant codon that lacks the wild-type nucleotide.
[0144] In some embodiments, the inserted nucleotide is comprised by a codon that encodes a different amino acid than the mutant codon.
[0145] In some embodiments, the inserted nucleotide is the wild-type nucleotide that is comprised by a wild-type codon.
[0146] In some embodiments, the inserted nucleotide is comprised by a codon that encodes a wildtype amino acid.
[0147] In some embodiments, the inserted nucleotide is comprised by a stop codon.
[0148] 16
[0149] P89339 2070WO (01288) In some embodiments, the first genomic sequence and / or the second genomic sequence comprise a wild-type exon sequence.
[0150] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence and / or the second genomic sequence.
[0151] In some embodiments, the inserted nucleotide is comprised by a mutant codon.
[0152] In some embodiments, the inserted nucleotide is comprised by a stop codon.
[0153] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an intron sequence that lacks at least one wild-type nucleotide.
[0154] In some embodiments, the intron sequence comprises a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence that lacks the wild-type nucleotide.
[0155] In some embodiments, the inserted nucleotide comprises the wild-type nucleotide.
[0156] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0157] In some embodiments, the first genomic sequence and / or the second genomic sequence comprise a wild-type intron sequence.
[0158] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence or the second genomic sequence.
[0159] In some embodiments, the inserted nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0160] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a non-gene protein coding sequence that lacks at least one wild-type nucleotide.
[0161] In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0162] In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0163] In some embodiments, the inserted nucleotide comprises the wild-type nucleotide.
[0164] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a wild-type non-gene protein coding sequence.
[0165] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence and / or the second genomic sequence.
[0166] In some embodiments, the first homology arm and the second homology arm are approximately the same length. In some embodiments, the first homology arm and the second
[0167] 17
[0168] P89339 2070WO (01288) homology arm are between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0169] In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0170] In some embodiments, the first homology arm and the second homology arm are between about 400 to 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0171] In some embodiments, the first homology arm and the second homology arm are about 500 base pairs in length.
[0172] In some embodiments, the first homology arm and the second homology arm are different lengths. In some embodiments, the first homology arm is longer than the second homology arm. In some embodiments, the second homology arm is longer than the first homology arm.
[0173] In some embodiments, the first homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments,
[0174] 18
[0175] P89339 2070WO (01288) the first homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0176] In some embodiments, the first homology arm is between about 200 to 800 base pairs in length. In some embodiments, the first homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0177] In some embodiments, the first homology arm is between about 400 to 600 base pairs in length. In some embodiments, the first homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0178] In some embodiments, the first homology arm is about 500 base pairs in length.
[0179] In some embodiments, the second homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the second homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the second homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0180] In some embodiments, the second homology arm is between about 200 to 800 base pairs in length. In some embodiments, the second homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the second homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0181] In some embodiments, the second homology arm is between about 400 to 600 base pairs in length. In some embodiments, the second homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in
[0182] 19
[0183] P89339 2070WO (01288) length. In some embodiments, the second homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0184] In some embodiments, the second homology arm is about 500 base pairs in length.
[0185] In some embodiments, the first homology arm has at least 95% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 96% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 97% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 98% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 99% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has 100% sequence homology to the first genomic sequence.
[0186] In some embodiments, the second homology arm has at least 95% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 96% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 97% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 98% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 99% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has 100% sequence homology to the second genomic sequence.
[0187] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 1- 10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the first homology arm.
[0188] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 50- 100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275-325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the first homology arm.
[0189] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between
[0190] 20
[0191] P89339 2070WO (01288) 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the first homology arm.
[0192] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700-1900, or 1800-2000 base pairs from the 3' end of the first homology arm.
[0193] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 5' end of the second homology arm.
[0194] The second homology arm can comprise the at least one nucleotide modification at a position between 1-10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the second homology arm.
[0195] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 50-100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275- 325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the second homology arm.
[0196] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the second homology arm.
[0197] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700- 1900, or 1800-2000 base pairs from the 3' end of the second homology arm.
[0198] In some embodiments, the exogenous polynucleotide comprises a nucleic acid sequence encoding a nuclease. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 5' upstream of the first homology arm. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 3' downstream of the second homology arm.
[0199] 21
[0200] P89339 2070WO (01288) In some embodiments, the nuclease is capable of binding and cleaving the genome of the cell to generate the double -strand break.
[0201] In some embodiments, the exogenous polynucleotide comprises a promoter that is operably linked to the nucleic acid sequence encoding the nuclease.
[0202] In some embodiments, the nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, or a compact TALEN.
[0203] In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo.
[0204] In some embodiments, the exogenous polynucleotide is an mRNA, a single-stranded DNA, or a double-stranded DNA.
[0205] In some embodiments, the exogenous polynucleotide is comprised by a viral genome. In some embodiments, the exogenous polynucleotide is comprised by a delivery vehicle.
[0206] In some embodiments, the delivery vehicle is a recombinant virus and the exogenous polynucleotide is comprised by a viral genome.
[0207] In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).
[0208] In some embodiments, the delivery vehicle is a lipid nanoparticle.
[0209] In some embodiments, the exogenous polynucleotide is an mRNA, wherein the mRNA is comprised by the lipid nanoparticle.
[0210] In another aspect, the disclosure provides method for genetically modifying a cell, the method comprising introducing into a cell: a) an exogenous polynucleotide; and b) a nuclease or a gene encoding a nuclease, wherein the nuclease is expressed in the cell; wherein the nuclease generates a double-strand break in the genome of the cell, wherein the exogenous polynucleotide comprises, from 5' to 3', a first homology arm and a second homology arm, wherein the first homology arm has homology to a first genomic sequence that is 5' upstream to the double-strand break, wherein the second homology arm has homology to a second genomic sequence that is 3' downstream to the double-strand break, wherein: i) the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence; ii) the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence; or iii) the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence, and the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence; and wherein the
[0211] 22
[0212] P89339 2070WO (01288) nucleotide modification is introduced into the first genomic sequence and / or the second genomic sequence by homologous recombination of the exogenous polynucleotide.
[0213] In some embodiments, the first genomic sequence is adjacent to the double-strand break.
[0214] In some embodiments, the first genomic sequence is not adjacent to the double-strand break.
[0215] In some embodiments, the second genomic sequence is adjacent to the double-strand break.
[0216] In some embodiments, the second genomic sequence is not adjacent to the double-strand break.
[0217] In some embodiments, the double-strand break comprises a 3' overhang.
[0218] In some embodiments, the double-strand break comprises a four base pair 3' overhang.
[0219] In some embodiments, the double-strand break comprises a blunt end.
[0220] In some embodiments, the double-strand break comprises a 5' overhang.
[0221] In some embodiments, the exogenous polynucleotide comprises, from 5' to 3', the first homology arm, a heterologous nucleic acid sequence, and the second homology arm, wherein the heterologous nucleic acid sequence is inserted into the genome at the double-strand break by homologous recombination.
[0222] In some embodiments, the double-strand break is positioned within an exon of a gene, an intron of a gene, or a non-protein coding region of a gene.
[0223] In some embodiments, the nucleotide modification comprises a mismatched nucleotide that differs from a corresponding endogenous nucleotide of the first genomic sequence or the second genomic sequence, wherein the corresponding endogenous nucleotide is replaced with the mismatched nucleotide following homologous recombination of the exogenous polynucleotide.
[0224] In some embodiments, the corresponding endogenous nucleotide is an A, and the mismatched nucleotide is a T, C, or G. In some embodiments, the endogenous nucleotide is a T, and the mismatched nucleotide is an A, C, or G. In some embodiments, the corresponding endogenous nucleotide is a C, and the mismatched nucleotide is an A, T, or G. In some embodiments, the corresponding endogenous nucleotide is a G, and the mismatched nucleotide is an A, T, or C.
[0225] In some embodiments, the first homology arm comprises between 1-20 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched
[0226] 23
[0227] P89339 2070WO (01288) nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the first homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the first homology arm comprises between 11-13 mismatched nucleotides. In some embodiments, the first homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the first homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the first homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the first homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the first homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the first homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the first homology arm comprises between 18-20 mismatched nucleotides.
[0228] In some embodiments, the first homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides.
[0229] In some embodiments, the first homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-5 mismatched nucleotides.
[0230] In some embodiments, the first homology arm comprises 1 mismatched nucleotide. In some embodiments, the first homology arm comprises 2 mismatched nucleotides. In some embodiments, the first homology arm comprises 3 mismatched nucleotides. In some embodiments, the first
[0231] 24
[0232] P89339 2070WO (01288) homology arm comprises 4 mismatched nucleotide. In some embodiments, the first homology arm comprises 5 mismatched nucleotides.
[0233] In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the second homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the second homology arm comprises between 11- 13 mismatched nucleotides. In some embodiments, the second homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the second homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the second homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the second homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the second homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the second homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the second homology arm comprises between 18-20 mismatched nucleotides.
[0234] In some embodiments, the second homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides.
[0235] In some embodiments, the second homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-5 mismatched
[0236] 25
[0237] P89339 2070WO (01288) nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-5 mismatched nucleotides.
[0238] In some embodiments, the second homology arm comprises 1 mismatched nucleotide. In some embodiments, the second homology arm comprises 2 mismatched nucleotides. In some embodiments, the second homology arm comprises 3 mismatched nucleotides. In some embodiments, the second homology arm comprises 4 mismatched nucleotides. In some embodiments, the second homology arm comprises 5 mismatched nucleotides. In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotides. In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0239] In some embodiments, the corresponding endogenous nucleotide is comprised by an exon.
[0240] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0241] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the mutant codon.
[0242] In some embodiments, the mismatched nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0243] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a wild-type amino acid.
[0244] In some embodiments, the mismatched nucleotide is comprised by a stop codon.
[0245] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0246] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the wild-type codon.
[0247] In some embodiments, the mismatched nucleotide is a mutant nucleotide comprised by a mutant codon.
[0248] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a wild-type amino acid.
[0249] In some embodiments, the mismatched nucleotide is comprised by a stop codon.
[0250] In some embodiments, the corresponding endogenous nucleotide is comprised by an intron. In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0251] In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0252] 26
[0253] P89339 2070WO (01288) In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0254] In some embodiments, the mismatched nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0255] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide.
[0256] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0257] In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0258] In some embodiments, the mismatched nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0259] In some embodiments, the corresponding endogenous nucleotide is comprised by a nongene protein coding sequence.
[0260] In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0261] In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0262] In some embodiments, the corresponding endogenous nucleotide is comprised by a mutant non-gene protein coding sequence.
[0263] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0264] In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0265] In some embodiments, the mismatched nucleotide differs from a wild-type nucleotide.
[0266] In some embodiments, the corresponding endogenous nucleotide is comprised by a wildtype non-gene protein coding sequence.
[0267] In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0268] In some embodiments, the nucleotide modification comprises a nucleotide deletion, wherein the first homology arm and / or the second homology arm lacks at least one corresponding endogenous nucleotide that is present in the first genomic sequence or the second genomic sequence, and wherein corresponding endogenous nucleotide is deleted from the genome following homologous recombination of the exogenous polynucleotide.
[0269] In some embodiments, the first homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some
[0270] 27
[0271] P89339 2070WO (01288) embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 17-29 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0272] In some embodiments, the first homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks
[0273] 28
[0274] P89339 2070WO (01288) between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0275] In some embodiments, the first homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0276] In some embodiments, the first homology arm lacks 1 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0277] In some embodiments, the first homology arm lacks 2 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0278] In some embodiments, the first homology arm lacks 3 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0279] In some embodiments, the first homology arm lacks 4 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0280] In some embodiments, the first homology arm lacks 5 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0281] In some embodiments, the second homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the
[0282] 29
[0283] P89339 2070WO (01288) second genomic sequence. In some embodiments, the second homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 17-29 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0284] In some embodiments, the second homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 5-7
[0285] 30
[0286] P89339 2070WO (01288) corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0287] In some embodiments, the second homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0288] In some embodiments, the second homology arm lacks 1 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0289] In some embodiments, the second homology arm lacks 2 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0290] In some embodiments, the second homology arm lacks 3 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0291] In some embodiments, the second homology arm lacks 4 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0292] In some embodiments, the second homology arm lacks 5 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0293] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0294] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0295] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an exon sequence.
[0296] In some embodiments, the exon sequence is a mutant exon sequence.
[0297] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0298] 31
[0299] P89339 2070WO (01288) In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a codon that encodes a different amino acid than the mutant codon.
[0300] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a wild-type codon.
[0301] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a codon that encodes a wild-type amino acid.
[0302] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a stop codon.
[0303] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0304] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a mutant codon that encodes a different amino acid than the wild-type codon.
[0305] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a codon that encodes the same amino acid as the wild-type codon.
[0306] In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a stop codon.
[0307] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an intron sequence.
[0308] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0309] In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0310] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0311] In some embodiments, the first homology arm and / or the second homology arm comprises a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0312] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide.
[0313] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0314] 32
[0315] P89339 2070WO (01288) In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide.
[0316] In some embodiments, the first homology arm and / or the second homology arm comprises a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0317] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an non-gene coding sequence.
[0318] In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0319] In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0320] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0321] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0322] In some embodiments, the non-gene coding sequence is a wild-type non-gene coding sequence.
[0323] In some embodiments, the first homology arm and / or the second homology arm does not comprise at least one corresponding wild-type nucleotide.
[0324] In some embodiments, the nucleotide modification comprises a nucleotide insertion, wherein the first homology arm and / or the second homology arm comprises at least one inserted nucleotide that is not present in the first genomic sequence or the second genomic sequence, and wherein the inserted nucleotide is introduced into the genome following homologous recombination of the exogenous polynucleotide.
[0325] In some embodiments, the first homology arm comprises between 1-20 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 1-3 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-4 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-6 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 5-7 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 6-8 inserted nucleotides that are not
[0326] 33
[0327] P89339 2070WO (01288) present in the first genomic sequence. In some embodiments, the first homology arm comprises between 7-9 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 8-10 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 9-11 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 10-12 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 11-13 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 12-14 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 13-15 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 14-16 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 15-17 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 16-18 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 17-29 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 18-20 inserted nucleotides that are not present in the first genomic sequence.
[0328] In some embodiments, the first homology arm comprises between 1-10 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 1-3 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-4 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-6 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 5-7 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 6-8 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 7-9 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 8-10 inserted nucleotides that are not present in the first genomic sequence.
[0329] 34
[0330] P89339 2070WO (01288) In some embodiments, the first homology arm comprises between 1-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-5 inserted nucleotides that are not present in the first genomic sequence.
[0331] In some embodiments, the first homology arm comprises 1 inserted nucleotide that is not present in the first genomic sequence.
[0332] In some embodiments, the first homology arm comprises 2 inserted nucleotides that are not present in the first genomic sequence.
[0333] In some embodiments, the first homology arm comprises 3 inserted nucleotides that are not present in the first genomic sequence.
[0334] In some embodiments, the first homology arm comprises 4 inserted nucleotides that are not present in the first genomic sequence.
[0335] In some embodiments, the first homology arm comprises 5 inserted nucleotides that are not present in the first genomic sequence.
[0336] In some embodiments, the second homology arm comprises between 1-20 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 1-3 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-4 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 4- 6 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 5-7 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 6- 8 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 7-9 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 8- 10 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 9-11 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 10-12 inserted nucleotides that are not present in the second genomic sequence. In some
[0337] 35
[0338] P89339 2070WO (01288) embodiments, the second homology arm comprises between 11-13 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 12-14 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 13-15 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 14-16 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 15-17 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 16-18 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 17-29 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 18-20 inserted nucleotides that are not present in the second genomic sequence.
[0339] In some embodiments, the second homology arm comprises between 1-10 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 1-3 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-4 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 4- 6 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 5-7 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 6- 8 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 7-9 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 8- 10 inserted nucleotides that are not present in the second genomic sequence.
[0340] In some embodiments, the second homology arm comprises between 1-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 2-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments, the second homology arm comprises between 3-5 inserted nucleotides that are not present in the second genomic sequence. In some embodiments,
[0341] 36
[0342] P89339 2070WO (01288) the second homology arm comprises between 4-5 inserted nucleotides that are not present in the second genomic sequence.
[0343] In some embodiments, the second homology arm comprises 1 inserted nucleotide that is not present in the second genomic sequence.
[0344] In some embodiments, the second homology arm comprises 2 inserted nucleotides that are not present in the second genomic sequence.
[0345] In some embodiments, the second homology arm comprises 3 inserted nucleotides that are not present in the second genomic sequence.
[0346] In some embodiments, the second homology arm comprises 4 inserted nucleotides that are not present in the second genomic sequence.
[0347] In some embodiments, the second homology arm comprises 5 inserted nucleotides that are not present in the second genomic sequence.
[0348] In some embodiments, the first genomic sequence and / or the second genomic sequence lacks at least one wild-type nucleotide.
[0349] In some embodiments, the inserted nucleotide is the wild-type nucleotide.
[0350] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an exon sequence that lacks at least one wild-type nucleotide.
[0351] In some embodiments, the exon sequence comprises a mutant codon that lacks the wild-type nucleotide.
[0352] In some embodiments, the inserted nucleotide is comprised by a codon that encodes a different amino acid than the mutant codon.
[0353] In some embodiments, the inserted nucleotide is the wild-type nucleotide that is comprised by a wild-type codon.
[0354] In some embodiments, the inserted nucleotide is comprised by a codon that encodes a wildtype amino acid.
[0355] In some embodiments, the inserted nucleotide is comprised by a stop codon.
[0356] In some embodiments, the first genomic sequence and / or the second genomic sequence comprise a wild-type exon sequence.
[0357] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence and / or the second genomic sequence.
[0358] In some embodiments, the inserted nucleotide is comprised by a mutant codon.
[0359] In some embodiments, the inserted nucleotide is comprised by a stop codon.
[0360] 37
[0361] P89339 2070WO (01288) In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an intron sequence that lacks at least one wild-type nucleotide.
[0362] In some embodiments, the intron sequence comprises a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence that lacks the wild-type nucleotide.
[0363] In some embodiments, the inserted nucleotide comprises the wild-type nucleotide.
[0364] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0365] In some embodiments, the first genomic sequence and / or the second genomic sequence comprise a wild-type intron sequence.
[0366] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence or the second genomic sequence.
[0367] In some embodiments, the inserted nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0368] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a non-gene protein coding sequence that lacks at least one wild-type nucleotide.
[0369] In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0370] In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0371] In some embodiments, the inserted nucleotide comprises the wild-type nucleotide.
[0372] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a wild-type non-gene protein coding sequence.
[0373] In some embodiments, the inserted nucleotide is not comprised by the first genomic sequence and / or the second genomic sequence.
[0374] In some embodiments, the first homology arm and the second homology arm are approximately the same length.
[0375] In some embodiments, the first homology arm and the second homology arm are between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900,
[0376] 38
[0377] P89339 2070WO (01288) or at least about 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0378] In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0379] In some embodiments, the first homology arm and the second homology arm are between about 400 to 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0380] In some embodiments, the first homology arm and the second homology arm are about 500 base pairs in length.
[0381] In some embodiments, the first homology arm and the second homology arm are different lengths. In some embodiments, the first homology arm is longer than the second homology arm. In some embodiments, the second homology arm is longer than the first homology arm.
[0382] In some embodiments, the first homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the first homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0383] 39
[0384] P89339 2070WO (01288) In some embodiments, the first homology arm is between about 200 to 800 base pairs in length. In some embodiments, the first homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0385] In some embodiments, the first homology arm is between about 400 to 600 base pairs in length. In some embodiments, the first homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0386] In some embodiments, the first homology arm is about 500 base pairs in length.
[0387] In some embodiments, the second homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the second homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the second homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0388] In some embodiments, the second homology arm is between about 200 to 800 base pairs in length. In some embodiments, the second homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the second homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0389] In some embodiments, the second homology arm is between about 400 to 600 base pairs in length. In some embodiments, the second homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the second homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0390] In some embodiments, the second homology arm is about 500 base pairs in length.
[0391] 40
[0392] P89339 2070WO (01288) In some embodiments, the first homology arm has at least 95% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 96% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 97% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 98% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 99% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has 100% sequence homology to the first genomic sequence.
[0393] In some embodiments, the second homology arm has at least 95% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 96% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 97% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 98% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 99% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has 100% sequence homology to the second genomic sequence.
[0394] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 1- 10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the first homology arm.
[0395] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 50- 100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275-325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the first homology arm.
[0396] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the first homology arm.
[0397] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the first homology arm.
[0398] 41
[0399] P89339 2070WO (01288) The first homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700-1900, or 1800-2000 base pairs from the 3' end of the first homology arm.
[0400] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 5' end of the second homology arm.
[0401] The second homology arm can comprise the at least one nucleotide modification at a position between 1-10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the second homology arm.
[0402] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 50-100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275- 325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the second homology arm.
[0403] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the second homology arm.
[0404] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700- 1900, or 1800-2000 base pairs from the 3' end of the second homology arm.
[0405] In some embodiments, the nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, or a compact TALEN.
[0406] In some embodiments, the exogenous polynucleotide comprises a nucleic acid sequence encoding the nuclease.
[0407] In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 5' upstream of the first homology arm. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 3' downstream of the second homology arm.
[0408] In some embodiments, the exogenous polynucleotide comprises a promoter that is operably linked to the nucleic acid sequence encoding the nuclease.
[0409] 42
[0410] P89339 2070WO (01288) In some embodiments, the exogenous polynucleotide is introduced into the cell by a recombinant virus, wherein the recombinant virus comprises the exogenous polynucleotide in its viral genome, and wherein the recombinant virus is contacted with the cell.
[0411] In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).
[0412] In some embodiments, the exogenous polynucleotide is introduced into the cell by non- viral delivery.
[0413] In some embodiments, the exogenous polynucleotide is comprised by lipid nanoparticles that are contacted with the cell.
[0414] In some embodiments, the exogenous polynucleotide is an mRNA.
[0415] In some embodiments, the gene encoding the nuclease is introduced into the cell by a recombinant virus, wherein the recombinant virus comprises the gene encoding the nuclease in its viral genome, and wherein the recombinant virus is contacted with the cell.
[0416] In some embodiments, the recombinant virus is a recombinant AAV that is contacted with the cell.
[0417] In some embodiments, the gene encoding the nuclease is introduced into the cell by non- viral delivery.
[0418] In some embodiments, the gene encoding the nuclease is comprised by lipid nanoparticles that are contacted with the cell.
[0419] In some embodiments, the gene encoding the nuclease is an mRNA.
[0420] In some embodiments, the exogenous polynucleotide is introduced into the cell by a recombinant virus that comprises the exogenous polynucleotide in its viral genome, and wherein the gene encoding the nuclease is comprised by lipid nanoparticles, wherein the recombinant virus and the lipid nanoparticles are contacted with the cell.
[0421] In some embodiments, the recombinant virus is a recombinant AAV.
[0422] In some embodiments, the gene encoding the nuclease is an mRNA.
[0423] In some embodiments, the exogenous polynucleotide is introduced into the cell by a first recombinant virus that comprises the exogenous polynucleotide in its viral genome, and wherein the gene encoding the nuclease is introduced into the cell by a second recombinant virus that comprises the gene encoding the nuclease in its viral genome, wherein the first recombinant nuclease and the second recombinant nucleases are contacted with the cell.
[0424] In some embodiments, the first recombinant virus and the second recombinant virus are each a recombinant AAV.
[0425] 43
[0426] P89339 2070WO (01288) In some embodiments, the exogenous polynucleotide is comprised by a first population of lipid nanoparticles, and wherein the gene encoding the nuclease is comprised by a second population of lipid nanoparticles, wherein the first population of lipid nanoparticles and the second population of lipid nanoparticles are contacted with the cell.
[0427] In some embodiments, the exogenous polynucleotide is a single-stranded DNA or a doublestranded DNA.
[0428] In some embodiments, the gene encoding the nuclease is an mRNA.
[0429] In some embodiments, the exogenous polynucleotide is comprised by a first population of non-viral particles, and wherein the gene encoding the nuclease is comprised by a second population of non-viral particles, wherein the first population of non-viral particles and the second population of non-viral particles are contacted with the cell.
[0430] In some embodiments, the exogenous polynucleotide and the gene encoding the nuclease are introduced into the cell by a recombinant virus that comprises the exogenous polynucleotide and the gene encoding the nuclease in its viral genome, wherein the recombinant virus is contacted with the cell.
[0431] In some embodiments, the recombinant virus is a recombinant AAV.
[0432] In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo.
[0433] In some embodiments, the method is a method for genetically modifying a target cell in a subject, wherein the exogenous polynucleotide and the nuclease, or the gene encoding the nuclease, are delivered to the target cell in the subject, and wherein the nucleotide modification is introduced into the first genomic sequence and / or the second genomic sequence by homologous recombination of the exogenous polynucleotide.
[0434] In some embodiments, the method is a method for treating a subject in need thereof, wherein the subject is administered an effective amount of a) the exogenous polynucleotide; and b) the nuclease, or the gene encoding the nuclease; wherein the exogenous polynucleotide and the nuclease, or the gene encoding the nuclease, are delivered to a target cell in the subject, and wherein the nucleotide modification is introduced into the first genomic sequence and / or the second genomic sequence by homologous recombination of the exogenous polynucleotide.
[0435] BRIEF DESCRIPTION OF THE FIGURES
[0436] Figure 1A provides an exemplary graphic depiction of a nuclease -initiated HDR base editing in a gene, wherein the DSB is generated within an exon denoted as exon 1, and the first and
[0437] 44
[0438] P89339 2070WO (01288) second genomic sequences are within the same cleaved exon. In this figure, there are two mutant nucleotides or codons relative to a wild type sequence (shown as gray bars) that are corrected by introducing the wild type nucleotide or codon in the second homology arm (shown as black bars). Following homologous recombination these wild type nucleotides or codons are introduced into the second genomic sequence, thus restoring a wild type protein.
[0439] Figure IB provides an exemplary graphic depiction of a nuclease-initiated HDR base editing in a gene, wherein the DSB is generated within an exon denoted as exon 1, and the first and second genomic sequences are within the same cleaved exon. In this figure, the endogenous exon 1 has a deletion causing the overall gene coding sequence to be reduced. The repair construct has a nucleotide insertion present within the second homology arm that has homology to the second genomic sequence. Following homologous recombination this nucleotide insertion is introduced into the second genomic sequence. In this example, the wild type coding sequence of exon 1 can be restored following the nucleotide insertion.
[0440] Figure 1C provides an exemplary graphic depiction of a nuclease-initiated HDR base editing in a gene, wherein the DSB is generated within an exon denoted as exon 1, and the first and second genomic sequences are within the same cleaved exon. In this figure, the endogenous exon 1 has a nucleotide expansion causing the overall gene coding sequence to be increased. The repair construct has a nucleotide deletion present within the second homology arm that has homology to the second genomic sequence. Following homologous recombination this nucleotide deletion is introduced into the second genomic sequence. In this example, the wild type coding sequence of exon 1 can be restored following the nucleotide deletion.
[0441] Figure ID provides an exemplary graphic depiction of a nuclease-initiated HDR base editing in a gene, wherein the DSB is generated at the exon / intron boundary where the first genomic sequence is present within exon 1 and the second genomic sequence is present with the intron. In this figure, the repair construct has a heterologous sequence at the 5’ portion of the second homology arm that has homology to the second genomic sequence. Following homologous recombination this heterologous sequence is introduced to the 3 ’ portion of exon 1.
[0442] Figure 2 provides a graphic depiction of seven different repair constructs that were made and tested in Example 1. In this figure, an engineered meganuclease TRC 1-2L.2307 is used to generate a double strand break within the TRAC locus of a T cell thereby disrupting expression of this gene. Four of the repair constructs contain a 5’ homology arm (LHA) that have either 1, 3, 5, or 10 base pair mismatches compared to the endogenous gene sequence that otherwise has homology to the corresponding endogenous genomic sequence. Three of the repair constructs contain a 3’
[0443] 45
[0444] P89339 2070WO (01288) homology arm (RHA) that have either 1, 3, or 5 base pair mismatches compared to the endogenous gene sequence that otherwise has homology to the 3’ homology to the corresponding endogenous genomic sequence. Each of these repair constructs contain a GFP coding sequence driven off of a Jet promoter for detection of homologous recombination of these repair constructs.
[0445] Figure 3 provides a bar graph denoting the percentage of repair construct knock-in (KI) of cells that have had the TRAC locus knocked-out (KO).
[0446] Figure 4A provides a chart showing the percentage of mismatched base pair incorporation into the TRAC locus following homologous recombination of the repair constructs of Example 1 (see Figure 3).
[0447] Figure 4B provides a bar graph showing the percentage alignment of mismatched base pair incorporation into the TRAC locus following homologous recombination of the repair constructs of Example 1 (see Figure 3).
[0448] Figure 5A provides a graphic depiction of repair construct that was made and tested in Example 2. In this figure, an engineered meganuclease having specificity for a site within the TGBRII locus was used to generate a double strand break within the 5’UTR of the TGFBRII locus of a T cell. A repair construct was made that incorporates a single base pair mismatch that mutates the ATG start codon of the TGFBRII gene to GTG thereby abrogating expression of the gene.
[0449] Figure 5B provides a graphic depiction of a repair construct that was made and tested in Example 2. In this figure, an engineered meganuclease having specificity for a site within the TGBRII locus was used to generate a double strand break within the 5’UTR of the TGFBRII locus of a T cell. A repair construct was made that deletes the 3-base pair ATG start codon of the TGFBRII gene thereby abrogating expression of the gene.
[0450] Figure 5C provides a graphic depiction of a repair construct that was made and tested in Example 2. In this figure, an engineered meganuclease having specificity for a site within the TGBRII locus was used to generate a double strand break within the 5’UTR of the TGFBRII locus of a T cell. A repair construct was made that inserts a 3-base pair TGA start codon immediately after the ATG start codon of the TGFBRII gene thereby abrogating expression of the gene immediately after the first methionine codon.
[0451] Figure 6A provides a flow cytometry plot highlighting expression of the TGFBRII gene following introduction of the TGFBRII meganuclease of Example 2 alone. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0452] 46
[0453] P89339 2070WO (01288) Figure 6B provides a flow cytometry plot showing the control levels of TGFBRII gene expression following introduction of the TGFBRII GTG mutation repair template of Example 2 alone. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0454] Figure 6C provides a flow cytometry plot showing the levels of TGFBRII gene expression following introduction of the TGBRII meganuclease and the TGFBRII GTG mutation repair template of Example 2. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0455] Figure 6D provides a flow cytometry plot showing the levels of TGFBRII gene expression following introduction of the TGBRII meganuclease and a wild-type coding TGFBRII repair template where no changes to the coding sequence have been made. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0456] Figure 7A provides a flow cytometry plot showing the control levels of TGFBRII gene expression following introduction of the TGFBRII ATG deletion repair template of Example 2 alone. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0457] Figure 7B provides a flow cytometry plot showing the levels of TGFBRII gene expression following introduction of the TGBRII meganuclease and the TGFBRII ATG deletion repair template of Example 2. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0458] Figure 7C provides a flow cytometry plot showing the control levels of TGFBRII gene expression following introduction of the TGFBRII TGA stop codon insertion repair template of Example 2 alone. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0459] Figure 7D provides a flow cytometry plot showing the levels of TGFBRII gene expression following introduction of the TGBRII meganuclease and the TGFBRII TGA stop codon insertion repair template of Example 2. The percentage of cells not expressing TGFBRII is indicated above the line in the chart.
[0460] Figure 8 provides a graphic depiction of five different repair constructs that were made and tested in Example 3 to demonstrate the ability of HDR base editing to be able to change every base in a gene. In this figure, an engineered meganuclease having specificity for a site within the 5’UTR of the TGBRII locus was used to generate a double strand break within the TGFBRII locus of a T cell. The first repair construct is the WT repair construct where no change to the sequence is made. The remaining four repair constructs contain a 3 ’ homology arm that have degenerate nucleotides at position -1 (C) relative to the ATG start codon, position 1 of the ATG (A), position 2 of the ATG
[0461] 47
[0462] P89339 2070WO (01288) (T), or position 3 of the ATG (G) of the TGFBRII gene. This degenerate nucleotide allows for any substitution of A, T, G, or C to occur at the indicated position. In the case of position 1, 2, and 3 a substitution to the non-wild type sequence will result in knock out of the TGFBRII gene. A substitution at position -1 will have little or no effect on TGFBRII gene expression since it does not change the start codon sequence.
[0463] Figure 9 provides charts showing the roughly equal percentage of degenerate nucleotide incorporation in the AAV6 plasmid cloning (top graphs) and in the packaged AAV6 containing the repair constructs (bottom graphs).
[0464] Figure 10A provides flow cytometry plots highlighting expression of the TGFBRII gene following introduction of the TGFBRII meganuclease and a wild type repair construct (upper left panel) or the repair constructs changing position 3 of the ATG (G)>N (upper right panel), position 1 of the ATG (A)>N (lower left panel), or position 2 of the ATG (T)>N (lower right panel). The percentage of TGFBRII knock out (KO) is shown in each panel.
[0465] Figure 10B in the upper panel, provides pie charts showing the percentage of nucleotide modification at each position within the TGFBRII locus relative to the ATG start codon. From left to right, the first pie chart shows at position -1 bases that were changed from the wild type (C) relative to the ATG start codon. The bar graph below this pie chart indicates out of the bases that were changed, which proportion were changed from C>A, C>T, or C>G. The second pie chart shows at position 1 bases that were changed from the wild type (A) relative to the ATG start codon. The bar graph below this second pie chart indicates out of the bases that were changed which proportion were changed from A>T, A>G, or A>C. The third pie chart shows at position 2 bases that were changed from the wild type (T) relative to the ATG start codon. The bar graph below this third pie chart indicates out of the bases that were changed which proportion were changed from T>A, T>G, or T>C. The fourth pie chart shows at position 3 bases that were changed from the wild type (G) relative to the ATG start codon. The bar graph below this third pie chart indicates out of the bases that were changed which proportion were changed from G>A, G>T, or G>C.
[0466] BRIEF DESCRIPTION OF THE SEQUENCES
[0467] SEQ ID NO: 1 sets forth the nucleic acid sequence of a wild type repair construct of Example 1.
[0468] SEQ ID NO: 2 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 5' homology arm has a 1 base pair mismatch.
[0469] 48
[0470] P89339 2070WO (01288) SEQ ID NO: 3 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 5' homology arm has a 3 base pair mismatch.
[0471] SEQ ID NO: 4 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 5' homology arm has a 5 base pair mismatch.
[0472] SEQ ID NO: 5 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 5' homology arm has a 10 base pair mismatch.
[0473] SEQ ID NO: 6 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 3' homology arm has a 1 base pair mismatch.
[0474] SEQ ID NO: 7 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 3' homology arm has a 3 base pair mismatch.
[0475] SEQ ID NO: 8 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 3' homology arm has a 5 base pair mismatch.
[0476] SEQ ID NO: 9 sets forth the nucleic acid sequence of a repair construct of Example 1, wherein the 3' homology arm has a 10 base pair mismatch.
[0477] DETAILED DESCRIPTION OF THE INVENTION
[0478] 1 References and Definitions
[0479] The patent and scientific literature referred to herein establishes knowledge that is available to those of skill in the art. The issued US patents, allowed applications, published foreign applications, and references, including GenBank database sequences, which are cited herein are hereby incorporated by reference to the same extent as if each was specifically and individually indicated to be incorporated by reference.
[0480] The present disclosure can be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art. For example, features illustrated with respect to one embodiment can be incorporated into other embodiments, and features illustrated with respect to a particular embodiment can be deleted from that embodiment. In addition, numerous variations and additions to the embodiments suggested herein will be apparent to those skilled in the art in light of the instant disclosure, which do not depart from the instant disclosure.
[0481] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure
[0482] 49
[0483] P89339 2070WO (01288) belongs. The terminology used in the description of the disclosure herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure.
[0484] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference herein in their entirety.
[0485] As used herein, “a,” “an,” or “the” can mean one or more than one. For example, “a” cell can mean a single cell or a multiplicity of cells.
[0486] As used herein, unless specifically indicated otherwise, the word “or” is used in the inclusive sense of “and / or” and not the exclusive sense of “either / or.”
[0487] As used herein, the term “recombinant” or “engineered” with respect to a protein means having an altered amino acid sequence as a result of the application of genetic engineering techniques to nucleic acids that encode the protein and cells or organisms that express the protein. With respect to a nucleic acid, the term “recombinant” or “engineered” means having an altered nucleic acid sequence as a result of the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning technologies; transfection, transformation, and other gene transfer technologies; homologous recombination; site- directed mutagenesis; and gene fusion. In accordance with this definition, a protein having an amino acid sequence identical to a naturally-occurring protein but produced by cloning and expression in a heterologous host, is not considered recombinant or engineered.
[0488] As used herein, the term “modified nucleotides” refers to one or more nucleotides present in a nucleic acid sequence (e.g., a first and / or second homology arm) that differ from one or more corresponding, endogenous nucleotides in a genomic nucleic acid sequence. One or more modified nucleotides may result in one or more modified codons.
[0489] As used herein, the term “modified codons” or “altered codons,” used interchangeably, refers to one or more codons present in a nucleic acid sequence (e.g., a first and / or second homology arm) that differ from one or more corresponding endogenous codons in a genomic nucleic acid sequence. In some cases, the one or more modified / altered codons in the nucleic acid sequence encode the same amino acid as the one or more corresponding endogenous codons. In other cases, the modified / altered codons encode a different amino acid than the one or more corresponding endogenous codons.
[0490] As used herein, the term “nucleotide modification” refers to one or more nucleotide that differ between the polynucleotide sequence of a homology arm and the corresponding, genomic sequence to which the homology arm sequence has homology. The nucleotide modification is determined by comparing the genomic sequence present within a genome following integration of a
[0491] 50
[0492] P89339 2070WO (01288) homology arm sequence into the genome through homologous recombination with an original, corresponding genomic sequence prior to integration of the homology arm into the cell genome. Exemplary, non-limiting, nucleotide modifications include mismatched nucleotides, nucleotide deletions, and nucleotide insertions.
[0493] As used herein, the term “coding sequence” refers to a nucleotide sequence of a gene, which is comprised of codons, that determines the sequence of amino acids in a protein.
[0494] As used herein, the term “codon” refers to a sequence of three nucleotides (a trinucleotide) that forms a unit of genomic information encoding a particular amino acid or signaling the termination of protein synthesis (stop codon). A codon can be altered by removing one or more than one nucleotide from a nucleotide sequence. There are 64 different codons: 61 specify amino acids and 3 are used as stop codons. The coding sequence begins with a start codon and ends with a stop codon. The coding sequence may comprise any number of codons between the start and stop codon. The start codon is the first codon of a messenger RNA (mRNA) transcript translated by a ribosome. Start and stop codon sequences are known. The most commonly understood start codon is the three -nucleotide sequence of AUG (i.e., encoded by the DNA sequence of ATG). Alternative start codons are less common, but do exist and would be understood to a person having skill in the art. Within a coding sequence, a new codon begins every three nucleotides, starting with the first nucleotide of the start codon. For example, the start codon (i.e., the first codon) comprises nucleotide residues 1, 2, and 3 of the coding sequence; the second codon comprises nucleotide residues 4, 5, and 6 of the coding sequence; and the third codon comprises nucleotide residues 7, 8, and 9.
[0495] As used herein, the term “wild-type codon” refers to a codon present in a wild-type gene or that is present in a polynucleotide sequence that corresponds to the wild-type codon found in a wild-type gene. The wild-type codon encodes a wild-type amino acid leading to expression of a wild-type protein.
[0496] As used herein, the term “mutant codon” refers to a codon present in a mutant gene that differs in sequence from the corresponding wild-type codon found in a wild-type gene. In some cases, the mutant codon can encode a wild-type amino acid leading to expression of a wild-type protein. In other cases, the mutant codon can encode a mutant amino acid that differs from the corresponding wild-type amino acid. A mutant gene can comprise, for example, one or more mutant codons. A mutant gene comprising a mutant codon can, in some cases, not express a protein, express a mutant protein that is not full length, and / or express a mutant protein is not functional relative to the wild-type protein.
[0497] 51
[0498] P89339 2070WO (01288) A “codon that encodes a different amino acid than a mutant codon” refers to a trinucleotide sequence that encodes an amino acid that differs from an amino acid encoded by a mutant codon. A codon can differ from a mutant codon at any nucleotide residue. In some embodiments, a codon that encodes a different amino acid than a mutant codon is a wild-type codon. In some instances, a codon that encodes a different amino acid than a mutant codon can be a second mutant codon comprising a different trinucleotide sequence than a first mutant codon.
[0499] A “codon that encodes a different amino acid than a wild-type codon” refers to a trinucleotide sequence that encodes an amino acid that differs from the amino acid encoded by a wild-type codon. A codon can differ from a wild-type codon at any nucleotide residue. In some embodiments, a codon that encodes a different amino acid than a wild-type codon is a mutant codon.
[0500] A “codon that encodes a wild-type amino acid” refers to trinucleotide sequence that encodes a wild-type amino acid.
[0501] As used herein, the term “wild-type” refers to the allele (i.e., polynucleotide or polynucleotide sequence) within an allele population of the same type of gene that is associated with a non-diseased state. In some embodiments, a wild-type allele is incorporated into the genome to correct a diseased or disorder. In such instances, a wild-type allele can replace a mutant allele. The term “wild-type” can also refer to the epigenetic state (e.g., methylation or acetylation state) of a polynucleotide or allele within a gene locus. It is understood that the epigenetic state of an allele or entire gene locus influences gene expression and cell state, and abnormal epigenetic patterns have been linked to many diseases, such as neurological disorders and cardiovascular disease. As such, a wild-type epigenetic state refers to an epigenetic state of a polynucleotide within a gene locus that is associated with a non-diseased state. The term “wild-type” can also refer to a cell, an organism, and / or a subject which possesses a wild-type allele of a particular gene, or a cell, an organism, and / or a subject used for comparative purposes.
[0502] As used herein, the term “mutant” refers to an allele having one or more mutations and / or substitutions relative to a corresponding wild-type allele. It would be understood that a mutant allele may be associated with a disease and / or disorder. As such, the term mutant can refer to an allele that is associated with a disease and / or a disorder in a subject (i.e., a diseased state). The term mutant can also refer to an allele that differs from the wild-type allele and is not associated with a disease and / or a disorder in a subject (i.e., a diseased state). In some embodiments, a mutant allele is related and / or associated with a diseased state, and a wild-type allele is related and / or associated with a non-diseased state. In some embodiments, a mutant allele is not related and / or associated
[0503] 52
[0504] P89339 2070WO (01288) with a diseased state, and a wild-type allele is related and / or associated with a non-diseased state. A mutant polynucleotide may have an epigenetic state (e.g., methylation or acetylation state) that differs from a wild-type polynucleotide. In some embodiments, a mutant allele is replaced with a wild-type allele to correct a disease and / or disorder in a subject. The terms “wild-type” and “mutant” can refer to any allele (polynucleotide) or polypeptide disclosed herein, including for instance a nucleotide, codon, exon sequence, intron sequence, such as splice donor sequence, splice acceptor sequence, splice enhancer sequence, non-gene coding sequence, and amino acid.
[0505] As used herein, the term “lacks at least one wild-type nucleotide” refers to a polynucleotide sequence (i.e., an allele or entire gene locus) that does not comprise at least one wild-type nucleotide. A locus that lacks at least one wild-type nucleotide is a mutant locus.
[0506] As used herein, the term “original” refers to an allele or gene locus present within a genomic sequence on the genome prior to homologous recombination and integration of homology arms with a corresponding endogenous genomic sequence. Any locus reference to herein can be referred to in its original state within a genomic sequence. An original locus can be replaced by a replacement locus through homologous recombination the homology arms described herein comprising a replacement locus. A nucleotide, codon, and / or amino acid can be referred to in its original state prior to recombination of the homologous arms with. For instance, an original nucleotide, codon, or amino acid within a genomic sequence can be replaced with a replacement nucleotide, codon, or amino acid, respectively. In some instances, the original locus is a wild-type locus. In some embodiments, the original locus is a mutant locus.
[0507] As used herein, the term “replacement locus” or “replacement allele” refers to a locus or allele present within a genomic sequence on the genome following integration of homology arms with their corresponding endogenous homologous genomic sequence. Following integration of homology arms, a replacement locus or allele present in a first and / or second homology is incorporated and present on the genome. A replacement locus allele replaces an original locus or allele that was present on the genome prior to homologous recombination. A replacement nucleic acid sequence may be identical to the original nucleic acid sequence; however, the replacement nucleic acid sequence may comprise one or more modified nucleotides and / or modified codons as compared to the original nucleic acid sequence. In some embodiments, the replacement locus is a wild-type locus. In some embodiments, the replacement locus is a mutant locus.
[0508] As used herein, the term “locus” or “gene locus” refers to a specific fixed position within a genomic sequence. The genomic sequence can be a portion or region of a genome of a cell. A locus
[0509] 53
[0510] P89339 2070WO (01288) can be comprised within a gene or a non-gene sequence of a genome. A gene locus comprises two alleles that may be the same or different, which comprises at least one nucleotide.
[0511] As used herein, the term “mutant protein” refers to a protein having one or more amino acids present in the polypeptide that differ from the corresponding, wild-type amino acid. That is, the mutant protein may comprise one or more “mutant amino acids.” As used herein, the term “mutant amino acid” refers to an amino acid present in a polypeptide that differs from the corresponding wild-type polypeptide. A polypeptide sequence (i.e., a protein) having a mutant amino acid is considered a variant of the wild-type polypeptide sequence.
[0512] As used herein, the term “heterologous nucleic acid sequence” refers to a nucleic acid sequence that is designed to be inserted into the genome of a cell, an organism, and / or a subject, preferably via homologous recombination. Following homologous recombination, the heterologous nucleic acid sequence is inserted into the genome of a cell, an organism, and / or a subject.
[0513] As used herein, the term “gene regulatory element” refers to a nucleic acid sequence that controls or modulates gene expression of a downstream gene. A gene regulatory element is typically operably linked to a polynucleotide sequence of a gene encoding sequence and can be located adjacent and 5’ upstream to a gene it controls.
[0514] As used herein the term “promoter” refers to a region of DNA upstream of a gene where relevant proteins, such as RNA polymerase and transcription factors, bind to initiate transcription of that gene.
[0515] As used herein the term “gene enhancer” refers to a regulatory DNA sequence that, when bound by specific proteins called transcription factors, enhances the transcription of an associated gene.
[0516] As used herein, the term “mutant nucleotide” refers to a nucleotide present in a polynucleotide sequence that differs from a corresponding nucleotide present in a wild-type polynucleotide sequence. A polynucleotide sequence having a mutant nucleotide is considered a variant of the wild-type polynucleotide sequence. The mutant nucleotide may be part of an endogenous genomic sequence or as part of a nucleotide modification present in a homology arm described herein.
[0517] As used herein, the term “wild-type nucleotide” refers to a nucleotide present in a polynucleotide sequence that is the same as the nucleotide from the corresponding wild-type polynucleotide sequence. The wild-type nucleotide may be part of an endogenous genomic sequence or as part of a nucleotide modification present in a homology arm described herein.
[0518] 54
[0519] P89339 2070WO (01288) As used herein with respect to both amino acid sequences and nucleic acid sequences, the terms “percent identity,” “sequence identity,” “percentage similarity,” “sequence similarity” and the like refer to a measure of the degree of similarity of two sequences based upon an alignment of the sequences that maximizes similarity between aligned amino acid residues or nucleotides, and which is a function of the number of identical or similar residues or nucleotides, the number of total residues or nucleotides, and the presence and length of gaps in the sequence alignment. A variety of algorithms and computer programs are available for determining sequence similarity using standard parameters. As used herein, sequence similarity is measured using the BLASTp program for amino acid sequences and the BLASTn program for nucleic acid sequences, both of which are available through the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov / ), and are described in, for example, Altschul et al. (1990), J. Mol. Biol. 215:403-410; Gish and States (1993), Nature Genet. 3:266-272; Madden et al. (1996), Meth. Enzymol.266: 131-141; Altschul et al. (1997), Nucleic Acids Res. 25:33 89-3402); Zhang et al. (2000), J. Comput. Biol. 7( l-2):203- 14. As used herein, percent similarity of two amino acid sequences is the score based upon the following parameters for the BLASTp algorithm: word size=3; gap opening penalty=-l l; gap extension penalty=-l; and scoring matrix=BLOSUM62. As used herein, percent similarity of two nucleic acid sequences is the score based upon the following parameters for the BLASTn algorithm: word size=l l; gap opening penalty=-5; gap extension penalty=-2; match reward=l; and mismatch penalty=-3.
[0520] As used herein, the term “corresponding to” with respect to modifications of two proteins or amino acid sequences, is used to indicate that a specified modification in the first protein is a substitution of the same amino acid residue as in the modification in the second protein, and that the amino acid position of the modification in the first protein corresponds to or aligns with the amino acid position of the modification in the second protein when the two proteins are subjected to standard sequence alignments (e.g., using the BLASTp program). Thus, the modification of residue “X” to amino acid “A” in the first protein will correspond to the modification of residue “Y” to amino acid “A” in the second protein if residues X and Y correspond to each other in a sequence alignment, and despite the fact that X and Y may be different numbers.
[0521] The methods described herein utilize a first and a second homology arm that have homology to a first and second genomic sequence and have one or more nucleotide modifications to alter the genome of a cell, an organism, and / or a subject through HDR. As such, the nucleotide modification present within the homology arms may comprise one or more mutant nucleotides not normally present in a wild-type gene; one or more mutant codons not normally present in a wild-
[0522] 55
[0523] P89339 2070WO (01288) type gene; one or more deleted nucleotides normally present in a wild-type gene; and / or one or more deleted codons normally present in a wild-type gene.
[0524] As used herein, the term “homologous recombination” (HR) or “homology-directed repair” (HDR), used interchangeably, refers to the natural, cellular process in which a double-stranded DNA-break is repaired using a homologous DNA sequence as the repair template (see, e.g., Cahill et al. (2006), Front. Biosci. 11: 1958-1976). As used herein, the term “non-homologous endjoining” (NHEJ) refers to the natural, cellular process in which a double-stranded DNA-break is repaired by the direct joining of two non-homologous DNA-break is repaired by the direct joining of two non-homologous DNA segments (see, e.g., Cahill et al. (2006), Front. Biosci. 11: 1958- 1976). DNA repair by NHEJ is error prone and frequently results in the untemplated addition or deletion of DNA sequences at the site of repair.
[0525] As used herein, the term “exogenous polynucleotide” refers to a polynucleotide that originates from outside of a cell, an organism, and / or a subject. An exogeneous polynucleotide may be integrated, in whole or in part, into the genome of a cell, and organism, and / or a subject. In some embodiments, the exogenous polynucleotide is integrated into the genome via homologous recombination.
[0526] To perform HDR, an exogenous polynucleotide is introduced into a cell, an organism, and / or a subject, wherein the exogenous polynucleotide comprises, from 5’ to 3’, a first homology arm and a second homology arm. In some instances, the exogenous polynucleotide comprises a heterologous nucleic acid sequence located between the first homology arm and the second homology arm, whereby the exogenous polynucleotide comprises, from 5’ to 3’, the first homology arm, a heterologous nucleic acid sequence, and the second homology arm.
[0527] As used herein, a “homology arm” refers to a sequence having homology to a genomic nucleic acid sequence. A homology arm can comprise at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide that is present within the genomic sequence, which enables the at least one corresponding modification to be incorporated into the genome of a cell, an organism, and / or a subject. A homology arm can flank the 5’ and 3’ ends of a nucleic acid molecule, i.e., a heterologous nucleic acid sequence, which promotes integration of a heterologous nucleic acid sequence into the genome of a cell, an organism, and / or a subject.
[0528] As used herein, the term “first homology arm” refers to a homology arm that has homology to a first genomic sequence. A first homology arm can flank the 5’ end of a heterologous nucleic acid sequence. As used herein, the term “first genomic sequence” refers to a region of the genome to which the first homology arm shares homology. The first genomic sequence is upstream of a
[0529] 56
[0530] P89339 2070WO (01288) double strand break, preferably a double strand break generated by an engineered nuclease. The first genomic sequence (and therefore the first homology arm) may be located immediately upstream of a double strand break, or may be located, for example, 500 bp, 1000 bp, 1.5 kb, 2kb, etc. upstream of the double strand break.
[0531] As used herein, the term “second homology arm” refers to a homology arm that has homology to a second genomic sequence. A second homology arm can flank the 3’ end of a heterologous nucleic acid sequence. As used herein, the term “second genomic sequence” refers a region of the genome to which the second homology arm shares homology. The second genomic sequence is located downstream of a double strand break, preferably a double strand break generated by an engineered nuclease. The second genomic sequence (and therefore the second homology arm) may be located immediately downstream of a double strand break, or may be located, for example, 500 bp, 1000 bp, 1.5kb, 2kb, 5kb, lOkb, etc. downstream of the double strand break.
[0532] The first and second homology arms may be approximately the same length. As used herein, the term “approximately the same length” refers to a first homology arm and a second homology arm that are within 10% of length of each other as measured by the number of base pairs (nucleotides). Alternatively, the first and second homology arms may be different lengths. As used herein, the term “different lengths” refers to a first homology arm and a second homology arm that have a greater than 10% difference in length of each other as measured by the number of base pairs (nucleotides). In general, homology arms can have a length of at least 50 base pairs, 100 base pairs, 500 base pairs, 1000 base pairs, 2000 base pairs, 3000 base pairs, or more. For example, but not by way of limitation, a first homology arm having a length of 500 base pairs and a second homology arm having a length of 600 base pairs would be classified as having approximately the same length. A first homology arm having a length of 500 base pairs and a second homology arm having a length of 1000 base pairs would be classified as having different lengths.
[0533] As used herein the term “nucleotide” refers to the basic building block of nucleic acids. A nucleotide consists of a sugar molecule (either ribose in RNA or deoxyribose in DNA) attached to a phosphate group and a nitrogen-containing base. The bases used in DNA are adenine (A), cytosine (C), guanine (G) and thymine (T). In RNA, the base uracil (U) takes the place of thymine. DNA and RNA molecules are polymers made up of long chains of nucleotides.
[0534] As used herein the term “corresponding endogenous nucleotide” refers to a nucleotide present within a genomic sequence that is modified following integration of a homology arm sequence into a genome through homologous recombination with a homologous genomic sequence.
[0535] 57
[0536] P89339 2070WO (01288) As used herein, the term “mismatched nucleotide” refers to a nucleotide modification wherein a nucleotide present within an original genomic sequence is replaced by a different nucleotide in a replacement genomic sequence that is incorporated into the genome following integration of a homology arm with its corresponding, homologous sequence on the genome.
[0537] As used herein, the term “nucleotide deletion” refers a nucleotide modification wherein a nucleotide present within an original genomic sequence is lacking (e.g., removed, or absent (i.e., deleted) from a replacement genomic sequence that is incorporated into the genome following integration of a homology arm with its corresponding, homologous sequence on the genome. A first and / or second homology arm can lack a corresponding endogenous nucleotide that is present in a first and / or second genomic sequence, respectively.
[0538] As used herein, the term “nucleotide insertion” refers a nucleotide modification wherein a nucleotide that is not present within an original genomic sequence is introduced into the replacement genomic sequence that is incorporated into the genome following integration of a homology arm with its corresponding, homologous sequence on the genome. A nucleotide that has been introduced into a sequence is referred to as an “inserted nucleotide.” An inserted nucleotide can be added within any sequence of a genome, such as an intron sequence or exon sequence, or non-gene coding sequence.
[0539] As used herein, the term “splice donor sequence” refers to the sequence at the beginning of an intron in a gene that marks the start of an intron and is required for splicing. The splice donor sequence is almost always GT, but can also be GC-AG, GG-AG, GT-TG, GT-CG, or CT-AG (Burset M, Seledtsov IA, Solovyev VV. Analysis of canonical and non-canonical splice sites in mammalian genomes. Nucleic Acids Res. 2000 Nov l;28(21):4364-75, the contents of which are incorporated herein in their entirety). A “mutant splice donor sequence” differs from a wild-type splice donor sequence by at least one nucleotide modification. The nucleotide modification may be a nucleotide mismatch, a nucleotide deletion, and / or a nucleotide insertion. The presence of a nucleotide modification may disrupt the function of the splice donor sequence. A mutant modification can be located anywhere within a splice donor sequence. For instance, a mutant nucleotide may be inserted between the two nucleotides of the splice donor sequence. Alternatively, a mismatch nucleotide may replace a nucleotide of the splice donor sequence, or a splice donor sequence nucleotide may be deleted. As used herein, the term “splice acceptor site” refers to the sequence located at the 3' end of the intron and terminates the intron. The splice acceptor site is most commonly AG. Commonly, the splice acceptor sequence is preceded by a polypyrimidine tract, which is a region high in pyrimidines (C and U). A “mutant splice acceptor
[0540] 58
[0541] P89339 2070WO (01288) sequence” differs from a wild-type splice acceptor sequence by at least one nucleotide modification. The nucleotide modification may be a nucleotide mismatch, a nucleotide deletion, and / or a nucleotide insertion. The presence of a nucleotide modification may disrupt the function of the splice acceptor sequence. For instance, a mutant nucleotide may be inserted between the two nucleotides of the splice acceptor sequence. Alternatively, a mismatch nucleotide may replace a nucleotide of the splice acceptor sequence, or a splice acceptor sequence nucleotide may be deleted.
[0542] As used herein, the term “splice enhancer sequence” refers to sequences required for accurate splice site recognition and the control of alternative splicing. Pre-mRNA splicing enhancer elements are short RNA sequences capable of activating weak splice sites in nearby introns. These elements can activate heterologous pre-mRNAs and can function at a distance as great as 500 nucleotides from the affected intron. Both constitutive and regulated splicing enhancers contain binding sites for SR proteins, a family of essential splicing factors containing one or more RNA recognition motifs (RRM) and an arginine / serine (RS)-rich domain. RRM is required for RNA binding, and the RS domain is required for protein-protein interactions. Mechanistic studies of splicing enhancer function are consistent with a recruitment model in which SR proteins activate splicing by binding to enhancers and recruiting the splicing machinery to the adjacent intron. A “mutant splice enhancer sequence” refers to a sequence with at least one nucleotide modification compared to a wild-type splice enhancer sequence.
[0543] As used herein, the term “non-gene protein coding sequence” or “non-protein coding region” refers to portions of an organism’s genome that do not code for amino acids, the building blocks of proteins. Some non-coding DNA sequences are known to serve functional roles, such as in the regulation of gene expression, such as gene regulatory elements, promoters, and gene enhancers. Additional regions of DNA sequences that do not encode amino acids include intronic regions. Other areas of non-coding DNA have no known function. A mutant non-gene protein coding sequence refers to sequences which comprise one or more nucleotide that differ from a wild-type non-gene protein coding sequence.
[0544] As used herein, the term “flanked” or “flanking” in relation to a particular nucleic acid sequence refers to a sequence (i.e., a flanking sequence) being on each side of a particular nucleic acid sequence. For example, a first homology arm and a second homology arm can be flanking sequences to a heterologous nucleic acid sequence.
[0545] As used herein, the terms “transfected” or “transformed” or “transduced” or “nucleofected” refer to a process by which exogenous nucleic acid is transferred or introduced into the host cell. A “transfected” or “transformed” or “transduced” cell is one which has been transfected, transformed,
[0546] 59
[0547] P89339 2070WO (01288) or transduced with exogenous nucleic acid. The cell includes the primary subject cell and its progeny.
[0548] As used herein, the term “double-strand break” refers to a site in the genome (e.g., a cleavage site) in which both strands of the DNA are severed. A double -stranded break as described herein occurs within the genome of a cell and is generated by a nuclease. Exemplary nucleases that can generated double-strand breaks and can be used are known and described herein.
[0549] As used herein, the term “blunt end” refers to a DNA fragment with no overhanging bases at either end. It is created when a restriction enzyme or nuclease cuts a DNA molecule straight across at the same site on both strands, resulting in two blunt ends.
[0550] As used herein, the term “5’ overhang” refers to a stretch of single-stranded bases at the end of a DNA molecule that ends with a 5’ phosphate.
[0551] As used herein, the term “3 ’ overhang” refers to a single-stranded DNA sequence at the end of a DNA molecule where the unpaired nucleotides have a 3' hydroxyl group, meaning the last nucleotide in the strand is at the 3' position, creating a protruding single -stranded region on the DNA molecule.
[0552] The methods described herein use a site-directed nuclease that cleaves a site in the genome to generate a double strand break. As used herein, the terms “cleave” or “cleavage” refer to the hydrolysis of phosphodiester bonds within the backbone of a recognition sequence within a target sequence that results in a double-stranded break within the target sequence, referred to herein as a “cleavage site.”
[0553] As used herein, the term “nuclease” refers to a naturally-occurring or engineered enzyme, which cleaves a phosphodiester bond within a polynucleotide chain, thereby generating a double strand break. Examples of nucleases include, but are not limited to, engineered meganucleases, CRISPR system nucleases, TALE nucleases (TALENs), zinc finger nucleases, compact TALENs, and megaTALs.
[0554] As used herein, the terms “recognition sequence” or “recognition site” refers to a DNA sequence that is bound and cleaved by a nuclease. In the case of a meganuclease, a recognition sequence comprises a pair of inverted, 9 base pair “half sites” which are separated by four base pairs. In the case of a single-chain meganuclease, the N-terminal domain of the protein contacts a first half-site and the C-terminal domain of the protein contacts a second half-site. Cleavage by a meganuclease produces four base pair 3' overhangs. “Overhangs,” or “sticky ends” are short, single-stranded DNA segments that can be produced by endonuclease cleavage of a doublestranded DNA sequence. In the case of meganucleases and single-chain meganucleases derived
[0555] 60
[0556] P89339 2070WO (01288) from I-Crel, the overhang comprises bases 10-13 of the 22 base pair recognition sequence. In the case of a compact TALEN, the recognition sequence comprises a first CNNNGN sequence that is recognized by the I-TevI domain, followed by a non-specific spacer 4-16 base pairs in length, followed by a second sequence 16-22 bp in length that is recognized by the TAL-effector domain (this sequence typically has a 5' T base). Cleavage by a compact TALEN produces two base pair 3' overhangs. In the case of a CRISPR nuclease, the recognition sequence is the sequence, typically 16-24 base pairs, to which the guide RNA binds to direct cleavage. Full complementarity between the guide sequence and the recognition sequence is not necessarily required to effect cleavage. Cleavage by a CRISPR nuclease can produce blunt ends (such as by a class 2, type II CRISPR nuclease) or overhanging ends (such as by a class 2, type V CRISPR nuclease), depending on the CRISPR nuclease. In those embodiments wherein a Cpfl CRISPR nuclease is utilized, cleavage by the CRISPR complex comprising the same will result in 5' overhangs and in certain embodiments, 5 nucleotide 5' overhangs. Each CRISPR nuclease enzyme also requires the recognition of a PAM (protospacer adjacent motif) sequence that is near the recognition sequence complementary to the guide RNA. The precise sequence, length requirements for the PAM, and distance from the target sequence differ depending on the CRISPR nuclease enzyme, but PAMs are typically 2-5 base pair sequences adjacent to the target / recognition sequence. PAM sequences for particular CRISPR nuclease enzymes are known in the art (see, for example, U.S. Patent No. 8,697,359 and U.S. Publication No. 20160208243, each of which is incorporated by reference in its entirety) and PAM sequences for novel or engineered CRISPR nuclease enzymes can be identified using methods known in the art, such as a PAM depletion assay (see, for example, Karvelis et al. (2017) Methods 121-122:3-8, which is incorporated herein in its entirety). In the case of a zinc finger, the DNA binding domains typically recognize an 18-bp recognition sequence comprising a pair of nine base pair “half-sites” separated by 2-10 base pairs and cleavage by the nuclease creates a blunt end or a 5' overhang of variable length (frequently four base pairs).
[0557] As used herein, the term “meganuclease” refers to an endonuclease that binds doublestranded DNA at a recognition sequence that is greater than 12 base pairs. A meganuclease can be an endonuclease that is derived from I-Crel and can refer to an engineered variant of I-Crel that has been modified relative to natural I-Crel with respect to, for example, DNA-binding specificity, DNA cleavage activity, DNA-binding affinity, or dimerization properties. Methods for producing such modified variants of I-Crel are known in the art (e.g., WO 2007 / 047859, incorporated by reference in its entirety). A meganuclease as used herein binds to double-stranded DNA as a heterodimer. A meganuclease may also be a “single-chain meganuclease” in which a pair of
[0558] 61
[0559] P89339 2070WO (01288) meganuclease DNA-binding domains is joined into a single polypeptide using a peptide linker. The term “homing endonuclease” is synonymous with the term “meganuclease.” Meganucleases used in the presently disclosed methods and compositions are substantially non-toxic when expressed in the targeted cells as described herein such that cells can be transfected and maintained at 37°C without observing deleterious effects on cell viability or significant reductions in meganuclease cleavage activity.
[0560] As used herein, the term “TALE nuclease” or “TALEN” refers to an endonuclease comprising a DNA-binding domain comprising a plurality of TAL domain repeats fused to a nuclease domain or an active portion thereof from an endonuclease or exonuclease, including but not limited to a restriction endonuclease, homing endonuclease, S 1 nuclease, mung bean nuclease, pancreatic DNAse I, micrococcal nuclease, and yeast HO endonuclease. See, for example, Christian et al. (2010) Genetics 186:757-761, which is incorporated by reference in its entirety. Nuclease domains useful for the design of TALENs include those from a Type Ils restriction endonuclease, including but not limited to FokI, FoM, StsI, Hhal, Hindlll, Nod, BbvCI, EcoRI, Bgll, and AlwI. Additional Type Ils restriction endonucleases are described in International Publication No. WO 2007 / 014275, which is incorporated by reference in its entirety. In some embodiments, the nuclease domain of the TALEN is a FokI nuclease domain or an active portion thereof. TAL domain repeats can be derived from the TALE (transcription activator-like effector) family of proteins used in the infection process by plant pathogens of the Xanthomonas genus. TAL domain repeats are 33-34 amino acid sequences with divergent 12th and 13th amino acids. These two positions, referred to as the repeat variable dipeptide (RVD), are highly variable and show a strong correlation with specific nucleotide recognition. Each base pair in the DNA target sequence is contacted by a single TAL repeat with the specificity resulting from the RVD. In some embodiments, the TALEN comprises 16-22 TAL domain repeats. DNA cleavage by a TALEN requires two DNA recognition regions (i.e., “half-sites”) flanking a nonspecific central region (i.e., the “spacer”). The term “spacer” in reference to a TALEN refers to the nucleic acid sequence that separates the two nucleic acid sequences recognized and bound by each monomer constituting a TALEN. The TAL domain repeats can be native sequences from a naturally-occurring TALE protein or can be redesigned through rational or experimental means to produce a protein that binds to a pre-determined DNA sequence (see, for example, Boch et al. (2009) Science 326(5959): 1509- 1512 and Moscou and Bogdanove (2009) Science 326(5959): 1501, each of which is incorporated by reference in its entirety). See also, U.S. Publication No. 20110145940 and International Publication No. WO 2010 / 079430 for methods for engineering a TALEN to recognize and bind a
[0561] 62
[0562] P89339 2070WO (01288) specific sequence and examples of RVDs and their corresponding target nucleotides. In some embodiments, each nuclease (e.g., FokI) monomer can be fused to a TAL effector sequence that recognizes and binds a different DNA sequence, and only when the two recognition sites are in close proximity do the inactive monomers come together to create a functional enzyme. It is understood that the term “TALEN” can refer to a single TALEN protein or, alternatively, a pair of TALEN proteins (i.e., a left TALEN protein and a right TALEN protein) which bind to the upstream and downstream half-sites adjacent to the TALEN spacer sequence and work in concert to generate a cleavage site within the spacer sequence. Given a predetermined DNA locus or spacer sequence, upstream and downstream half-sites can be identified using a number of programs known in the art (Komel Labun; Tessa G. Montague; James A. Gagnon; Summer B. Thyme; Eivind Valen. (2016). CHOPCHOP v2: a web tool for the next generation of CRISPR genome engineering. Nucleic Acids Research; doi: 10.1093 / nar / gkw398; Tessa G. Montague; Jose M. Cruz; James A. Gagnon; George M. Church; Eivind Valen. (2014). CHOPCHOP: a CRISPR / Cas9 and TALEN web tool for genome editing. Nucleic Acids Res. 42. W401-W407). It is also understood that a TALEN recognition sequence can be defined as the DNA binding sequence (i.e., half-site) of a single TALEN protein or, alternatively, a DNA sequence comprising the upstream half-site, the spacer sequence, and the downstream half-site.
[0563] As used herein, the term “compact TALEN” refers to an endonuclease comprising a DNA- binding domain with one or more TAL domain repeats fused in any orientation to any portion of the I-TevI homing endonuclease or any of the endonucleases listed in Table 2 in U.S. Application No. 20130117869 (which is incorporated by reference in its entirety), including but not limited to Mmel, EndA, Endl, I-BasI, I-TevII, I-TevIII, I-Twol, MspI, Mval, NucA, and NucM. Compact TALENs do not require dimerization for DNA processing activity, alleviating the need for dual target sites with intervening DNA spacers. In some embodiments, the compact TALEN comprises 16-22 TAL domain repeats.
[0564] As used herein, the term “megaTAL” refers to a single-chain endonuclease comprising a transcription activator-like effector (TALE) DNA binding domain with an engineered, sequencespecific homing endonuclease.
[0565] As used herein, the terms “CRISPR nuclease” or “CRISPR system nuclease” refers to a CRISPR (clustered regularly interspaced short palindromic repeats)-associated (Cas) endonuclease or a variant thereof, such as Cas9, that associates with a guide RNA that directs nucleic acid cleavage by the associated endonuclease by hybridizing to a recognition site in a polynucleotide. In certain embodiments, the CRISPR nuclease is a class 2 CRISPR enzyme. In some of these
[0566] 63
[0567] P89339 2070WO (01288) embodiments, the CRISPR nuclease is a class 2, type II enzyme, such as Cas9. In other embodiments, the CRISPR nuclease is a class 2, typeV enzyme, such as Cpfl . The guide RNA comprises a direct repeat and a guide sequence (often referred to as a spacer in the context of an endogenous CRISPR system), which is complementary to the target recognition site. In certain embodiments, the CRISPR system further comprises a tracrRNA (trans-activating CRISPR RNA) that is complementary (fully or partially) to the direct repeat sequence (sometimes referred to as a tracr-mate sequence) present on the guide RNA. In particular embodiments, the CRISPR nuclease can be mutated with respect to a corresponding wild-type enzyme such that the enzyme lacks the ability to cleave one strand of a target polynucleotide, functioning as a nickase, cleaving only a single strand of the target DNA. Non-limiting examples of CRISPR enzymes that function as a nickase include Cas9 enzymes with a D10A mutation within the RuvC I catalytic domain, or with a H840A, N854A, or N863A mutation. Given a predetermined DNA locus, recognition sequences can be identified using a number of programs known in the art (Komel Labun; Tessa G. Montague; James A. Gagnon; Summer B. Thyme; Eivind Valen. (2016). CHOPCHOP v2: a web tool for the next generation of CRISPR genome engineering. Nucleic Acids Research; doi: 10.1093 / nar / gkw398; Tessa G. Montague; Jose M. Cruz; James A. Gagnon; George M. Church; Eivind Valen. (2014). CHOPCHOP: a CRISPR / Cas9 and TALEN web tool for genome editing. Nucleic Acids Res. 42. W401-W407).
[0568] As used herein, the terms “zinc finger nuclease” or “ZFN” refers to a chimeric protein comprising a zinc finger DNA-binding domain fused to a nuclease domain from an endonuclease or exonuclease, including but not limited to a restriction endonuclease, homing endonuclease, S 1 nuclease, mung bean nuclease, pancreatic DNAse I, micrococcal nuclease, and yeast HO endonuclease. Nuclease domains useful for the design of zinc finger nucleases include those from a Type Ils restriction endonuclease, including but not limited to FokI, FoM, and StsI restriction enzyme. Additional Type Ils restriction endonucleases are described in International Publication No. WO 2007 / 014275, which is incorporated by reference in its entirety. The structure of a zinc finger domain is stabilized through coordination of a zinc ion. DNA binding proteins comprising one or more zinc finger domains bind DNA in a sequence -specific manner. The zinc finger domain can be a native sequence or can be redesigned through rational or experimental means to produce a protein which binds to a pre-determined DNA sequence ~18 base pairs in length, comprising a pair of nine base pair half-sites separated by 2-10 base pairs. See, for example, U.S. Pat. Nos. 5,789,538, 5,925,523, 6,007,988, 6,013,453, 11,311,574, and International Publication Nos. WO 95 / 19431, WO 96 / 06166, WO 98 / 53057, WO 98 / 54311, WO 00 / 27878, WO 01 / 60970, WO
[0569] 64
[0570] P89339 2070WO (01288) 01 / 88197, and WO 02 / 099084, each of which is incorporated by reference in its entirety. By fusing this engineered protein domain to a nuclease domain, such as FokI nuclease, it is possible to target DNA breaks with genome-level specificity. The selection of target sites, zinc finger proteins and methods for design and construction of zinc finger nucleases are known to those of skill in the art and are described in detail in U.S. Publications Nos. 20030232410, 20050208489, 2005064474, 20050026157, 20060188987 and International Publication No. WO 07 / 014275, each of which is incorporated by reference in its entirety. In the case of a zinc finger, the DNA binding domains typically recognize an 18-bp recognition sequence comprising a pair of nine base pair “half-sites” separated by a 2-10 base pair “spacer sequence”, and cleavage by the nuclease creates a blunt end or a 5' overhang of variable length (frequently four base pairs). It is understood that the term “zinc finger nuclease” can refer to a single zinc finger protein or, alternatively, a pair of zinc finger proteins (i.e., a left ZFN protein and a right ZFN protein) that bind to the upstream and downstream half-sites adjacent to the zinc finger nuclease spacer sequence and work in concert to generate a cleavage site within the spacer sequence. Given a predetermined DNA locus or spacer sequence, upstream and downstream half-sites can be identified using a number of programs known in the art (Mandell JG, Barbas CF 3rd. Zinc Finger Tools: custom DNA-binding domains for transcription factors and nucleases. Nucleic Acids Res. 2006 Jul 1 ;34 (Web Server issue):W516-23). It is also understood that a zinc finger nuclease recognition sequence can be defined as the DNA binding sequence (i.e., half-site) of a single zinc finger nuclease protein or, alternatively, a DNA sequence comprising the upstream half-site, the spacer sequence, and the downstream half-site.
[0571] As used herein, the term “specificity” means the ability of a nuclease to bind and cleave double-stranded DNA molecules only at a particular sequence of base pairs referred to as the recognition sequence, or only at a particular set of recognition sequences. The set of recognition sequences will share certain conserved positions or sequence motifs but may be degenerate at one or more positions. A highly-specific nuclease is capable of cleaving only one or a very few recognition sequences. Specificity can be determined by any method known in the art.
[0572] As used herein, the terms “target site” or “target sequence” refers to a region of the chromosomal DNA of a cell comprising a recognition sequence for a nuclease.
[0573] As used herein, the term “recognition half-site,” “recognition sequence half-site,” or simply “half-site” means a nucleic acid sequence in a double-stranded DNA molecule that is recognized and bound by a monomer of a homodimeric or heterodimeric meganuclease or by one subunit of a single-chain meganuclease or by one subunit of a single-chain meganuclease, or by a monomer of a TALEN or zinc finger nuclease.
[0574] 65
[0575] P89339 2070WO (01288) As used herein, the term “5' portion” when referring to a nuclease recognition sequence is intended to mean the nucleotides of the recognition sequence that are 5' upstream of a cleavage site generated by an engineered nuclease. Similarly, the term “3' portion” when referring to a nuclease recognition sequence is intended to mean the nucleotides of the recognition sequence that are 3' downstream of a cleavage site generated by an engineered nuclease. By way of example, where an engineered nuclease, such as a CRISPR / Cas9 nuclease system, generates a blunt end cleavage site within a recognition sequence, the 5' portion of the recognition sequence comprises the nucleotides of the sequence that are 5' upstream of the cleavage site, while the 3' portion of the recognition sequence comprises the nucleotides of the sequence that are 3' downstream of the cleavage site. For nucleases that generate 5' or 3' overhangs (e.g., engineered meganucleases), each 5' portion or 3' portion can, in some examples, include only the nucleotides that are 5' upstream or 3' downstream of the cleavage site, respectively.
[0576] In other examples, each 5' portion or 3' portion can include the nucleotides that are 5' upstream or 3' downstream of the cleavage site, respectively, as well as the nucleotides of the 5' or 3' overhang. By way of example, an I-Crel-derived engineered meganuclease generates a cleavage site between nucleotide positions 13 and 14 of its 22 base pair recognition sequence, wherein positions 9-13 represent the 4 base pair 3' overhang. Thus, in some examples, the 5' portion of the sequence can comprise the nucleotides of positions 1-13, and the 3' portion of the sequence can comprise the nucleotides of positions 14-22. In other examples, the 5' portion of the sequence can comprise the nucleotides of positions 1-13, and the 3' portion of the sequence can comprise the nucleotides of positions 10-22, which includes the 4 base pair overhang and the 9 base pairs of the 3' half site. In some cases, the inclusion or exclusion of the overhang nucleotides as components of a homology arm may be used to affect homology directed repair. It is understood that in examples wherein a 5' portion of a recognition sequence and a 3' portion of a recognition sequence are positioned adjacent to one another for the purpose of generating a complete recognition sequence, only the portion that normally includes the overhang base pairs will comprise the overhang nucleotides.
[0577] As used herein, the term “operably linked” is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a nucleic acid sequence encoding a nuclease and a regulatory sequence (e.g., a promoter) is a functional link that allows for expression of the nucleic acid sequence encoding the nuclease. Operably linked elements may be contiguous or non-contiguous. When used to refer to the joining of two protein coding regions, by operably linked is intended that the coding regions are in the same reading frame.
[0578] 66
[0579] P89339 2070WO (01288) As used herein, the term “promoter” or “regulatory sequence” refers to a nucleic acid sequence which is required for expression of a gene product operably linked to the promoter or regulatory sequence. In some instances, this sequence may be the core promoter sequence and in other instances, this sequence may also include an enhancer sequence and other regulatory elements which are required for expression of the gene product. The promoter / regulatory sequence may, for example, be one which expresses the gene product in a tissue specific manner.
[0580] As used herein, the term “delivery vehicle” refers to a delivery mechanism for contacting a cell, an organism, and / or a subject with one or more exogenous polynucleotides. Delivery vehicles may be, for example, recombinant viruses, such as an adeno-associated virus (AAV), or lipid nanoparticles.
[0581] As used herein, the term “adeno-associated virus,” “adeno-associated viral particles,” “adeno-associated viral vectors,” or “AAV” or “AAV particles” or “AAV vectors” are used interchangeably and refer to viruses and vectors derived from the Parvoviridcie family. Adeno- associated viruses can infect both dividing and quiescent cells. Adeno-associated viruses of any serotype can be used in the presently disclosed methods and compositions, as well as variants thereof. As used herein, the term “serotype” refers to a distinct variant within a species of virus that is determined based on the viral cell surface antigens. Known serotypes of AAV include, for example, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, and AAV11 (Weitzman and Linden (2011) In Snyder and Moullier Adeno-associated virus methods and protocols. Totowa, NJ: Humana Press). Other serotypes of AAV are known, and a person of skill in the art could determine their usefulness in the present disclosure, including, for example, AAV 13, AAV14, AAV15, AAV16, AAV.rh8, AAV.rhlO, AAV.rh20, AAV.rh39, AAV.rh74, AAV.rh79, AAV.RHM4-1, AAV.hu37, AAV.hu68, AAV.Anc80, AAV.Anc90L65, AAV.7m8, AAV. PHP. B, AAV2.5, AAV2tYF, AAV3B, AAV.LK03, AAV.HsCl, AAV.HSC2, AAV.HSC3, AAV.HSC4, AAV.HSC5, AAV.HSC6, AAV.HSC7, AAV.HSC9, AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAVMYO, MYOAAV and AAV.HSC16 serotypes, such as those described in PCT Publication No. WO 2022 / 133051, incorporated herein by reference. An AAV particle may be of a single serotype or a combination of serotypes, or derivatives thereof. Generally, adeno-associated viruses comprise a transgene or portion thereof that is flanked by parvoviral or AAV inverted terminal repeat sequences (ITRs). Such AAV vectors can be replicated and packaged into infectious viral particles when present in a host cell that is expressing AAV replication (rep) and capsid (cap) gene products. Methods of
[0582] 67
[0583] P89339 2070WO (01288) producing recombinant AAV particles are described in the art, including, for example, PCT Publication No. WO 2022 / 133051.
[0584] As used herein, the term “inverted terminal repeat” or “ITR” refers to regions found at the 5' and 3' termini of an AAV genome. The ITRs are about 145 nt each that flank a transgene or portion thereof within an AAV genome. The ITRs are self-complementary and organized so that an energetically stable intramolecular duplex forming a T-shaped hairpin may be formed. These hairpin structures function as an origin for viral DNA replication, serving as primers for the cellular DNA polymerase complex. The ITRs also aid in concatamer formation in the nucleus and integration into the genome. Sequences of AAV-associated ITRs are known in the art, for example those disclosed by Yan et al., J. Virol. 79(l):364-379 (2005). ITR sequences that find use herein may be full length, wild-type AAV ITRs or fragments thereof that retain functional capability, or may be sequence variants of full-length, wild-type AAV ITRs that are capable of functioning in cis as origins of replication. AAV ITRs useful in the presently disclosed methods and compositions may derive from any known AAV serotype.
[0585] As used herein, the term “D sequence” refers to D-sequence, a stretch of nucleotides, which in some cases can be 20 nucleotides in length, that are associated with an AAV inverted terminal repeat but do not play a role in hairpin formation. The D sequence has been shown to play a role in the life cycle of AAV in a number of ways. For example, the D sequence acts as the packaging signal for AAV. Further, the first 10 nucleotides of a D sequence have been shown to be necessary for AAV DNA replication (see, Kwon et al., Human Gene Therapy (2020), Vol. 31 (9-10): 565- 574).
[0586] As used herein, the term “lipid nanoparticle” refers to a lipid composition having a typically spherical structure with an average diameter between 10 and 1000 nanometers. In some formulations, lipid nanoparticles can comprise at least one cationic lipid, at least one non-cationic lipid, and at least one conjugated lipid. Lipid nanoparticles known in the art that are suitable for encapsulating nucleic acids, such as mRNA, are contemplated for use in the presently disclosed methods. A nucleic acid, such as an mRNA, may be encapsulated in the lipid portion of the lipid nanoparticle or aqueous space enveloped by some or all of the lipid portion of the lipid nanoparticle. This affords protection from enzymatic degradation or other undesired effects induced by a cell, an organism, and / or a subject contacted with the lipid nanoparticle. Lipid nanoparticles may further be conjugated with a targeting moiety, such as an antibody or a ligand, to direct the lipid nanoparticle to a target cell or target tissue.
[0587] 68
[0588] P89339 2070WO (01288) As used herein, the term “therapeutically effective amount” refers to an amount sufficient to effect beneficial or desirable biological and / or clinical results, thereby treating a subject.
[0589] As used herein, the term “treat,” “treating,” or “treatment” refers to the administration of a pharmaceutical composition disclosed herein, comprising, for example an exogenous polynucleotide encoding a replacement nucleic acid sequence and an exogenous polynucleotide encoding a nuclease, to a subject having a disease, disorder, or condition. For example, the subject may have a disease characterized by a pathogenic allele, wherein the treatment corrects the pathogenic allele to a wild-type allele.
[0590] As used herein, the recitation of a numerical range for a variable is intended to convey that the present disclosure may be practiced with the variable equal to any of the values within that range. Thus, for a variable which is inherently discrete, the variable can be equal to any integer value within the numerical range, including the end-points of the range. Similarly, for a variable which is inherently continuous, the variable can be equal to any real value within the numerical range, including the end-points of the range. As an example, and without limitation, a variable which is described as having values between 0 and 2 can take the values 0, 1 or 2 if the variable is inherently discrete, and can take the values 0.0, 0.1, 0.01, 0.001, or any other real values =0 and =2 if the variable is inherently continuous.
[0591] 2 Principle of the Invention
[0592] The present disclosure is based, in part, on the discovery that a novel nuclease-initiated homology directed repair (HDR) method can be used to predictably modify an endogenous genomic sequence at any locus within regions downstream (or upstream) of a double strand break. The regions, for example, can be any region within 10 kilobases upstream (e.g., 5’) or downstream (e.g., 3’) of the double-strand break. In particular, it was discovered herein that perfect homology is not required for nuclease mediated HDR and thus certain nucleotide modifications (e.g., nucleotide changes, nucleotide insertions, and nucleotide deletions) can be introduce to the genome.
[0593] Provided herein are, inter alia, exogenous polynucleotide sequences that find use in introducing at least one nucleotide modification within a genomic sequence. The genomic sequence can be a region or portion of a genome of a subject, such as a mammal or a plant. The exogenous polynucleotides comprise homology arms that are specifically designed to have homology with genomic sequences located at specified distances from a double strand break. Aspects of the homology arms, such as length and homology with genomic sequences, as well as specified distances of the homologous sequence from a double-strand break are described herein. The
[0594] 69
[0595] P89339 2070WO (01288) provided compositions and methods can be used to modify any locus of a genomic sequence with individual nucleotide residue precision.
[0596] The exogenous polynucleotides comprise a first and a second homology arm. To introduce a nucleotide modification into a genomic sequence, the first and / or second homology arm are designed to comprise at least one nucleotide modification relative to the corresponding endogenous genomic sequence. For instance, the first homology arm comprises at least one nucleotide modification relative to its corresponding endogenous genomic sequence (i.e., original un-modified first genomic sequence) and / or the second homology arm comprises at least one nucleotide modification relative to its corresponding endogenous genomic sequence (i.e., original un-modified second genomic sequence). A nucleotide modification can be introduced into a first genomic sequence (i.e., first nucleotide modification) and / or a second genomic sequence (i.e., second nucleotide modification) by engineering the first and / or second homology arms to comprise at least one nucleotide modification according to the disclosure. Following integration of the first and second homology arm into the genome at sites of their corresponding endogenous sequences, any nucleotide modification present within the first and / or second homology arm will be introduced and integrated into the first and / or second genomic sequence.
[0597] A genomic sequence (i.e., original genomic sequence) can be modified at any location. The modification site can include any sequence within the genome (i.e., locus). A modified genomic sequence comprises at least one nucleotide modification compared to the original genomic sequence. The at least one nucleotide modification is introduced by integration of homology arms having homology with their corresponding endogenous sequence through HDR. A nucleotide modification can comprise the modification of any endogenous polynucleotide sequence of a genomic sequence.
[0598] One aspect of the compositions and methods included in the present disclosure is the particularly effective ability to modify genomic sequences with precision. Such precision allows a genomic sequence to be modified at the individual nucleotide level at any site of the genome. That is, a nucleotide modification can comprise modification of a single nucleotide reside. A modification can also include more than one nucleotide, as described herein.
[0599] As described herein, a locus can comprise any nucleotide sequence, such as a nucleotide residue, a codon, or a larger partition of a genomic sequence, such as an exon, or an intron, and / or a region thereof, such as a splice donor sequence, a splice acceptor sequence, and / or a splice enhancer sequence. A locus can also comprise a non-gene protein coding sequence, such as a gene regulatory element, including a promoter and / or a gene enhancer.
[0600] 70
[0601] P89339 2070WO (01288) A nucleotide modification can involve any modification involving at least one nucleotide difference between an original and a replacement genomic sequence. For example, a nucleotide modification can involve a nucleotide mismatch, nucleotide deletion, and / or nucleotide insertion. Moreover, a modification can include, for example: replacement, removal, and / or insertion of at least one nucleotide. A nucleotide can be replaced with a different nucleotide, as described herein. In some instances, a single nucleotide is removed (i.e., deleted), while in other instances, multiple nucleotides are removed. Multiple nucleotides can be removed, for example, from a gene having a mutation involving multiple mutant nucleotides (e.g., a nucleotide repeat). It is also possible to insert at least one nucleotide into an original genomic sequence. In any of the modifications described, homology arms of an exogenous polynucleotide are designed to comprise the modification such that it is introduced into the genome at the site corresponding with the homologous genomic sequence. As such, the homology arms differ from their corresponding endogenous genomic sequence by any nucleotide modification being introduced (i.e., at least one nucleotide modification), but otherwise are the same.
[0602] Recombination of homology arms with their corresponding endogenous sequences can also be particularly useful for inserting larger nucleic acid sequences into the genome. In such instances, a heterologous nucleic acid sequence can be comprised within the exogenous polynucleotide between the first homology arm and second homology arm. As such, the heterologous nucleic acid sequence can be integrated into any location of the genome described herein by specifically designing the flanking homology arms to comprise sequences having homology to sequences flanking the desired insertion site. Following integration of homology arms into their corresponding sequence homology sites, the heterologous nucleic acid sequence will be present on the replacement genome between the homology sites.
[0603] 3, Methods of Nuclease-Initiated Homology Directed Repair and Modification
[0604] Site-specific nucleases generate a double strand break in the genome of a cell, which can result in permanent modification of the genome via homologous recombination with a DNA sequence. The use of nucleases to induce a double-strand break in a target locus stimulates homologous recombination of DNA sequences that are flanked by sequences having homology to the genomic target; a process known as homology-directed repair (HDR). The natural mechanism of HDR can be utilized to integrate nucleotide modifications, such as the nucleotide mismatches, nucleotide deletions, and / or nucleotide insertions described herein, at a specific site in the genome. HDR can also be used to integrate exogenous DNA, such as a heterologous nucleic acid sequence
[0605] 71
[0606] P89339 2070WO (01288) described herein. The exogenous polynucleotides described herein comprise homology arms that have homology to genomic sequences flanking a double-strand break. As such, when used in combination with a nuclease, which generates a double-strand break in a cell genome, the exogenous polynucleotides can be incorporated into the genome through homologous recombination of a first homology arm and a second homology arm, which comprise a homology to a first genomic sequence and a second genomic sequence, respectively flanking a double-strand break.
[0607] 3 , 1 Nucleases
[0608] The presently disclosed methods and compositions utilize nucleases to cleave nuclease recognition sequences, generating a double strand break, thereby allowing the integration of an exogenous nucleic acid sequence, i.e., a replacement nucleic acid sequence described herein, into a genomic locus via HDR. Non-limiting examples of nucleases useful in the present disclosure include engineered meganucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), compact TALENs, megaTALs, and CRISPR system nucleases.
[0609] Engineered Meganucleases
[0610] In one aspect, the methods and compositions described herein use an engineered meganuclease to generate a DSB in the genome. Meganucleases are homing endonucleases that recognize between 12 to 40 base pairs, thus conferring stringent site specificity. Meganucleases can be engineered to alter DNA-binding specificity, DNA cleavage activity, DNA-binding affinity, or dimerization properties. A non-limiting example of a meganuclease is the I-Crel meganuclease, and engineered meganucleases derived from the I-Crel meganuclease that are modified (e.g., as described above) to alter DNA-binding specificity, DNA cleavage activity, and / or DNA-binding affinity. Methods for producing modified I-Crel meganucleases are known in the art, see, e.g., PCT Publication Nos. WO 2007 / 047859, WO 2017 / 062439, WO 2017 / 062451, and WO 2019 / 200122, incorporated herein by reference. I-Crel-derived engineered meganucleases cleave a 22 base pair recognition sequence in DNA that comprises two half-sites, each comprising nine base pairs, separated by a four base pair center sequence. Cleavage by I-Crel-derived engineered meganucleases occurs between position 13 and 14 of the recognition sequence and results in the formation of a cleavage site having a four base pair, 3’ overhang on each strand. More specifically, on each strand, a cleavage site is formed that comprises positions 1-13 on the 5’ end and positions
[0611] 72
[0612] P89339 2070WO (01288) 14-22 on the 3’ end. The presence of a 3’ overhang is believed to enhance the efficiency of homologous recombination, and thereby contribute to the efficiency of HDR repair and replacement. Thus, in some particular embodiments, the double strand break utilized in HDR repair and replacement comprises a 3’ overhang, and particularly a four base pair 3’ overhang.
[0613] Nucleases referred to as megaTALs are single-chain endonucleases comprising a transcription activator-like effector (TALE) DNA binding domain with an engineered, sequencespecific homing endonuclease.
[0614] Zinc Finger Nucleases (ZFNs)
[0615] In some embodiments, the methods and compositions described herein utilize a zinc finger nuclease (ZFN) to generate a DSB in the genome. Zinc finger nucleases (ZFNs) ZFNs are chimeric proteins comprising a zinc finger DNA-binding domain fused to a nuclease domain derived from an endonuclease or exonuclease, such as the Type II restriction endonuclease Fokl. The zinc finger DNA-binding domain can be a native sequence or can be redesigned through rational or experimental means to produce a protein which binds to a pre-determined DNA sequence of approximately 18 base pairs in length. By fusing the zinc finger DNA-binding domain to the nuclease domain, the resultant chimeric ZFN can be used to cleave the genome at the target site, i.e., the pre-determined DNA sequence. The engineering of ZFNs, particularly the zinc finger DNA-binding domain is further described in U.S. Pat. Nos. 6,453,242; 6,534,261; 6,599,692; 6,503,717; 6,689,558; 7,030,215; 6,794,136; 7,067,317; 7,262,054; 7,070,934; 7,361,635; and 7,253,273, and U.S. Patent Application No. US 2023 / 0295563, all incorporated herein by reference.
[0616] Transcription Activator-Like Effector Nucleases (TALENs)
[0617] In one aspect, the methods and compositions described herein use a transcription activatorlike effector nuclease (TAEEN) to generate a DSB in the genome. TAL-effector nucleases (TALENs) are chimeric proteins comprising a site-specific DNA-binding domain, a TAL effector, fused to an endonuclease or exonuclease, such as the Type II restriction endonuclease Fokl. The DNA-binding domain of TALENs comprises TAL domain repeats, which are 33-34 amino acid sequences with divergent 12th and 13th amino acids. These two positions, referred to as the repeat variable dipeptide (RVD), are highly variable and show a strong correlation with specific nucleotide recognition. Each base pair in the DNA target sequence is contacted by a single TAL repeat with the specificity resulting from the RVD. As the tandem array of TAL-effector domains
[0618] 73
[0619] P89339 2070WO (01288) each recognize a single DNA base pair, the TALEN can be engineered to cleave the genome at the target site conferred by the DNA-binding domain.
[0620] Clustered Regularly Interspaced Short Palindromic Repeats (CRISP R)-Cas Nucleases
[0621] In one aspect, the methods and compositions described herein use a CRISPR-Cas system- nuclease to generate a DSB in the genome. CRISPR / Cas systems comprise two components: (1) a CRISPR nuclease; and (2) a guide RNA (gRNA) comprising a ~20 nucleotide targeting sequence that confers site-specificity to the nuclease. The CRISPR system may further comprise a transactivating RNA (tracrRNA) which complexes with the guide RNA. The gRNA and the tracrRNA, if present, may be the same or separate molecules, referred to as a single guide RNA (sgRNA). Methods for engineering gRNAs to confer target site -specificity are known in the art, and further described in, for example, PCT Publication No. WO 2016 / 182959, incorporated herein by reference.
[0622] There are a number of CRISPR nucleases known in the art that can be used to generate DSB in the genome of a cell, a subject, and / or an organism (see, e.g., Xu and Li (2020) Comput Struct Biotechnol J. 18:2401-2415). In some embodiments, the nuclease is a Cas9 nuclease. Cas9 nucleases comprise a recognition domain (REC) and a nuclease domain. The REC domain is important for recognition of the gRNA, thereby conferring site-specificity. The nuclease domain comprises a RuvC domain, a HNH domain, and a PAM-interacting domain. The RuvC domain is structurally similar to retroviral integrase superfamily members and cleaves a single strand, e.g., the non-complementary strand targeted by the gRNA. The HNH domain shares structural similarity with HNH endonucleases, and cleaves a single strand, e.g., the complementary strand targeted by the gRNA. The PAM-interacting domain interacts with the PAM targeted by the gRNA. Cas9 nucleases are further described in PCT Publication No. WO 2016 / 182959. Non-limiting examples of Cas nucleases Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, CaslO, Cpfl, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, and Cmr6.
[0623] 3 ,2 Exogenous Polynucleotides and Modification Constructs
[0624] The presently disclosed methods for generating a genetically modified cell utilize exogenous polynucleotides comprising homology arms (i.e., modification constructs) to modify a region of the genome (i.e., genomic sequence) of a cell, an organism, and / or a subject. The homology arms comprise a polynucleotide sequence having homology to genomic sequences
[0625] 74
[0626] P89339 2070WO (01288) allowing them to be integrated with their corresponding homologous sequence through recombination.
[0627] The exogenous polynucleotide modification constructs generally comprise the following:
[0628] 5’ [First homology arm] -[second homology arm] 3’.
[0629] The first and / or second homology arms, having homology to a first and second genomic sequence, respectively, may comprise one or more nucleotide modification, such as mismatched nucleotides, nucleotide deletion, and / or nucleotide insertion.
[0630] In some embodiments, the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence.
[0631] In some embodiments, the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence.
[0632] In some embodiments, the first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence, and the second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence.
[0633] The first and second homology arms, which have homology to a first and second genomic sequence, respectively, allow for recombination with the chromosome via HDR. Upon integration of the first and second homology arm into the corresponding endogenous homologous genomic sequences, a nucleotide modification present within the first and / or second homology arm is incorporated into the genome of a cell. The methods disclosed herein are capable of generating a genetically modified cell comprising in its genome a double-stranded break that was generated by a nuclease. Also comprised within the cell is an exogenous polynucleotide comprising from 5’ to 3’, a first homology arm and a second homology arm, wherein the first homology arm has homology to a first genomic sequence that is 5' upstream of the double-strand break, wherein the second homology arm has homology to a second genomic sequence that is 3' downstream of the doublestrand break.
[0634] In some embodiments, the first genomic sequence is adjacent to the double-strand break. In some embodiments, the first genomic sequence is not adjacent to the double-strand break.
[0635] In some embodiments, the second genomic sequence is adjacent to the double-strand break. In some embodiments, the second genomic sequence is not adjacent to the double-strand break.
[0636] 75
[0637] P89339 2070WO (01288) In some embodiments, the double-strand break comprises a 3’ overhang. In some embodiments, the double-strand break comprises a four base pair 3’ overhang. In some embodiments, the double-strand break comprises a blunt end. In some embodiments, the doublestrand break comprises a 5' overhang. In some embodiments, the exogenous polynucleotide comprises, from 5’ to 3’, the first homology arm, a heterologous nucleic acid sequence, and the second homology arm.
[0638] The double-strand break can be positioned within an exon of a gene, an intron of a gene, or a non-protein coding region of a gene.
[0639] In one aspect, the first and second homology arms are approximately the same length. In some embodiments, the first homology arm and the second homology arm are between about 100 to 2000 base pairs in length.
[0640] In some embodiments, the first homology arm and the second homology arm are at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the first homology arm and the second homology arm are about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0641] In some embodiments, the first homology arm and the second homology arm are between about 200 to 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0642] In some embodiments, the first homology arm and the second homology arm are between about 400 to 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm and the second homology arm are between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0643] 76
[0644] P89339 2070WO (01288) In some embodiments, the first homology arm and the second homology arm are about 500 base pairs in length.
[0645] In one aspect, the first and second homology arms are different lengths. In some embodiments, the first homology arm is longer than the second homology arm. In some embodiments, the second homology arm is longer than the first homology arm.
[0646] In some embodiments, the first homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the first homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the first homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600 to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0647] In some embodiments, the first homology arm is between about 200 to 800 base pairs in length. In some embodiments, the first homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the first homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0648] In some embodiments, the first homology arm is between about 400 to 600 base pairs in length. In some embodiments, the first homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the first homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0649] In some embodiments, the first homology arm is about 500 base pairs in length.
[0650] In some embodiments, the second homology arm is between about 100 to 2000 base pairs in length. In some embodiments, the second homology arm is at least about 100, at least about 200, at least about 300, at least about 400, at least about 500, at least about 600, at least about 700, at least about 800, at least about 900, at least about 1000, at least about 1100, at least about 1200, at least about 1300, at least about 1400, at least about 1500, at least about 1600, at least about 1700, at least about 1800, at least about 1900, or at least about 2000 base pairs in length. In some embodiments, the second homology arm is about 100 to 300, 200 to 400, 300 to 500, 400 to 600, 500 to 700, 600
[0651] 77
[0652] P89339 2070WO (01288) to 800, 700 to 900, 800 to 1000, 900 to 1100, 1000 to 1200, 1100 to 1300, 1200 to 1400, 1300 to 1500, 1400 to 1600, 1500 to 1700, 1600 to 1800, 1700 to 1900, or 1800 to 2000 base pairs in length.
[0653] In some embodiments, the second homology arm is between about 200 to 800 base pairs in length. In some embodiments, the second homology arm is at least 200, at least 300 at least 400, at least 500, at least 600, at least 700, or at least 800 base pairs in length. In some embodiments, the second homology arm is between about 200 to 300, 250 to 350, 300 to 400, 350 to 450, 400 to 500, 450 to 550, 500 to 600, 550 to 650, 600 to 700, 650 to 750, or 700 to 800 base pairs in length.
[0654] In some embodiments, the second homology arm is between about 400 to 600 base pairs in length. In some embodiments, the second homology arm is at least at least 400, at least 425, at least 450, at least 475, at least 500, at least 525, at least 550, at least 575, or at least 600 base pairs in length. In some embodiments, the second homology arm is between about 400 to 450, 425 to 475, 450 to 500, 475 to 525, 500 to 550, 525 to 575, or 550 to 600 base pairs in length.
[0655] In some embodiments, the second homology arm is about 500 base pairs in length.
[0656] In some embodiments, the first homology arm has between at least 80% and 100% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 95% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 96% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 97% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 98% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has at least 99% sequence homology to the first genomic sequence. In some embodiments, the first homology arm has 100% sequence homology to the first genomic sequence.
[0657] The first homology arm can have a nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence. In such instances, a first homology arm can comprise a polynucleotide sequence that, other than a nucleotide modification, is the same polynucleotide sequence as the first genomic sequence. In some embodiments, the first homology arm comprises 1 nucleotide modification that differs from the first genomic sequence. In some embodiments, the first homology arm comprises 2 nucleotide modifications that differ from
[0658] 78
[0659] P89339 2070WO (01288) the first genomic sequence. In some embodiments, the first homology arm comprises 3 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 4 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 5 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 6 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 7 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 8 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 9 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 10 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 11 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 12 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 13 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 14 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 15 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 16 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 17 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 18 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 19 nucleotide modifications that differ from the first genomic sequence. In some embodiments, the first homology arm comprises 20 nucleotide modifications that differ from the first genomic sequence.
[0660] In some embodiments, the second homology arm has between at least 80% and 100% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 95% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 96% sequence homology to the second genomic sequence. In
[0661] 79
[0662] P89339 2070WO (01288) some embodiments, the second homology arm has at least 97% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 98% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has at least 99% sequence homology to the second genomic sequence. In some embodiments, the second homology arm has 100% sequence homology to the second genomic sequence.
[0663] The second homology arm can have a nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence. In such instances, a second homology arm can comprise a polynucleotide sequence that, other than a nucleotide modification, is the same polynucleotide sequence as the second genomic sequence. In some embodiments, the second homology arm comprises 1 nucleotide modification that differs from the second genomic sequence. In some embodiments, the second homology arm comprises 2 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 3 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 4 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 5 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 6 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 7 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 8 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 9 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 10 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 11 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 12 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 13 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 14 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 15 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 16 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 17 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm
[0664] 80
[0665] P89339 2070WO (01288) comprises 18 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 19 nucleotide modifications that differ from the second genomic sequence. In some embodiments, the second homology arm comprises 20 nucleotide modifications that differ from the second genomic sequence.
[0666] The first homology arm can have a nucleotide modification that differs from a corresponding endogenous nucleotide of the first genomic sequence. The first homology arm can have at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the first genomic sequence.
[0667] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 1- 10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the first homology arm.
[0668] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 50- 100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275-325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the first homology arm.
[0669] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the first homology arm.
[0670] In some embodiments, the first homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the first homology arm. The first homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700-1900, or 1800-2000 base pairs from the 3' end of the first homology arm.
[0671] The second homology arm can have a nucleotide modification that differs from a corresponding endogenous nucleotide of the second genomic sequence. The second homology arm can have at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of the second genomic sequence.
[0672] 81
[0673] P89339 2070WO (01288) In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 1-10, 5-15, 10-20, 15-25, 20-30, 25-35, 30-40, 35-45, or 40-50 base pairs from the 3' end of the second homology arm.
[0674] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 50-100, 75-125, 100-150, 125-175, 150-200, 175-225, 200-250, 225-257, 250-300, 275- 325, 300-350, 325-375, 350-400, 375-425, 400-450, 425-475, or 450-500 base pairs from the 3' end of the second homology arm.
[0675] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 500-600, 550-650, 600-700, 650-750, 700-800, 750-850, 800-900, 850-950, or 900-1000 base pairs from the 3' end of the second homology arm.
[0676] In some embodiments, the second homology arm comprises the at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of the second homology arm. The second homology arm can comprise the at least one nucleotide modification at a position between 1000-1200, 1100-1300, 1200-1400, 1300-1500, 1400-1600, 1500-1700, 1600-1800, 1700- 1900, or 1800-2000 base pairs from the 3' end of the second homology arm.
[0677] The methods disclosed herein are useful for generating a genetically modified cell comprising, inter alia, an exogenous polynucleotide. In some embodiments, the exogenous polynucleotide comprises a nucleic acid sequence encoding a nuclease. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 5' upstream of the first homology arm. In some embodiments, the nucleic acid sequence encoding the nuclease is positioned 3' downstream of the second homology arm.
[0678] In some embodiments, the nuclease is capable of binding and cleaving the genome of the cell to generate the double -strand break.
[0679] In some embodiments, the exogenous polynucleotide comprises a promoter that is operably linked to the nucleic acid sequence encoding the nuclease.
[0680] The nuclease can be any nuclease known to a person having skill in the art. The nuclease is capable of binding and cleaving the genome of the cell. In some embodiments, the nuclease is an
[0681] 82
[0682] P89339 2070WO (01288) engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, or a compact TALEN.
[0683] In some embodiments, the cell is in vitro. In some embodiments, the cell is in vivo.
[0684] In some embodiments, the exogenous polynucleotide is an mRNA, a single-stranded DNA, or a double-stranded DNA.
[0685] In some embodiments, the exogenous polynucleotide is comprised by a viral genome.
[0686] In some embodiments, the exogenous polynucleotide is comprised by a delivery vehicle.
[0687] In some embodiments, the delivery vehicle is a recombinant virus and the exogenous polynucleotide is comprised by a viral genome.
[0688] In some embodiments, the recombinant virus is a recombinant adeno-associated virus (AAV).
[0689] In some embodiments, the delivery vehicle is a lipid nanoparticle.
[0690] In some embodiments, the exogenous polynucleotide is an mRNA, wherein the mRNA is comprised by the lipid nanoparticle.
[0691] 3,2, 1 Mismatched Nucleotide
[0692] In some embodiments, the nucleotide modification comprises a mismatched nucleotide that differs from a corresponding endogenous nucleotide of the first genomic sequence or the second genomic sequence. The mismatched nucleotide can be present in the first homology, the second homology arm, or both.
[0693] The corresponding endogenous nucleotide can be comprised within a gene coding sequence or a non-gene coding sequence in the genome. In some embodiments, the corresponding endogenous nucleotide is comprised within an exon or a gene coded intron sequence.
[0694] The mismatched nucleotide is a nucleotide that differs from the corresponding endogenous nucleotide that is present within the genomic sequence. In some embodiments, the corresponding endogenous nucleotide is an A, and the mismatched nucleotide is a T, C, or G. In some embodiments, the corresponding endogenous nucleotide is a T, and the mismatched nucleotide is an A, C, or G. In some embodiments, the corresponding endogenous nucleotide is a C, and the mismatched nucleotide is an A, T, or G. In some embodiments, the corresponding endogenous nucleotide is a G, and the mismatched nucleotide is an A, T, or C.
[0695] In some embodiments, the first homology arm comprises between 1-20 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched
[0696] 83
[0697] P89339 2070WO (01288) nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the first homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the first homology arm comprises between 11-13 mismatched nucleotides. In some embodiments, the first homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the first homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the first homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the first homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the first homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the first homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the first homology arm comprises between 18-20 mismatched nucleotides.
[0698] In some embodiments, the first homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the first homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the first homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the first homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the first homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the first homology arm comprises between 8-10 mismatched nucleotides.
[0699] In some embodiments, the first homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 2-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the first homology arm comprises between 4-5 mismatched nucleotides.
[0700] In some embodiments, the first homology arm comprises 1 mismatched nucleotides.
[0701] 84
[0702] P89339 2070WO (01288) In some embodiments, the first homology arm comprises 2 mismatched nucleotides. In some embodiments, the first homology arm comprises 3 mismatched nucleotides. In some embodiments, the first homology arm comprises 4 mismatched nucleotides. In some embodiments, the first homology arm comprises 5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 1-20 mismatched nucleotides. In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 9-11 mismatched nucleotides. In some embodiments, the second homology arm comprises between 10-12 mismatched nucleotides. In some embodiments, the second homology arm comprises between 11- 13 mismatched nucleotides. In some embodiments, the second homology arm comprises between 12-14 mismatched nucleotides. In some embodiments, the second homology arm comprises between 13-15 mismatched nucleotides. In some embodiments, the second homology arm comprises between 14-16 mismatched nucleotides. In some embodiments, the second homology arm comprises between 15-17 mismatched nucleotides. In some embodiments, the second homology arm comprises between 16-18 mismatched nucleotides. In some embodiments, the second homology arm comprises between 17-29 mismatched nucleotides. In some embodiments, the second homology arm comprises between 18-20 mismatched nucleotides.
[0703] In some embodiments, the second homology arm comprises between 1-10 mismatched nucleotides. In some embodiments, the second homology arm comprises between 1-3 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-4 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-6 mismatched nucleotides. In some embodiments, the second homology arm comprises between 5-7 mismatched nucleotides. In some embodiments, the second homology arm comprises between 6-8 mismatched nucleotides. In some embodiments, the second homology arm comprises between 7-9 mismatched
[0704] 85
[0705] P89339 2070WO (01288) nucleotides. In some embodiments, the second homology arm comprises between 8-10 mismatched nucleotides.
[0706] In some embodiments, the second homology arm comprises between 1-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 2-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 3-5 mismatched nucleotides. In some embodiments, the second homology arm comprises between 4-5 mismatched nucleotides.
[0707] In some embodiments, the second homology arm comprises 1 mismatched nucleotides.
[0708] In some embodiments, the second homology arm comprises 2 mismatched nucleotides.
[0709] In some embodiments, the second homology arm comprises 3 mismatched nucleotides.
[0710] In some embodiments, the second homology arm comprises 4 mismatched nucleotides.
[0711] In some embodiments, the second homology arm comprises 5 mismatched nucleotides.
[0712] The corresponding endogenous nucleotide with the first genomic sequence or the second genomic sequence can be a mutant nucleotide.
[0713] In some embodiments, the mismatched nucleotide is a wild-type nucleotide. For instance, when the corresponding endogenous nucleotide within the first genomic sequence and / or the second genomic sequence is a mutant nucleotide, the mismatched nucleotide within the first homology arm or second homology arm, respectively, is a wild-type nucleotide. The nucleotide sequence comprised within the first and / or second homology arm will comprise the same sequence as the first and / or second genomic sequence, other than the one or more mismatched nucleotide.
[0714] In some embodiments, the corresponding endogenous nucleotide is comprised within a gene sequence. The gene sequence can be a gene coding sequence. In some embodiments, the corresponding endogenous nucleotide is comprised by an exon.
[0715] When the endogenous nucleotide comprised by an exon is a mutant nucleotide, the nucleotide can be comprised by a mutant codon. In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0716] A homology arm can comprise a sequence of a corresponding endogenous gene coding sequence, or a portion thereof, comprising the mutant codon. The sequence comprised within the homology arm will differ from the gene coding sequence of the genomic sequence by one or more mismatched nucleotide. In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the mutant codon of the endogenous genomic sequence. That is, the presence of a mismatched nucleotide in place of a corresponding endogenous nucleotide, when integrated into the endogenous sequence, can alter the codon of the gene coding
[0717] 86
[0718] P89339 2070WO (01288) sequence comprising the endogenous nucleotide. The altered codon can encode for a different amino acid when the mismatched nucleotide is incorporated into the genomic sequence following recombination of the homology arm comprising the mismatched nucleotide.
[0719] In some embodiments, the mismatched nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0720] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a wild-type amino acid. When the homology arm is integrated into the genome, the replacement of the mutant nucleotide with the mismatched nucleotide present within the homology arm results in a wild-type codon within the genomic sequence. In some instances, the wild-type codon encodes a wild-type amino acid. Thus, integration of a homology arm with the corresponding endogenous sequence can replace a mutant nucleotide comprised within a mutant codon with a wild-type nucleotide comprised within a wild-type codon, whereby the wild-type codon encodes a wild-type amino acid.
[0721] The replacement of an endogenous corresponding nucleotide with a mismatched nucleotide that is present on a homology arm can also alter the codon comprising the endogenous corresponding nucleotide such that it becomes a stop codon, following integration of the homology arm. A stop codon can be generated at any location of a corresponding endogenous codon. Nucleotide sequences of a stop codon are well known in the art. A skilled artisan would readily understand how to alter one or more corresponding endogenous nucleotide within a genomic sequence with one or more mismatched nucleotides within a homology arm to generate a stop codon within an exon sequence following integration of the homology arm.
[0722] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon. A wild-type codon encodes for a wild-type amino acid.
[0723] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes a different amino acid than the wild-type codon. The mismatched nucleotide can be present at any location within the codon such that it encodes a different amino acid than the wild-type codon when integrated at the corresponding position within the endogenous sequence.
[0724] In some embodiments, the mismatched nucleotide is a mutant nucleotide comprised by a mutant codon. A mutant codon may or may not encode for a different amino acid compared to a wild-type codon.
[0725] In some embodiments, the mismatched nucleotide is comprised by a codon that encodes the wild-type amino acid.
[0726] 87
[0727] P89339 2070WO (01288) In some embodiments, the mismatched nucleotide is comprised by a stop codon. In such instances, a stop codon can replace a wild-type codon within a gene coding sequence following integration of a homology arm. The presence of the one or more mismatched nucleotide within the homology arm serves to replace one or more corresponding endogenous nucleotide following integration onto the genome within the genomic sequence. The replacement, or alteration, of a wild-type codon with a stop codon can result in the production of a shortened, or truncated, polypeptide sequence being generated from the respective coding sequence. A stop codon can replace any wild-type codon within a gene coding sequence to generate any size truncated polypeptide.
[0728] In some embodiments, the corresponding endogenous nucleotide is comprised by an intron. To replace an endogenous nucleotide within an intron sequence of a gene, the first homology arm and / or second homology arm can be engineered with a mismatched nucleotide that differs from a corresponding endogenous nucleotide. Homology arms that comprise a mismatched nucleotide in place of a corresponding endogenous nucleotide may comprise a wild-type intron sequence or a portion thereof, or a mutant intron sequence or a portion thereof. The mismatched nucleotide can be a wild-type nucleotide or a mutant nucleotide.
[0729] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide. For example, a mutant gene sequence may have a mutant nucleotide within an intron sequence of a gene sequence. Removal of a mutant nucleotide from an intron sequence can restore a wild-type intron sequence. The removal of a mutant nucleotide can involve replacement with a mismatched nucleotide, such as a wild-type nucleotide. The mutant nucleotide can be replaced with any nucleotide that restores the wild-type intron sequence. The presence of a mutant nucleotide within an intron sequence can disrupt expression of the gene (e.g., disrupt primary and mature mRNA transcription), resulting in disrupted RNA expression and / or polypeptide production. The mutant nucleotide may be a single nucleotide polymorphism that is replaced with a mismatched nucleotide. In some instances, more than one endogenous mutant nucleotide is replaced with a mismatched nucleotide. For example, an intron sequence may comprise a nucleotide repeat mutant. Replacement of a mutant nucleotide in an intron sequence can restore a wild-type intron sequence.
[0730] In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0731] In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0732] In some embodiments, the mismatched nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0733] 88
[0734] P89339 2070WO (01288) In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide.
[0735] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0736] In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0737] In some embodiments, the mismatched nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0738] In some embodiments, the corresponding endogenous nucleotide is comprised by a nongene protein coding sequence. In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element. In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0739] In some embodiments, the corresponding endogenous nucleotide is comprised by a mutant non-gene protein coding sequence.
[0740] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0741] In some embodiments, the mismatched nucleotide is a wild-type nucleotide.
[0742] In some embodiments, the mismatched nucleotide differs from a wild-type nucleotide.
[0743] In some embodiments, the corresponding endogenous nucleotide is comprised by a wildtype non-gene protein coding sequence.
[0744] In some embodiments, the mismatched nucleotide is a mutant nucleotide.
[0745] The exogenous polynucleotides described herein can be used in the genetic modification of a cell, when introduced into a cell also having a nuclease or gene encoding a nuclease capable of generating a double-stranded break in the cell genome. When expressed in the cell, the nuclease generates a double -stranded break in the genome of the cell allowing for integration of the exogenous polynucleotide into the genome via homologous recombination of the homology arms with their corresponding endogenous homologous sequences, which are designed to flank the double-stranded break, as described herein. For instance, a first genomic sequence and / or a second genomic sequence can be adjacent to the double-stranded break. In some instances, a first genomic sequence and / or a second genomic sequence are not adjacent to the double-stranded break.
[0746] Following homologous recombination, a nucleotide modification that is present within a first homology arm and / or a second homology arm is introduced into the first genomic sequence and / or second genomic sequence, respectively by homologous recombination.
[0747] Disclosed herein are methods of introducing an exogenous polynucleotide into a cell to replace an endogenous nucleotide from an original genomic sequence, such as a single mutant nucleotide or a nucleotide repeat. Following homologous recombination of the exogenous
[0748] 89
[0749] P89339 2070WO (01288) polynucleotide, a mismatched nucleotide present in a homology arm, which is integrated into the genome as a replacement genomic sequence, replaces a corresponding endogenous nucleotide present in an original genomic sequence. Replacement of a mutant nucleotide, for example, within a mutant gene sequence with a wild-type nucleotide can restore a wild-type gene sequence. For instance, a mutant nucleotide may be a single nucleotide polymorphism that alters the coding sequence or expression of a gene and / or a polypeptide encoded by the gene, whereby its replacement with a mismatched nucleotide restores a wild-type coding sequence. Alternatively, the replaced nucleotide may be a wild-type nucleotide within a wild-type gene sequence that is replaced with a mismatched nucleotide so that the replacement genomic sequence comprises a mutant gene sequence. A nucleotide in intron or exon of a gene coding sequence, or from a nongene coding sequence of a genome, can be replaced. Nucleotide replacement using the methods disclosed herein can alter a codon or corresponding amino acid encoded by a codon in an original genomic sequence. In some instances, more than one endogenous nucleotide is replaced from a genomic sequence. For example, a mutant genomic sequence may comprise a mutant endogenous nucleotide that is not present within a wild-type gene (i.e., mutant nucleotide(s)), and / or that is associated with a disease and / or disorder (i.e., a “diseased state”), such as a single nucleotide polymorphism. A gene may comprise a nucleotide repeat mutant associated with a disease or disorder that can be replaced, such that the mutant gene is restored to a wild-type gene in the replacement genomic sequence. A nucleotide is replaced following homologous recombination of the homology arms of an exogenous polynucleotide resulting in a replacement genomic sequence (which comprises a replacement locus) replacing the original genomic sequence. Without wishing to be bound by any theory, a disease and / or disorder may result from a disruption in the structure and / or function of a polypeptide encoded by a mutant allele comprising a mutant nucleotide, resulting in altered gene expression and the production of a mutant polypeptide. A mutant polypeptide exhibits decreased activity relative to wild-type polypeptide, resulting in a diseased state. For example, a polypeptide encoded by a mutant allele may comprise one or more than one mutant amino acid, which alter protein function. A mutant allele may encode a truncated polypeptide that is incapable of performing its wild-type function. The presence of a mutant nucleotide may introduce a pre-mature stop codon into a coding sequence, whereby no polypeptide is encoded, or a truncated version of the original full-length polypeptide is encoded by the coding sequence. Any of the above scenarios could result from the presence of a mutant nucleotide within a gene coding sequence. Replacing a mutant nucleotide in an original sequence can correct any of
[0750] 90
[0751] P89339 2070WO (01288) the situations described above following replacement of the corresponding genomic sequences with the sequences of the homology arms through homologous recombination.
[0752] To replace an endogenous nucleotide from an original genomic sequence, a first homology arm and / or second homology arm are designed to comprise a mismatched nucleotide that differs from a corresponding endogenous nucleotide of a first genomic sequence and / or a second genomic sequence, respectively. In such instances, an endogenous nucleotide that is present within a first genomic sequence and / or a second genomic sequence is replaced with a mismatched nucleotide following homologous recombination of the exogenous polynucleotide. The exogenous polynucleotide introduced into the cell to generate a genetically modified cell can include any exogenous polynucleotide disclosed herein. For instance, the first and / or second homology arms comprised within the exogenous polynucleotide can include any first and / or second homology arms of the present disclosure, whereby following integration of the homology arms into the genome through homologous recombination, the replacement genomic sequence can comprise any modification relative to an original genomic sequence described herein.
[0753] 3,2,2 Nucleotide Deletion
[0754] A nucleotide modification of the presently disclosed compositions and methods can comprise a nucleotide deletion. For instance, a first homology arm, having homology to a first genomic sequence, can lack at least one corresponding endogenous nucleotide that is present within the first genomic sequence. When a nucleotide modification within a first homology arm comprises a nucleotide deletion, it would be understood that that the first homology arm is lacking a corresponding endogenous nucleotide present in the first genomic sequence. Similarly, when the nucleotide modification within a second homology arm comprises a nucleotide deletion, it would be understood that that the second homology arm is lacking a corresponding endogenous nucleotide present in the second genomic sequence. The first and / or second homology arm can comprise a deletion of more than one nucleotide (i.e., the first and / or second arms are lacking more than one corresponding endogenous nucleotide present in the first and / or second genomic sequence, respectively).
[0755] In some embodiments, the first homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some
[0756] 91
[0757] P89339 2070WO (01288) embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 17-29 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0758] In some embodiments, the first homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks
[0759] 92
[0760] P89339 2070WO (01288) between 4-6 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0761] In some embodiments, the first homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the first genomic sequence. In some embodiments, the first homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the first genomic sequence.
[0762] In some embodiments, the first homology arm lacks 1 corresponding endogenous nucleotide that is present in the first genomic sequence.
[0763] In some embodiments, the first homology arm lacks 2 corresponding endogenous nucleotide that are present in the first genomic sequence.
[0764] In some embodiments, the first homology arm lacks 3 corresponding endogenous nucleotide that are present in the first genomic sequence.
[0765] In some embodiments, the first homology arm lacks 4 corresponding endogenous nucleotide that are present in the first genomic sequence.
[0766] In some embodiments, the first homology arm lacks 5 corresponding endogenous nucleotide that are present in the first genomic sequence.
[0767] In some embodiments, the second homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the
[0768] 93
[0769] P89339 2070WO (01288) second genomic sequence. In some embodiments, the second homology arm lacks between 5-7 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 9-11 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 10-12 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 11-13 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 12-14 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 13-15 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 14-16 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 15-17 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 16-18 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 17-29 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 18-20 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0770] In some embodiments, the second homology arm lacks between 1-10 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 1-3 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-4 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-6 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 5-7
[0771] 94
[0772] P89339 2070WO (01288) corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 6-8 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 7-9 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 8-10 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0773] In some embodiments, the second homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 2-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 3-5 corresponding endogenous nucleotides that are present in the second genomic sequence. In some embodiments, the second homology arm lacks between 4-5 corresponding endogenous nucleotides that are present in the second genomic sequence.
[0774] In some embodiments, the second homology arm lacks 1 corresponding endogenous nucleotide that is present in the second genomic sequence.
[0775] In some embodiments, the second homology arm lacks 2 corresponding endogenous nucleotide that are present in the second genomic sequence.
[0776] In some embodiments, the second homology arm lacks 3 corresponding endogenous nucleotide that are present in the second genomic sequence.
[0777] In some embodiments, the second homology arm lacks 4 corresponding endogenous nucleotide that are present in the second genomic sequence.
[0778] In some embodiments, the second homology arm lacks 5 corresponding endogenous nucleotide that are present in the second genomic sequence.
[0779] A corresponding endogenous nucleotide that is absent from (i.e., lacking or missing from) a homology arm can be a mutant nucleotide.
[0780] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0781] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide.
[0782] A gene can comprise more than one mutant nucleotide. A mutant nucleotide can be present at any location within a gene. For instance, a mutant nucleotide can be present in either the noncoding or coding sequence of a gene, or both. A mutant nucleotide can be present in within an intron or exon. When a mutant nucleotide is present within an exon sequence, the exon sequence is
[0783] 95
[0784] P89339 2070WO (01288) understood as a mutant exon sequence. When a mutant nucleotide is present within an intron sequence, the intron sequence is understood as a mutant intron sequence.
[0785] A genomic sequence can comprise more than one mutant nucleotide. As such, a homology arm can lack more than one corresponding endogenous nucleotide that is present in the genomic sequence. For instance, a genomic sequence can comprise a nucleotide repeat sequence. The term “nucleotide repeat” refers to a sequence of nucleotides repeated a number of times in tandem. Many genetic disorders are known to be associated with nucleotide repeats, as described in, Wallace SE, Bean LJH. Resources for Genetics Professionals — Genetic Disorders Caused by Nucleotide Repeat Expansions and Contractions. 2017 Mar 14 [Updated 2022 Oct 20], the contents of which are incorporated herein by reference in their entirety. A nucleotide repeat mutant can be located within any region of a gene, such as the coding sequence.
[0786] A mutant nucleotide comprised within a gene coding sequence of a genomic sequence (e.g., an exon) can be comprised within a mutant codon. A mutant codon may or may not encode for a mutant amino acid. A mutant amino acid refers to an amino acid that differs from the corresponding endogenous amino acid encoded by the corresponding wild-type codon of the wildtype gene sequence. A mutant codon encoding a mutant amino acid will result in the production of a mutant polypeptide. A mutant polypeptide can differ from a wild-type polypeptide by one or more amino acid residue. A skilled artisan would readily understand that a mutant codon comprising one or more mutant nucleotide may encode a mutant or a wild-type amino acid. The exogenous polynucleotides can be used to alter the amino acid sequence of a mutant polypeptide encoded by an endogenous mutant gene coding sequence by removing one or more corresponding endogenous nucleotides from a genomic sequence (e.g., the mutant gene sequence encoding the mutant polypeptide). A skilled artisan would understand that the homology arms of the exogenous polynucleotide are designed based upon the mutant polypeptide encoding endogenous sequence and the specific sequence edits required to correct the mutant polypeptide. As described above, specific edits can include removal of a mutant polynucleotide, removal of a mutant codon, and / or removal of a mutant amino acid. The exogenous polynucleotide may alter one or more codon or amino acid of a corresponding endogenous sequence, by removing at least one mutant nucleotide.
[0787] It would be understood that removal of the one or more than one mutant nucleotide from a variant of the wild-type polynucleotide sequence could restore the variant sequence to a wild-type polynucleotide sequence.
[0788] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an exon sequence. In some embodiments, the exon sequence is a mutant exon sequence.
[0789] 96
[0790] P89339 2070WO (01288) In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
[0791] To remove one or more nucleotide that is present in a mutant gene sequence comprised in a genomic sequence, the first homology arm and / or second homology arm can be engineered to comprise a polynucleotide sequence corresponding to a region around the mutant nucleotide of the endogenous sequence, but without (i.e., lacking) the one or more mutant nucleotide. That is, the first homology arm and / or second homology arm can comprise a polynucleotide sequence that specifically lacks a corresponding endogenous nucleotide that is present in the first and / or second genomic sequence, respectively. In such instances, the homology arm(s) comprise a polynucleotide sequence that, other than lacking the endogenous nucleotide(s) that is to be removed, corresponds to the endogenous genomic nucleotide sequence. An endogenous nucleotide sequence comprised by a genomic sequence can comprise a gene coding sequence. As such, the homology arm can comprise a variant of the corresponding gene coding sequence that differs by the missing corresponding endogenous nucleotide(s). Either or both homology arms may lack an endogenous nucleotide present in a corresponding endogenous genomic sequence, the absence of which results in an altered coding sequence of a gene when the region of the homology arm lacking the nucleotide is inserted into the endogenous genomic sequence. An altered gene coding sequence refers to a change in the position of one or more than one individual nucleotide residue relative to the first nucleotide of the start codon (i.e., position 1 of the coding sequence) such that the codons of the coding sequence are different after the nucleotide alteration compared to before the nucleotide alteration. Such alterations can result in either a reduction of the number of encoded amino acids by at least one or a frame shift. A homology arm that lacks a number of nucleotides that is evenly divisible by 3 (e.g., 3, 6, 9) when compared to the endogenous coding sequence will result in one or more amino acids being removed from the encoded sequence, but the original open reading frame of the gene will remain the same. For example, a mutant endogenous gene that has an additional mutant proline at amino acid position 34 of the mutant endogenous gene can be corrected to the wild type coding sequence by designing either the first or second homology arm to lack the three base pair coding sequence that encodes that proline (e.g., a CCG codon). In this case, the overall open reading frame will remain the same, however, the total length of the encoded protein will be reduced by one amino acid closer to the start codon.
[0792] Alternatively, such alterations may result in a change in the frame of the endogenous coding sequence. For instance, deletion of a nucleotide from a coding sequence can result in one or more nucleotide located 3’ to the deleted nucleotide being shifted to a position closer to the start codon.
[0793] 97
[0794] P89339 2070WO (01288) Moreover, a nucleotide that is located at position 101 of an endogenous coding sequence (i.e., position 2 within the thirty-fourth codon of the endogenous coding sequence) can be shifted to position 100 (i.e., position 1 within the thirty-fourth codon of the coding sequence) when the nucleotide at position 100 in the endogenous sequence is removed. Deletion of more than one nucleotide within a coding sequence can shift nucleotides located 3’ to the deleted nucleotide more than one position. The positional shift of a nucleotide will correspond with the number of nucleotides that are removed from an endogenous coding sequence. The deletion of five nucleotides located at position 100-104 within the coding sequence of a mutant gene sequence, can shift the nucleotide of position 105, and all nucleotides located 3’ to position 105 (e.g., position 106, 107, 108, etc.), five positions closer to the start codon; thus, the nucleotide at position 105 (i.e., the first position of the thirty-fifth codon) is located at position 100 (i.e., position 1 within the thirty-fourth codon of the coding sequence) after deletion the nucleotides at position 100 to 104 of the endogenous coding sequence. This type of frameshift can result in a premature stop codon, restoration of the wild-type coding sequence, or other alteration of the protein sequence encoded by the polynucleotide compared to the mutant endogenous sequence prior to the deletion. As a frameshift can alter the position of one or more nucleotide, codons can be altered by deletion of nucleotides. A codon may be altered to encode a different amino acid. For instance, following the integration of a homology arm lacking a mutant nucleotide comprised by a mutant codon in the endogenous sequence, the mutant codon may be altered to another mutant codon encoding a different amino acid. Alternatively, the mutant codon may be altered to a wild-type codon encoding a different amino acid. Codon sequences and the corresponding amino acids they encode are well established and would be readily understood by a skilled artisan. Thus, in some embodiments described herein the first homology arm and / or the second homology arm does not comprise the mutant nucleotide resulting in a coding sequence that has a codon that encodes a different amino acid than the mutant codon.
[0795] The removal of an endogenous nucleotide from a genomic sequence can remove a mutant codon from an endogenous coding sequence. For instance, the mutant codon can be altered to a wild-type codon. Conversion of a codon (e.g., alteration of a mutant codon to a wild-type codon) refers to the process of when a codon (e.g., mutant codon) of a genomic sequence is altered to a differ codon (e.g., wild-type codon) when the region of the homology arm lacking the one or more corresponding endogenous nucleotide is inserted into the endogenous genomic sequence. One or two nucleotides present within a mutant codon of the endogenous genomic sequence may be present within the newly generated wild-type codon. A skilled artisan would readily understand that
[0796] 98
[0797] P89339 2070WO (01288) the process of determining if a codon present in an endogenous genomic sequence would be altered following integration of a homology arm that does not comprise an endogenous nucleotide involves comparing the corresponding codons in the genomic sequence and homology arm (i.e., comparing the corresponding codons in the genomic sequence prior to and after integration of a homology arm). A corresponding codon of a coding sequence is located the same distance from the start codon. When a mutant codon is altered to a wild-type codon, both codons will be located at the same codon position within a coding sequence relative to the start codon. In some embodiments described herein, up to 20 endogenous nucleotides can be absent from a homology arm to alter at least one mutant codon to a wild-type codon following integration of the homology arm into the endogenous genomic sequence. Thus, in some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a wild-type codon. In some other embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a codon that encodes a wild-type amino acid.
[0798] The removal of one or more nucleotide from a coding sequence can alter the amino acid composition of a polypeptide encoded by the coding sequence following integration of a homology arm. In some instances, removal of a nucleotide at any position within a codon alters whether the codon encodes an amino acid. The removal of a nucleotide can alter the codon from a stop codon to an amino acid-encoding codon. Alternatively, removal of a nucleotide can alter the codon from an amino acid encoding codon to a stop codon. It would be understood that removal of one or more nucleotide from a codon, in addition to altering the codon comprising the nucleotide, may alter all codons within the coding sequence located 3’ to the removed nucleotide (i.e., a frame shift may occur). A mutant polypeptide encoded by an endogenous mutant gene sequence may comprise a mutant amino acid sequence at a portion of the polypeptide and / or may have the endogenous wildtype stop codon removed. This may be due to the presence of a mutant nucleotide that causes a frameshift in an otherwise wild-type coding sequence. The mutant polypeptide may be a truncated protein due to a pre-mature stop codon in the coding sequence. Alternatively, the mutant polypeptide may lack a stop codon that is present in the corresponding codon position of the wildtype gene coding sequence. When the one or more mutant nucleotide that is not present in the homology arm is comprised within a mutant codon in the corresponding endogenous sequence, the frameshift in the coding sequence following integration of the homology arm into the corresponding endogenous sequence can be such that the mutant codon is altered to encode a stop codon. Thus, in some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and comprises a stop codon.
[0799] 99
[0800] P89339 2070WO (01288) In some instances, when a genomic sequence comprises a mutant nucleotide that is comprised within a mutant codon encoding a pre-mature stop codon, the homology arm does not comprise the mutant nucleotide and comprises a wild-type codon. Thus, when the region of the homology arm is recombined with the corresponding endogenous genomic region, the stop codon is removed in the endogenous genomic sequence and the wild-type coding sequence is restored. Alternatively, a mutant polypeptide may comprise one or more than one mutant amino acid inserted within a corresponding wild-type polypeptide sequence. For example, the wild-type and mutant polypeptide differ by only the presence of one or more mutant amino acids. To remove a mutant amino acid from a mutant polypeptide, a homology arm can be designed to lack a mutant codon. In such instances, the homology arm(s) comprise a polynucleotide sequence that, other than lacking the endogenous nucleotide(s) of the mutant codon, corresponds to the endogenous nucleotide sequence. Thus, in some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide and the mutant stop codon is removed from the resulting coding sequence.
[0801] In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
[0802] The absence of a corresponding endogenous nucleotide in a homology arm comprising homology to a genomic sequence can alter a wild-type codon comprising the endogenous nucleotide following homologous recombination of the homology arm with the genomic sequence. As described above, removal of a nucleotide, such as a wild-type nucleotide, from a gene coding sequence may alter the codon comprising the removed nucleotide and / or the gene coding sequence. The impact of a nucleotide deletion on its corresponding codon is dependent, in part, on the nucleotide present in the corresponding position of the removed nucleotide. An altered codon within the replacement genomic sequence present on the genome following integration of a homology arm may encode a different amino acid than the corresponding codon in the original genomic sequence (i.e., original codon). Removal of a wild-type nucleotide from a wild-type codon may result in an altered codon that is a mutant codon. The mutant codon may encode the same amino acid as the wild-type codon, or it may encode a different amino acid, depending on the nucleotide present at the position of the removed wild-type nucleotide within the replacement codon compared to the original codon. In some instances, the altered replacement codon encodes a wild-type amino acid. In some instances, the altered replacement codon encodes a mutant amino acid. In some instances, the amino acid encoded by the altered replacement codon is different than the amino acid encoded by the original wild-type codon. One having skill in the art would readily
[0803] 100
[0804] P89339 2070WO (01288) understand that the amino acid encoded by a codon will be dependent upon the nucleotides and their locations within a codon. An endogenous nucleotide present within a genomic sequence and absent from a homology arm can be a wild-type nucleotide comprised within a gene coding sequence (e.g., an exon), and as such can be comprised within a wild-type codon. Thus, in some embodiments, the first homology arm and / or the second homology arm does not comprise the wildtype nucleotide and comprises a mutant codon that encodes a different amino acid than the wildtype codon. In some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a codon that encodes the same amino acid as the wild-type codon.
[0805] A skilled artisan would readily understand that determining if wild-type codon of an original coding sequence is an altered codon in a replacement coding sequence following homologous recombination of a homology arm into a genome involves comparing the corresponding codons in the original genomic sequence and replacement genomic sequence. As a homology arm comprises the homologous sequence that will be incorporated into the genome following homologous recombination (i.e., replacement genomic sequence). A corresponding codon of a coding sequence is located the same distance from the start codon. In a particular instance, a wild-type nucleotide is removed from a genomic sequence (i.e., the first and / or second homology arm does not comprise the one or more wild-type nucleotide) to introduce a stop codon within the coding sequence. That is, the original wild-type codon that comprised the wild-type nucleotide is altered to encode a stop codon following integration of the homology arm which lacks the wild-type nucleotide. To introduce a stop codon into a coding sequence, a homology arm can be engineered to lack a nucleotide in a codon such that the codon in the homology arm comprises three nucleotides encoding a stop codon. Nucleotide sequences of a codon encoding a stop codon are well known in the art. The insertion of a stop codon may generate a shortened, or truncated protein. Thus, in some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide and comprises a stop codon.
[0806] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises an intron sequence. In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide. When a mutant nucleotide is present within a non-protein coding sequence of a gene, the presence of the mutant nucleotide may disrupt expression of the gene (e.g., disrupt primary and mature mRNA), resulting in disrupted RNA expression and / or polypeptide production. The mutant nucleotide may be comprised for instance in an intron sequence of a gene (i.e., a mutant intron sequence). The mutant intron sequence may comprise one or more than one nucleotide that is
[0807] 101
[0808] P89339 2070WO (01288) not present within a wild-type intron sequence, and the presence of which is associated with a diseased state.
[0809] In some embodiments, the mutant nucleotide is comprised by a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence. For instance, a mutant nucleotide may be located between two nucleotides of the splice donor sequence site within a mutant splice donor sequence. A mutant nucleotide can be present anywhere within a splice donor sequence, splice acceptor sequence, or splice enhancer sequence. The mutant nucleotide may disrupt the function of the splice donor sequence, splice acceptor sequence, and / or splice enhancer sequence, such that efficacy and / or efficiency of splicing of the coding sequence. In some instances, splicing of the coding sequence is abolished when a mutant nucleotide is present in a splice acceptor sequence, splice donor sequence, and / or splice enhancer sequence.
[0810] The homology arms can lack a corresponding endogenous nucleotide of a genomic sequence that is present within an intron sequence. For example, a mutant gene sequence may have a mutant nucleotide within an intron sequence. A mutant intron sequence may comprise a single nucleotide polymorphism or a nucleotide repeat mutation. Removal of a mutant nucleotide from an intron sequence can correct the mutant intron sequence to a wild-type intron sequence. In some instances, more than one endogenous nucleotide is removed from an intron sequence of a genomic sequence, and thus is absent from a homology arm having homology to a first genomic sequence comprising the intron sequence. The exogenous polynucleotide can be engineered to remove one or more than one nucleotide from an intron sequence of a gene within a genomic sequence. That is, the first homology arm and / or second homology arm can comprise a nucleotide sequence that is specifically lacking a mutant nucleotide within an intron sequence, such that following homologous recombination, a wild-type intron sequence is integrated as a replacement intron sequence into the genomic sequence in place of the original, mutant intron sequence. To remove a mutant nucleotide, the first homology arm and / or second homology arm is engineered to lack a mutant nucleotide present within the mutant splice acceptor sequence, splice donor sequence, and / or splice enhancer sequence, while having the remaining sequence intact of the mutant splice acceptor sequence, splice donor sequence, and / or splice enhancer sequence, respectively. As such, in some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide. In some embodiments, the first homology arm and / or the second homology arm comprises a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0811] 102
[0812] P89339 2070WO (01288) In some embodiments, the corresponding endogenous nucleotide is a wild-type nucleotide. For instance, a wild-type nucleotide may be removed from an intron sequence of a gene such that a mutant replacement intron sequence is integrated in its place on the genome, following homologous recombination of a homology arm lacking the wild-type nucleotide.
[0813] In some embodiments, the wild-type nucleotide is comprised by a wild-type splice donor sequence, a wild-type splice acceptor sequence, or a wild-type splice enhancer sequence.
[0814] In a particular instance, a first homology arm and / or second homology arm comprise a sequence having homology to an endogenous, original intron sequence, but lack a wild-type nucleotide present within the original intron sequence. The original intron sequence present on the genome may be a wild-type intron sequence. Alternatively, the original intron sequence present on the genome may be a mutant intron sequence. Removal of a wild-type nucleotide from an intron sequence can alter the intron sequence to a mutant intron sequence. In some instances, more than one nucleotide is removed from an intron sequence of a genomic sequence, and thus is absent from a homology arm having homology to a genomic sequence comprising the intron sequence. The exogenous polynucleotide can be engineered to remove any wild-type nucleotide from an intron sequence of a gene within a genomic sequence. That is, the first homology arm and / or second homology arm can comprise a nucleotide sequence that is specifically lacking a mutant nucleotide within an intron sequence, such that following homologous recombination, a mutant intron sequence wherein a wild-type nucleotide has been deleted is integrated as a replacement intron sequence into the genomic sequence in place of the original, wild-type intron sequence. To remove a wild-type nucleotide, the first homology arm and / or second homology arm is engineered to lack a wild-type nucleotide. In some embodiments, the first homology arm and / or second homology arm is engineered to lack a wild-type nucleotide present within the mutant splice acceptor sequence, splice donor sequence, and / or splice enhancer sequence, while having the remaining sequence intact of the mutant splice acceptor sequence, splice donor sequence, and / or splice enhancer sequence, respectively. As such, in some embodiments, the first homology arm and / or the second homology arm does not comprise the wild-type nucleotide.
[0815] In some embodiments, the first homology arm and / or the second homology arm comprises a mutant splice donor sequence, a mutant splice acceptor sequence, or a mutant splice enhancer sequence.
[0816] In some embodiments, the first genomic sequence and / or the second genomic sequence comprises a non-gene coding sequence. In some embodiments, the non-gene protein coding sequence comprises a gene regulatory element.
[0817] 103
[0818] P89339 2070WO (01288) In some embodiments, the gene regulatory element comprises a promoter or a gene enhancer.
[0819] In some embodiments, the corresponding endogenous nucleotide is a mutant nucleotide.
[0820] When a mutant nucleotide is present within a non-gene protein coding sequence, the presence of the mutant nucleotide may disrupt expression of a gene (e.g., disrupt expression amount and / or control) resulting in disrupted RNA expression and / or polypeptide production. The mutant nucleotide may be comprised for instance in a gene regulatory element of a gene (i.e., a mutant gene regulatory element). The mutant gene regulatory element sequence may comprise a mutant nucleotide that is not present within a wild-type sequence, and the presence of which is associated with a diseased state. Removal of a mutant nucleotide from a non-gene protein coding sequence can restore the coding sequence to a wild-type non-gene protein coding sequence.
[0821] In some embodiments, the first homology arm and / or the second homology arm does not comprise the mutant nucleotide. The exogenous polynucleotide can be engineered to remove a mutant nucleotide from a non-gene protein coding sequence within a genomic sequence. That is, the first homology arm and / or second homology arm can comprise a nucleotide sequence that is specifically lacking a mutant nucleotide within a mutant non-gene protein coding sequence, such that following homologous recombination, a wild-type non-gene protein coding sequence is integrated as a replacement sequence into the genomic sequence in place of the original, mutant sequence.
[0822] Alternatively, a wild-type nucleotide may be removed from non-gene protein coding sequence of a gene such that a mutant replacement sequence (i.e, mutant non-gene protein coding sequence) is integrated in its place on the genome, following homologous recombination of a homology arm lacking the wild-type nucleotide. In such instances, the non-gene coding sequence is a wild-type non-gene coding sequence.
[0823] The exogenous polynucleotide can be engineered to remove any wild-type nucleotide from a non-gene protein coding sequence of a gene within a genomic sequence. The wild-type nucleotide can be comprised with a gene regulatory element, for example, a promoter or gene enhancer. The first homology arm and / or second homology arm can comprise a nucleotide sequence having homology to a non-gene protein coding sequence of a gene within a genomic sequence. The first homology arm and / or second homology can specifically lack a wild-type nucleotide within a wildtype non-gene protein coding sequence. In some embodiments, the first homology arm and / or the second homology arm does not comprise at least one corresponding wild-type nucleotide. Following homologous recombination of the first and / or second homology arm, a nucleotide
[0824] 104
[0825] P89339 2070WO (01288) sequence having homology to a non-gene protein coding sequence of a gene within a genomic sequence, but lacking a wild-type nucleotide is integrated into the genomic sequence as a replacement non-gene protein coding sequence in place of the original, wild-type non-gene protein coding sequence.
[0826] The exogenous polynucleotides can be used in the genetic modification of a cell, when introduced into a cell also having a nuclease or gene encoding a nuclease capable of generating a double-stranded break in the cell genome. When expressed in the cell, the nuclease generates a double-stranded break in the genome of the cell allowing for integration of the exogenous polynucleotide into the genome via homologous recombination of the homology arms with their corresponding endogenous homologous sequences, which are designed to flank the double-stranded break, as described herein. For instance, a first genomic sequence and / or a second genomic sequence can be adjacent to the double -stranded break. In some instances, a first genomic sequence and / or a second genomic sequence are not adjacent to the double-stranded break.
[0827] Following homologous recombination, a nucleotide modification that is present within a first homology arm and / or a second homology arm is introduced into the first genomic sequence and / or second genomic sequence, respectively by homologous recombination.
[0828] Disclosed herein are methods of introducing an exogenous polynucleotide into a cell to remove an endogenous nucleotide from an original genomic sequence, such as a single mutant nucleotide or a nucleotide repeat. Following homologous recombination of the exogenous polynucleotide, an endogenous nucleotide present in an original genomic sequence and which is lacking in a homology arm, is deleted in the replacement genomic sequence. Removal of a mutant nucleotide, for example, within a mutant gene sequence can restore a wild-type gene sequence. For instance, a mutant nucleotide may be a single nucleotide polymorphism that alters the coding sequence or expression of a gene and / or a polypeptide encoded by the gene, whereby its removal restores a wild-type coding sequence. Alternatively, the nucleotide may be a wild-type nucleotide within a wild-type gene sequence that is removed so that the replacement genomic sequence comprises a mutant gene sequence. A nucleotide can be removed from an intron or exon of a gene coding sequence, or from a non-gene coding sequence of a genome. Nucleotide deletion using the methods disclosed herein can alter a codon or corresponding amino acid encoded by a codon in an original genomic sequence. In some instances, more than one endogenous nucleotide is removed from a genomic sequence. For example, a mutant genomic sequence may comprise an endogenous nucleotide that is not present within a wild-type gene (i.e., mutant nucleotide(s)), and / or that is associated with a disease and / or disorder (i.e., a “diseased state”), such as a single nucleotide
[0829] 105
[0830] P89339 2070WO (01288) polymorphism. A gene may comprise a nucleotide repeat mutant associated with a disease or disorder that can be removed, such that the mutant gene is restored to a wild-type gene in the replacement genomic sequence. A nucleotide can be deleted following homologous recombination of the homology arms of an exogenous polynucleotide resulting in a replacement genomic sequence (which comprises a replacement locus) replacing the original genomic sequence. Without wishing to be bound by any theory, a disease and / or disorder may result from a disruption in the structure and / or function of a polypeptide encoded by a mutant allele comprising a mutant nucleotide, resulting in altered gene expression and the production of a mutant polypeptide. A mutant polypeptide exhibits decreased activity relative to wild-type polypeptide, resulting in a diseased state. For example, a polypeptide encoded by a mutant allele may comprise one or more than one mutant amino acid, which alter protein function. A mutant allele may encode a truncated polypeptide that is incapable of performing its wild-type function. The presence of a mutant nucleotide may introduce a pre-mature stop codon into a coding sequence, whereby no polypeptide is encoded, or a truncated version of the original full-length polypeptide is encoded by the coding sequence. Any of the above scenarios could result from the presence of a mutant nucleotide within a gene coding sequence. The absence of a corresponding endogenous nucleotide of a genomic sequence from a homology arm can function to correct any of the situations described above following replacement of the corresponding genomic sequences with the sequences of the homology arms through homologous recombination.
[0831] To delete an endogenous nucleotide from an original genomic sequence, a first homology arm and / or second homology arm are designed to lack at least one corresponding endogenous nucleotide of a first genomic sequence and / or a second genomic sequence, respectively. In such instances, an endogenous nucleotide that is present within a first genomic sequence and / or a second genomic sequence is deleted from the genome following homologous recombination of the exogenous polynucleotide. The exogenous polynucleotide introduced into the cell to generate a genetically modified cell can include any exogenous polynucleotide disclosed herein. For instance, the first and / or second homology arms comprised within the exogenous polynucleotide can include any first and / or second homology arms of the present disclosure, whereby following integration of the homology arms into the genome through homologous recombination, the replacement genomic sequence can comprise any modification relative to an original genomic sequence described herein.
[0832] 3,2,3 Nucleotide Insertion
[0833] 106
[0834] P89339 2070WO (01288) A nucleotide modification of the presently disclosed compositions and methods can comprise a nucleotide insertion. In some embodiments, the nucleotide modification comprises a nucleotide insertion, wherein the first homology arm and / or the second homology arm comprises at least one inserted nucleotide that is not present in the first genomic sequence or the second genomic sequence. For instance, a first homology arm, having homology to a first genomic sequence, can comprise at least one inserted nucleotide that is not present in the first genomic sequence. When a nucleotide modification within a first homology arm comprises a nucleotide insertion, it would be understood that that the first homology arm comprises an inserted nucleotide that is not present in the first genomic sequence. Similarly, when the nucleotide modification within a second homology arm comprises a nucleotide insertion, it would be understood that that the second homology arm comprises an inserted nucleotide that is not present in the second genomic sequence. The first and / or second homology arm can comprise an insertion of more than one nucleotide (i.e., the first and / or second arms comprise more than one nucleotide that is not present in the first and / or second genomic sequence, respectively).
[0835] In some embodiments, the first homology arm comprises between 1-20 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 1-3 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 2-4 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 3-5 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homology arm comprises between 4-6 inserted nucleotides that are not present in the first genomic sequence. In some embodiments, the first homolog...
Claims
CLAIMSWhat is claimed is:
1. A cell comprising an exogenous polynucleotide, wherein said cell comprises in its genome a double-strand break, wherein said exogenous polynucleotide comprises, from 5’ to 3’, a first homology arm and a second homology arm, wherein said first homology arm has homology to a first genomic sequence that is 5 ’ upstream of said double-strand break; wherein said second homology arm has homology to a second genomic sequence that is 3 ’ downstream of said double-strand break, and wherein: a) said first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said first genomic sequence; b) said second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said second genomic sequence; or c) said first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said first genomic sequence and said second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said second genomic sequence.
2. The cell of claim 1, wherein said first genomic sequence is adjacent to said double-strand break.
3. The cell of claim 1, wherein said first genomic sequence is not adjacent to said doublestrand break.
4. The cell of any one of claims 1-3, wherein said second genomic sequence is adjacent to said double-strand break.
5. The cell of any one of claims 1-3, wherein said second genomic sequence is not adjacent to said double-strand break.152P89339 2070WO (01288)6. The cell of any one of claims 1-6, wherein said double-strand break comprises a 3’ overhang.
7. The cell of claim 6, wherein said double-strand break comprises a four base pair 3’ overhang.
8. The cell of any one of claims 1-7, wherein said exogenous polynucleotide comprises, from 5’ to 3’, said first homology arm, a heterologous nucleic acid sequence, and said second homology arm.
9. The cell of any one of claims 1-8, wherein said double-strand break is positioned within an exon of a gene, an intron of a gene, or a non-protein coding region of a gene.
10. The cell of any one of claims 1-8, wherein said nucleotide modification comprises a mismatched nucleotide that differs from a corresponding endogenous nucleotide of said first genomic sequence or said second genomic sequence.
11. The cell of claim 10, wherein a) said corresponding endogenous nucleotide is an A, and said mismatched nucleotide is a T, C, or G; b) said corresponding endogenous nucleotide is a T, and said mismatched nucleotide is an A, C, or G; c) said corresponding endogenous nucleotide is a C, and said mismatched nucleotide is an A, T, or G; or d) said corresponding endogenous nucleotide is a G, and said mismatched nucleotide is an A, T, or C.
12. The cell of claim 10 or 11, wherein said first homology arm comprises between 1-20 mismatched nucleotides.
13. The cell of any one of claims 10-12, wherein said first homology arm comprises between 1- 5 mismatched nucleotides.153P89339 2070WO (01288)14. The cell of any one of claims 10-13, wherein said second homology arm comprises between 1-20 mismatched nucleotides.
15. The cell of any one of claims 10-14, wherein said second homology arm comprises between 1-5 mismatched nucleotides.
16. The cell of any one of claims 10-15, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
17. The cell of claim 16, wherein said mismatched nucleotide is a wild-type nucleotide.
18. The cell of any one of claims 10-15, wherein said corresponding endogenous nucleotide is comprised by an exon.
19. The cell of claim 18, wherein said corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
20. The cell of claim 19, wherein said mismatched nucleotide is a wild-type nucleotide comprised by a wild-type codon.
21. The cell of claim 18, wherein said corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.22 The cell of claim 21, wherein said mismatched nucleotide is a mutant nucleotide comprised by a mutant codon.
23. The cell of any one of claims 10 -15, wherein said corresponding endogenous nucleotide is comprised by an intron.
24. The cell of claim 23, wherein said corresponding endogenous nucleotide is a mutant nucleotide.154P89339 2070WO (01288)25. The cell of claim 24, wherein said mismatched nucleotide is a wild-type nucleotide.
26. The cell of claim 23, wherein said corresponding endogenous nucleotide is a wild-type nucleotide.
27. The cell of claim 26, wherein said mismatched nucleotide is a mutant nucleotide.
28. The cell of any one of claims 10-15, wherein said corresponding endogenous nucleotide is comprised by a non-gene protein coding sequence.
29. The cell of claim 28, wherein said corresponding endogenous nucleotide is comprised by a mutant non-gene protein coding sequence.
30. The cell of claim 29, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
31. The cell of claim 29 or claim 30, wherein said mismatched nucleotide is a wild-type nucleotide.
32. The cell of claim 28, wherein said corresponding endogenous nucleotide is comprised by a wild-type non-gene protein coding sequence.
33. The cell of claim 32, wherein said mismatched nucleotide is a mutant nucleotide.
34. The cell of any one of claims 1-9, wherein said nucleotide modification comprises a nucleotide deletion, wherein said first homology arm and / or said second homology arm lacks at least one corresponding endogenous nucleotide that is present in said first genomic sequence or said second genomic sequence.
35. The cell of claim 34, wherein said first homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in said first genomic sequence.155P89339 2070WO (01288)36. The cell of claim 34 or claim 35, wherein said first homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in said first genomic sequence.
37. The cell of any one of claims 34-36, wherein said second homology arm lacks between 1- 20 corresponding endogenous nucleotides that are present in said second genomic sequence.
38. The cell of any one of claims 34-37, wherein said second homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in said second genomic sequence.
39. The cell of any one of claims 34-38, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
40. The cell of claim 39, wherein said first homology arm and / or said second homology arm does not comprise said mutant nucleotide.
41. The cell of any one of claims 1-9, wherein said nucleotide modification comprises a nucleotide insertion, wherein said first homology arm and / or said second homology arm comprises at least one inserted nucleotide that is not present in said first genomic sequence or said second genomic sequence.
42. The cell of claim 41, wherein said first homology arm comprises between 1-20 inserted nucleotides that are not present in said first genomic sequence.
43. The cell of claim 41 or claim 42, wherein said first homology arm comprises between 1-5 inserted nucleotides that are not present in said first genomic sequence.
44. The cell of any one of claims 41-43, wherein said second homology arm comprises between 1-20 inserted nucleotides that are not present in said second genomic sequence.
45. The cell of any one of claims 41-44, wherein said second homology arm comprises between 1-5 inserted nucleotides that are not present in said second genomic sequence.156P89339 2070WO (01288)46. The cell of any one of claims 41-45, wherein said first genomic sequence and / or said second genomic sequence lacks at least one wild-type nucleotide.
47. The cell of claim 46, wherein said inserted nucleotide is said wild-type nucleotide.
48. The cell of any one of claims 1-47, wherein said first homology arm and said second homology arm are approximately the same length.
49. The cell of claim 48, wherein said first homology arm and said second homology arm are between about 100 to 2000 base pairs in length.
50. The cell of claim 48 or claim 49, wherein said first homology arm and said second homology arm are between about 200 to 800 base pairs in length.
51. The cell of any one of claims 48-50, wherein said first homology arm and said second homology arm are between about 400 to 600 base pairs in length.
52. The cell of any one of claims 48-51, wherein said first homology arm and said second homology arm are about 500 base pairs in length.
53. The cell of any one of claims 1-47, wherein said first homology arm and said second homology arm are different lengths.
54. The cell of claim 53, wherein said first homology arm is longer than said second homology arm.
55. The cell of claim 53, wherein said second homology arm is longer than said first homology arm.
56. The cell of any one of claims 53-55, wherein said first homology arm is between about 100 to 2000 base pairs in length.157P89339 2070WO (01288)57. The cell of any one of claims 53-56, wherein said first homology arm is between about 200 to 800 base pairs in length.
58. The cell of any one of claims 53-57, wherein said first homology arm is between about 400 to 600 base pairs in length.
59. The cell of any one of claims 53-58, wherein said first homology arm is about 500 base pairs in length.
60. The cell of any one of claims 53-59, wherein said second homology arm is between about 100 to 2000 base pairs in length.
61. The cell of any one of claims 53-60, wherein said second homology arm is between about 200 to 800 base pairs in length.
62. The cell of any one of claims 53-61, wherein said second homology arm is between about 400 to 600 base pairs in length.
63. The cell of any one of claims 53-62, wherein said second homology arm is about 500 base pairs in length.
64. The cell of any one of claims 1-63, wherein said first homology arm has at least 95% sequence homology to said first genomic sequence.
65. The cell of any one of claims 1-64, wherein said second homology arm has at least 95% sequence homology to said second genomic sequence.
66. The cell of any one of claims 1-65, wherein said first homology arm comprises said at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of said first homology arm.158P89339 2070WO (01288)67. The cell of any one of claims 1-65, wherein said first homology arm comprises said at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of said first homology arm.
68. The cell of any one of claims 1-65, wherein said first homology arm comprises said at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of said first homology arm.
69. The cell of any one of claims 1-65, wherein said first homology arm comprises said at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of said first homology arm.
70. The cell of any one of claims 1-69, wherein said second homology arm comprises said at least one nucleotide modification at a position between 1-50 base pairs from the 5' end of said second homology arm.
71. The cell of any one of claims 1-69, wherein said second homology arm comprises said at least one nucleotide modification at a position between 50-500 base pairs from the 5' end of said second homology arm.
72. The cell of any one of claims 1-69, wherein said second homology arm comprises said at least one nucleotide modification at a position between 500-1000 base pairs from the 5' end of said second homology arm.
73. The cell of any one of claims 1-69, wherein said second homology arm comprises said at least one nucleotide modification at a position between 1000-2000 base pairs from the 5' end of said second homology arm.
74. The cell of any one of claims 1-73, wherein said exogenous polynucleotide comprises a nucleic acid sequence encoding a nuclease.
75. The cell of claim 74, wherein said nuclease is capable of binding and cleaving the genome of said cell to generate said double-strand break.159P89339 2070WO (01288)76. The cell of claim 74 or claim 75, wherein said exogenous polynucleotide comprises a promoter that is operably linked to said nucleic acid sequence encoding said nuclease.
77. The cell of any one of claims 74-76, wherein said nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, or a compact TALEN.
78. The cell of any one of claims 1-77, wherein said cell is in vitro.
79. The cell of any one of claims 1-77, wherein said cell is in vivo.
80. The cell of any one of claims 1-79, wherein said exogenous polynucleotide is an mRNA, a single-stranded DNA, or a double-stranded DNA.
81. The cell of any one of claims 1-79, wherein said exogenous polynucleotide is comprised by a viral genome.
82. The cell of any one of claims 1-81, wherein said exogenous polynucleotide is comprised by a delivery vehicle.
83. The cell of claim 82, wherein said delivery vehicle is a recombinant virus and said exogenous polynucleotide is comprised by a viral genome.
84. The cell of claim 83, wherein said recombinant virus is a recombinant adeno-associated virus (AAV).
85. The cell of claim 82, wherein said delivery vehicle is a lipid nanoparticle.
86. The cell of claim 83, wherein said exogenous polynucleotide is an mRNA, wherein said mRNA is comprised by said lipid nanoparticle.
87. A method for genetically modifying a cell, said method comprising introducing into a cell:160P89339 2070WO (01288)a) an exogenous polynucleotide; and b) a nuclease or a gene encoding a nuclease, wherein said nuclease is expressed in said cell; wherein said nuclease generates a double-strand break in the genome of said cell, wherein said exogenous polynucleotide comprises, from 5' to 3', a first homology arm and a second homology arm, wherein said first homology arm has homology to a first genomic sequence that is 5' upstream to said double-strand break, wherein said second homology arm has homology to a second genomic sequence that is 3' downstream to said double-strand break, wherein: i) said first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said first genomic sequence; ii) said second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said second genomic sequence; or iii) said first homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said first genomic sequence, and said second homology arm comprises at least one nucleotide modification that differs from at least one corresponding endogenous nucleotide of said second genomic sequence; and wherein said nucleotide modification is introduced into said first genomic sequence and / or said second genomic sequence by homologous recombination of said exogenous polynucleotide.
88. The method of claim 87, wherein said first genomic sequence is adjacent to said doublestrand break.
89. The method of claim 87, wherein said first genomic sequence is not adjacent to said doublestrand break.
90. The method of any one of claims 87-89, wherein said second genomic sequence is adjacent to said double-strand break.
91. The method of any one of claims 87-89, wherein said second genomic sequence is not adjacent to said double-strand break.161P89339 2070WO (01288)92. The method of any one of claims 87-91, wherein said double-strand break comprises a 3' overhang.
93. The method of claim 92, wherein said double-strand break comprises a four base pair 3' overhang.
94. The method of any one of claims 87-93, wherein said exogenous polynucleotide comprises, from 5' to 3', said first homology arm, a heterologous nucleic acid sequence, and said second homology arm, wherein said heterologous nucleic acid sequence is inserted into the genome at said double-strand break by homologous recombination.
95. The method of any one of claims 87-94, wherein said double-strand break is positioned within an exon of a gene, an intron of a gene, or a non-protein coding region of a gene.
96. The method of any one of claims 1-95, wherein said nucleotide modification comprises a mismatched nucleotide that differs from a corresponding endogenous nucleotide of said first genomic sequence or said second genomic sequence, wherein said corresponding endogenous nucleotide is replaced with said mismatched nucleotide following homologous recombination of said exogenous polynucleotide.
97. The method of claim 96, wherein a) said corresponding endogenous nucleotide is an A, and said mismatched nucleotide is a T, C, or G; b) said corresponding endogenous nucleotide is a T, and said mismatched nucleotide is an A, C, or G; c) said corresponding endogenous nucleotide is a C, and said mismatched nucleotide is an A, T, or G; or d) said corresponding endogenous nucleotide is a G, and said mismatched nucleotide is an A, T, or C.
98. The method of claim 96 or claim 97, wherein said first homology arm comprises between 1-20 mismatched nucleotides.162P89339 2070WO (01288)99. The method of any one of claims 96-98, wherein said first homology arm comprises between 1-5 mismatched nucleotides.
100. The method of any one of claims 96-99, wherein said second homology arm comprises between 1-20 mismatched nucleotides.
101. The method of any one of claims 96-100, wherein said second homology arm comprises between 1-5 mismatched nucleotides.
102. The method of any one of claims 96-101, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
103. The method of claim 102, wherein said mismatched nucleotide is a wild-type nucleotide.
104. The method of any one of claims 96-101, wherein said corresponding endogenous nucleotide is comprised by an exon.
105. The method of claim 104, wherein said corresponding endogenous nucleotide is a mutant nucleotide comprised by a mutant codon.
106. The method of claim 105, wherein said mismatched nucleotide is a wild-type nucleotide comprised by a wild-type codon.
107. The method of claim 104, wherein said corresponding endogenous nucleotide is a wild-type nucleotide comprised by a wild-type codon.
108. The method of claim 107, wherein said mismatched nucleotide is a mutant nucleotide comprised by a mutant codon.
109. The method of any one of claims 96-101, wherein said corresponding endogenous nucleotide is comprised by an intron.163P89339 2070WO (01288)110. The method of claim 109, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
111. The method of claim 110, wherein said mismatched nucleotide is a wild-type nucleotide.
112. The method of claim 109, wherein said corresponding endogenous nucleotide is a wild-type nucleotide.
113. The method of claim 112, wherein said mismatched nucleotide is a mutant nucleotide.
114. The method of any one of claims 96-101, wherein said corresponding endogenous nucleotide is comprised by a non-gene protein coding sequence.
115. The method of claim 114, wherein said corresponding endogenous nucleotide is comprised by a mutant non-gene protein coding sequence.116 The method of claim 115, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
117. The method of claim 115 or claim 116, wherein said mismatched nucleotide is a wild-type nucleotide.
118. The method of claim 114, wherein said corresponding endogenous nucleotide is comprised by a wild-type non-gene protein coding sequence.
119. The method of claim 118, wherein said mismatched nucleotide is a mutant nucleotide.
120. The method of any one of claims 87-95, wherein said nucleotide modification comprises a nucleotide deletion, wherein said first homology arm and / or said second homology arm lacks at least one corresponding endogenous nucleotide that is present in said first genomic sequence or said second genomic sequence, and wherein corresponding endogenous nucleotide is deleted from the genome following homologous recombination of said exogenous polynucleotide.164P89339 2070WO (01288)121. The method of claim 120, wherein said first homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in said first genomic sequence.
122. The method of claim 120 or claim 121, wherein said first homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in said first genomic sequence.
123. The method of any one of claims 120-122, wherein said second homology arm lacks between 1-20 corresponding endogenous nucleotides that are present in said second genomic sequence.
124. The method of any one of claims 120-123, wherein said second homology arm lacks between 1-5 corresponding endogenous nucleotides that are present in said second genomic sequence.
125. The method of any one of claims 120-124, wherein said corresponding endogenous nucleotide is a mutant nucleotide.
126. The method of claim 125, wherein said first homology arm and / or said second homology arm does not comprise said mutant nucleotide.
127. The method of any one of claims 87-95, wherein said nucleotide modification comprises a nucleotide insertion, wherein said first homology arm and / or said second homology arm comprises at least one inserted nucleotide that is not present in said first genomic sequence or said second genomic sequence, and wherein said inserted nucleotide is introduced into the genome following homologous recombination of said exogenous polynucleotide.
128. The method of claim 127, wherein said first homology arm comprises between 1-20 inserted nucleotides that are not present in said first genomic sequence.
129. The method of claim 127 or claim 128, wherein said first homology arm comprises between 1-5 inserted nucleotides that are not present in said first genomic sequence.165P89339 2070WO (01288)130. The method of any one of claims 127-129, wherein said second homology arm comprises between 1-20 inserted nucleotides that are not present in said second genomic sequence.
131. The method of any one of claims 127-130, wherein said second homology arm comprises between 1-5 inserted nucleotides that are not present in said second genomic sequence.
132. The method of any one of claims 127-131, wherein said first genomic sequence and / or said second genomic sequence lacks at least one wild-type nucleotide.
133. The method of claim 132, wherein said inserted nucleotide is said wild-type nucleotide.
134. The method of any one of claims 87-133, wherein said first homology arm and said second homology arm are approximately the same length.
135. The method of claim 134, wherein said first homology arm and said second homology arm are between about 100 to 2000 base pairs in length.
136. The method of claim 134 or claim 135, wherein said first homology arm and said second homology arm are between about 200 to 800 base pairs in length.
137. The method of any one of claims 134-136, wherein said first homology arm and said second homology arm are between about 400 to 600 base pairs in length.
138. The method of any one of claims 134-137, wherein said first homology arm and said second homology arm are about 500 base pairs in length.
139. The method of any one of claims 87-133, wherein said first homology arm and said second homology arm are different lengths.
140. The method of claim 139, wherein said first homology arm is longer than said second homology arm.166P89339 2070WO (01288)141. The method of claim 139, wherein said second homology arm is longer than said first homology arm.
142. The method of any one of claims 139-141, wherein said first homology arm is between about 100 to 2000 base pairs in length.
143. The method of any one of claims 139-142, wherein said first homology arm is between about 200 to 800 base pairs in length.
144. The method of any one of claims 139-143, wherein said first homology arm is between about 400 to 600 base pairs in length.
145. The method of any one of claims 139-144, wherein said first homology arm is about 500 base pairs in length.
146. The method of any one of claims 139-145, wherein said second homology arm is between about 100 to 2000 base pairs in length.
147. The method of any one of claims 139-146, wherein said second homology arm is between about 200 to 800 base pairs in length.
148. The method of any one of claims 139-147, wherein said second homology arm is between about 400 to 600 base pairs in length.
149. The method of any one of claims 139-148, wherein said second homology arm is about 500 base pairs in length.
150. The method of any one of claims 87-149, wherein said first homology arm has at least 95% sequence homology to said first genomic sequence.
151. The method of any one of claims 87-150, wherein said first homology arm has at least 95% sequence homology to said first genomic sequence.167P89339 2070WO (01288)152. The method of any one of claims 87-151, wherein said first homology arm comprises said at least one nucleotide modification at a position between 1-50 base pairs from the 3' end of said first homology arm.
153. The method of any one of claims 87-151, wherein said first homology arm comprises said at least one nucleotide modification at a position between 50-500 base pairs from the 3' end of said first homology arm.
154. The method of any one of claims 87-151, wherein said first homology arm comprises said at least one nucleotide modification at a position between 500-1000 base pairs from the 3' end of said first homology arm.
155. The method of any one of claims 87-151, wherein said first homology arm comprises said at least one nucleotide modification at a position between 1000-2000 base pairs from the 3' end of said first homology arm.
156. The method of any one of claims 87-155, wherein said second homology arm comprises said at least one nucleotide modification at a position between 1-50 base pairs from the 5' end of said second homology arm.
157. The method of any one of claims 87-155, wherein said second homology arm comprises said at least one nucleotide modification at a position between 50-500 base pairs from the 5' end of said second homology arm.
158. The method of any one of claims 87-155, wherein said second homology arm comprises said at least one nucleotide modification at a position between 500-1000 base pairs from the 5' end of said second homology arm.
159. The method of any one of claims 87-155, wherein said second homology arm comprises said at least one nucleotide modification at a position between 1000-2000 base pairs from the 5' end of said second homology arm.168P89339 2070WO (01288)160. The method of any one of claims 87-159, wherein said nuclease is an engineered meganuclease, a CRISPR-system nuclease, a zinc finger nuclease (ZFN), a TALEN, or a compact TALEN.
161. The method of any one of claims 87-160, wherein said exogenous polynucleotide comprises a nucleic acid sequence encoding said nuclease.
162. The method of claim 161, wherein said exogenous polynucleotide comprises a promoter that is operably linked to said nucleic acid sequence encoding said nuclease.
163. The method of any one of claims 87-162, wherein said exogenous polynucleotide is introduced into said cell by a recombinant virus, wherein said recombinant virus comprises said exogenous polynucleotide in its viral genome, and wherein said recombinant virus is contacted with said cell.
164. The method of claim 163, wherein said recombinant virus is a recombinant adeno- associated virus (AAV).
165. The method of any one of claims 87-162, wherein said exogenous polynucleotide is introduced into said cell by non-viral delivery.
166. The method of claim 165, wherein said exogenous polynucleotide is comprised by lipid nanoparticles that are contacted with said cell.
167. The method of any one of claims 87-166, wherein said gene encoding said nuclease is introduced into said cell by a recombinant virus, wherein said recombinant virus comprises said gene encoding said nuclease in its viral genome, and wherein said recombinant virus is contacted with said cell.
168. The method of claim 167, wherein said recombinant virus is a recombinant AAV that is contacted with said cell.169P89339 2070WO (01288)169. The method of any one of claims 87-166, wherein said gene encoding said nuclease is introduced into said cell by non-viral delivery.
170. The method of claim 169, wherein said gene encoding said nuclease is comprised by lipid nanoparticles that are contacted with said cell.
171. The method of any one of claims 87-162, wherein said exogenous polynucleotide is introduced into said cell by a recombinant virus that comprises said exogenous polynucleotide in its viral genome, and wherein said gene encoding said nuclease is comprised by lipid nanoparticles, wherein said recombinant virus, and said lipid nanoparticles are contacted with said cell.
172. The method of claim 171, wherein said recombinant virus is a recombinant AAV.
173. The method of any one of claims 87-162, wherein said exogenous polynucleotide is introduced into said cell by a first recombinant virus that comprises said exogenous polynucleotide in its viral genome, and wherein said gene encoding said nuclease is introduced into said cell by a second recombinant virus that comprises said gene encoding said nuclease in its viral genome, wherein said first recombinant nuclease and said second recombinant nucleases are contacted with said cell.
174. The method of claim 173, wherein said first recombinant virus and said second recombinant virus are each a recombinant AAV.
175. The method of any one of claims 87-162, wherein said exogenous polynucleotide is comprised by a first population of lipid nanoparticles, and wherein said gene encoding said nuclease is comprised by a second population of lipid nanoparticles, wherein said first population of lipid nanoparticles, and said second population of lipid nanoparticles are contacted with said cell.
176. The method of any one of claims 87-162, wherein said exogenous polynucleotide is comprised by a first population of non-viral particles, and wherein said gene encoding said nuclease is comprised by a second population of non-viral particles, wherein said first population of non-viral particles and said second population of non-viral particles are contacted with said cell.170P89339 2070WO (01288)177. The method of any one of claims 87-162, wherein said exogenous polynucleotide and said gene encoding said nuclease are introduced into said cell by a recombinant virus that comprises said exogenous polynucleotide and said gene encoding said nuclease in its viral genome, wherein said recombinant virus is contacted with said cell.
178. The method of claim 177, wherein said recombinant virus is a recombinant AAV.
179. The method of any one of claims 87-178, wherein said cell is in vitro.
180. The method of any one of claims 87-178, wherein said cell is in vivo.
181. The method of any one of claims 87-178, wherein said method is a method for genetically modifying a target cell in a subject, wherein said exogenous polynucleotide and said nuclease, or said gene encoding said nuclease, are delivered to said target cell in said subject, and wherein said nucleotide modification is introduced into said first genomic sequence and / or said second genomic sequence by homologous recombination of said exogenous polynucleotide.
182. The method of any one of claims 87-178, wherein said method is a method for treating a subject in need thereof, wherein said subject is administered an effective amount of a) said exogenous polynucleotide; and b) said nuclease, or said gene encoding said nuclease; wherein said exogenous polynucleotide and said nuclease, or said gene encoding said nuclease, are delivered to a target cell in said subject, and wherein said nucleotide modification is introduced into said first genomic sequence and / or said second genomic sequence by homologous recombination of said exogenous polynucleotide.171P89339 2070WO (01288)
Citation Information
Patent Citations
Methods and compositions for targeted cleavage and recombination
US11311574B2
Methods and compositions for using zinc finger endonucleases to enhance homologous recombination
US20030232410A1
Use of chimeric nucleases to stimulate gene targeting
US20050026157A1
Methods and compositions for targeted cleavage and recombination
US20050064474A1
Targeted chromosomal mutagenasis using zinc finger nucleases
US20050208489A1