Serpina-modulating systems and methods
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TESSERA THERAPEUTICS INC
- Filing Date
- 2025-09-12
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for integrating longer nucleic acid sequences into genomes are inefficient and lack site specificity, and current treatments for alpha-1 antitrypsin deficiency (AATD) are inadequate, particularly in correcting the SERPINA1 PiZ mutation and addressing both liver and lung disease manifestations.
A gene modifying system comprising a reverse transcriptase (RT) domain and a StlCas9 domain, along with a template RNA with a variant gRNA scaffold, is used to correct the SERPINA1 PiZ mutation by inserting, altering, or deleting sequences, thereby modulating alpha-1 antitrypsin (AAT) activity.
The system effectively corrects the SERPINA1 PiZ mutation, potentially restoring normal AAT levels and activity, thereby mitigating liver and lung disease symptoms associated with AATD.
Smart Images

Figure US2025046291_07052026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: 2017469-0049SERPINA-MODULATING SYSTEMS AND METHODSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 694,517, filed September 13, 2024, U.S. Provisional Patent Application No. 63 / 694,731, filed September 13, 2024, and U.S. Provisional Patent Application No. 63 / 804,434, filed May 12, 2025, the titles of all of which are “SERPINA-MODULATING SYSTEMS AND METHODS” and the contents of which are incorporated herein by reference in their entirety.BACKGROUND
[0002] Integration of a nucleic acid of interest into a genome occurs at low frequency and with little site specificity, in the absence of a specialized protein to promote the insertion event. Some existing approaches, like CRISPR / Cas9, are more suited for small edits that rely on host repair pathways and are less effective at integrating longer sequences. Other existing approaches, like Cre / loxP, require a first step of inserting a loxP site into the genome and then a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, altering, or deleting sequences of interest in a genome.
[0003] Alpha- 1 antitrypsin deficiency (AATD) is characterized by low circulating levels of alpha- 1 antitrypsin (AAT). AAT is produced primarily in liver cells and secreted into the blood, but it is also made by other cell types including lung epithelial cells and certain white blood cells. AAT inhibits several serine proteases secreted by inflammatory cells (most notably neutrophil elastase [NE], proteinase 3, and cathepsin G) and thus protects organs, such as the lung, from protease-induced damage, especially during periods of inflammation.
[0004] The two most common clinical variants of AAT are E264V (PiS) and E342K (PiZ) alleles. The clinical E342K (PiZ) mutation (also referred to as the Z mutation) is caused by a single base pair substitution in the SERPINAgene (referred to as the Z allele) and results in a glutamic acid to lysine mutation at position 342 of AAT. Inheritance of the Z allele is autosomal codominant and more than half of AATD patients harbor at least one copy of the Z allele.
[0005] The E342K mutation leads to structurally unstable and / or inactive AAT-Z protein that causes toxicity in the liver and is inactive in the lungs. The E342K mutation is located at the hinge between the beta sheet and the Reactive Center Loop (RCL) of the AATPage l of 39612985085vlAttorney Docket No.: 2017469-0049 protein and causes a loop-sheet dimer that can extend to form long chains of loop-sheet polymers. These polymers form aggregates that accumulate inside the rough endoplasmic reticulum of hepatocytes during translation and are therefore not secreted into the bloodstream. Consequently, circulating AAT levels in individuals homozygous for the Z allele (PiZZ) are markedly reduced; only approximately 15% of mutant AAT-Z protein folds correctly and is secreted by the cell. An additional consequence of the Z mutation is that the secreted AAT-Z protein has reduced activity compared to wild-type protein, with 40% to 80% of normal antiprotease activity (American thoracic society / European respiratory society, Am J Respir Crit Care Med. 2003; 168(7): 818-900; and Ogushi et al. J Clin Invest . 1987;80(5): 1366-74, herein incorporated by reference in their entirety).
[0006] There are two disease phenotypes associated with the PiZZ genotype. A gain- of-function phenotype presents as the accumulation of polymerized AAT-Z protein in hepatocytes results in a gain-of-function cytotoxicity that can result in cellular stress, inflammation, fibrosis, cirrhosis, hepatocellular carcinoma (HCC), and neonatal liver disease in 12% of patients. This accumulation may spontaneously remit but can be fatal in a small number of children. A loss-of-function phenotype results from the reduced systemic levels of AAT that lead to increased protease digestion of connective tissue in the lower airway. Excess protease-digestion of the connective tissues and alveolar linings deteriorates lung elasticity and pulmonary function, leading to emphysema, a hallmark of Chronic Obstructive Pulmonary Disease (COPD). This effect is severe in PiZZ individuals and typically manifests in middle age, resulting in a decline in quality of life and shortened lifespan (mean 68 years of age) (Tanash et al. Int J Chron Obstruct Pulm Dis. 2016; 11 : 1663-9, herein incorporated by reference in its entirety). The effect is more pronounced in PiZZ individuals who smoke, resulting in an even further shortened lifespan (58 years)(Piitulainen and Tanash, COPD 2015; 12(l):36-41 , herein incorporated by reference in its entirety). PiZZ individuals account for the majority of patients with clinically relevant AATD lung disease.
[0007] A milder form of AATD is associated with the SZ genotype in which a patient has a Z-allele and an S-allele. The S allele is associated with somewhat reduced levels of circulating AAT but the AAT S-protein is not hepatotoxic. Accordingly, the SZ genotype is associated with clinically significant lung disease but not liver disease. Fregonese and Stolk, Orphanet J Rare Dis. 2008; 33: 16. As with the PiZZ genotype, the deficiency of circulating AAT in subjects with the SZ genotype results in dysregulated protease activity that degrades lung tissue over time and can result in emphysema, particularly in smokers.Page 2 of 39612985085vlAttorney Docket No.: 2017469-0049
[0008] While limited treatment options for AATD exist, there is currently no cure. A small fraction of newborn patients and patients having advanced stage liver disease undergo liver transplant. The current standard of care for AATD patients who have or show signs of significant or developing lung disease is augmentation therapy or protein replacement therapy. Augmentation therapy involves administration (weekly infusion) of a human AAT protein concentrate purified from pooled healthy donor plasma. Although infusions of plasma protein have been shown to improve survival or slow the rate of emphysema progression, augmentation therapy is often insufficient under challenging conditions (e.g., active lung infection). Augmentation therapy also fails to restore the normal physiological regulation of AAT in patients and efficacy has been difficult to demonstrate. In addition, augmentation therapy does not remedy liver disease driven by the toxic gain-of-function of the Z allele. Accordingly, there is a need for new and more effective treatments for AATD.SUMMARY OF THE INVENTION
[0009] This disclosure relates to novel compositions, systems, and methods for altering a genome at one or more locations in a host cell, tissue, or subject, in vivo or in vitro. The disclosure provides, for instance, gene modifying systems that comprise a gene modifying polypeptide comprising a reverse transcriptase (RT) domain and a StlCas9 domain, and a template RNA comprising a variant gRNA scaffold that has been engineered for improved performance, e.g., when used in concert with the StlCas9 domain. The disclosure also provides gene modifying systems that are capable of modulating (e.g., inserting, altering, or deleting sequences of interest) alpha- 1 antitrypsin (AAT) activity and methods of treating alpha-1 antitrypsin deficiency (AATD) by administering one or more such systems to alter a genomic sequence at a single nucleotide to correct the SERPINA1 PiZ mutation that causes AATD.
[0010] In some embodiments, the present disclosure provides a system for modifying DNAto correct a human SERPINA1 gene mutation that causes AATD comprising (a) a nucleic acid encoding a gene modifying polypeptide capable of target primed reverse transcription, the polypeptide comprising (i) a reverse transcriptase domain and (ii) a StlCas9 nickase that binds DNA and has endonuclease activity, and (b) a template RNA comprising (i) a gRNA spacer that is complementary to a first portion of the human SERPINA1 gene, (ii) a gRNA scaffold that binds the polypeptide, (iii) a heterologous object sequence comprising a mutation region to correct the SERPINA1 gene mutation, and (iv) a primer binding site (PBS) sequence comprising at least 3, 4, 5, 6, 7, or 8 bases of 100% homology to a target Page 3 of 39612985085vlAttorney Docket No.: 2017469-0049DNA strand at the 3' end of the template RNA. In some embodiments, the SERPINA1 gene may comprise an E342K mutation (also referred to as a PiZ mutation). A template RNA sequence may comprise a sequence described herein, e.g., in Table 14, Table 19, Table 21, Table 22, or Table 23.
[0011] A gRNA spacer may comprise at least 15 bases of 100% homology to a target DNA at the 5 ' end of the template RNA. A template RNA may further comprise a PBS sequence comprising at least 5 bases of at least 80% homology to a target DNA strand. A template RNA may comprise one or more chemical modifications.
[0012] Domains of a gene modifying polypeptide may be joined by a peptide linker. A polypeptide may comprise one or more peptide linkers. A gene modifying polypeptide may further comprise a nuclear localization signal. A polypeptide may comprise more than one nuclear localization signal, e.g., multiple adjacent nuclear localization signals or one or more nuclear localization signals in different regions of the polypeptide, e.g., one or more nuclear localization signals in the N-terminus of the polypeptide and one or more nuclear localization signals in the C-terminus of the polypeptide. A nucleic acid encoding a gene modifying polypeptide may encode one or more intein domains.
[0013] Introduction of a system of the present disclosure into a target cell may result in insertion of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs of exogenous DNA. Introduction of a system of the present disclosure into a target cell may result in a deletion, wherein the deletion is less than 2, 3, 4, 5, 10, 50, or 100 base pairs of genomic DNA upstream or downstream of an insertion. Introduction of a system of the present disclosure into a target cell may result in substitution, e.g., substitution of 1, 2, or 3 nucleotides, e.g., consecutive nucleotides.
[0014] A heterologous object sequence may be at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 base pairs.
[0015] In some embodiments, the present disclosure provides a pharmaceutical composition comprising a system described herein and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle. In some embodiments, the present disclosure provides a pharmaceutical composition comprising a system described herein and multiple pharmaceutically acceptablePage 4 of 39612985085vlAttorney Docket No.: 2017469-0049 excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle, e.g., where the system described above is delivered by two distinct excipients or carriers, e.g., two lipid nanoparticles, two viral vectors, or one lipid nanoparticle and one viral vector. A viral vector may be an adeno-associated virus (AAV).
[0016] In some embodiments, the disclosure relates to a host cell (e.g., a mammalian cell, e.g., a human cell) comprising a system described herein.
[0017] In some embodiments, the present disclosure provides a method of correcting a mutation in the human SERPINA1 gene in a cell, tissue or subject, the method comprising administering a system described herein to the cell, tissue or subject, wherein optionally the correction of the mutant SERPINA1 gene comprises an amino acid substitution of K342E (i.e., reversing the pathogenic E342K mutation). A system described herein may be introduced in vivo, in vitro, ex vivo, or in situ. A nucleic acid of a system described herein may be integrated into the genome of a host cell. In some embodiments, a nucleic acid of a system described herein is not integrated into the genome of a host cell. In some embodiments, a heterologous object sequence is inserted at only one target site in a host cell genome. A heterologous object sequence may be inserted at two or more target sites in a host cell genome, e.g., at the same corresponding site in two homologous chromosomes or at two different sites on the same or different chromosomes. A heterologous object sequence may encode a mammalian polypeptide, or a fragment or a variant thereof. Components of a system described herein may be delivered on 1, 2, 3, 4, or more distinct nucleic acid molecules. A system of the present disclosure may be introduced into a host cell by electroporation or by using at least one vehicle selected from a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle.
[0018] In some embodiments, the present disclosure provides a gene modifying system comprising: (a) gene modifying polypeptide-encoding polynucleotide comprising, (i) a 5’ untranslated region (UTR) sequence, (ii) a polynucleotide sequence encoding a first nuclear localization signal (NLS), (iii) a polynucleotide sequence encoding an Stl Cas9 domain, (iv) a polynucleotide encoding a peptide linker, (v) a polynucleotide encoding a reverse transcriptase (RT) domain, (vi) a polynucleotide encoding a second NLS, (vii) a 3’ UTR sequence, and (viii) a poly(A) tail sequence; and (b) a template ribonucleotide comprising, (i) a gRNA spacer, (ii) a gRNA scaffold, (iii) a heterologous object sequence comprising a mutation region for correcting a mutation in the human SERPINA1 gene, and Page 5 of 39612985085vlAttorney Docket No.: 2017469-0049(iv) a primer binding site (PBS) sequence complementary to a portion of the human SERPINA1 gene.
[0019] In some embodiments, an RT domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a an amino acid sequence according to SEQ ID NO: 20, SEQ ID NO: 54, SEQ ID NO: 46, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 47, SEQ ID NO: 61, or SEQ ID NO: 62.
[0020] In some embodiments, a 5’ UTR sequence comprises a nucleotide sequence at least at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 13.
[0021] In some embodiments, a polynucleotide encoding the first NLS encodes a modified c-myc NLS and comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 85.
[0022] In some embodiments, a polynucleotide encoding the STI Cas9 domain encodes a modified Stl Cas9 domain comprising an N955A mutation, and wherein the polynucleotide encoding the STI Cas9 domain comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 115.
[0023] In some embodiments, a polynucleotide encoding the linker comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 80, SEQ ID NO: 81, or SEQ ID NO: 82.
[0024] In some embodiments, a polynucleotide encoding the second NLS comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 86.Page 6 of 39612985085vlAttorney Docket No.: 2017469-0049
[0025] In some embodiments, a 3’ UTR sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 14.
[0026] In some embodiments, a poly (A) tail sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 22.
[0027] In some embodiments, a gene modifying polypeptide-encoding polynucleotide further comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 21.
[0028] In some embodiments, a gRNA spacer comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 1 or SEQ ID NO: 6.
[0029] In some embodiments, a gRNA scaffold comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 100, SEQ ID NO: 110, SEQ ID NO: 101, or SEQ ID NO: 111.
[0030] In some embodiments, a heterologous object sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 3, SEQ ID NO: 8, SEQ ID NO: 102, SEQ ID NO: 112, SEQ ID NO: 103, or SEQ ID NO: 113.
[0031] In some embodiments, a PBS sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 4, or SEQ ID NO: 9.Page 7 of 39612985085vlAttorney Docket No.: 2017469-0049
[0032] In some embodiments, the present disclosure provides a gene modifying system comprising: (a) a gene modifying polypeptide comprising, (i) a first nuclear localization signal (NLS), (ii) an Stl Cas9 domain, (iii) a peptide linker, (iv) a reverse transcriptase (RT) domain, and (v) a second NLS; and (b) a template ribonucleotide comprising, (i) a gRNA spacer, (ii) a gRNA scaffold, (iii) a heterologous object sequence comprising a mutation region for correcting a mutation in the human SERPINA1 gene, and (iv) a primer binding site (PBS) sequence complementary to a portion of the human SERPINA1 gene.
[0033] In some embodiments, a first NLS comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 15.
[0034] In some embodiments, an Stl Cas9 domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 12.
[0035] In some embodiments, a peptide linker comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 19, SEQ ID NO: 48, or SEQ ID NO: 49.
[0036] In some embodiments, an RT domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 20, SEQ ID NO: 46, or SEQ ID NO: 47.
[0037] In some embodiments, a second NLS comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 17.
[0038] In some embodiments, a gene modifying system further comprises a lipid delivery vehicle comprising: (i) an ionizable lipid; (ii) a sterol; (iii) a non-cationic lipid; andPage 8 of 39612985085vlAttorney Docket No.: 2017469-0049(iv) a conjugated lipid. In some embodiments, a lipid delivery vehicle comprises: (i) 47.5% (w / w) Lipid 1; (ii) 40% (w / w) cholesterol; (iii) 10% (w / w) DSPC; and (iv) 2.5% DMG- PEG2000. In some embodiments, a lipid delivery vehicle has an N / P ratio of 6.
[0039] In some embodiments, the present disclosure provides a pharmaceutical composition comprising the gene modifying system of any of the preceding embodiments and a pharmaceutically acceptable excipient or carrier.
[0040] In some embodiments, the present disclosure provides a method for modifying a human SERPINA1 gene in a cell comprising contacting the cell with the gene modifying system of pharmaceutical composition of any one of the preceding embodiments, thereby modifying the human SERPINA1 gene.
[0041] In some embodiments, the present disclosure provides a method of treating a subject having a disease or condition associated with a mutation in the human SERPINA1 gene comprising administering to the subject the gene modifying system of pharmaceutical composition of any one of the preceding embodiments, thereby treating the subject having the disease or condition.
[0042] Features of the compositions or methods can include one or more of the following enumerated embodiments.Enumerated Embodiments
[0043] 1. A template RNA (tgRNA) comprising from 5 ’ to 3 ’ :(1) a gRNA spacer comprising the nucleic acid sequence of SEQ ID NO: 1, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(2) a gRNA scaffold comprising the nucleic acid sequence of SEQ ID NO: 2, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(3) a heterologous object sequence comprising the nucleic acid sequence of SEQ ID NO: 3, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto; and(4) a primer binding site (PBS) sequence comprising the nucleic acid sequence of SEQ ID NO: 4, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto.Page 9 of 39612985085vlAttorney Docket No.: 2017469-0049
[0044] 2. A template RNA (tgRNA) comprising from 5 ’ to 3 ’ :(1) a gRNA spacer;(2) a gRNA scaffold;(3) a heterologous object sequence; and(4) a primer binding site (PBS) sequence; wherein the template RNA comprises the nucleic acid sequence of a template RNA of SEQ ID NO: 5, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0045] 3. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 1.
[0046] 4. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 2.
[0047] 5. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 3.
[0048] 6. The template RNA of any one of the preceding embodiments, wherein the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 4.
[0049] 7 The template RNA of any one of the preceding embodiments, wherein:(i) the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 1;(ii) the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 2;(iii) the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 3; and(iv) the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 4.
[0050] 8. The template RNA of any one of the preceding embodiments, wherein template RNA comprises the nucleic acid sequence of SEQ ID NO: 5.
[0051] 9. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer consists of the nucleic acid sequence of SEQ ID NO: 1.
[0052] 10. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold consists of the nucleic acid sequence of SEQ ID NO: 2.Page 10 of 39612985085vlAttorney Docket No.: 2017469-0049
[0053] 11. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence consists of the nucleic acid sequence of SEQ ID NO: 3.
[0054] 12. The template RNA of any one of the preceding embodiments, wherein the primer binding site (PBS) sequence consists of the nucleic acid sequence of SEQ ID NO: 4.
[0055] 13. The template RNA of any one of the preceding embodiments, which does not comprise any nucleotides situated between the gRNA spacer and the gRNA scaffold.
[0056] 14. The template RNA of any one of the preceding embodiments, which does not comprise any nucleotides situated between the gRNA scaffold and the heterologous object sequence.
[0057] 15. The template RNA of any one of the preceding embodiments, which does not comprise any nucleotides situated between the heterologous object sequence and the PBS sequence.
[0058] 16. The template RNA of any one of the preceding embodiments, wherein:(i) the gRNA spacer consists of the nucleic acid sequence of SEQ ID NO: 1;(ii) the gRNA scaffold consists of the nucleic acid sequence of SEQ ID NO: 2;(iii) the heterologous object sequence consists of the nucleic acid sequence of SEQ ID NO: 3; and(iv) the primer binding site (PBS) sequence consists of the nucleic acid sequence of SEQ ID NO: 4.
[0059] 17. The template RNA of any one of the preceding embodiments, wherein template RNA consists of the nucleic acid sequence of SEQ ID NO: 5.
[0060] 18. The template RNA of any one of the preceding embodiments, wherein the template RNA comprises one or more chemically modified nucleotides.
[0061] 19. The template RNA of any one of the preceding embodiments, wherein the template RNA comprises one or more 2’-O-methyl chemically modified nucleotides.
[0062] 20. The template RNA of any one of the preceding embodiments, wherein the template RNA comprises one or more 2’-fluoro chemically modified nucleotides.
[0063] 21. The template RNA of any one of the preceding embodiments, wherein the template RNA comprises one or more phosphorothioate linkages.Page 11 of 39612985085vlAttorney Docket No.: 2017469-0049
[0064] 22. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer comprises one or more chemically modified nucleotides.
[0065] 23. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises one or more chemically modified nucleotides.
[0066] 24. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises one or more chemically modified nucleotides.
[0067] 25. The template RNA of any one of the preceding embodiments, wherein thePBS sequence comprises one or more chemically modified nucleotides.
[0068] 26. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer comprises a chemically modified nucleotide at one or more (e.g., 2 or 3) of positions 1, 2, and 3 relative to SEQ ID NO: 1.
[0069] 27. The template RNA of any one of the preceding embodiments, wherein the first three nucleotides at the 5’ end of the template RNA are chemically modified.
[0070] 28. The template RNA of any one of the preceding embodiments, wherein thePBS sequence comprises a chemically modified nucleotide at one or more (e.g., 2 or 3) of positions -6, -7, and -8 relative to SEQ ID NO: 4.
[0071] 29. The template RNA of any one of the preceding embodiments, wherein the last three nucleotides at the 3’ end of the template RNA are chemically modified.
[0072] 30. The template RNA of any one of embodiments 27-29, wherein the chemically modified nucleotide is a 2’-O-methyl chemically modified nucleotide.
[0073] 31. The template RNA any one of embodiments 27-30, wherein the last three nucleotides at the 3’ end of the template RNA are linked to their 5’ adjacent nucleotides by phosphorothioate linkages.
[0074] 32. The template RNA any one of embodiments 27-31, wherein the first three nucleotides at the 5’ end of the template RNA are linked to their 3’ adjacent nucleotides by phosphorothioate linkages.
[0075] 33. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises: a) a Repeat: anti-repeat (RAR) region, wherein optionally the RAR region comprises a RAR lower stem, a RAR upper stem, and an RAR loop (e.g., a tetraloop);Page 12 of 39612985085vlAttorney Docket No.: 2017469-0049 b) a stem-loop 1 (SL1) region that is 3’ of the RAR region, and c) optionally, a stem loop 2 (SL2) region that is optionally 3’ of the SL1 region; wherein the gRNA scaffold comprises one or more chemically modified nucleotides in one or both of the RAR region or the SL1 region.
[0076] 34. The template RNA of embodiment 33, wherein the one or more chemically modified nucleotides are situated in the tetraloop region, optionally wherein the chemically modified nucleotides are 2’-O-methyl nucleotides.
[0077] 35. The template RNA of embodiment 33 or 34, wherein the one or more chemically modified nucleotides are situated in the three nucleotides adjacent (e.g., 5’ adjacent) to the RAR region, optionally wherein the chemically modified nucleotides are 2’- O-methyl nucleotides.
[0078] 36. The template RNA of any one of embodiments 33-35, wherein the one or more chemically modified nucleotides are situated in the three nucleotides preceding the SL1 region, optionally wherein the chemically modified nucleotides are 2’-O-methyl nucleotides.
[0079] 37. The template RNA of any one of embodiments embodiment 33-36, wherein the one or more chemically modified nucleotides are situated in the SL1 region, optionally wherein the chemically modified nucleotides are 2’ -Fluoro nucleotides.
[0080] 38. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at one or more of (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or all of ) positions 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 42, 43, 44, 45, 46, 47, 48, 49, 50, 53, 54, and 55, relative to SEQ ID NO: 2.
[0081] 39. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 12 through 15 (if present), relative to SEQ ID NO: 2.
[0082] 40. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 16 through 18 (if present), relative to SEQ ID NO: 2.
[0083] 41. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 19 through 21 (if present), relative to SEQ ID NO: 2.Page 13 of 39612985085vlAttorney Docket No.: 2017469-0049
[0084] 42. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 22 through 25 (if present), relative to SEQ ID NO: 2.
[0085] 43. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 26 through 29 (if present), relative to SEQ ID NO: 2.
[0086] 44. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 12 through 29 (if present), relative to SEQ ID NO: 2.
[0087] 45. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 42 through 44 (if present), relative to SEQ ID NO: 2.
[0088] 46. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 45 through 47 (if present), relative to SEQ ID NO: 2.
[0089] 47. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 48 through 50 (if present), relative to SEQ ID NO: 2.
[0090] 48. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a chemically modified nucleotide at each of positions 53 through 55 (if present), relative to SEQ ID NO: 2.
[0091] 49. The template RNA of embodiment 38, wherein the chemically modified nucleotide at position 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 42, 43, or 44 (if present) is a 2’-O-methyl chemically modified nucleotide.
[0092] 50. The template RNA of embodiment 38, wherein the chemically modified nucleotide at position 45, 46, 47, 48, 49, 50, 53, 54, or 55 (if present) is a 2’-fluoro chemically modified nucleotide.
[0093] 51. The template RNA of any one of the preceding embodiments, wherein the scaffold comprises a 2’-O-methyl chemically modified nucleotide at each of positions 12 through 29, a 2’-O-methyl chemically modified nucleotide at each of positions 42 through 44,Page 14 of 39612985085vlAttorney Docket No.: 2017469-0049 a 2’ -fluoro chemically modified nucleotide at each of positions 45 through 50, and a 2’ -fluoro chemically modified nucleotide at each of positions 53 through 55, relative to SEQ ID NO: 2.
[0094] 52. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises a chemically modified nucleotide at one or more (e.g., 2 or 3) of positions +4, +5, and +6, relative to SEQ ID NO: 3.
[0095] 53. The template RNA of embodiment 52, wherein the chemically modified nucleotide is a 2’ -fluoro chemically modified nucleotide.
[0096] 54. The template RNA of any one the preceding embodiments, wherein the heterologous object sequence comprises one or more (e.g., 1, 2, or 3) 2’ -fluoro chemically modified nucleotides.
[0097] 55. The template RNA of embodiment 54, wherein the heterologous object sequence comprises a plurality of 2’-fluoro chemically modified nucleotides (e.g., 2, 3, 4, or 5 2’-fluoro chemically modified nucleotides) positioned adjacent to each other.
[0098] 56. The template RNA of embodiment 55, wherein the heterologous object sequence comprises three 2’ -fluoro chemically modified nucleotides positioned adjacent to each other.
[0099] 57. The template RNA of embodiment 55, wherein the plurality of 2’-fluoro chemically modified nucleotides are at least 1, 2, 3, 4, or 5 nucleotides from the 3’ end of the heterologous object sequence.
[0100] 58. The template RNA of embodiment 57, wherein the plurality of 2’-fluoro chemically modified nucleotides are at least 3 nucleotides from the 3’ end of the heterologous object sequence.
[0101] 59. The template RNA of embodiment 57, wherein the plurality of 2’ -fluoro chemically modified nucleotides are exactly 3 nucleotides from the 3’ end of the heterologous object sequence.
[0102] 60. The template RNA of any one of the preceding embodiments, which does not comprise a 2’ -fluoro chemically modified nucleotide at position -1 of the PBS sequence or at position +1 of the heterologous object sequence.
[0103] 61. The template RNA of any one of the preceding embodiments, which does not comprise a 2’ -fluoro chemically modified nucleotide at positions -2 or -1 of the PBS sequence or at positions +1 or +2 of the heterologous object sequence.Page 15 of 39612985085vlAttorney Docket No.: 2017469-0049
[0104] 62. The template RNA of embodiment 33-61, wherein the chemically modified nucleotide in the heterologous object sequence region is a 2’ -fluoro chemically modified nucleotide, optionally wherein the chemically modified nucleotide is situated in the nucleic acid sequence of UCG of the heterologous object sequence.
[0105] 63. The template RNA of any one of the preceding embodiments, wherein the template RNA comprises a chemically modified nucleotide at one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38 or all of) of positions 1, 2, 3, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 63, 64, 65, 66, 67, 68, 69, 70, 71, 74, 75, 76, 90, 91, 92, 101, 102, and 103 relative to SEQ ID NO: 5.
[0106] 64. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 6, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.
[0107] 65. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 7, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.
[0108] 66. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 8, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.
[0109] 67. The template RNA of any one of the preceding embodiments, wherein the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 9, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.
[0110] 68. The template RNA of any one of the preceding embodiments, wherein:(i) the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 6, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto;(ii) the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 7, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto;Page 16 of 39612985085vlAttorney Docket No.: 2017469-0049(iii) the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 8, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto; and(iv) the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 9, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.[OHl] 69. The template RNA of any one of the preceding embodiments, wherein template RNA comprises the nucleic acid sequence of a template RNA of SEQ ID NO: 10, or a sequence having no more than 1, 2, or 3, chemical modification differences relative thereto.
[0112] 70. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 6.
[0113] 71. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 7.
[0114] 72. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 8.
[0115] 73. The template RNA of any one of the preceding embodiments, wherein the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 9.
[0116] 74. The template RNA of any one of the preceding embodiments, wherein:(i) the gRNA spacer comprises the nucleic acid sequence of SEQ ID NO: 6;(ii) the gRNA scaffold comprises the nucleic acid sequence of SEQ ID NO: 7;(iii) the heterologous object sequence comprises the nucleic acid sequence of SEQ ID NO: 8; and(iv) the primer binding site (PBS) sequence comprises the nucleic acid sequence of SEQ ID NO: 9.
[0117] 75. The template RNA of any one of the preceding embodiments, wherein template RNA comprises the nucleic acid sequence of a template RNA of SEQ ID NO: 10.
[0118] 76. The template RNA of any one of the preceding embodiments, wherein the gRNA spacer consists of the nucleic acid sequence of SEQ ID NO: 6.Page 17 of 39612985085vlAttorney Docket No.: 2017469-0049
[0119] 77. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold consists of the nucleic acid sequence of SEQ ID NO: 7.
[0120] 78. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence consists of the nucleic acid sequence of SEQ ID NO: 8.
[0121] 79. The template RNA of any one of the preceding embodiments, wherein the primer binding site (PBS) sequence consists of the nucleic acid sequence of SEQ ID NO: 9.
[0122] 80. The template RNA of any one of the preceding embodiments, wherein:(i) the gRNA spacer consists of the nucleic acid sequence of SEQ ID NO: 6;(ii) the gRNA scaffold consists of the nucleic acid sequence of SEQ ID NO: 7;(iii) the heterologous object sequence consists of the nucleic acid sequence of SEQ ID NO: 8; and(iv) the primer binding site (PBS) sequence consists of the nucleic acid sequence of SEQ ID NO: 9.
[0123] 81. The template RNA of any one of the preceding embodiments, wherein template RNA consists of the nucleic acid sequence of a template RNA of SEQ ID NO: 10.
[0124] 82. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold has a deletion of part or all of Stem loop 2, as compared to a wild-type StlCas9 scaffold, optionally wherein the wild-type StlCas9 scaffold comprises the nucleic acid sequence of SEQ ID NO: 45.
[0125] 83. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a mutation in the tetraloop as compared to a wild-type StlCas9 scaffold.
[0126] 84. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a tetraloop comprising the nucleic acid sequence of SEQ ID NO: 11, or a sequence having no more than 1, 2, or 3 sequence alterations (e.g., substitutions or deletions) relative thereto.
[0127] 85. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a tetraloop comprising the nucleic acid sequence of SEQ ID NO: 11.Page 18 of 39612985085vlAttorney Docket No.: 2017469-0049
[0128] 86. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold comprises a tetraloop consisting of the nucleic acid sequence of SEQ ID NO: 11.
[0129] 87. The template RNA of any one of the preceding embodiments, wherein the gRNA scaffold has a length no more than 70, 68, 66, 64, or 62 nucleotides.
[0130] 88. The template RNA of any one of the preceding embodiments, wherein the heterologous object sequence comprises the nucleic acid sequence of CC at position 1 and 2 of SEQ ID NO: 3.
[0131] 89. The template RNA of any one of the preceding embodiments, wherein the template RNA has a length of about 90, about 95, about 100, about 101, about 102, about 103, about 104, about 105, about 108, about 110, about 115, or about 120 nucleotides.
[0132] 90. The template RNA of any one of the preceding embodiments, wherein the template RNA has a length of 90-95, 95-100, 100-105, 105-110, or 110-120 nucleotides.
[0133] 91. The template RNA according to any one of the preceding embodiments, wherein the mutation region comprises a first region (e.g., a first nucleotide) designed to correct a pathogenic mutation in the SERPINA1 gene and a second region (e.g., a second nucleotide) designed to inactivate a PAM sequence (e.g., a “PAM-kill” mutation).
[0134] 92. The gene modifying system of any one of the preceding embodiments, which further comprises a second strand-targeting gRNA.
[0135] 93. The gene modifying system of embodiment 92, wherein the second strandtargeting gRNA has a “PAM-in orientation” with the template RNA.
[0136] 94. The gene modifying system of embodiment 92, wherein the second strandtargeting gRNA has a “PAM-out orientation” with the template RNA of the gene modifying system.
[0137] 95. The template RNA according to any one of the preceding embodiments, wherein the mutation introduced by the system is a K342E mutation (e.g., to correct a pathogenic E342K mutation) of the SERPINA1 gene.
[0138] 96. An mRNA comprising: a 5’ cap chosen from m7(3'OMeG)(5')ppp(5')(2'OMeA)pG or m7(3’OMeG)(5’)ppp(‘5)m6(2’OMeA)pG; andPage 19 of 39612985085vlAttorney Docket No.: 2017469-0049 a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a Cas domain, a linker, and an RT domain.
[0139] 97. The mRNA of embodiment 96, wherein the Cas domain is an StlCas9 domain.
[0140] 98. An mRNA comprising: a 5’ cap chosen from m7GpppN2’OmeN, m7(3'OMeG)(5')ppp(5')(2'OMeA)pG, or m7(3’OMeG)(5’)ppp(‘5)m6(2’OMeA)pG; and a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a StlCas9 domain, a linker, and an RT domain.
[0141] 99. An mRNA comprising: a WPRE element having the nucleic acid sequence of SEQ ID NO: 21, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, provided that position 412 is G; and a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a StlCas9 domain, a linker, and an RT domain.
[0142] 100. An mRNA comprising: a 5’ UTR comprising the nucleic acid sequence of SEQ ID NO: 13 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a StlCas9 domain, a linker, and an RT domain.
[0143] 101. The mRNA of any one of embodiments 96-100, wherein the mRNA further comprises a polyA region comprising the nucleic acid sequence of SEQ ID NO: 22 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0144] 102. An mRNA comprising: a polyA region comprising the nucleic acid sequence of SEQ ID NO: 22; and a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a StlCas9 domain, a linker, and an RT domain.Page 20 of 39612985085vlAttorney Docket No.: 2017469-0049
[0145] 103. The mRNA of any one of embodiments 99-102, wherein the mRNA further comprises a 5’ cap chosen from m7GpppN2’OmeN, m7(3'OMeG)(5')ppp(5')(2'OMeA)pG , or m7(3’OMeG)(5’)ppp(‘5)m6(2’OMeA)pG.
[0146] 104. The mRNA of any one of embodiments 96-98 or 100-103, wherein the mRNA further comprises a WPRE element having the nucleic acid sequence of SEQ ID NO: 21, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, provided that residue 412 is G.
[0147] 105. The mRNA of any one of embodiments 96-99, 102, or 103, wherein the mRNA further comprises a 5’ UTR comprising the nucleic acid sequence of SEQ ID NO: 13 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0148] 106. An mRNA comprising: a 5’ cap chosen from m7GpppN2’OmeN, m7(3'OMeG)(5')ppp(5')(2'OMeA)pG, or m7(3’OMeG)(5’)ppp(‘5)m6(2’OMeA)pG; and a region encoding a gene modifying polypeptide, wherein the gene modifying polypeptide comprises a Cas domain, a linker, and an RT domain, wherein:(i) the Cas domain is an StlCas9 domain;(ii) the mRNA comprises a WPRE element having the nucleic acid sequence of SEQ ID NO: 21, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, optionally provided that residue 412 is G;(iii) the mRNA comprises a 5’ UTR comprising the nucleic acid sequence of SEQ ID NO: 13; and / or(iv) the mRNA comprises a polyA region comprising the nucleic acid sequence of SEQ ID NO: 22.
[0149] 107. The mRNA of any one of embodiments 96-106, wherein the StlCas9 domain encoded by the mRNA comprises the amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.Page 21 of 39612985085vlAttorney Docket No.: 2017469-0049
[0150] 108. The mRNA of any one of embodiments 96-107, wherein the StlCas9 domain encoded by the mRNA comprises the amino acid sequence of SEQ ID NO: 12, wherein residue 622 of the StlCas9 domain is not an asparagine.
[0151] 109. The mRNA of any one of embodiments 96-108, wherein the StlCas9 domain encoded by the mRNA comprises the amino acid sequence of SEQ ID NO: 12, wherein residue 622 of the StlCas9 domain is an alanine.
[0152] 110. The mRNA of any one of embodiments 96-109, wherein the linker encoded by the mRNA comprises the amino acid sequence of SEQ ID NO: 19, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto.
[0153] 111. The mRNA of any one of embodiments 96-110, which comprises the amino acid sequence of SEQ ID NO: 18, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto, situated between the StlCas9 domain and the RT domain.
[0154] 112. The mRNA of any one of embodiments 96-111, wherein the RT domain encoded by the mRNA is an AVIRE RT domain.
[0155] 113. The mRNA of any one of embodiments 96-112, wherein the RT domain encoded by the mRNA comprises the amino acid sequence of SEQ ID NO: 20, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0156] 114. The mRNA of any one of embodiments 96-113, wherein the mRNA encodes:(i) the amino acid sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto;(ii) the amino acid sequence of SEQ ID NO: 19, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions or deletions) relative thereto; and(iii) the amino acid sequence of SEQ ID NO: 20, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0157] 115. The mRNA of embodiment 114, wherein (i), (ii), and (iii), are within the same polypeptide.Page 22 of 39612985085vlAttorney Docket No.: 2017469-0049
[0158] 116. The mRNA of embodiment 114 or 115, wherein the mRNA encodes a polypeptide comprising the amino acid sequence of SEQ ID NO: 24, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0159] 117. The mRNA of any one of embodiments 114-116, wherein the mRNA comprises the nucleic acid sequence of SEQ ID NO: 23, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0160] 118. The mRNA of any one of embodiments 96-117, wherein the mRNA further encodes a nuclear localization sequence (NLS).
[0161] 119. The mRNA of embodiment 118, wherein the NLS comprises the amino acid sequence of SEQ ID NO: 16, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto, wherein optionally the NLS is situated N- terminal of the Cas domain.
[0162] 120. The mRNA of embodiment 118 or 119, wherein the NLS comprises the amino acid sequence of SEQ ID NO: 15, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto.
[0163] 121. The mRNA of any one of embodiments 118-120, wherein the mRNA further encodes a second NLS.
[0164] 122. The mRNA of embodiment 121, wherein the second NLS comprises the amino acid sequence of SEQ ID NO: 17, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto, wherein optionally the second NLS is situated C-terminal of the RT domain.
[0165] 123. The mRNA of any one of embodiments 96-122, wherein the mRNA comprises a 5’ UTR.
[0166] 124. The mRNA of embodiment 123, wherein the 5’ UTR comprises the nucleic acid sequence of SEQ ID NO: 13, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto.
[0167] 125. The mRNA of embodiment 123 or 124, wherein the 5’ UTR consists of the nucleic acid sequence of SEQ ID NO: 13.
[0168] 126. The mRNA of any one of embodiments 96-125, wherein the mRNA comprises a 3’ UTR.Page 23 of 39612985085vlAttorney Docket No.: 2017469-0049
[0169] 127. The mRNA of embodiment 126, wherein the 3’ UTR comprises the nucleic acid sequence of SEQ ID NO: 14, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto.
[0170] 128. The mRNA of embodiment 126 or 127, wherein the 3’ UTR consists of the nucleic acid sequence of SEQ ID NO: 14.
[0171] 129. The mRNA of any one of embodiments 96-128, wherein the mRNA further encodes an expression element.
[0172] 130. The mRNA of embodiment 129, wherein the expression element is aWPRE element.
[0173] 131. The mRNA of embodiment 129 or 130, wherein the WPRE element comprises the nucleic acid sequence of SEQ ID NO: 21, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0174] 132. The mRNA of any one of embodiments 129-131, wherein position 412 of the WPRE element is not a T.
[0175] 133. The mRNA of any one of embodiments 129-132, wherein position 412 of the WPRE element is a G.
[0176] 134. The mRNA of embodiment 129-133, wherein the WPRE element consists of the nucleic acid sequence of SEQ ID NO: 21.
[0177] 135. The mRNA of any one of embodiments 96-134, wherein the mRNA further comprises a polyA region, e.g., a polyAtail, e.g., at the 3 ’ end of the mRNA.
[0178] 136. The mRNA of embodiment 135, wherein the polyA tail is at least 5, at least 10, at least 20, at least 30, at least 50, at least 70, at least 100, at least 120, at least 125, at least 127, at least 129, at least 131, or at least 135 nucleotides in length.
[0179] 137. The mRNA of embodiment 135 or 136, wherein the polyA region comprises the nucleic acid sequence of SEQ ID NO: 22, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0180] 138. The mRNA of any one of embodiments 135-137, wherein the polyA region consists of the nucleic acid sequence of SEQ ID NO: 22.
[0181] 139. An mRNA encoding a gene modifying polypeptide comprising from 5’ to3’:Page 24 of 39612985085vlAttorney Docket No.: 2017469-0049(a) a 5’ UTR of SEQ ID NO: 13, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(b) a region encoding a first NLS of SEQ ID NO: 15, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(c) a region encoding a Cas domain of SEQ ID NO: 12, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto;(d) a region encoding a linker of SEQ ID NO: 18, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(e) a region encoding a reverse transcriptase (RT) domain of SEQ ID NO: 20, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto;(f) a region encoding a second NLS of SEQ ID NO: 17, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(g) a 3’ UTR of SEQ ID NO: 14, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a sequence having no more than 1, 2, or 3, sequence alterations (e.g., substitutions) relative thereto;(h) an expression element of SEQ ID NO: 21, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and(i) a poly(A) tail of SEQ ID NO: 22, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0182] 140. An mRNA encoding a gene modifying polypeptide comprising from 5’ to3’:(a) a 5’ UTR of SEQ ID NO: 13;(b) a region encoding a first NLS of SEQ ID NO: 15;(c) a region encoding a Cas domain of SEQ ID NO: 12;(d) a region encoding a linker of SEQ ID NO: 18;Page 25 of 39612985085vlAttorney Docket No.: 2017469-0049(e) a region encoding a reverse transcriptase (RT) domain of SEQ ID NO: 20;(f) a region encoding a second NLS of SEQ ID NO: 17;(g) a 3’ UTR of SEQ ID NO: 14;(h) an expression element of SEQ ID NO: 21; and(i) a poly(A) tail of SEQ ID NO: 22.
[0183] 141. The mRNA of any one of embodiments 96-140, wherein the mRNA comprises one or more chemically modified nucleotides.
[0184] 142. The mRNA of embodiment 141, wherein one or more uridine residues of the mRNA are Nl-methyl-pseudouri dine residues.
[0185] 143. The mRNA of embodiment 141 or 142, wherein all uridine residues of the mRNA are Nl-methyl-pseudouri dine residues.
[0186] 144. The mRNA of any one of embodiments 96-143, wherein the mRNA has a length of at least 1000, at least 2000, at least 3000, at least 4000, at least 5000, or at least 6000 nucleotides.
[0187] 145. The mRNA of any one of embodiments 96-144, wherein the mRNA has a length of about 1000, about 2000, about 3000, about 4000, about 5000, about 6000, about 6300, about 6400, about 6500, or about 7000 nucleotides.
[0188] 146. The mRNA of any one of embodiments 96-145, wherein the mRNA has a length of 1000-2000, 2000-3000, 3000-4000, 4000-5000, 5000-6000, 6000-7000 nucleotides.
[0189] 147. The mRNA of any one of embodiments 96-146, wherein the mRNA has a length of no more than about 10000, about 9000, about 8000, about 7000, or about 6500 nucleotides.
[0190] 148. The mRNA of any one of embodiments 96-147, wherein the StlCas9 encoded by the mRNA comprises a nickase domain.
[0191] 149. The mRNA of any one of embodiments 129-134, wherein the mRNA has an enhanced nuclear export activity compared to an otherwise similar mRNA that does not have an WPRE element.Page 26 of 39612985085vlAttorney Docket No.: 2017469-0049
[0192] 150. The mRNA of any one of embodiments 135-138, wherein the polypeptide encoded by the mRNA has an enhanced rewriting activity compared to an otherwise similar mRNA that does not have a polyA domain.
[0193] 151. The mRNA of any one of embodiments 123-128, wherein the polypeptide encoded by the mRNA has an enhanced rewriting activity compared to an otherwise similar mRNA that does not have a UTR.
[0194] 152. A gene modifying system comprising:(a) a template RNA of any one of embodiments 1-95; and(b) an mRNA encoding a gene modifying polypeptide, the gene modifying polypeptide comprising:(1) a Cas domain;(2) a linker; and(3) a reverse transcriptase (RT) domain.
[0195] 153. A gene modifying system comprising:(a) a template RNA (tgRNA) comprising, from 5’ to 3’ :(1) a gRNA spacer;(2) a gRNA scaffold;(3) a heterologous object sequence comprising a mutation region to correct a mutation in a second portion of the human SERPINA1 gene; and(4) a primer binding site (PBS) sequence complementary to a third portion of human SERPINA1; and(b) an mRNA of any one of embodiments 96-151; optionally wherein the gRNA spacer is complementary to a first portion of the human SERPINA1 gene.
[0196] 154. A gene modifying system comprising:(a) a template RNA of any one of embodiments 1-95; and(b) an mRNA of any one of embodiments 96-151.
[0197] 155. A gene modifying system comprising:Page 27 of 39612985085vlAttorney Docket No.: 2017469-0049(a) a template RNA (tgRNA) comprising, from 5’ to 3’ :(1) a gRNA spacer complementary to a first portion of the human SERPINA1 gene;(2) a gRNA scaffold;(3) a heterologous object sequence comprising a mutation region to correct a mutation in a second portion of the human SERPINA1 gene; and(4) a primer binding site (PBS) sequence complementary to a third portion of human SERPINA1; and(b) an mRNA encoding a gene modifying polypeptide, the gene modifying polypeptide comprising:(1) a Cas domain;(2) a linker; and(3) a reverse transcriptase (RT) domain; wherein the mRNA comprises a 5’ cap chosen from m7GpppN2’OmeN, m7(3'OMeG)(5')ppp(5')(2'OMeA)pG, or m7(3’OMeG)(5’)ppp(‘5)m6(2’OMeA)pG.
[0198] 156. A gene modifying system comprising:(a) a template RNA (tgRNA) comprising, from 5’ to 3’ :(1) a gRNA spacer complementary to a first portion of the human SERPINA1 gene;(2) a gRNA scaffold;(3) a heterologous object sequence comprising a mutation region to correct a mutation in a second portion of the human SERPINA1 gene; and(4) a primer binding site (PBS) sequence complementary to a third portion of human SERPINA1; and(b) an mRNA encoding a gene modifying polypeptide, the gene modifying polypeptide comprising:(1) a Cas domain;(2) a linker; andPage 28 of 39612985085vlAttorney Docket No.: 2017469-0049(3) a reverse transcriptase (RT) domain; wherein the mRNA comprises a 5’ UTR, and wherein the 5’UTR comprises the nucleic acid sequence of SEQ ID NO: 13.
[0199] 157. A gene modifying system comprising:(a) a template RNA (tgRNA) comprising, from 5’ to 3’ :(1) a gRNA spacer complementary to a first portion of the human SERPINA1 gene;(2) a gRNA scaffold;(3) a heterologous object sequence comprising a mutation region to correct a mutation in a second portion of the human SERPINA1 gene; and(4) a primer binding site (PBS) sequence complementary to a third portion of human SERPINA1; and(b) an mRNA encoding a gene modifying polypeptide, the gene modifying polypeptide comprising:(1) a Cas domain;(2) a linker; and(3) a reverse transcriptase (RT) domain; wherein the mRNA comprises a polyA region (e.g., a polyA tail), and wherein the polyA region is situated at the 3’ of the mRNA, and optionally wherein the polyA region comprises the nucleic acid sequence of SEQ ID NO: 22.
[0200] 158. The system of any one of embodiments 152-157, wherein the ratio of the template RNA to the mRNA (e.g., weightweight) is about 1 :0.5, about 1 :0.8, about 1 :0.9, about 1 : 1, about 1 : 1.1, about 1 : 1.2, about 1 : 1.5, or about 1 :2.
[0201] 159. The system of embodiment 158, wherein the ratio of the template RNA to the mRNA is about 1 :1.
[0202] 160. The system of embodiment 158, wherein the ratio of the template RNA to the mRNA is about 1 :0.5.
[0203] 161. The system of any one of embodiments 152-160, wherein the system further comprises a second nick gRNA.Page 29 of 39612985085vlAttorney Docket No.: 2017469-0049
[0204] 162. The gene modifying system of any one of embodiments 152-161, wherein the gene modifying system comprises a second strand-targeting gRNA, and wherein correction of the mutation in a population of target cells is increased relative to a population of target cells treated with a gene modifying system comprising a template RNA without a second strand-targeting gRNA.
[0205] 163. The system of any one of embodiments 152-162, which is capable of specifically cleaving a target nucleic acid.
[0206] 164. The system of any one of embodiments 152-163, which is capable of introducing an edit specified by the heterologous object sequence in a target nucleic acid.
[0207] 165. The gene modifying system of any one of embodiments 152-164, wherein introduction of the system into a target cell results in a mutation that causes the restoration of the function of the SERPINA gene.
[0208] 166. The gene modifying system of any one of embodiments 152-165, wherein correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of target nucleic acids.
[0209] 167. The gene modifying system any one of embodiments 152-166, wherein correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of target cells.
[0210] 168. The template RNA, mRNA, or system of any one of the preceding embodiments, wherein the template RNA and / or the mRNA are formulated in a lipid nanoparticle (LNP).
[0211] 169. The template RNA, mRNA, or system of any one of the preceding embodiments, wherein the template RNA, the mRNA, and / or the second nick gRNA are formulated in an LNP.
[0212] 170. The system of embodiment 168 or 169, wherein the LNP composition is formulated at a N:P of about 4: 1 to about 6: 1, e.g., about 4: 1, about 4.5: 1, about 5: 1, about 5.5: 1, or about 6: 1.
[0213] 171. The system of any embodiment 170, wherein the LNP composition is formulated at a N:P of about 4: 1.
[0214] 172. The system of embodiment 170, wherein the LNP composition is formulated at a N:P of about 6: 1.Page 30 of 39612985085vlAttorney Docket No.: 2017469-0049
[0215] 173. A pharmaceutical composition, comprising the template RNA of any one of embodiments 1-95, the mRNA of any one of embodiments 96-151, or the system of any one of embodiments 152-172, and a pharmaceutically acceptable excipient or carrier.
[0216] 174. The pharmaceutical composition of embodiment 173, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle (LNP).
[0217] 175. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising the system, template RNA, or mRNA of any one of the preceding embodiments.
[0218] 176. A method of making the nucleic acid or template RNA of any one of embodiments 1-151, the method comprising synthesizing the template R.NA / / / vitro (e.g., by in vitro transcription or solid-state synthesis).
[0219] 177. A method for modifying a target site (e.g., a target site in the humanSERPINA1 gene) in a cell, the method comprising contacting the cell with the system of any one of embodiments 152-172, or DNA encoding the same, or the pharmaceutical composition of embodiment 173 or 174, thereby modifying the target site.
[0220] 178. A method for treating a subject having a disease or condition associated with a mutation in a gene (e.g., the human SERPINA1 gene), the method comprising administering to the subject the system of any one of embodiments 152-172, or DNA encoding the same, or the pharmaceutical composition of embodiment 173 or 174, thereby treating the subject having a disease or condition.
[0221] 179. The method of embodiment 178, wherein the disease or condition is alpha- 1 antitrypsin deficiency (AATD).
[0222] 180. The method of embodiment 178 or 179, wherein the subject has aE342K mutation.
[0223] 181. A method for treating a subj ect having AATD, the method comprising administering to the subject the system of any one of embodiments 152-172, or DNA encoding the same, or the pharmaceutical composition of embodiment 173 or 174, thereby treating the subject having AATD.
[0224] 182. The method of any one of embodiments 176-181, wherein the contacting occurs ex vivo, e.g., wherein the cell’s or subject’s DNA is modified ex vivo.Page 31 of 39612985085vlAttorney Docket No.: 2017469-0049
[0225] 183. The method of any one of embodiments 176-182, wherein the contacting occurs in vivo, e.g., wherein the cell’s or subject’s DNA is modified in vivo.
[0226] 184. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide, and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 89, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0227] 185. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 89, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0228] 186. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0229] 187. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,Page 32 of 39612985085vlAttorney Docket No.: 2017469-0049(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0230] 188. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 89, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0231] 189. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 89, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0232] 190. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.Page 33 of 39612985085vlAttorney Docket No.: 2017469-0049
[0233] 191. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0234] 192. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 89, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0235] 193. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to 89, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0236] 194. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%,Page 34 of 39612985085vlAttorney Docket No.: 2017469-0049 at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0237] 195. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 24 or SEQ ID NO: 28, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0238] 196. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 91, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0239] 197. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 91, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0240] 198. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 87, and wherein,Page 35 of 39612985085vlAttorney Docket No.: 2017469-0049(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0241] 199. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 87, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0242] 200. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 90, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0243] 201. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 91, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0244] 202. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, orPage 36 of 39612985085vlAttorney Docket No.: 2017469-0049100% sequence identity to an amino acid sequence according to SEQ ID NO: 87, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0245] 203. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 87, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0246] 204. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 91, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0247] 205. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 91, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0248] 206. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,Page 37 of 39612985085vlAttorney Docket No.: 2017469-0049(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to 87, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0249] 207. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 87, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0250] 208. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 93, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0251] 209. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 93, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.Page 38 of 39612985085vlAttorney Docket No.: 2017469-0049
[0252] 210. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0253] 211. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 5 or SEQ ID NO: 10.
[0254] 212. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 93, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0255] 213. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 93, and wherein,Page 39 of 39612985085vlAttorney Docket No.: 2017469-0049(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0256] 214. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0257] 215. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 105 or SEQ ID NO: 107.
[0258] 216. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 93, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0259] 217. A gene modifying system comprising: (A) a nucleic acid encoding a gene modifying polypeptide and (B) a template RNA, wherein,Page 40 of 39612985085vlAttorney Docket No.: 2017469-0049(A) the nucleic acid encoding the gene modifying polypeptide comprises a nucleotide sequence according to SEQ ID NO: 93, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0260] 218. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.
[0261] 219. A gene modifying system comprising: (A) a gene modifying polypeptide and (B) a template RNA, wherein,(A) the gene modifying polypeptide comprises an amino acid sequence according to SEQ ID NO: 88, and wherein,(B) the template RNA comprises a nucleotide sequence according to SEQ ID NO: 106 or SEQ ID NO: 108.BRIEF DESCRIPTION OF THE DRAWINGS
[0262] FIG. 1A and FIG. IB are diagrams depicting a gene modifying system as described herein. FIG. 1A is a diagram showing a gene modifying polypeptide, which comprises a Cas nickase domain (e.g., StlCas9) and a reverse transcriptase domain (RT domain) which are linked by a linker. FIG. IB is a diagram showing a template RNA which comprises, from 5’ to 3’, a gRNA spacer, a gRNA scaffold, a heterologous object sequence, and a primer binding site sequence (PBS sequence). The heterologous object sequence can comprise a mutation region that comprises one or more sequence differences relative to the target site. The heterologous object sequence can also comprise a pre-edit homology region and a post-edit homology region, which flank the mutation region. Without wishing to be bound by theory, it is thought that the gRNA spacer of the template RNA binds to the second strand of a target site in the genome, and the gRNA scaffold of the template RNA binds to thePage 41 of 39612985085vlAttorney Docket No.: 2017469-0049 gene modifying polypeptide, e.g., localizing the gene modifying polypeptide to the target site in the genome. It is thought that the Cas domain of the gene modifying polypeptide nicks the target site (e.g., the first strand of the target site), e.g., allowing the PBS sequence to bind to a sequence adjacent to the site to be altered on the first strand of the target site. It is thought that the RT domain of the gene modifying polypeptide uses the first strand of the target site that is bound to the complementary sequence comprising the PBS sequence of the template RNA as a primer and the heterologous object sequence of the template RNA as a template to, e.g., polymerize a sequence complementary to the heterologous object sequence. Without wishing to be bound by theory, it is thought that reverse transcription can then proceed through the pre-edit homology region, then through the mutation region, and then through the post-edit homology region, thereby producing a DNA strand comprising a mutation specified by the heterologous object sequence.
[0263] FIG. 2A shows a bar graph of the rewriting performance of StlCas9-based gene modifying systems comprising exemplary template RNAs in primary mouse hepatocytes. FIG. 2B shows a bar graph of the % INDEL levels of the same gene modifying systems evaluated in FIG. 2A.
[0264] FIG. 3A shows bar graphs of the rewriting performance of StlCas9-based gene modifying systems comprising an exemplary template RNA at various dosages following administration to mice modified to carry hSERPINA 1*E342K (PiZ) encoding alpha- 1 -antitrypsin (Al AT) protein. FIG. 3B shows bar graphs of the % INDEL levels of the same gene modifying systems evaluated in FIG. 3A.
[0265] FIG. 4A-and FIG. 4B show graphs of the rewriting performance of an StlCas9-based gene modifying system comprising an exemplary template RNA in the liver of hSERPINAl E342K mice. FIG. 4C shows a graph of the % INDEL levels of the same gene modifying systems evaluated in FIGs. 4A and 4B.
[0266] FIG. 5A shows a graph of the human Al AT levels in serum from hSERPINAl E342K mice treated with an StlCas9-based gene modifying systems comprising an exemplary template RNA at various dosages. FIG. 5B shows a correlation plot of Al AT levels and rewriting levels in FIG. 4.
[0267] FIG. 6A is a diagram illustrating the positions of an exemplary reference wild-type StlCas9 scaffold sequence.Page 42 of 39612985085vlAttorney Docket No.: 2017469-0049
[0268] FIG. 6B is a diagram illustrating the positions of the reference a mutant StlCas9 scaffold sequence according to SEQ ID NO: 2.
[0269] FIG. 7A is a bar graph showing the rewriting performance of StlCas9-based gene modifying systems formulated in various LNPs administered to mice modified to carry hSEPPINAl*E342K (PiZ) encoding alpha- 1 -antitrypsin (Al AT) protein. FIG. 7B is a bar graph showing the % INDEL levels of the same gene modifying systems evaluated in FIG. 7A.
[0270] FIG. 8 A, FIG. 8B, and FIG. 8C are bar graphs showing the rewriting performance of StlCas9-based gene modifying systems formulated in various LNPs administered, at various dosages, to non-human primates modified to carry hSEPPINAl*E342K (PiZ) encoding alpha- 1 -antitrypsin (Al AT) protein. FIG. 8A shows the rewriting performance of said gene modifying systems formulated in Lipid 1 LNP or Lipid 2 LNP. FIG. 8B and FIG. 8C shows the rewriting performance of said gene modifying systems formulated in Lipid 1 LNP at two different dosages by measuring rewriting via gDNA or cDNA.
[0271] FIG. 9 is a bar graph showing the rewriting performance of an exemplary gene modifying system formulated in exemplary LNPs following administration to non- human primates at a dose of 1.5 mg / kg.
[0272] FIG. 10A is a bar graph showing the rewriting efficiency in the livers of mice administered with exemplary gene modifying systems comprising different gene modifying polypeptides and different template RNAs comprising scaffold chemical modifications in combination with fluoro modifications at the heterologous object sequence.
[0273] FIG. 10B is a bar graph showing the % indel levels in the livers of mice administered with the exemplary gene modifying systems evaluated in FIG. 10A.Definitions
[0274] The term “expression cassette,” as used herein, refers to a nucleic acid construct comprising nucleic acid elements sufficient for the expression of the nucleic acid molecule of the instant invention.
[0275] A “gRNA spacer”, as used herein, refers to a portion of a nucleic acid that has complementarity to a target nucleic acid and can, together with a gRNA scaffold, target a Cas protein to the target nucleic acid.Page 43 of 39612985085vlAttorney Docket No.: 2017469-0049
[0276] A “gRNA scaffold”, as used herein, refers to a portion of a nucleic acid that can bind a Cas protein and can, together with a gRNA spacer, target the Cas protein to the target nucleic acid. In some embodiments, the gRNA scaffold comprises a crRNA sequence, tetraloop, and tracrRNA sequence.
[0277] A “variant gRNA scaffold”, as used herein, refers to gRNA scaffold having a non-naturally occurring sequence. In some embodiments, the variant sequence comprises one or more substitutions relative to the closest naturally occurring sequence. In some embodiments, the variant sequence comprises one or more insertions relative to the closest naturally occurring sequence. In some embodiments, the variant sequence comprises one or more deletions relative to the closest naturally occurring sequence.
[0278] The term “StlCas9 scaffold,” as used herein, refers to a gRNA scaffold that can bind an StlCas9 protein and can, together with a gRNA spacer, target the StlCas9 protein to the target nucleic acid. In some embodiments, an StlCas9 scaffold comprises a crRNA sequence, tetraloop, and tracrRNA sequence.
[0279] In some embodiments, an StlCas9 scaffold comprises a full length wild-type sequence. In some embodiments, an StlCas9 scaffold comprises a sequence with at least 80%, 85%. 90%, 95%, 96%, 97%, 98%, or 99% identity to the sequence of: GUCGUUGUACUCUGCGCGGUAACGCGCAGAAGCUACAACGAUAAGGCUUCAUG CCGAAAUCA (SEQ ID NO: 2). In some embodiments, an StlCas9 scaffold comprises a sequence identical to SEQ ID NO: 2. In some embodiments, an StlCas9 scaffold comprises an insertion, deletion, or substitution relative to a reference sequence of SEQ ID NO: 2. In some embodiments, an StlCas9 scaffold comprises a chemically modified nucleotide.
[0280] The term “StlCas9 domain” refers to a Cas9 domain having at least 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity to the amino acid sequence of SEQ ID NO: 12. In some embodiments, the StlCas9 domain has nickase activity. A StlCas9 domain can comprise a naturally occurring StlCas9 amino acid sequence, or a variant thereof. In some embodiments, a gene modifying polypeptide comprising an StlCas9 domain is used together with a compatible template RNA comprising a variant gRNA scaffold described herein.
[0281] A “gene modifying polypeptide”, as used herein, refers to a polypeptide comprising a retroviral reverse transcriptase, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acidPage 44 of 39612985085vlAttorney Docket No.: 2017469-0049 sequence identity to a retroviral reverse transcriptase, which is capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., in a mammalian host cell, such as a genomic DNA molecule in the host cell). In some embodiments, the gene modifying polypeptide is capable of integrating the sequence substantially without relying on host machinery. In some embodiments, the gene modifying polypeptide integrates a sequence into a random position in a genome, and in some embodiments, the gene modifying polypeptide integrates a sequence into a specific target site. In some embodiments, a gene modifying polypeptide includes one or more domains that, collectively, facilitate 1) binding the template nucleic acid, 2) binding the target DNA molecule, and 3) facilitate integration of the at least a portion of the template nucleic acid into the target DNA. Gene modifying polypeptides include both naturally occurring polypeptides as well as engineered variants of the foregoing, e.g., having one or more amino acid substitutions to the naturally occurring sequence. Gene modifying polypeptides also include heterologous constructs, e.g., where one or more of the domains recited above are heterologous to each other, whether through a heterologous fusion (or other conjugate) of otherwise wild-type domains, as well as fusions of modified domains, e.g., by way of replacement or fusion of a heterologous sub-domain or other substituted domain. Exemplary gene modifying polypeptides, and systems comprising them and methods of using them, that can be used in the methods provided herein are described, e.g., in PCT / US2021 / 020948, which is incorporated herein by reference with respect to gene modifying polypeptides that comprise a retroviral reverse transcriptase domain. In some embodiments, a gene modifying polypeptide integrates a sequence into a gene. In some embodiments, a gene modifying polypeptide integrates a sequence into a sequence outside of a gene. A “gene modifying system,” as used herein, refers to a system comprising a gene modifying polypeptide and a template nucleic acid.
[0282] The term “domain” as used herein refers to a structure of a biomolecule that contributes to a specified function of the biomolecule. A domain may comprise a contiguous region (e.g., a contiguous sequence) or distinct, non-contiguous regions (e.g., non-contiguous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, an endonuclease domain, a DNA binding domain, a reverse transcription domain; an example of a domain of a nucleic acid is a regulatory domain, such as a transcription factor binding domain. In some embodiments, a domain (e.g., a Cas domain) can comprise two or more smaller domains (e.g., a DNA binding domain and an endonuclease domain).Page 45 of 39612985085vlAttorney Docket No.: 2017469-0049
[0283] As used herein, the term “exogenous”, when used with reference to a biomolecule (such as a nucleic acid sequence or polypeptide) means that the biomolecule was introduced into a host genome, cell or organism by the hand of man. For example, a nucleic acid that is as added into an existing genome, cell, tissue or subject using recombinant DNA techniques or other methods is exogenous to the existing nucleic acid sequence, cell, tissue or subject.
[0284] As used herein, “first strand” and “second strand”, as used to describe the individual DNA strands of target DNA, distinguish the two DNA strands based upon which strand the reverse transcriptase domain initiates polymerization, e.g., based upon where target primed synthesis initiates. The first strand refers to the strand of the target DNA upon which the reverse transcriptase domain initiates polymerization, e.g., where target primed synthesis initiates. The second strand refers to the other strand of the target DNA. First and second strand designations do not describe the target site DNA strands in other respects; for example, in some embodiments the first and second strands are nicked by a polypeptide described herein, but the designations ‘first’ and ‘second’ strand have no bearing on the order in which such nicks occur.
[0285] The term “heterologous,” as used herein to describe a first element in reference to a second element means that the first element and second element do not exist in nature disposed as described. For example, a heterologous polypeptide, nucleic acid molecule, construct or sequence refers to (a) a polypeptide, nucleic acid molecule or portion of a polypeptide or nucleic acid molecule sequence that is not native to a cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been altered or mutated relative to its native state, or (c) a polypeptide or nucleic acid molecule with an altered expression as compared to the native expression levels under similar conditions. For example, a heterologous regulatory sequence (e.g., promoter, enhancer) may be used to regulate expression of a gene or a nucleic acid molecule in a way that is different than the gene or a nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA binding domain of a polypeptide or nucleic acid encoding a DNA binding domain of a polypeptide) may be disposed relative to other domains or may be a different sequence or from a different source, relative to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may exist in a native host cell genome but may have an altered expression level or have a differentPage 46 of 39612985085vlAttorney Docket No.: 2017469-0049 sequence or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous to a host cell or host genome but instead may have been introduced into a host cell by transformation (e.g., transfection, electroporation), wherein the added molecule may integrate into the host genome or can exist as extra-chromosomal genetic material either transiently (e.g., mRNA) or semi-stably for more than one generation (e.g., episomal viral vector, plasmid or other self-replicating vector).
[0286] As used herein, “insertion” of a sequence into a target site refers to the net addition of DNA sequence at the target site, e.g., where there are new nucleotides in the heterologous object sequence with no cognate positions in the unedited target site. In some embodiments, a nucleotide alignment of the PBS sequence and heterologous object sequence to the target nucleic acid sequence would result in an alignment gap in the target nucleic acid sequence.
[0287] As used herein, a “deletion” generated by a heterologous object sequence in a target site refers to the net deletion of DNA sequence at the target site, e.g., where there are nucleotides in the unedited target site with no cognate positions in the heterologous object sequence. In some embodiments, a nucleotide alignment of the PBS sequence and heterologous object sequence to the target nucleic acid sequence would result in an alignment gap in the molecule comprising the PBS sequence and heterologous object sequence.
[0288] The term “inverted terminal repeats” or “ITRs” as used herein refers to AAV viral cis-elements named so because of their symmetry. These elements promote efficient multiplication of an AAV genome. It is hypothesized that the minimal elements for ITR function are a Rep-binding site (RBS; 5 -GCGCGCTCGCTCGCTC-3 for AAV2; SEQ ID NO: 27) and a terminal resolution site (TRS; 5 -AGTTGG-3' for AAV2; SEQ ID NO: 41) plus a variable palindromic sequence allowing for hairpin formation. According to the present invention, an ITR comprises at least these three elements (RBS, TRS, and sequences allowing the formation of a hairpin). In addition, in the present invention, the term “ITR” refers to ITRs of known natural AAV serotypes (e.g. ITR of a serotype 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or 11 AAV), to chimeric ITRs formed by the fusion of ITR elements derived from different serotypes, and to functional variants thereof. “Functional variant” refers to a sequence presenting a sequence identity of at least 80%, 85%, 90%, preferably of at least 95% with a known ITR and allowing multiplication of the sequence that includes said ITR in the presence of Rep proteins.Page 47 of 39612985085vlAttorney Docket No.: 2017469-0049
[0289] The term “mutation region,” as used herein, refers to a region in a template RNA having one or more sequence difference relative to the corresponding sequence in a target nucleic acid. The sequence difference may comprise, for example, a substitution, insertion, frameshift, or deletion.
[0290] The term “mutated” when applied to nucleic acid sequences means that nucleotides in a nucleic acid sequence are inserted, deleted, or changed compared to a reference (e.g., native) nucleic acid sequence. A single alteration may be made at a locus (a point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. A nucleic acid sequence may be mutated by any method known in the art.
[0291] “Nucleic acid molecule” refers to both RNA and DNA molecules including, without limitation, complementary DNA (“cDNA”), genomic DNA (“gDNA”), and messenger RNA (“mRNA”), and also includes synthetic nucleic acid molecules, such as those that are chemically synthesized or recombinantly produced, such as RNA templates, as described herein. The nucleic acid molecule can be double-stranded or single-stranded, circular, or linear. If single-stranded, the nucleic acid molecule can be the sense strand or the antisense strand. Unless otherwise indicated, and as an example for all sequences described herein under the general format “SEQ ID NO:,” or “nucleic acid comprising SEQ ID NO: 1” refers to a nucleic acid, at least a portion which has either (i) the sequence of SEQ ID NO: 1, or (ii) a sequence complimentary to SEQ ID NO: 1. The choice between the two is dictated by the context in which SEQ ID NO: 1 is used. For instance, if the nucleic acid is used as a probe, the choice between the two is dictated by the requirement that the probe be complementary to the desired target. Nucleic acid sequences of the present disclosure may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with an analog, inter-nucleotide modifications such as uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (for example, phosphorothioates, phosphorodithioates, etc.), pendant moieties, (for example, polypeptides), intercalators (for example, acridine, psoralen, etc.), chelators, alkylators, and modified linkages (for example, alpha anomeric nucleic acids, etc.). Also included are chemically modified bases (see, for example, Table 16), backbones (see, forPage 48 of 39612985085vlAttorney Docket No.: 2017469-0049 example, Table 17), and modified caps (see, for example, Table 18). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to a designated sequence via hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages substitute for phosphate linkages in the backbone of a molecule, e.g., peptide nucleic acids (PNAs). Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structure such as modifications found in “locked” nucleic acids (LNAs). In various embodiments, the nucleic acids are in operative association with additional genetic elements, such as tissue-specific expression-control sequence(s) (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences), as well as additional elements, such as inverted repeats (e.g., inverted terminal repeats, such as elements from or derived from viruses, e.g., AAV ITRs) and tandem repeats, inverted repeats / direct repeats, homology regions (segments with various degrees of homology to a target DNA), untranslated regions (UTRs) (5', 3', or both 5' and 3' UTRs), and various combinations of the foregoing. The nucleic acid elements of the systems provided by the invention can be provided in a variety of topologies, including single-stranded, double-stranded, circular, linear, linear with open ends, linear with closed ends, and particular versions of these, such as doggybone DNA (dbDNA), closed-ended DNA (ceDNA).
[0292] As used herein, a “gene expression unit” is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence.Operably linked DNA sequences may be contiguous or non-contiguous. Where necessary to join two protein-coding regions, operably linked sequences may be in the same reading frame.
[0293] The terms “host genome” or “host cell”, as used herein, refer to a cell and / or its genome into which protein and / or genetic material has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell and / or genome, but to the progeny of such a cell and / or the genome of the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, butPage 49 of 39612985085vlAttorney Docket No.: 2017469-0049 are still included within the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line or may be a host cell or host genome in living tissue or an organism. In some instances, a host cell may be an animal cell or a plant cell, e.g., as described herein. In certain instances, a host cell may be a mammalian cell, a human cell, avian cell, reptilian cell, bovine cell, horse cell, pig cell, goat cell, sheep cell, chicken cell, or turkey cell. In certain instances, a host cell may be a com cell, soy cell, wheat cell, or rice cell.
[0294] As used herein, “operative association” describes a functional relationship between two nucleic acid sequences, such as a 1) promoter and 2) a heterologous object sequence, and means, in such example, the promoter and heterologous object sequence (e.g., a gene of interest) are oriented such that, under suitable conditions, the promoter drives expression of the heterologous object sequence. For instance, a template nucleic acid carrying a promoter and a heterologous object sequence may be single-stranded, e.g., either the (+) or (-) orientation. An “operative association” between the promoter and the heterologous object sequence in this template means that, regardless of whether the template nucleic acid will be transcribed in a particular state, when it is in the suitable state (e.g., is in the (+) orientation, in the presence of required catalytic factors, and NTPs, etc.), it is accurately transcribed. Operative association applies analogously to other pairs of nucleic acids, including other tissue-specific expression control sequences (such as enhancers, repressors and microRNA recognition sequences), IR / DR, ITRs, UTRs, or homology regions and heterologous object sequences or sequences encoding a retroviral RT domain.
[0295] The term “primer binding site sequence” or “PBS sequence,” as used herein, refers to a portion of a template RNA capable of binding to a region comprised in a target nucleic acid sequence. In some instances, a PBS sequence is a nucleic acid sequence comprising at least 3, 4, 5, 6, 7, or 8 bases with 100% identity to the region comprised in the target nucleic acid sequence. In some embodiments the primer region comprises at least 5, 6, 7, 8 bases with 100% identity to the region comprised in the target nucleic acid sequence. Without wishing to be bound by theory, in some embodiments when a template RNA comprises a PBS sequence and a heterologous object sequence, the PBS sequence binds to a region comprised in a target nucleic acid sequence, allowing a reverse transcriptase domain to use that region as a primer for reverse transcription, and to use the heterologous object sequence as a template for reverse transcription.Page 50 of 39612985085vlAttorney Docket No.: 2017469-0049
[0296] As used herein, the term “position” with respect to an StlCas9 scaffold refers to the nucleotide of the StlCas9 scaffold that aligns with the corresponding nucleotide of the reference sequence. Alignments of nucleic acid or polypeptide sequences can be performed by using a sequence analysis tool such as Basic Local Alignment Search Tool (BLAST), for instance NIH megablast using default parameters.
[0297] In some embodiments, a position of an StlCas9 scaffold can be identified by providing an alignment of the StlCas9 scaffold (query sequence) to a reference sequence of GUCGUUGUACUCUGCGCGGUAACGCGCAGAAGCUACAACGAUAAGGCUUCAUG CCGAAAUCA (SEQ ID NO: 2; an exemplary scaffold region compatible with StlCas9, see e g , FIG 6B) or to a reference sequence ofGUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAU CAACACCCUGUCAUUUUAUGGCAGGGUGUUUU (SEQ ID NO: 45; a full length wildtype sequence, see e.g., FIG 6A), and identifying the position in the query sequence that corresponds to the position in the reference sequence.
[0298] For example, in an StlCas9 scaffold consisting of the sequence of SEQ ID NO: 2 except that a single new nucleotide is inserted just 5’ of the 5’ most G, the G is still position 1.
[0299] As another example, in an StlCas9 scaffold consisting of the sequence of SEQ ID NO: 2 except that a sequence of n nucleotides is inserted between the G of position 1 and the U of position 2, nucleotides 3’ of the insert maintain their original position number. For example, the U of position 2 is still position 2 rather than position n+2. A nucleotide that is inserted relative to the reference sequence need not be assigned a position number.
[0300] Similarly, as used herein, the term “position” with respect to a spacer sequence refers to the nucleotide number counting from the 5’ end of the spacer sequence. For instance, position 1 relative to SEQ ID NO: 1 refers to the 5’ most nucleotide of SEQ ID NO: 1, position 2 relative to SEQ ID NO: 1 refers to the nucleotide of SEQ ID NO: 1 that is adjacent to position 1, and so on.
[0301] As used herein, the term “position” with respect to a template RNA refers to the nucleotide number counting from the 5’ end of the template RNA sequence.
[0302] As used herein, the term “position” with respect to a heterologous object sequence refers to the nucleotide of the heterologous object sequence, counting from the junction of the heterologous object sequence and the PBS sequence in a 3’-to-5’ direction.Page 51 of 39612985085vlAttorney Docket No.: 2017469-0049For example, the most 3’ nucleotide of the heterologous object sequence is position +1, and the second most 3’ nucleotide of heterologous object sequence is position +2.
[0303] As used herein, the term “position” with respect to a PBS sequence refers to the nucleotide of the PBS sequence, counting from the junction of the heterologous object sequence and the PBS sequence, counting in a 5’-to-3’ direction. For example, the most 5’ nucleotide of the PBS sequence is position -1, and the adjacent nucleotide in the PBS sequence is position -2.
[0304] As used herein, a “stem-loop sequence” refers to a nucleic acid sequence (e.g., RNA sequence) with sufficient self-complementarity to form a stem-loop, e.g., having a stem comprising at least two (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs, and a loop with at least three (e.g., four) base pairs. The stem may comprise mismatches or bulges.
[0305] As used herein, a “tissue-specific expression-control sequence” means nucleic acid elements that increase or decrease the level of a transcript comprising the heterologous object sequence in a target tissue in a tissue-specific manner, e.g., preferentially in on-target tissue(s), relative to off-target tissue(s). In some embodiments, a tissue-specific expressioncontrol sequence preferentially drives or represses transcription, activity, or the half-life of a transcript comprising the heterologous object sequence in the target tissue in a tissue-specific manner, e.g., preferentially in an on-target tissue(s), relative to an off-target tissue(s). Exemplary tissue-specific expression-control sequences include tissue-specific promoters, repressors, enhancers, or combinations thereof, as well as tissue-specific microRNA recognition sequences. Tissue specificity refers to on-target (tissue(s) where expression or activity of the template nucleic acid is desired or tolerable) and off-target (tissue(s) where expression or activity of the template nucleic acid is not desired or is not tolerable). For example, a tissue-specific promoter drives expression preferentially in on-target tissues, relative to off-target tissues. In contrast, a microRNA that binds the tissue-specific microRNA recognition sequences is preferentially expressed in off-target tissues, relative to on-target tissues, thereby reducing expression of a template nucleic acid in off-target tissues. Accordingly, a promoter and a microRNA recognition sequence that are specific for the same tissue, such as the target tissue, have contrasting functions (promote and repress, respectively, with concordant expression levels, i.e., high levels of the microRNA in off-target tissues and low levels in on-target tissues, while promoters drive high expression in on-target tissues and low expression in off-target tissues) with regard to the transcription, activity, or half-life of an associated sequence in that tissue.Page 52 of 39612985085vlAttorney Docket No.: 2017469-0049
[0306] Unless specified otherwise, the following numbering system will be adhered to for describing the position of nucleotides having chemical modifications in the PBS and / or heterologous object sequence of a template RNA. The positions of nucleotides in the heterologous object sequence are numbered +1, +2, +3, and so on, starting from the 3 ’-most end of the heterologous object sequence. The positions of nucleotides in the PBS sequence are numbered -1, -2, -3, and so on, starting from the 5 ’-most end of the PBS sequence. Thus, positions +1 and -1 are directly adjacent to each other.
[0307] The term “chemically modified nucleotide,” as used herein, refers to a nucleotide comprising one or more structural differences relative to the canonical ribonucleotides (i.e., G, U, C, and A). A chemically modified nucleotide may have (relative to a canonical nucleotide) a chemically modified nucleobase, a chemically modified sugar, a chemically modified phosphodiester linkage, or a combination thereof. In some embodiments, a chemically modified nucleotide is a 2'-O-methyl nucleotide, e.g., 2'-O- methyl-Adenosine, 2'-O-methyl-Cytidine, 2'-O-methyl-Guanosine, or 2'-O-methyl-Uridine. No particular process of making is implied; for instance, a chemically modified nucleotide can be produced directly by chemical synthesis, or by covalently modifying a canonical nucleotide.
[0308] The term “chemical modification”, as used herein, refers to a structural difference of a chemical modified nucleotide relative to the canonical ribonucleotides (i.e., G, U, C, and A). A chemical modification may comprise a modification resulting in a chemically modified nucleobase, a chemically modified sugar, a chemically modified phosphodiester linkage, or a combination thereof. In some embodiments, a chemical modification is 2'-O- methylation or 2’ -fluoro modification. No particular process of making is implied; for instance, a chemical modification can be produced directly by chemical synthesis, or by covalently modifying a canonical nucleotide.
[0309] It is understood that, herein, when a nucleic acid molecule is said to comprise the nucleic acid sequence of a given SEQ ID NO, and if the sequence of the SEQ ID NO specifies particular chemical modifications (such as phosphorothioate, 2’-0-Me ribose, or 2’F ribose), the nucleic acid molecule comprises those chemical modifications at the position(s) indicated. Sometimes a sequence is said to have one or more chemical modification difference(s) relative to the nucleic acid sequence of a given SEQ ID NO; several types of chemical modification differences are possible. For instance, the chemical modification difference may be replacement of a chemically modified moiety with its canonical equivalent Page 53 of 39612985085vlAttorney Docket No.: 2017469-0049(e.g., replacement of a phosphorothioate moiety with a phosphate moiety, or replacement of a 2’-0-Me ribose with a canonical ribose). Alternatively, the chemical modification difference may be replacement of a canonical nucleotide with a chemically modified one (e.g., replacement of a nucleotide having a phosphate moiety with a nucleotide having a phosphorothioate moiety, or replacement of a nucleotide having a canonical ribose with a nucleotide having a 2’-0-Me ribose). Alternatively, the chemical modification difference may be replacement of one chemical modification with a different chemical modification (e.g., replacement of 2’-0-Me ribose with 2’F ribose). When a nucleic acid molecule is said to comprise the nucleic acid sequence of a given SEQ ID NO, and if the SEQ ID NO does not specify any chemical modifications, then the nucleic acid molecule may have canonical nucleotides and / or one or more chemically modified nucleotides.DETAILED DESCRIPTION
[0310] This disclosure relates to methods for treating alpha- 1 antitrypsin deficiency (AATD) and compositions for targeting, editing, modifying or manipulating a DNA sequence at one or more locations in a DNA sequence in a cell, tissue or subject, e.g., in vivo or in vitro. In some embodiments, a modification may include, e.g., a substitution.
[0311] The present disclosure further provides methods for treating AATD using reverse transcriptase-based systems for altering a genomic DNA sequence of interest, e.g., by inserting, deleting, or substituting one or more nucleotides into / from a sequence of interest.
[0312] The disclosure provides, in part, methods for treating AATD using a gene modifying system comprising a gene modifying polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, a gene modifying system can be used to introduce an alteration into a target site in a genome. In some embodiments, the gene modifying polypeptide component comprises a writing domain (e.g., a reverse transcriptase domain), a DNA-binding domain, and an endonuclease domain (e.g., nickase domain). In some embodiments, a template nucleic acid (e.g., template RNA) comprises a sequence (e.g., a gRNA spacer) that binds a target site in the genome (e.g., that binds to a second strand of the target site), a sequence (e.g., a gRNA scaffold) that binds the gene modifying polypeptide component, a heterologous object sequence, and a PBS sequence. Without wishing to be bound by any particular theory, it is thought that a template nucleic acid (e.g., template RNA) binds to the second strand of a target site in the genome, and binds to a gene modifying polypeptide component (e.g., localizing the polypeptidePage 54 of 39612985085vlAttorney Docket No.: 2017469-0049 component to the target site in the genome). It is thought that the endonuclease (e.g., nickase) of a gene modifying polypeptide component cuts the target site (e.g., the first strand of the target site), e.g., allowing a PBS sequence to bind to a sequence adjacent to the target site to be altered on the first strand of the target site. It is thought that the writing domain (e.g., reverse transcriptase domain) of a gene modifying polypeptide component uses the first strand of a target site that is bound to the complementary sequence comprising the PBS sequence of a template nucleic acid as a primer and a heterologous object sequence of the template nucleic acid as a template to, e.g., polymerize a sequence complementary to the heterologous object sequence. Without wishing to be bound by any particular theory, it is thought that selection of an appropriate heterologous object sequence can result in substitution, deletion, and / or insertion of one or more nucleotides at a target site.Gene modifying systems
[0313] In some embodiments, a gene modifying system described herein comprises: (A) a gene modifying polypeptide or a nucleic acid encoding the gene modifying polypeptide, wherein the gene modifying polypeptide comprises (i) a reverse transcriptase domain, and either (x) an endonuclease domain that contains DNA binding functionality or (y) an endonuclease domain and separate DNA binding domain; and (B) a template RNA. A gene modifying polypeptide, in some embodiments, acts as a substantially autonomous protein machine capable of integrating a template nucleic acid sequence into a target DNA molecule (e.g., in a mammalian host cell, such as a genomic DNA molecule in the host cell), substantially without relying on host machinery. For example, a gene modifying polypeptide may comprise a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, a DNA-binding function may involve an RNA component that directs a gene modifying polypeptide to a DNA sequence, e.g., a gRNA spacer. In other embodiments, a gene modifying polypeptide may comprise a reverse transcriptase domain and an endonuclease domain. An RNA template element of a gene modifying system may be heterologous to a gene modifying polypeptide element and provides an object sequence to be inserted (reverse transcribed) into a host genome. In some embodiments, a gene modifying polypeptide is capable of target primed reverse transcription (TPRT). In some embodiments, a gene modifying polypeptide is capable of second-strand synthesis.
[0314] In some embodiments a gene modifying system is combined with a second polypeptide. In some embodiments, a second polypeptide may comprise an endonuclease domain. In some embodiments, a second polypeptide may comprise a polymerase domain, Page 55 of 39612985085vlAttorney Docket No.: 2017469-0049 e.g., a reverse transcriptase domain. In some embodiments, a second polypeptide may comprise a DNA-dependent DNA polymerase domain. In some embodiments, a second polypeptide aids in completion of a genome edit, e.g., by contributing to second-strand synthesis or DNA repair resolution.
[0315] A functional gene modifying polypeptide can be made up of unrelated DNA binding, reverse transcription, and endonuclease domains. This modular structure allows combining of functional domains: e.g., dCas9 (DNA binding); e.g., Avian reticuloendotheliosis virus (AVIRE), porcine endogenous retrovirus (PERV), or baboon endogenous virus M (BAEVM) reverse transcriptase (reverse transcription); FokI (endonuclease). In some embodiments, multiple functional domains may arise from a single protein, e.g., Cas9 or Cas9 nickase (DNA binding, endonuclease).
[0316] In some embodiments, a gene modifying polypeptide includes one or more domains that, collectively, facilitate 1) binding a template nucleic acid, 2) binding a target DNA molecule, and 3) integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, a gene modifying polypeptide is an engineered polypeptide that comprises one or more amino acid substitutions to a corresponding naturally occurring sequence. In some embodiments, a gene modifying polypeptide comprises two or more domains that are heterologous relative to each other, e.g., through a heterologous fusion (or other conjugate) of otherwise wild-type domains, or as fusions of modified domains, e.g., by way of replacement or fusion of a heterologous sub-domain or other substituted domain. For instance, in some embodiments, one or more of: an RT domain is heterologous to a DNA- binding domain (DBD); a DBD is heterologous to an endonuclease domain; or an RT domain is heterologous to an endonuclease domain.
[0317] In some embodiments, a template RNA molecule for use in a system of the present disclosure comprises, from 5' to 3' (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous object sequence; and (4) a primer binding site (PBS) sequence. In some embodiments, a gRNA spacer is about 18 to 22 nucleotides in length (e.g., about 20 or about 21 nucleotides in length). In some embodiments, a gRNA scaffold comprises one or more hairpin loops, e.g., 1, 2, or 3 loops for associating a template RNA with a Cas domain, e.g., a nickase Cas9 domain. In some embodiments, a gRNA scaffold comprises the sequence, from 5' to 3', GUCGUUGUACUCUGCGCGGUAACGCGCAGAAGCUACAACGAUAAGGCUUCAUG CCGAAAUCA (SEQ ID NO: 2). In some embodiments, a heterologous object sequence is, Page 56 of 39612985085vlAttorney Docket No.: 2017469-0049 e.g., 7-74, e.g., 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, or 70-80 nucleotides or, 80-90 nucleotides in length. In some embodiments, a first (i.e., 5’-most) base of a heterologous object sequence is not C. In some embodiments, a PBS sequence that binds a target priming sequence after nicking occurs is e.g., 3-20 nucleotides, e.g., 7-15 nucleotides, e.g., 12-14 nucleotides in length. In some embodiments, a PBS sequence has 40-60% GC content.
[0318] In some embodiments, a second gRNA associated with a system of the present disclosure may help drive complete integration. In some embodiments, a second gRNA may target a location that is 0-200 nucleotides away from a first-strand nick, e.g., 0-50, 50-100, 100-200 nucleotides away from the first-strand nick. In some embodiments, a second gRNA can only bind its target sequence after an edit is made, e.g., the gRNA binds a sequence present in a heterologous object sequence, but not in an initial target sequence.
[0319] In some embodiments, a gene modifying system described herein is used to make an edit in hepatocytes, HEK293, K562, U2OS, or HeLa cells. In some embodiment, a gene modifying system is used to make an edit in primary cells, e.g., primary liver cells, or primary lung cells.
[0320] In some embodiments, a gene modifying polypeptide as described herein comprises a reverse transcriptase or RT domain (e.g., as described herein) that comprises an AVIRE RT sequence or variant thereof. In some embodiments, a gene modifying polypeptide as described herein comprises an RT domain comprising the sequence of SEQ ID NO: 20, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity thereof.
[0321] In some embodiments, a gene modifying polypeptide as described herein comprises a reverse transcriptase or RT domain (e.g., as described herein) that comprises a PERV RT sequence or variant thereof. In some embodiments, a gene modifying polypeptide as described herein comprises an RT domain comprising the sequence of SEQ ID NO: 46, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity thereof.
[0322] In some embodiments, a gene modifying polypeptide as described herein comprises a reverse transcriptase or RT domain (e.g., as described herein) that comprises a BAEVM RT sequence or variant thereof. In some embodiments, a gene modifying polypeptide as described herein comprises an RT domain comprising the sequence of SEQ IDPage 57 of 39612985085vlAttorney Docket No.: 2017469-0049NO: 47, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity thereof.
[0323] In some embodiments, an endonuclease domain (e.g., as described herein) comprises a Cas9 domain (e.g., StlCas9, e.g., comprising an N622 A mutation (e.g., in StlCas9)). In some embodiments, a gene modifying polypeptide as described herein comprises a Cas9 domain comprising the sequence of SEQ ID NO: 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity thereof.
[0324] In some embodiments, a heterologous object sequence (e.g., of a system as described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600- 700, 700-800, 800-900, 900-1000, or more, nucleotides in length.
[0325] In some embodiments, RT and endonuclease domains are joined by a flexible linker. In some embodiments, a linker comprises the amino acid sequence GGWQAAESYEVGG (SEQ ID NO: 18), or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 18.
[0326] In some embodiments, a linker comprises the amino acid sequence WQAAESYEV (SEQ ID NO: 19), or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19. In some embodiments, a linker comprises the amino acid sequence AEIKYDGV (SEQ ID NO:48), or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 48. In some embodiments, a linker comprises the amino acid sequence EAAAKEAAAKEAAAKEAAAKEAAAK (SEQ ID NO:49), or a sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 49. In some embodiments, an endonuclease domain is N-terminal relative to an RT domain. In some embodiments, an endonuclease domain is C-terminal relative to an RT domain.Page 58 of 39612985085vlAttorney Docket No.: 2017469-0049
[0327] In some embodiments, a system of the present disclosure incorporates a heterologous object sequence into a target site by target-primed reverse transcription (TPRT), e.g., as described herein.
[0328] In some embodiments, a gene modifying polypeptide as described herein comprises, e.g., from N terminus to C terminus: (i) an StlCas9 domain; (ii) a linker; and (iii) an AVIRE, PERV, or BAEVM RT domain.
[0329] In some embodiments, a gene modifying polypeptide as described herein comprises, e.g., from N-terminus to C-terminus: (i) an StlCas9 domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an StlCas9 domain having an amino acid sequence according to SEQ ID NO: 12; (ii) a linker comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a linker having an amino acid sequence according to SEQ ID NO: 19; and (iii) an AVIRE RT domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an AVIRE RT domain having an amino acid sequence according to SEQ ID NO: 20.
[0330] In some embodiments, a gene modifying polypeptide as described herein comprises, e.g., from N-terminus to C-terminus: (i) an StlCas9 domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an StlCas9 domain having an amino acid sequence according to SEQ ID NO: 12; (ii) a linker comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a linker having an amino acid sequence according to SEQ ID NO: 48; and (iii) a PERV RT domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a PERV RT domain having an amino acid sequence according to SEQ ID NO: 46.Page 59 of 39612985085vlAttorney Docket No.: 2017469-0049
[0331] In some embodiments, a gene modifying polypeptide as described herein comprises, e.g., from N-terminus to C-terminus: (i) an StlCas9 domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an StlCas9 domain having an amino acid sequence according to SEQ ID NO: 12; (ii) a linker comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a linker having an amino acid sequence according to SEQ ID NO: 49; and (iii) a BAEVM RT domain comprising an amino acid sequence at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a PERV RT domain having an amino acid sequence according to SEQ ID NO: 47.
[0332] In some embodiments, a gene modifying system is capable of producing an insertion of at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides (and optionally no more than 500, no more than 400, no more than 300, no more than 200, or no more than 100 nucleotides) in a target site. In some embodiments, a gene modifying system is capable of producing an insertion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides (and optionally no more than 500, no more than 400, no more than 300, no more than 200, or no more than 100 nucleotides) in a target site. In some embodiments, a gene modifying system is capable of producing an insertion of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, at least 1, at least 1.5, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 5.5, at least 6, at least 6.5, at least 7, at least 7.5, at least 8, at least 8.5, at least 9, at least 9.5 or at least 10 kilobases (and optionally no more than 1, 5, 10, or 20 kilobases) into a target site. In some embodiments, a gene modifying system is capable of producing a deletion of at least 81, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 nucleotides (and optionally no more than 500, no more than 400, no more than 300, or no more than 200 nucleotides). In somePage 60 of 39612985085vlAttorney Docket No.: 2017469-0049 embodiments, a gene modifying system is capable of producing a deletion of at least 81, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 nucleotides (and optionally no more than 500, no more than 400, no more than 300, or no more than 200 nucleotides). In some embodiments, a gene modifying system is capable of producing a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, or at least 200 nucleotides (and optionally no more than 500, no more than 400, no more than 300, or no more than 200 nucleotides). In some embodiments, a gene modifying system is capable of producing a deletion of at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, at least 1, at least 1.5, at least 2, at least 2.5, at least 3, at least 3.5, at least 4, at least 4.5, at least 5, at least 5.5, at least 6, at least 6.5, at least 7, at least 7.5, at least 8, at least 8.5, at least 9, at least 9.5 or at least 10 kilobases (and optionally no more than 1, no more than 5, no more than 10, or no more than 20 kilobases). In some embodiments, a gene modifying system is capable of producing a substitution of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 or more nucleotides in a target site. In some embodiments, a gene modifying system is capable of producing a substitution of 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15- 20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotides in a target site.
[0333] In some embodiments, a substitution is a transition mutation. In some embodiments, a substitution is a transversion mutation. In some embodiments, a substitution converts an adenine to a thymine, an adenine to a guanine, an adenine to a cytosine, a guanine to a thymine, a guanine to a cytosine, a guanine to an adenine, a thymine to a cytosine, a thymine to an adenine, a thymine to a guanine, a cytosine to an adenine, a cytosine to a guanine, or a cytosine to a thymine.
[0334] In some embodiments, an insertion, deletion, substitution, or combination thereof, increases or decreases expression (e.g., transcription or translation) of a gene. In some embodiments, an insertion, deletion, substitution, or combination thereof, increases orPage 61 of 39612985085vlAttorney Docket No.: 2017469-0049 decreases expression (e.g., transcription or translation) of a gene by altering, adding, or deleting sequences in a promoter or enhancer, e.g. sequences that bind transcription factors. In some embodiments, an insertion, deletion, substitution, or combination thereof alters translation of a gene (e.g., alters an amino acid sequence), inserts or deletes a start or stop codon, or alters or fixes the translation frame of a gene. In some embodiments, an insertion, deletion, substitution, or combination thereof alters splicing of a gene, e.g., by inserting, deleting, or altering a splice acceptor or donor site. In some embodiments, an insertion, deletion, substitution, or combination thereof alters transcript or protein half-life. In some embodiments, an insertion, deletion, substitution, or combination thereof alters protein localization in the cell (e.g., from the cytoplasm to a mitochondria, from the cytoplasm into the extracellular space (e.g., adds a secretion tag)). In some embodiments, an insertion, deletion, substitution, or combination thereof alters (e.g., improves) protein folding (e.g., to prevent accumulation of misfolded proteins). In some embodiments, an insertion, deletion, substitution, or combination thereof, alters, increases, decreases the activity of a gene, e.g., a protein encoded by the gene.
[0335] Exemplary gene modifying polypeptides, systems comprising the same, and methods of using the same are described, e.g., in PCT / US2021 / 020948, which is incorporated herein by reference with respect to retroviral RT domains, including the amino acid and nucleic acid sequences therein.
[0336] Exemplary gene modifying polypeptides and retroviral RT domain sequences are also described, e.g., in International Application No. PCT / US21 / 20948 filed March 4, 2021, e.g., at Table 30, Table 31, and Table 44 therein; the entire application is incorporated by reference herein with respect to retroviral RTs, e.g., in said sequences and Tables. Accordingly, a gene modifying polypeptide described herein may comprise an amino acid sequence according to any of the Tables mentioned in this paragraph, or a domain thereof (e.g., a retroviral RT domain), or a functional fragment or variant of any of the foregoing, or an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0337] In some embodiments, a polypeptide for use in any of the systems described herein can be a molecular reconstruction or ancestral reconstruction based upon the aligned polypeptide sequence of multiple homologous proteins. In some embodiments, a reverse transcriptase domain for use in any of the systems described herein can be a molecular reconstruction or an ancestral reconstruction, or can be modified at particular residues, based upon alignments of reverse transcriptase domains from the same or different sources. APage 62 of 39612985085vlAttorney Docket No.: 2017469-0049 skilled artisan can, based on the Accession numbers provided herein, align polypeptides or nucleic acid sequences, e.g., by using routine sequence analysis tools as Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Molecular reconstructions can be created based upon sequence consensus, e.g. using approaches described in Ivies et al., Cell 1997, 501 - 510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99.A. Gene Modifying Polypeptides
[0338] In some embodiments, a gene modifying polypeptide possesses the functions of DNA target site binding, template nucleic acid (e.g., template RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., template RNA) writing (e.g., reverse transcription or RT). In some embodiments, each function is contained within a distinct domain. In some embodiments, a function may be attributed to two or more domains (e.g., two or more domains, together, exhibit the functionality). In some embodiments, two or more domains may have the same or similar function (e.g., two or more domains each independently have DNA-binding functionality, e.g., for two different DNA sequences). In some embodiments, one or more domains may be capable of enabling one or more functions, e.g., a Cas9 domain enabling both DNA binding and target site cleavage. In some embodiments, domains are all located within a single polypeptide. In some embodiments, a first domain is in one polypeptide and a second domain is in a second polypeptide. For example, in some embodiments, sequences may be split between a first polypeptide and a second polypeptide, e.g., wherein the first polypeptide comprises a reverse transcriptase (RT) domain and wherein the second polypeptide comprises a DNA-binding domain and an endonuclease domain, e.g., a nickase domain. As a further example, in some embodiments, a first polypeptide and a second polypeptide each comprise a DNA binding domain (e.g., a first DNA binding domain and a second DNA binding domain). In some embodiments, a first polypeptide and a second polypeptide may be brought together post-translationally via a split- intein to form a single gene modifying polypeptide.
[0339] In some embodiments, a gene modifying polypeptide described herein comprises an StlCas9 domain. An StlCas9 domain can comprise a naturally occurring StlCas9 amino acid sequence, or a variant thereof. In some embodiments, an StlCas9 domain is a nickase. In some embodiments, an StlCas9 domain comprises a sequence according to SEQ ID NO: 12. In some embodiments, an StlCas9 domain comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at leastPage 63 of 39612985085vlAttorney Docket No.: 2017469-004992%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to an StlCas9 domain having an amino acid sequence according to SEQ ID NO: 12. In some embodiments, a gene modifying polypeptide comprising an StlCas9 domain is used together with a compatible template RNA comprising a variant gRNA scaffold described herein.
[0340] In some embodiments, a gene modifying polypeptide described herein comprises (e.g., a system described herein comprises a gene modifying polypeptide that comprises): 1) a Cas domain (e.g., a Cas nickase domain, e.g., a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain of Table D of International Application WO / 2023 / 039424 (which is herein incorporated by reference in its entirety), or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto, wherein the RT domain is C-terminal of the Cas domain; and a linker is disposed between the RT domain and the Cas domain, wherein the linker has a sequence from the same row of Table D of International Application WO / 2023 / 039424 as the RT domain, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, a at least 94%, t least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0341] In some embodiments, an RT domain has a sequence with 100% sequence identity to an RT domain of Table D of International Application WO / 2023 / 039424 and a linker has a sequence with 100% sequence identity to a linker sequence from the same row of Table D of International Application WO / 2023 / 039424 as the RT domain.
[0342] In some embodiments, a Cas domain comprises a sequence described in Table 5, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0343] In some embodiments, a gene modifying polypeptide comprises, from N- terminus to C-terminus, one or more (e.g., 1, 2, 3, 4, 5, or all 6) of an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA binding domain, a linker, an RT domain, and / or a second NLS. In some embodiments, a gene modifying polypeptide comprises an N-terminal methionine residue.Page 64 of 39612985085vlAttorney Docket No.: 2017469-0049
[0344] In some embodiments, a gene modifying polypeptide comprises a first NLS comprising an amino acid sequence of MPAAKRVKLDGG (SEQ ID NO: 15), or a sequence having no more than 1, no more than 2, or no more than 3 sequence differences thereof. In some embodiments, a gene modifying polypeptide comprises a first NLS comprising the amino acid sequence of PAAKRVKLDGG (SEQ ID NO: 16), or a sequence having no more than 1, no more than 2, or no more than 3 sequence differences thereof. In some embodiments, a gene modifying polypeptide comprises a second NLS comprising the amino acid sequence of KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 17), or a sequence having no more than 1, no more than 2, or no more than 3 sequence differences thereof.
[0345] In some embodiments, a gene modifying polypeptide comprises: a GG amino acid sequence between a Cas domain and a linker; an AG amino acid sequence between an RT domain and a second NLS; and / or a GG amino acid sequence between the linker and the RT domain.
[0346] In some embodiments, a gene modifying polypeptide comprises a sequence of SEQ ID NO: 25 which comprises a first NLS and a Cas domain, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a gene modifying polypeptide comprises a sequence of SEQ ID NO: 25 but does not comprise an N-terminal methionine residue. In some embodiments, a gene modifying polypeptide comprises a sequence of SEQ ID NO: 26 which comprises a reverse transcriptase (RT) domain and a second NLS, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.Exemplary N-terminal NLS-Cas9 domainMPAAKRVKLDGGSDLVLGLDIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNLVRRT NRQGRRLARRKKHRRVRLNRLFEESGLITDFTKISINLNPYQLRVKGLTDELSNEELFI ALKNMVKHRGISYLDDASDDGNSSVGDYAQIVKENSKQLETKTPGQIQLERYQTYG QLRGDFTVEKDGKKHRLINVFPTSAYRSEALRILQTQQEFNPQITDEFINRYLEILTGK RKYYHGPGNEKSRTDYGRYRTSGETLDNIFGILIGKCTFYPDEFRAAKASYTAQEFNL LNDLNNLTVPTETKKLSKEQKNQIINYVKNEKAMGPAKLFKYIAKLLSCDVADIKGY RIDKSGKAEIHTFEAYRKMKTLETLDIEQMDRETLDKLAYVLTLNTEREGIQEALEHE FADGSF SQKQVDELVQFRKANS SIFGKGWHNF S VKLMMELIPELYETSEEQMTILTRL GKQKTTSSSNKTKYIDEKLLTEEIYNPVVAKSVRQAIKIVNAAIKEYGDFDNIVIEMARPage 65 of 39612985085vlAttorney Docket No.: 2017469-0049ETNEDDEKKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQLATKIRL WHQQGERCLYTGKTISIHDLINNSNQFEVDHILPLSITFDDSLANKVLVYATAAQEKG QRTPYQALDSMDDAWSFRELKAFVRESKTLSNKKKEYLLTEEDISKFDVRKKFIERN LVDTRYASRVVLNALQEHFRAHKIDTKVSVVRGQFTSQLRRHWGIEKTRDTYHHHA VDALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVD TLKSKEFEDSILFSYQVDSKFNRKISDATIYATRQAKVGKDKADETYVLGKIKDIYTQ DGYDAFMKIYKKDKSKFLMYRHDPQTFEKVIEPILENYPNKQINEKGKEVPCNPFLK YKEEHGYIRKYSKKGNGPEIKSLKYYDSKLGNHIDITPKDSNNKVVLQSVSPWRADV YFNKTTGKYEILGLKYADLQFEKGTGTYKISQEKYNDIKKKEGVDSDSEFKFTLYKNDLLLVKDTETKEQQLFRFLSRTMPKQKHYVELKPYDKQKFEGGEALIKVLGNVANSG QCKKGLGKSNISIYKVRTDVLGNQHIIKNEGDKPKLDF(SEQ ID NO: 25)Exemplary C-terminal sequence comprising an NLSTAPLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPV RVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLRE VNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIFAFEWADAE EGESGQLTWTRLPQGFKNSPTLFNEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQ AACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQI PVPKTKRQVREFLGKIGYCRLFIPGFAELAQPLYAATRPGNDPLVWGEKEEEAFQSLK LALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAG WPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSLTSTRPDLTDQPLAQAEATLF TDGSSYIRDGKRYAGAAVVTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNI YTDSRYAFATLHVHGMIYRERGWLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKG HQKDDAPTSTGNRRADEVAREVAIRPLSTQATISKRTADGSEFEKRTADGSEFESPKKK AKVE(SEQ ID NO: 26) / . RT Domains / Writing Domains
[0347] In some embodiments, a writing domain of a gene modifying system of the present disclosure possesses reverse transcriptase activity and is also referred to as a reverse transcriptase domain (an RT domain). In some embodiments, an RT domain comprises an RT catalytic portion and an RNA-binding region (e.g., a region that binds a template RNA).
[0348] In some embodiments, a nucleic acid encoding a reverse transcriptase is altered from its natural sequence to have altered codon usage, e.g., improved for human cells.In some embodiments, a reverse transcriptase domain is a heterologous reverse transcriptase from a retrovirus. In some embodiments, an RT domain has been mutated from its original amino acid sequence, e.g., has at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100 substitutions. InPage 66 of 39612985085vlAttorney Docket No.: 2017469-0049 some embodiments, an RT domain is derived from the RT of a retrovirus. In some embodiments, an RT domain is derived from an Avian reticuloendotheliosis virus (AVIRE) (e.g., UniProtKB accession: P03360) RT. In some embodiments, an RT domain is derived from a porcine endogenous retrovirus (PERV) (e.g., UniProtKB accession: Q4VFZ2) RT. In some embodiments, an RT domain is derived from a baboon endogenous virus (BAEVM) (e.g., UniProtKB accession: Pl 0272) RT.
[0349] In some embodiments, a retroviral reverse transcriptase (RT) domain exhibits enhanced stringency of target-primed reverse transcription (TPRT) initiation, e.g., relative to an endogenous RT domain. In some embodiments, an RT domain initiates TPRT when the 3 nucleotides in a target site immediately upstream of a first strand nick, e.g., genomic DNA priming of an RNA template, have at least 66% or 100% complementarity to 3 nucleotides of homology in the RNA template. In some embodiments, an RT domain initiates TPRT when there are less than 5 nucleotides mismatched (e.g., less than 1, less than 2, less than 3, less than 4, or less than 5 nucleotides mismatched) between a template RNA and a target DNA priming reverse transcription. In some embodiments, an RT domain is modified such that the stringency for mismatches in priming the TPRT reaction is increased, e.g., wherein an RT domain does not tolerate any mismatches or tolerates fewer mismatches in a priming region relative to a wild-type (e.g., unmodified) RT domain.
[0350] Naturally heterodimeric RT domains may, in some embodiments, also be functional as homodimers. In some embodiments, dimeric RT domains are expressed as fusion proteins, e.g., as homodimeric fusion proteins or heterodimeric fusion proteins. In some embodiments, the RT function of a gene modifying system of the present disclosure is fulfilled by multiple RT domains. In some embodiments, multiple RT domains are fused or separate, e.g., may be on the same polypeptide or on different polypeptides.
[0351] In some embodiments, a gene modifying system described herein comprises an integrase domain, e.g., wherein the integrase domain may be part of an RT domain. In some embodiments, an RT domain (e.g., as described herein) comprises an integrase domain. In some embodiments, an RT domain (e.g., as described herein) lacks an integrase domain. In some embodiments, an RT domain comprises an integrase domain that has been inactivated by mutation or deleted. In some embodiments, a gene modifying system described herein comprises an RNase H domain, e.g., wherein the RNase H domain may be part of an RT domain. In some embodiments, an RNase H domain is not part of an RT domain and is covalently linked to the RT domain via a flexible linker. In somePage 67 of 39612985085vlAttorney Docket No.: 2017469-0049 embodiments, an RT domain (e.g., as described herein) comprises an RNase H domain, e.g., an endogenous RNAse H domain or a heterologous RNase H domain. In some embodiments, an RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, an RT domain (e.g., as described herein) comprises an RNase H domain that has been added, deleted, mutated, or swapped for a heterologous RNase H domain. In some embodiments, a gene modifying polypeptide comprises an inactivated endogenous RNase H domain. In some embodiments, an endogenous RNase H domain is genetically removed such that it is not included in a gene modifying polypeptide, e.g., the endogenous RNase H domain is partially or completely truncated. In some embodiments, mutation of an RNase H domain yields a polypeptide exhibiting lower RNase activity, e.g., as determined by the methods described in Kotewicz et al. Nucleic Acids Res 16(l):265-277 (1988) (incorporated herein by reference in its entirety), e.g., lower by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% compared to an otherwise similar domain without the mutation. In some embodiments, RNase H activity is abolished.
[0352] In some embodiments, an RT domain is mutated to increase fidelity compared to an otherwise similar domain without the mutation. For instance, in some embodiments, a YADD (SEQ ID NO: 50) or YMDD (SEQ ID NO: 51) motif in an RT domain (e.g, in a reverse transcriptase) is replaced with YVDD (SEQ ID NO: 52). In some embodiments, replacement of a YADD (SEQ ID NO: 50), YMDD (SEQ ID NO: 51), or YVDD (SEQ ID NO: 52) motif results in higher fidelity in retroviral reverse transcriptase activity (e.g., as described in Jamburuthugoda and Eickbush J Mol Biol 2011; incorporated herein by reference in its entirety).
[0353] In certain embodiments, a gene modifying polypeptide comprises the amino acid sequence of an RT domain sequence from an AVIRE RT domain. In some embodiments, an RT domain amino acid sequence comprises one or more point mutations at an amino acid position of an RT domain as listed in columns 1 and 2 of Table 1, or an amino acid position corresponding thereto.Table 1. Positions that can be mutated in exemplary AVIRE RT domainsPage 68 of 39612985085vlAttorney Docket No.: 2017469-0049
[0354] In some embodiments, an RT domain for use in a gene modifying system described herein comprises an RT domain from AVIRE, PERV or BAEVM. In some embodiments, an RT domain comprises one or more mutations as listed in Table 2 below compared to a wild-type or canonical RT domain (e.g., from AVIRE, PERV or BAEVM). In some embodiments, an RT domain comprises one, two, three, four, five, or six of the mutations listed in the corresponding row of Table 2 below.Table 2. Exemplary RT domain mutations (relative to corresponding wild-type sequences)Page 69 of 39612985085vlAttorney Docket No.: 2017469-0049
[0355] In some embodiments, an RT domain comprises the amino acid sequence of an RT domain of an AVIRE RT (e.g., an AVIRE P03360 sequence, e.g., SEQ ID NO: 53), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity thereto. In some embodiments, an RT domain comprises the amino acid sequence of an AVIRE RT further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, G330P, L605W, T306K, and W313F, or a corresponding position in a homologous RT domain. In some embodiments, an RT domain comprises the amino acid sequence of an AVIRE RT further comprising one, two, or three mutations selected from the group consisting of D200N, G330P, and L605W, or a corresponding position in a homologous RT domain.
[0356] In some embodiments, an RT domain comprises the amino acid sequence of an RT domain of a PERV RT (e g., a PERV Q4VFZ2 sequence, e g., SEQ ID NO: 55 or SEQ ID NO: 56), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, an RT domain comprises the amino acid sequence of a PERV RT further comprising one, two, three, four, or five mutations selected from the group consisting of D196N, E326P, L599W, T302K, and W309F, or a corresponding position in a homologous RT domain. In some embodiments, an RT domain comprises the amino acid sequence of a PERV RT further comprising one, two, or three mutations selected from the group consisting of D196N, E326P, and L599W, or a corresponding position in a homologous RT domain.
[0357] In some embodiments, an RT domain comprises the amino acid sequence of an RT domain of a BAEVM RT (e.g., an BAEVM P 10272 sequence, e.g., SEQ ID NO: 60), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity thereto. In some embodiments, an RT domain comprises the amino acid sequence of a BAEVM RT further comprising one, two, three, four, or five mutations selected from the group consisting of D198N, E328P, L602W, T304K, and W311F, or a corresponding position in a homologous RT domain. In some embodiments, an RT domain comprises the amino acid sequence of a BAEVM RT further comprising one, two, or three mutations selected from the group consisting of D198N, E328P, and L602W, or a corresponding position in a homologous RT domain.
[0358] In some embodiments, a gene modifying polypeptide described herein comprises an RT domain having an amino acid sequence according to Table 3, or a sequence Page 70 of 39612985085vlAttorney Docket No.: 2017469-0049 having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide described herein comprises an RT domain encoded by a nucleic acid sequence according to Table 4, or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a nucleic acid described herein encodes an RT domain having an amino acid sequence according to Table 3, or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.Table 3. Exemplary reverse transcriptase domains from retroviruses (amino acid sequences)Page 71 of 39612985085vlAttorney Docket No.: 2017469-0049Page 72 of 39612985085vlAttorney Docket No.: 2017469-0049Page 73 of 39612985085vlAttorney Docket No.: 2017469-0049Page 74 of 39612985085vlAttorney Docket No.: 2017469-0049Table 4. Exemplary reverse transcriptase domains from retroviruses (nucleotide sequences)Page 75 of 39612985085vlAttorney Docket No.: 2017469-0049Page 76 of 39612985085vlAttorney Docket No.: 2017469-0049Page 77 of 39612985085vlAttorney Docket No.: 2017469-0049Page 78 of 39612985085vlAttorney Docket No.: 2017469-0049
[0359] In some embodiments, an RT domain described herein is modified, for example by site-specific mutation. In some embodiments, an RT domain is engineered to have improved properties. In some embodiments, an RT domain may be engineered to have lower error rates as compared to a reference RT domain, e.g., as described in WO200 1068895, incorporated herein by reference. In some embodiments, the reverse transcriptase domain may be engineered to be more thermostable as compared to a reference RT domain. In some embodiments, an RT domain may be engineered to be more processive as compared to a reference RT domain. In some embodiments, an RT domain may be engineered to have improved tolerance to inhibitors as compared to a reference RT domain. In some embodiments, an RT domain may be engineered to be faster as compared to a reference RT domain. In some embodiments, an RT domain may be engineered to better tolerate modified nucleotides in an RNA template as compared to a reference RT domain. In some embodiments, an RT domain may be engineered to be capable of inserting modified DNA nucleotides. In some embodiments, an RT domain is engineered to bind a template RNA.
[0360] In some embodiments, a gene modifying polypeptide comprising an RT domain from a retroviral reverse transcriptase domain may comprise one or more mutations from a wild-type sequence that may improve features of the RT, e.g., thermostability, processivity, and / or template binding.
[0361] In some embodiments, a gene modifying polypeptide comprises the RT domain from a retroviral reverse transcriptase domain comprising an amino acid sequence according to SEQ ID NO: 20, or an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0362] In some embodiments, a gene modifying polypeptide comprises the RT domain from a retroviral reverse transcriptase domain comprising an amino acid sequence according to SEQ ID NO: 46, or an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0363] In some embodiments, a gene modifying polypeptide comprises the RT domain from a retroviral reverse transcriptase domain comprising an amino acid sequence according to SEQ ID NO: 47, or an amino acid sequence having at least 75%, at least 80%, atPage 79 of 39612985085vlAttorney Docket No.: 2017469-0049 least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0364] In some embodiments, a writing domain (e.g., an RT domain) comprises an RNA-binding domain, e.g., that specifically binds to an RNA sequence. In some embodiments, a template RNA comprises an RNA sequence that is specifically bound by an RNA-binding domain of a writing domain (e.g., an RT domain).
[0365] In some embodiments, an RT domain only recognizes and reverse transcribes a specific template, e.g., a template RNA of a gene modifying system of the present disclosure. In some embodiments, a template RNA comprises a sequence or structure that enables recognition and reverse transcription by a reverse transcription domain. In some embodiments, a template RNA comprises a sequence or structure that enables association with an RNA-binding domain of a gene modifying polypeptide component of a gene modifying system described herein. In some embodiments, a gene modifying system of the present disclosure reverse transcribes a template comprising an association sequence over a template lacking an association sequence.
[0366] In some embodiments, a writing domain (e.g., an RT domain) may also comprise DNA-dependent DNA polymerase activity, e.g., comprise enzymatic activity capable of writing DNA into the genome from a template DNA sequence. In some embodiments, DNA-dependent DNA polymerization is employed to complete second-strand synthesis of a target site edit. In some embodiments, DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in a gene modifying polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by an RT domain that is also capable of DNA-dependent DNA polymerization, e.g., second-strand synthesis. In some embodiments, DNA-dependent DNA polymerase activity is provided by a second polypeptide of a gene modifying system of the present disclosure. In some embodiments, DNA-dependent DNA polymerase activity is provided by an endogenous host cell polymerase that is optionally recruited to a target site by a component of a gene modifying system of the present disclosure.
[0367] In some embodiments, an RT domain has a lower probability of premature termination rate ( off) in vitro relative to a reference RT domain. In some embodiments, a reference RT domain is a viral RT domain, e.g., the RT domain from AVIRE.Page 80 of 39612985085vlAttorney Docket No.: 2017469-0049
[0368] In some embodiments, an RT domain has a lower probability of premature termination rate ( off) in vitro of less than about 5 x 10'3 / nucleotides, less than about 5 x 10"4 / nucleotides, or less than about 5 x 10'6 / nucleotides, e.g., as measured on a 1094 nucleotide RNA. In some embodiments, an in vitro premature termination rate is determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (incorporated by reference herein its entirety).
[0369] In some embodiments, an RT domain is able to complete at least about 30% or 50% of integrations in cells. The percent of complete integrations can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites that comprise at least 98% of the expected integrated sequence) by the number of total (including substantially full-length and partial) integration events in a population of cells. In some embodiments, the integrations in cells is determined (e.g., across the integration site) using long-read amplicon sequencing, e.g., as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated by reference herein in its entirety).
[0370] In some embodiments, quantifying integrations in cells comprises counting the fraction of integrations that contain at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% of a DNA sequence corresponding to a template RNA (e.g., a template RNA having a length of at least 0.05, at least 0.1, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, at least 1, at least 1.5, at least 2, at least 3, at least 4, or at least 5 kb, e.g., a length between 0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, 1.0-1.2, 1.2-1.4, 1.4-1.6, 1.6-1.8, 1.8-2.0, 2-3, 3-4, or 4-5 kb).
[0371] In some embodiments, an RT domain is capable of polymerizing dNTPs in vitro. In some embodiments, an RT domain is capable of polymerizing dNTPs in vitro at a rate between 0.1 - 50 nucleotides / sec (e.g., between 0.1-1, 1-10, or 10-50 nucleotides / sec). In some embodiments, polymerization of dNTPs by an RT domain is measured by a singlemolecule assay, e.g., as described in Schwartz and Quake (2009) PNAS 106(48):20294-20299 (incorporated by reference in its entirety).
[0372] In some embodiments, an RT domain has an in vitro error rate (e.g., misincorporation of nucleotides) of between 1 x IO’3- 1 x IO’4or 1 x IO’4- 1 x 10’5substitutions / nucleotide, e.g., as described in Yasukawa et al. (2017) Biochem Biophys ResPage 81 of 39612985085vlAttorney Docket No.: 2017469-0049Commun 492(2): 147-153 (incorporated herein by reference in its entirety). In some embodiments, an RT domain has an error rate (e.g., misincorporation of nucleotides) in cells (e.g., HEK293T cells, primary liver, or primary lung cells) of between 1 x IO’3- 1 x IO’4or 1 x 10'4- 1 x 10'5substitutions / nucleotide, e.g., by long-read amplicon sequencing, e.g., as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated by reference herein in its entirety).
[0373] In some embodiments, an RT domain is capable of performing reverse transcription of a target RNA in vitro. In some embodiments, an RT domain requires a primer of at least 3 nucleotides to initiate reverse transcription of a template. In some embodiments, reverse transcription of a target RNA is determined by detection of cDNA from the target RNA (e.g., when provided with a ssDNA primer, e.g., which anneals to a target with at least 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides at the 3' end), e.g., as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (incorporated herein by reference in its entirety).
[0374] In some embodiments, an RT domain performs reverse transcription at least 5 or 10 times more efficiently (e.g., by cDNA production), e.g., when converting its RNA template to cDNA, for example, as compared to an RNA template lacking a protein binding motif (e.g., a 3' UTR). In some embodiments, efficiency of reverse transcription is measured as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2): 147-153 (incorporated by reference herein in its entirety).
[0375] In some embodiments, an RT domain specifically binds a specific RNA template with higher frequency (e.g., about 5-fold or about 10-fold higher frequency) than any endogenous cellular RNA, e.g., when expressed in cells (e.g., HEK293T cells, primary liver cells, or primary lung cells). In some embodiments, frequency of specific binding between an RT domain and a template RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11): 5490-5501 (incorporated herein by reference in its entirety).2. Template nucleic acid binding domain
[0376] In some embodiments, a gene modifying polypeptide contains regions capable of associating with a template nucleic acid (e.g., template RNA). In some embodiments, a template nucleic acid binding domain is an RNA binding domain. In some embodiments, an RNA binding domain is a modular domain that can associate with RNA molecules containing specific signatures, e.g., structural motifs. In some embodiments, a template nucleic acidPage 82 of 39612985085vlAttorney Docket No.: 2017469-0049 binding domain (e.g., RNA binding domain) is contained within an RT domain, e.g., the reverse transcriptase-derived component has a known signature for RNA preference.
[0377] In some embodiments, a template nucleic acid binding domain (e.g., RNA binding domain) is contained within a target DNA binding domain. For example, in some embodiments, a DNA binding domain is a CRISPR-associated protein that recognizes the structure of a template nucleic acid (e.g., a template RNA) comprising a gRNA. In some embodiments, a gene modifying polypeptide comprises a DNA-binding domain comprising a CRISPR-associated protein that associates with a gRNA scaffold that allows the DNA- binding domain to bind a target genomic DNA sequence. In some embodiments, a gRNA scaffold and a gRNA spacer is comprised within a template nucleic acid (e.g., template RNA), thus, in some embodiments a DNA-binding domain is also a template nucleic acid binding domain. In some embodiments, a gene modifying polypeptide possesses RNA binding function in multiple domains, e.g., can bind a gRNA structure in a CRISPR- associated DNA binding domain and an additional sequence, or structure in an RT domain.
[0378] In some embodiments, an RNA binding domain is capable of binding to a template RNA with greater affinity than a reference RNA binding domain. In some embodiments, a reference RNA binding domain is an RNA binding domain from Cas9 of S. pyogenes. In some embodiments, an RNA binding domain is capable of binding to a template RNA with an affinity between 100 pM - 10 nM (e.g., between 100 pM - 1 nM or 1 nM - 10 nM). In some embodiments, the affinity of an RNA binding domain for a template RNA is measured in vitro, e.g., by thermophoresis, e.g., as described in Asmari et al. Methods 146: 107-119 (2018) (incorporated by reference herein in its entirety). In some embodiments, the affinity of an RNA binding domain for a template RNA is measured in cells (e.g., by FRET or CLIP-Seq).
[0379] In some embodiments, an RNA binding domain is associated with a template RNA in vitro at a frequency at least about 5-fold or at least aboutlO-fold higher than with a scrambled RNA. In some embodiments, the frequency of association between an RNA binding domain and a template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501 (incorporated by reference herein in its entirety). In some embodiments, an RNA binding domain is associated with a template RNA in cells (e.g., in HEK293T cells, primary liver cells, or primary lung cells) at a frequency at least about 5-fold or at least about 10-fold higher than with a scrambled RNA. In some embodiments, the frequency of association between the RNAPage 83 of 39612985085vlAttorney Docket No.: 2017469-0049 binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019), supra.3. Endonuclease domains and DNA binding domains
[0380] In some embodiments, a gene modifying polypeptide possesses the function of DNA target site cleavage via an endonuclease domain. In some embodiments, a gene modifying polypeptide comprises a DNA binding domain, e.g., for binding to a target nucleic acid. In some embodiments, a domain (e.g., a Cas domain) of a gene modifying polypeptide comprises two or more smaller domains, e.g., a DNA binding domain and an endonuclease domain. It is understood that when a DNA binding domain (e.g., a Cas domain) is said to bind to a target nucleic acid sequence, in some embodiments, the binding is mediated by a gRNA.
[0381] In some embodiments, a domain has two or more functions. For example, in some embodiments, an endonuclease domain is also a DNA-binding domain. In some embodiments, an endonuclease domain is also a template nucleic acid (e.g., a template RNA) binding domain. For example, in some embodiments, a gene modifying polypeptide comprises a CRISPR-associated endonuclease domain that binds a template RNA comprising a gRNA, binds a target DNA sequence (e.g., with complementarity to a portion of the gRNA), and cuts the target DNA sequence. In some embodiments, an endonuclease domain or endonuclease / DNA-binding domain from a heterologous source can be used or can be modified (e.g., by insertion, deletion, or substitution of one or more residues) in a gene modifying system described herein.
[0382] In some embodiments, a nucleic acid encoding an endonuclease domain or endonuclease / DNA binding domain is altered from its natural sequence to have altered codon usage, e.g. improved for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9), a type-II restriction endonuclease (e.g., Fokl), a meganuclease (e.g., LScel), or other endonuclease domain.
[0383] In some embodiments, a DNA-binding domain of a gene modifying polypeptide described herein is selected, designed, or constructed for binding to a desired host DNA target sequence. In some embodiments, a DNA-binding domain of a gene modifying polypeptide is a heterologous DNA-binding element. In some embodiments, a heterologous DNA binding element is a zinc-finger element or a TAL effector element, e.g., aPage 84 of 39612985085vlAttorney Docket No.: 2017469-0049 zinc-finger or TAL polypeptide or functional fragment thereof. In some embodiments, a heterologous DNA binding element is a sequence-guided DNA binding element, such as Cas9, Cpfl, or other CRISPR-r elated protein that has been altered to have no endonuclease activity. In some embodiments, a heterologous DNA binding element retains endonuclease activity. In some embodiments, a heterologous DNA binding element retains partial endonuclease activity to cleave ssDNA, e.g., possesses nickase activity. In some embodiments, a heterologous DNA-binding domain comprises a Cas9 domain, a TAL domain, a ZF domain, a Myb domain, a combination thereof, or multiples thereof.
[0384] In some embodiments, a DNA-binding domain is modified, for example by site-specific mutation, increasing or decreasing DNA-binding elements (for example, number, and / or specificity of zinc fingers), etc., to alter DNA-binding specificity and affinity. In some embodiments, a nucleic acid sequence encoding a DNA binding domain is altered from its natural sequence to have altered codon usage, e.g., improved for human cell expression. In some embodiments, a DNA binding domain comprises one or more modifications relative to a wild-type DNA binding domain, e.g., a modification via directed evolution, e.g., phage-assisted continuous evolution (PACE).
[0385] In some embodiments, a DNA binding domain comprises a meganuclease domain (e.g., as described herein), or a functional fragment thereof. In some embodiments, a meganuclease domain possesses endonuclease activity, e.g., double-strand cleavage and / or nickase activity. In some embodiments, a meganuclease domain has reduced activity, e.g., lacks endonuclease activity, e.g., the meganuclease is catalytically inactive. In some embodiments, a catalytically inactive meganuclease is used as a DNA binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012), incorporated herein by reference in its entirety.
[0386] In some embodiments, a gene modifying polypeptide comprises a modification to a DNA-binding domain, e.g., relative to the wild-type polypeptide. In some embodiments, a DNA-binding domain comprises an addition, deletion, replacement, or modification to the amino acid sequence of the original DNA-binding domain. In some embodiments, a DNA-binding domain is modified to include a heterologous functional domain that binds specifically to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, a functional domain replaces at least a portion of (e.g., the entirety of) the prior DNA-binding domain of the gene modifying polypeptide. In some embodiments, a functional domain comprises a zinc finger (e.g., a zinc finger that specifically binds to thePage 85 of 39612985085vlAttorney Docket No.: 2017469-0049 target nucleic acid (e.g., DNA) sequence of interest). In some embodiments, a functional domain comprises a Cas domain (e.g., a Cas domain that specifically binds to the target nucleic acid (e.g., DNA) sequence of interest). In some embodiments, a Cas domain comprises a Cas9 or a mutant or variant thereof (e.g., as described herein). In some embodiments, a Cas domain is associated with a guide RNA (gRNA), e.g., as described herein. In some embodiments, a Cas domain is directed to a target nucleic acid (e.g., DNA) sequence of interest by a gRNA. In some embodiments, a Cas domain is encoded in the same nucleic acid (e.g., RNA) molecule as a gRNA. In some embodiments, a Cas domain is encoded in a different nucleic acid (e.g., RNA) molecule from a gRNA.
[0387] In some embodiments, a DNA binding domain is capable of binding to a target sequence (e.g., a dsDNA target sequence) with greater affinity than a reference DNA binding domain. In some embodiments, a reference DNA binding domain is a DNA binding domain from Cas9 of S. pyogenes. In some embodiments, a DNA binding domain is capable of binding to a target sequence (e.g., a dsDNA target sequence) with an affinity between 100 pM - 10 nM (e.g., between 100 pM-1 nM or 1 nM - 10 nM).
[0388] In some embodiments, the affinity of a DNA binding domain for a target sequence (e.g., a dsDNA target sequence) is measured in vitro, e.g., by thermophoresis, e.g., as described in Asmari et al. Methods 146: 107-119 (2018) (incorporated by reference herein in its entirety).
[0389] In some embodiments, a DNA binding domain is capable of binding to a target sequence (e.g., dsDNA target sequence), e.g, with an affinity between 100 pM - 10 nM (e.g., between 100 pM-1 nM or 1 nM - 10 nM) in the presence of a molar excess of scrambled sequence competitor dsDNA, e.g., of about 100-fold molar excess.
[0390] In some embodiments, a DNA binding domain is found associated with a target sequence (e.g., dsDNA target sequence) more frequently than any other sequence in the genome of a target cell, e.g., human target cell, e.g., as measured by ChlP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) Curr ProtocMol Biol Chapter 21 (incorporated herein by reference in its entirety). In some embodiments, a DNA binding domain is found associated with a target sequence (e.g., a dsDNA target sequence) at least about 5-fold or at least about 10-fold, more frequently than any other sequence in the genome of a target cell, e.g., as measured by ChlP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010), supra.Page 86 of 39612985085vlAttorney Docket No.: 2017469-0049
[0391] In some embodiments, an endonuclease domain has nickase activity and cleaves one strand of a target DNA. In some embodiments, nickase activity reduces the formation of double-stranded breaks at a target site. In some embodiments, an endonuclease domain creates a staggered nick structure in the first and second strands of a target DNA. In some embodiments, a staggered nick structure generates free 3’ overhangs at a target site. In some embodiments, free 3’ overhangs at a target site improve editing efficiency, e.g., by enhancing access and annealing of a 3’ homology region of a template nucleic acid. In some embodiments, a staggered nick structure reduces the formation of double-stranded breaks at a target site.
[0392] In some embodiments, an endonuclease domain cleaves both strands of a target DNA, e.g., results in blunt-end cleavage of a target with no ssDNA overhangs on either side of a cut-site. The amino acid sequence of an endonuclease domain of a gene modifying system described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of an endonuclease domain described herein, e.g., an endonuclease domain from Table 5.
[0393] In some embodiments, a heterologous endonuclease is Fokl or a functional fragment thereof. In some embodiments, a heterologous endonuclease is a Holliday junction resolvase or homolog thereof, such as the Holliday junction resolving enzyme from Sulfolobus solfataricus — Ssol Hje (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In some embodiments, a heterologous endonuclease is an endonuclease of a large fragment of a spliceosomal protein, such as Prp8 (Mahbub et al., Mobile DNA 8: 16, 2017).
[0394] In some embodiments, a heterologous endonuclease is derived from a CRISPR-associated protein, e.g., Cas9. In some embodiments, a heterologous endonuclease is engineered to have only ssDNA cleavage activity, e.g., only nickase activity, e.g., be a Cas9 nickase, e.g., StlCas9 with an N622 A mutation. Table 5 provides exemplary Cas proteins and mutations associated with nickase activity. In some embodiments, an endonuclease domain is modified, for example by site-specific mutation, to alter DNA endonuclease activity. In some embodiments, an endonuclease domain is modified to reduce DNA- sequence specificity, e.g., by truncation to remove domains that confer DNA-sequence specificity or mutation to inactivate regions conferring DNA-sequence specificity.Page 87 of 39612985085vlAttorney Docket No.: 2017469-0049
[0395] In some embodiments, an endonuclease domain has nickase activity and does not form double-stranded breaks. In some embodiments, an endonuclease domain forms single-stranded breaks at a higher frequency than double-stranded breaks, e.g., at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the breaks are singlestranded breaks, or less than 10%, less than 5%, less than 4%, less than 3%, less than 2%, or less than 1% of the breaks are double-stranded breaks. In some embodiments, an endonuclease forms substantially no double-stranded breaks. In some embodiments, an endonuclease does not form detectable levels of double-stranded breaks.
[0396] In some embodiments, an endonuclease domain has nickase activity that nicks the first strand of a target site DNA. In some embodiments, an endonuclease domain cuts the genomic DNA of a target site near to the site of alteration on the strand that will be extended by a writing domain (e.g., an RT domain). In some embodiments, an endonuclease domain has nickase activity that nicks the first strand of a target site DNA and does not nick the second strand of the target site DNA. For example, when a gene modifying polypeptide comprises a CRISPR-associated endonuclease domain having nickase activity, in some embodiments, said CRISPR-associated endonuclease domain nicks a target site DNA strand containing a PAM site (e.g., and does not nick the target site DNA strand that does not contain the PAM site). As a further example, when a gene modifying polypeptide comprises a CRISPR-associated endonuclease domain having nickase activity, in some embodiments, said CRISPR-associated endonuclease domain nicks a target site DNA strand that does not contain a PAM site (e.g., and does not nick the target site DNA strand that contains the PAM site).
[0397] In some embodiments, an endonuclease domain has nickase activity that nicks the first strand and the second strand of a target site DN. Without wishing to be bound by any particular theory, after a writing domain (e.g., an RT domain) of a gene modifying polypeptide described herein polymerizes (e.g., reverse transcribes) from a heterologous object sequence of a template nucleic acid (e.g., a template RNA), the cellular DNA repair machinery must repair the nick on the first DNA strand. The target site DNA now contains two different sequences for the first DNA strand: one corresponding to the original genomic DNA (e.g., having a free 5' end) and a second corresponding to that polymerized from the heterologous object sequence (e.g., having a free 3' end). It is thought that the two different sequences equilibrate with one another, first one hybridizing the second strand, then the other, and which sequence the cellular DNA repair apparatus incorporates into its repaired targetPage 88 of 39612985085vlAttorney Docket No.: 2017469-0049 site may be a stochastic process. Without wishing to be bound by any particular theory, it is thought that introducing an additional nick to the second-strand may bias the cellular DNA repair machinery to adopt the heterologous object sequence-based sequence more frequently than the original genomic sequence (Anzalone et al. Nature 576: 149-157 (2019)). In some embodiments, an additional nick is positioned at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 105, at least 110, at least 115, at least 120, at least 125, at least 130, at least 135, at least 140, at least145, or at least 150 nucleotides 5' or 3' of a target site modification (e.g., an insertion, deletion, or substitution) or to a nick on the first strand.
[0398] Alternatively or additionally, without wishing to be bound by theory, it is thought that an additional nick to the second strand may promote second-strand synthesis. In some embodiments, where a gene modifying system of the present disclosure has inserted or substituted a portion of a first strand, synthesis of a new sequence corresponding to the insertion / substitution in the second strand is necessary.
[0399] In some embodiments, a gene modifying polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain) and said domain nicks both the first strand and the second strand. For example, in some embodiments, an endonuclease domain may be a CRISPR-associated endonuclease domain, and a template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs nicking of the first strand and an additional gRNA spacer that directs nicking of the second strand. In some embodiments, a gene modifying polypeptide comprises a plurality of domains having endonuclease activity, and a first endonuclease domain nicks the first strand and a second endonuclease domain nicks the second strand (optionally, the first endonuclease domain does not (e.g., cannot) nick the second strand and the second endonuclease domain does not (e.g., cannot) nick the first strand).
[0400] In some embodiments, an endonuclease domain is capable of nicking a first strand and a second strand. In some embodiments, first and second strand nicks occur at the same position in a target site but on opposite strands. In some embodiments, a second strand nick occurs in a staggered location, e.g., upstream or downstream, from a first nick. In some embodiments, an endonuclease domain generates a target site deletion if a second strand nick is upstream of a first strand nick. In some embodiments, an endonuclease domain generates a target site duplication if a second strand nick is downstream of a first strand nick. In some Page 89 of 39612985085vlAttorney Docket No.: 2017469-0049 embodiments, an endonuclease domain generates no duplication and / or deletion if the first and second strand nicks occur in the same position of the target site. In some embodiments, the endonuclease domain has altered activity depending on protein conformation or RNA- binding status, e.g., which promotes the nicking of the first or second strand (e.g., as described in Christensen et al. PNAS 2006; incorporated by reference herein in its entirety).
[0401] In some embodiments, an endonuclease domain comprises a meganuclease, or a functional fragment thereof. In some embodiments, an endonuclease domain comprises a homing endonuclease, or a functional fragment thereof. In some embodiments, an endonuclease domain comprises a meganuclease from the LAGLID ADG, GIY-YIG, HNH, His-Cys Box, or PD-(DZE) XK families, or a functional fragment or variant thereof, e.g., which possess conserved amino acid motifs, e.g., as indicated in the family names. In some embodiments, an endonuclease domain comprises a meganuclease, or fragment thereof, chosen from, e.g., SmaMI (Uniprot F7WD42), I-Scel (Uniprot P03882), I- Anil (Uniprot P03880), I-Dmol (Uniprot P21505), LCrel (Uniprot P05725), I-TevI (Uniprot Pl 3299), I- Onul (Uniprot Q4VWW5), or I-Bmol (Uniprot Q9ANR6). In some embodiments, a meganuclease is naturally monomeric, e.g., I-Scel, I-TevI, or dimeric, e.g., LCrel, in its functional form. For example, LAGLID ADG meganucleases with a single copy of the LAGLID ADG motif generally form homodimers, whereas members with two copies of the LAGLID ADG motif are generally found as monomers. In some embodiments, a meganuclease that normally forms as a dimer is expressed as a fusion, e.g., two subunits are expressed as a single open reading frame (ORF) and, optionally, connected by a linker, e.g., an LCrel dimer fusion (Rodriguez-Fomes et al. Gene Therapy 2020; incorporated by reference herein in its entirety). In some embodiments, a meganuclease, or a functional fragment thereof, is altered to favor nickase activity for one strand of a double-stranded DNA molecule, e.g., LScel (K122I and / or K223I) (Niu et al. J Mol Biol 2008), LAnil (K227M) (McConnell Smith et al. PNAS 2009), LDmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, a meganuclease or functional fragment thereof possessing this preference for single-strand cleavage is used as an endonuclease domain, e.g., with nickase activity. In some embodiments, an endonuclease domain comprises a meganuclease, or a functional fragment thereof, which naturally targets or is engineered to target a safe harbor site, e.g., an LCrel targeting SH6 site (Rodriguez-F ornes et al., supra). In some embodiments, an endonuclease domain comprises a meganuclease, or a functional fragment thereof, with a sequence tolerant catalytic domain, e.g., LTevI recognizing thePage 90 of 39612985085vlAttorney Docket No.: 2017469-0049 minimal motif CNNNG (Kleinstiver et al. PNAS 2012). In some embodiments, a target sequence tolerant catalytic domain is fused to a DNA binding domain, e.g., to direct activity, e.g., by fusing I-TevI to: (i) zinc fingers to create Tev-ZFEs (Kleinstiver et al. PNAS 2012),(ii) other meganucleases to create MegaTevs (Wolfs et al. Nucleic Acids Res 2014), and / or(iii) Cas9 to create TevCas9 (Wolfs et al. PNAS 2016).
[0402] In some embodiments, an endonuclease domain comprises a restriction enzyme, e.g., a Type IIS or Type IIP restriction enzyme. In some embodiments, an endonuclease domain comprises a Type IIS restriction enzyme, e.g., FokI, or a fragment or variant thereof. In some embodiments, an endonuclease domain comprises a Type IIP restriction enzyme, e.g., PvuII, or a fragment or variant thereof. In some embodiments, a dimeric restriction enzyme is expressed as a fusion such that it functions as a single chain, e.g., a FokI dimer fusion (Minczuk et al. Nucleic Acids Res 36(12):3926-3938 (2008)).
[0403] The use of additional endonuclease domains is described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565 (2017), which is incorporated herein by reference in its entirety.
[0404] In some embodiments, a gene modifying polypeptide comprises a modification to an endonuclease domain, e.g., relative to a wild-type Cas protein. In some embodiments, an endonuclease domain comprises an addition, deletion, replacement, or modification to the amino acid sequence of a wild-type Cas protein. In some embodiments, an endonuclease domain is modified to include a heterologous functional domain that binds specifically to and / or induces endonuclease cleavage of a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, an endonuclease domain comprises a zinc finger. In some embodiments, an endonuclease domain comprising a Cas domain is associated with a guide RNA (gRNA), e.g., as described herein. In some embodiments, an endonuclease domain is modified to include a functional domain that does not target a specific target nucleic acid (e.g., DNA) sequence. In some embodiments, an endonuclease domain comprises a FokI domain.
[0405] In some embodiments, an endonuclease domain is associated with a target dsDNA in vitro at a frequency at least about 5-fold or at least about 10-fold higher than with a scrambled dsDNA. In some embodiments, an endonuclease domain is associated with a target dsDNA in vitro at a frequency at least about 5-fold or at least about 10-fold higher than with a scrambled dsDNA, e.g., in a cell (e.g., a HEK293T cell, a primary liver cell, or a primaryPage 91 of 39612985085vlAttorney Docket No.: 2017469-0049 lung cell). In some embodiments, the frequency of association between an endonuclease domain and a target DNA or scrambled DNA is measured by ChlP-seq, e.g., as described in He and Pu (2010) Curr Protoc Mol Biol Chapter 21 (incorporated by reference herein in its entirety).
[0406] In some embodiments, an endonuclease domain can catalyze the formation of a nick at a target sequence, e.g., to an increase of at least about 5-fold or at least about 10-fold relative to a non-target sequence (e.g., relative to any other genomic sequence in the genome of the target cell). In some embodiments, the level of nick formation is determined using NickSeq, e.g., as described in Elacqua et al. (2019) bioRxiv doi.org / 10.1101 / 867937 (incorporated herein by reference in its entirety).
[0407] In some embodiments, an endonuclease domain is capable of nicking DNA / / / vitro. In some embodiments, a nick results in an exposed base. In some embodiments, an exposed base can be detected using a nuclease sensitivity assay, e.g., as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res 23(19):3805-3809 (incorporated by reference herein in its entirety). In some embodiments, the level of exposed bases (e.g., detected by the nuclease sensitivity assay) is increased by at least about 10%, at least about 50%, or more relative to a reference endonuclease domain. In some embodiments, a reference endonuclease domain is an endonuclease domain from Cas9 of S. pyogenes.
[0408] In some embodiments, an endonuclease domain is capable of nicking DNA in a cell. In some embodiments, an endonuclease domain is capable of nicking DNA in a HEK293T cell. In some embodiments, an unrepaired nick that undergoes replication in the absence of Rad51 results in increased NHEJ rates at the site of the nick, which can be detected, e.g., by using a Rad51 inhibition assay, e.g., as described in Bothmer et al. (2017) Nat Commun 8: 13905 (incorporated by reference herein in its entirety). In some embodiments, NHEJ rates are increased above 0-5%. In some embodiments, NHEJ rates are increased to 20-70% (e.g., between 30%-60% or 40-50%), e.g., upon Rad51 inhibition.
[0409] In some embodiments, an endonuclease domain releases a target after cleavage. In some embodiments, release of a target is indicated indirectly by assessing for multiple turnovers by an enzyme, e.g., as described in Yourik at al. RNA 25(l):35-44 (2019) (incorporated herein by reference in its entirety) and shown in FIG. 2 therein. In some embodiments, the kexpof an endonuclease domain is 1 x 10'3- 1 x 10'5 min-1 as measured by such methods.Page 92 of 39612985085vlAttorney Docket No.: 2017469-0049
[0410] In some embodiments, an endonuclease domain has a catalytic efficiency kcai / Km) greater than about 1 x 108s'1M'1in vitro. In some embodiments, an endonuclease domain has a catalytic efficiency greater than about 1 x 105, greater than about 1 x 106, greater than about 1 x 107, or greater than about 1 x 108, s'1M'1in vitro. In some embodiments, catalytic efficiency is determined as described in Chen et al. (2018) Science 360(6387):436-439 (incorporated herein by reference in its entirety). In some embodiments, an endonuclease domain has a catalytic efficiency (W£m) greater than about 1 x 108s'1M'1in cells. In some embodiments, the endonuclease domain has a catalytic efficiency greater than about 1 x 105, greater than about 1 x 106, greater than about 1 x 107, or greater than about 1 x 108s'1M'1in cells.4. Gene modifying polypeptides comprising Cas domains
[0411] In some embodiments, a gene modifying polypeptide described herein comprises a Cas domain. In some embodiments, a Cas domain can direct a gene modifying polypeptide to a target site specified by a gRNA spacer, thereby modifying a target nucleic acid sequence in “cis.” In some embodiments, a gene modifying polypeptide is fused to a Cas domain. In some embodiments, a gene modifying polypeptide comprises a CRISPR / Cas domain (also referred to herein as a CRISPR-associated protein). In some embodiments, a CRISPR / Cas domain comprises a protein involved in the clustered regulatory interspaced short palindromic repeat (CRISPR) system, e.g., a Cas protein, and optionally binds a guide RNA, e.g., single guide RNA (sgRNA).
[0412] CRISPR systems are adaptive defense systems originally discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases termed CRISPR- associated or “Cas” endonucleases (e.g., Cas9 or Cpfl) to cleave foreign DNA. For example, in a typical CRISPR-Cas system, an endonuclease is directed to a target nucleotide sequence (e. g., a site in the genome that is to be sequence-edited) by sequence-specific, non-coding “guide RNAs” that target single- or double-stranded DNA sequences. Three classes (I-III) of CRISPR systems have been identified. The class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One class II CRISPR system includes a type II Cas endonuclease such as Cas9, a CRISPR RNA (“crRNA”), and a trans-activating crRNA (“tracrRNA”). The crRNA contains a “spacer” sequence, a typically about 20- nucleotide RNA sequence that corresponds to a target DNA sequence (“protospacer”). In the wild-type system, and in some engineered systems, crRNA also contains a region that binds to the tracrRNA to form a partially double-stranded structure that is cleaved by RNase III,Page 93 of 39612985085vlAttorney Docket No.: 2017469-0049 resulting in a crRNA / tracrRNA hybrid molecule. A crRNA / tracrRNA hybrid then directs the Cas endonuclease to recognize and cleave a target DNA sequence. A target DNA sequence is generally adjacent to a “protospacer adjacent motif’ (“PAM”) that is specific for a given Cas endonuclease and required for cleavage activity at a target site matching the spacer of the crRNA. CRISPR endonucleases identified from various prokaryotic species have unique PAM sequence requirements; examples of PAM sequences include 5 -NGG (Streptococcus pyogenes), 5'-NNAGAA (Streptococcus thermophilus CRISPR1; SEQ ID NO: 65), 5'- NGGNG (Streptococcus thermophilus CRISPR3; SEQ ID NO: 66), and 5'-NNNGATT (Neisseria meningiditis; SEQ ID NO: 67). Some endonucleases, e.g., Cas9 endonucleases, are associated with G-rich PAM sites, e.g., 5 -NGG), and perform blunt-end cleaving of the target DNA at a location 3 nucleotides upstream from (5' from) the PAM site. Another class II CRISPR system includes the type V endonuclease Cpfl, which is smaller than Cas9; examples include AsCpfl (from Acidaminococcus sp.) and LbCpfl (from Lachnospiraceae sp.). Cpfl -associated CRISPR arrays are processed into mature crRNAs without the requirement of a tracrRNA; in other words, a Cpfl system, in some embodiments, comprises only Cpfl nuclease and a crRNA to cleave a target DNA sequence. Cpfl endonucleases, are typically associated with T-rich PAM sites, e. g., 5'-TTN. Cpfl can also recognize a 5 -CTA PAM motif. Cpfl typically cleaves a target DNA by introducing an offset or staggered double-strand break with a 4- or 5-nucleotide 5' overhang, for example, cleaving a target DNA with a 5-nucleotide offset or staggered cut located 18 nucleotides downstream from (3' from) a PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complimentary strand; the 5-nucleotide overhang that results from such offset cleavage allows more precise genome editing by DNA insertion by homologous recombination than by insertion at blunt-end cleaved DNA. See, e.g., Zetsche et al. (2015) Cell, 163:759 - 771.
[0413] A variety of CRISPR associated (Cas) genes or proteins can be used in the technologies provided by the present disclosure and the choice of Cas protein will depend upon the particular conditions of the method. Specific examples of Cas proteins include class II systems including Casl, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, CaslO, Cpfl, C2C1, or C2C3. In some embodiments, a Cas protein, e.g., a Cas9 protein, may be from any of a variety of prokaryotic species. In some embodiments, a Cas protein, e.g., a Cas9 protein, is selected to recognize a particular protospacer-adjacent motif (PAM) sequence. In some embodiments, a DNA-binding domain or endonuclease domain includes a sequence targeting polypeptide, such as a Cas protein, e.g., Cas9. In some embodiments, a Cas protein, e.g., aPage 94 of 39612985085vlAttorney Docket No.: 2017469-0049Cas9 protein, may be obtained from a bacteria or archaea or synthesized using known methods. In some embodiments, a Cas protein may be from a gram-positive bacteria or a gram-negative bacteria. In some embodiments, a Cas protein may be from a Streptococcus (e.g., a S. pyogenes , or a S. thermophilus), a Francisella (e.g., an F. novicida), a Staphylococcus (e.g., an S. aureus), an Acidaminococcus (e.g., an Acidaminococcus sp. BV3L6), a Neisseria (e.g., an N. meningitidis), a Cryptococcus, a Corynebacterium, a Haemophilus, a Eubacterium, a Pasteurella, a Prevotella, a Veillonella, or a Marinobacter.
[0414] In some embodiments, a gene modifying polypeptide may comprise an amino acid sequence according to SEQ ID NO: 25, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto. In some embodiments, an amino acid sequence according to SEQ ID NO: 25, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto, is positioned at the N-terminal end of a gene modifying polypeptide. In some embodiments, an amino acid sequence of SEQ ID NO: 25, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto, is positioned within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids of the N- terminal end of a gene modifying polypeptide.
[0415] In some embodiments, a gene modifying polypeptide may comprise an amino acid sequence according to SEQ ID NO: 26, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto. In some embodiments, an amino acid sequence according to SEQ ID NO: 26, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto, is positioned at the C-terminal end of a gene modifying polypeptide. In some embodiments, an amino acid sequence according to SEQ ID NO: 26, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto, is positioned within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids of the C-terminal end of a gene modifying polypeptide.Page 95 of 39612985085vlAttorney Docket No.: 2017469-0049
[0416] In some embodiments, a gene modifying polypeptide may comprise a Cas domain comprising an amino acid sequence as listed in Table 5, or a functional fragment thereof, or a Cas domain compriaing an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% identity thereto. In some embodiments, a gene modifying polypeptide may comprise a Cas domain comprising an amino acid sequence as listed in Table 5 without an initial N-terminal methionine residue.Page 96 of 39612985085vlAttorney Docket No.: 2017469-0049Table 5. Sequences of Exemplary CRISPR / Cas Proteins, Species, and MutationsPage 97 of 39612985085vlAttorney Docket No.: 2017469-0049Page 98 of 39612985085vlAttorney Docket No.: 2017469-0049Page 99 of 39612985085vlAttorney Docket No.: 2017469-0049Page 100 of 39612985085vlAttorney Docket No.: 2017469-0049Page 101 of 39612985085vlAttorney Docket No.: 2017469-0049
[0417] In some embodiments, a gene modifying polypeptide may comprise a Cas domain encoded by a nucleotide sequence according toAGCGACCUUGUACUAGGCCUGGACAUCGGCAUCGGCAGCGUGGGCGUGGGCAU CCUGAACAAGGUGACCGGCGAGAUCAUCCACAAGAACAGCCGGAUCUUCCCCGCCGCCCAGGCCGAGAAUAACCUGGUGCGGCGGACAAACAGACAGGGACGGCGG UUAGCCCGGCGGAAAAAACACCGGCGGGUGCGGCUGAACAGACUCUUCGAGGAGAGCGGCCUGAUCACCGACUUCACCAAGAUCAGCAUCAACCUGAACCCCUACC AGCUGCGGGUGAAGGGCCUGACCGACGAGCUGAGCAACGAGGAGCUGUUCAUAGCCCUGAAGAACAUGGUGAAGCACCGGGGCAUCAGCUACCUGGACGACGCCA GCGACGACGGCAACAGCAGCGUGGGCGAUUACGCCCAGAUCGUGAAGGAGAACAGCAAGCAGCUGGAAACCAAGACCCCCGGCCAGAUCCAGCUGGAGCGGUACCA GACCUACGGCCAGCUGCGUGGCGACUUCACCGUGGAGAAGGACGGCAAGAAGCACCGGCUGAUCAACGUGUUUCCCACCAGCGCCUACAGAAGCGAGGCCCUGCGGAUCCUGCAGACCCAGCAGGAGUUCAACCCCCAGAUCACCGACGAGUUCAUCAA CCGGUACCUGGAGAUCCUGACCGGCAAGCGGAAGUACUACCACGGCCCCGGCAACGAAAAGAGCCGGACCGACUACGGCCGGUACCGGACCAGCGGCGAAACCCUG GACAACAUCUUCGGCAUCCUCAUCGGCAAGUGCACCUUCUACCCUGACGAGUUCCGGGCCGCCAAGGCCUCAUACACCGCCCAGGAGUUCAACCUGCUGAACGACC UGAACAACCUGACCGUGCCUACCGAGACUAAGAAGCUGAGCAAGGAGCAGAAGAACCAGAUCAUCAACUACGUGAAGAACGAGAAGGCCAUGGGCCCCGCCAAGCU GUUCAAGUACAUAGCCAAGCUGCUGAGCUGCGACGUGGCCGACAUCAAGGGCUACCGGAUCGACAAGAGCGGCAAGGCCGAGAUCCACACCUUCGAGGCCUACCGG AAGAUGAAAACCCUGGAAACCCUGGACAUCGAGCAGAUGGACCGGGAAACCCUGGACAAGCUGGCCUACGUGCUGACCCUGAACACCGAACGGGAGGGCAUCCAGG AAGCCCUGGAGCACGAGUUCGCCGACGGCAGCUUCAGCCAGAAGCAGGUAGACGAGCUGGUGCAGUUCCGGAAGGCCAACAGCAGCAUCUUCGGCAAGGGCUGGCA CAACUUCAGCGUGAAGCUGAUGAUGGAGCUGAUCCCCGAGCUGUACGAAACCAGCGAGGAGCAGAUGACCAUCCUGACCCGGCUGGGCAAGCAGAAAACCACCAGC AGCAGCAACAAGACCAAGUACAUCGACGAGAAGCUGCUGACCGAGGAGAUCUACAACCCCGUGGUGGCCAAGAGCGUGCGGCAGGCCAUCAAGAUCGUGAACGCCG CCAUCAAGGAGUACGGCGACUUCGACAACAUCGUGAUCGAGAUGGCCCGGGAGACAAACGAGGACGACGAGAAGAAGGCCAUCCAGAAGAUCCAGAAGGCCAACAA GGACGAGAAGGACGCCGCCAUGCUGAAGGCCGCCAACCAGUACAACGGCAAGGCCGAGCUGCCCCACAGCGUGUUCCACGGCCACAAGCAGCUGGCCACCAAGAUC CGGCUGUGGCACCAGCAGGGCGAGAGAUGCCUGUACACCGGCAAGACCAUCAGCAUCCACGACCUGAUCAACAACAGCAACCAGUUCGAGGUGGACCACAUCCUGC CACUGAGCAUCACCUUCGACGACAGCCUGGCCAACAAGGUGCUGGUGUACGCAACAGCCGCCCAGGAGAAGGGCCAGCGGACCCCAUACCAAGCCCUGGACAGCAU GGACGACGCCUGGAGCUUCCGGGAGCUGAAGGCCUUCGUGCGGGAGAGCAAGACCCUGAGCAACAAGAAGAAGGAGUACCUGCUGACCGAGGAGGACAUCAGCAAGUUCGACGUGCGGAAGAAGUUCAUCGAGCGGAACCUGGUGGACACCCGGUACGC CAGCCGGGUGGUGCUGAAUGCCCUGCAGGAGCACUUCCGGGCCCACAAGAUCGACACCAAGGUAAGCGUAGUGCGGGGCCAGUUCACCAGCCAGCUGCGGCGGCACUGGGGAAUCGAAAAGACCCGGGACACCUACCACCACCACGCCGUGGAUGCGCU GAUAAUCGCCGCCAGCAGCCAGCUGAACCUGUGGAAGAAGCAGAAGAACACCCUGGUGAGCUACAGCGAGGACCAGCUGCUGGACAUCGAAACCGGCGAGCUGAUC AGCGACGACGAGUACAAGGAGAGCGUGUUCAAGGCCCCCUACCAGCACUUCGUGGACACCUUAAAGAGCAAGGAGUUCGAGGACAGCAUCCUGUUCAGCUACCAGG UGGACAGCAAGUUCAACCGGAAGAUCAGCGACGCCACCAUCUACGCCACCCGGPage 102 of 39612985085vlAttorney Docket No.: 2017469-0049CAGGCCAAGGUGGGCAAGGACAAGGCCGACGAAACCUACGUGCUGGGCAAGAU CAAGGACAUCUACACCCAGGACGGCUACGACGCCUUCAUGAAGAUCUACAAGA AGGACAAGAGCAAGUUCCUGAUGUACCGGCACGACCCUCAGACCUUCGAGAAG GUGAUCGAGCCCAUCCUGGAGAACUACCCCAACAAGCAGAUCAACGAGAAGGG CAAGGAGGUGCCCUGCAACCCCUUCCUGAAGUACAAGGAGGAGCACGGCUACA UCCGGAAGUACAGCAAGAAGGGCAACGGCCCCGAGAUCAAGAGCCUGAAGUAC UACGACAGCAAGCUGGGCAACCACAUCGACAUCACCCCCAAGGACAGCAACAA CAAGGUGGUGCUGCAGAGCGUGAGCCCAUGGCGGGCCGACGUGUACUUCAACA AGACCACCGGCAAGUACGAGAUCCUGGGCCUGAAGUACGCCGACCUGCAGUUC GAGAAGGGCACCGGCACCUACAAGAUCAGCCAGGAGAAGUACAACGACAUCAA GAAGAAGGAGGGCGUGGACAGCGACAGCGAGUUCAAGUUCACCCUGUACAAG AACGACCUGCUGCUGGUGAAGGACACCGAAACCAAGGAGCAGCAGCUGUUUCG GUUCCUGAGCCGGACCAUGCCCAAGCAGAAGCACUACGUGGAGCUGAAGCCCU ACGACAAGCAGAAGUUCGAGGGCGGCGAGGCACUGAUCAAGGUGCUGGGCAA CGUGGCCAACAGCGGCCAGUGCAAGAAGGGCCUGGGCAAGAGCAACAUCAGCA UCUACAAGGUGCGGACCGACGUGCUGGGCAACCAGCACAUCAUCAAGAACGAG GGCGACAAGCCAAAGCUG (SEQ ID NO: 115) or a nucleotide sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity thereto.
[0418] In some embodiments, a Cas protein requires a protospacer adjacent motif (PAM) to be present in or adjacent to a target DNA sequence for the Cas protein to bind and / or function. In some embodiments, a PAM is or comprises, from 5' to 3', NGG, YG, NNGRRT (SEQ ID NO: 73), NNNRRT (SEQ ID NO: 74), NGA, TYCV (SEQ ID NO: 75), TATV (SEQ ID NO: 76), NTTN (SEQ ID NO: 77), or NNNGATT (SEQ ID NO: 78), where N stands for any nucleotide, Y stands for C or T, R stands for A or G, and V stands for A or C or G. In some embodiments, a Cas protein is a protein listed in Table 5. In some embodiments, a Cas protein comprises one or more mutations altering its PAM. In some embodiments, a Cas protein comprises E1369R, E1449H, and R1556A mutations or analogous substitutions to the amino acids corresponding to said positions. In some embodiments, a Cas protein comprises E782K, N968K, and R1015H mutations or analogous substitutions to the amino acids corresponding to said positions. In some embodiments, a Cas protein comprises DI 135V, R1335Q, and T1337R mutations or analogous substitutions to the amino acids corresponding to said positions. In some embodiments, a Cas protein comprises S542R and K607R mutations or analogous substitutions to the amino acids corresponding to said positions. In some embodiments, a Cas protein comprises S542R, K548V, and N552R mutations or analogous substitutions to the amino acids corresponding to said positions. Exemplary advances in the engineering of Cas enzymes to recognize altered PAM sequencesPage 103 of 39612985085vlAttorney Docket No.: 2017469-0049 are reviewed in Collias et al Nature Communications 12:555 (2021), incorporated herein by reference in its entirety.
[0419] In some embodiments, a Cas protein is catalytically active and cuts one or both strands of a target DNA site. In some embodiments, cutting a target DNA site is followed by formation of an alteration, e.g., an insertion or deletion, e.g., by the cellular repair machinery.
[0420] In some embodiments, a Cas protein is modified to deactivate or partially deactivate a nuclease, e.g., nuclease-deficient Cas9. Whereas wild-type Cas9 generates double-strand breaks (DSBs) at specific DNA sequences targeted by a gRNA, a number of CRISPR endonucleases having modified functionalities are available, for example: a “nickase” version of Cas9 that has been partially deactivated generates only a single-strand break; a catalytically inactive Cas9 (“dCas9”) does not cut target DNA. In some embodiments, dCas9 binding to a DNA sequence may interfere with transcription at that site by steric hindrance. In some embodiments, dCas9 binding to an anchor sequence may interfere with (e.g., decrease or prevent) genomic complex (e.g., ASMC) formation and / or maintenance. In some embodiments, a DNA-binding domain comprises a catalytically inactive Cas9, e.g., dCas9. In some embodiments, dCas9 comprises mutations in each endonuclease domain of the Cas protein, e.g., a N622 A mutation. In some embodiments, a catalytically inactive or partially inactive CRISPR / Cas domain comprises a Cas protein comprising one or more mutations.
[0421] In some embodiments, a catalytically inactive, e.g., dCas9, or partially deactivated Cas9 protein comprises a N622 mutation (e.g., N622A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive, e.g., dCas9, or partially deactivated Cas9 protein comprises a Dll mutation (e.g., Dll A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a H969 mutation (e.g., H969A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a N995 mutation (e.g., N995 A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises mutations at one, two, or three of positions Dll, H969,Page 104 of 39612985085vlAttorney Docket No.: 2017469-0049 and N995 (e.g., Dll A, H969A, and N995 A mutations) or analogous substitutions to the amino acids corresponding to said positions.
[0422] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a N622 mutation (e.g., a N622 A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a DIO mutation (e.g., a D10A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a H557 mutation (e.g., a H557 A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a DIO mutation (e.g., a D10A mutation) and a H557 mutation (e.g., a H557A mutation) or analogous substitutions to the amino acids corresponding to said positions.
[0423] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a D839 mutation (e.g., a D839A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a H840 mutation (e.g., a H840A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a N863 mutation (e.g., a N863 A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a DIO mutation (e.g., D10A), a D839 mutation (e.g., D839A), a H840 mutation (e.g., H840A), and a N863 mutation (e.g., N863A) or analogous substitutions to the amino acids corresponding to said positions.
[0424] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a E993 mutation (e.g., a E993 A mutation) or an analogous substitution to the amino acid corresponding to said position.
[0425] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a D917 mutation (e.g., a D917A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9Page 105 of 39612985085vlAttorney Docket No.: 2017469-0049 protein comprises a a E1006 mutation (e.g., a E1006A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a D1255 mutation (e.g., a D1255A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D917 mutation (e.g., D917A), a E1006 mutation (e.g., E1006A), and a D1255 mutation (e.g., D1255A) or analogous substitutions to the amino acids corresponding to said positions.
[0426] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a DI 6 mutation (e.g., a D16A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a D587 mutation (e.g., a D587A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a partially deactivated Cas domain has nickase activity. In some embodiments, a partially deactivated Cas9 domain is a Cas9 nickase domain. In some embodiments, the catalytically inactive Cas domain or dead Cas domain produces no detectable double strand break formation. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a H588 mutation (e.g., a H588A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein comprises a N611 mutation (e.g., a N611 A mutation) or an analogous substitution to the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D16 mutation (e.g., D16A), a D587 mutation (e.g., D587A), a H588 mutation (e.g., H588A), and a N611 mutation (e.g., N611A) or analogous substitutions to the amino acids corresponding to said positions.
[0427] In some embodiments, a DNA-binding domain or endonuclease domain may comprise a Cas molecule comprising or linked (e.g., covalently) to a gRNA (e.g., a template nucleic acid, e.g., template RNA, comprising a gRNA).
[0428] In some embodiments, an endonuclease domain or DNA binding domain comprises a Streptococcus thermophilus Cas9 (StlCas9) or a functional fragment or variant thereof. In some embodiments, an endonuclease domain or DNA binding domain comprises a modified StlCas9. In embodiments, a modified StlCas9 comprises a modification that altersPage 106 of 39612985085vlAttorney Docket No.: 2017469-0049 protospacer-adjacent motif (PAM) specificity. In some embodiments, a PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In some embodiments, an endonuclease domain or DNA binding domain comprises a Cas domain, e.g., a Cas9 domain. In some embodiments, an endonuclease domain or DNA binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas) domain. In some embodiments, an endonuclease domain or DNA binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas9) domain.
[0429] In some embodiments, an endonuclease domain or DNA binding domain comprises a Cas9 sequence, e.g., as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737; incorporated herein by reference. In some embodiments, an endonuclease domain or DNA binding domain comprises a HNH nuclease subdomain and / or a RuvCl subdomain of a Cas, e.g., Cas9, e.g., as described herein, or a variant thereof. In some embodiments, an endonuclease domain or DNA binding domain comprises a Cas polypeptide (e.g., enzyme), or a functional fragment thereof. In some embodiments, a Cas9 domain comprises one or more substitutions and / or one or more mutations.
[0430] In some embodiments, a Cas9 derivative with enhanced activity may be used in a gene modification polypeptide. In some embodiments, a Cas9 derivative may comprise mutations that improve activity of a HNH endonuclease domain, (see, e.g., Spencer and Zhang Sci Rep 7: 16836 (2017), the Cas9 derivatives and comprising mutations of which are incorporated herein by reference). In some embodiments, a Cas9 derivative may comprise one or more types of mutations described herein, e.g., PAM-modifying mutations, protein stabilizing mutations, activity enhancing mutations, and / or mutations partially or fully inactivating one or two endonuclease domains relative to the parental enzyme (e.g., one or more mutations to abolish endonuclease activity towards one or both strands of a target DNA, e.g., a nickase or catalytically dead enzyme). In some embodiments, a Cas9 enzyme used in a gene modifying system described herein may comprise one or more mutations that confer nickase activity toward the enzyme in addition to one or more mutations improving catalytic efficiency.5. Linkers
[0431] In some embodiments, a gene modifying polypeptide may comprise a linker, e.g., a peptide linker, e.g., a linker as described in Table 6. In some embodiments, a gene modifying polypeptide comprises, from N-terminus to C-terminus, a Cas domain (e.g., a CasPage 107 of 39612985085vlAttorney Docket No.: 2017469-0049 domain as listed in Table 5 or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto), a linker (e.g., a linker as listed in Table 6 or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto), and an RT domain (e.g., an RT domain as listed in Table 3 or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto). In some embodiments, a gene modifying polypeptide comprises a flexible linker between an endonuclease and an RT domain, e.g., a linker comprising an amino acid sequence according to SEQ ID NO: 18, or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises a flexible linker between an endonuclease and an RT domain, e.g., a linker comprising an amino acid sequence according to SEQ ID NO: 19, or a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto. In some embodiments, an RT domain of a gene modifying polypeptide may be located C-terminal to an endonuclease domain. In some embodiments, an RT domain of a gene modifying polypeptide may be located N-terminal to an endonuclease domain.Table 6. Exemplary linker sequences
[0432] In some embodiments, a gene modifying polypeptide comprises: (i) a linker comprising a linker sequence as listed in a row of Table 7, or an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%,Page 108 of 39612985085vlAttorney Docket No.: 2017469-0049 at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto; and (ii) an RT domain comprising an RT domain sequence as listed in the same row of Table 7, or an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto.Table 7. Selection of exemplary gene modifying polypeptidesPage 109 of 39612985085vlAttorney Docket No.: 2017469-0049
[0433] In some embodiments, a gene modifying polypeptide comprises: (i) a Cas domain comprising an Cas amino acid sequence as listed in a row of Table 8, or an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto; (ii) a linker comprising a linker amino acid sequence as listed in the same row of Table 8, or an amino acid sequence having at least 70%, at least 80%, at leastPage 110 of 39612985085vlAttorney Docket No.: 2017469-004985%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto; and (iii) an RT domain comprising an RT domain amino acid sequence as listed in the same row of Table 8, or an amino acid sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or 99% identity thereto.
[0434] In some embodiments, a gene modifying polypeptide comprises any Cas domain sequence listed in Table 5 and any RT domain sequence listed Table 3 where the Cas domain and RT domain are joined by any linker listed in Table 6.Page 111 of 39612985085vlAttorney Docket No.: 2017469-0049Table 8. Selection of exemplary gene modifying polypeptidePage 112 of 39612985085vlAttorney Docket No.: 2017469-0049Page 113 of 39612985085vlAttorney Docket No.: 2017469-0049Page 114 of 39612985085vlAttorney Docket No.: 2017469-0049Page 115 of 39612985085vlAttorney Docket No.: 2017469-0049Page 116 of 39612985085vlAttorney Docket No.: 2017469-00496. Localization sequences
[0435] In some embodiments, an RNA encoding a gene modifying polypeptide comprises an intracellular localization sequence, e.g., a nuclear localization sequence (NLS). In some embodiments, a gene modifying polypeptide comprises an NLS having an amino acid sequence according to SEQ ID NO: 15, SEQ ID NO: 16, and / or SEQ ID NO: 17, or an NLS having an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises an NLS having an amino acid sequence according to SEQ ID NO: 15 and / or SEQ ID NO: 17, or an NLS having an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises an NLS having an amino acid sequence according to SEQ ID NO: 16 and / or SEQ ID NO: 17, or an NLS having an amino acid sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0436] A nuclear localization sequence may be an RNA sequence that promotes the import of an RNA into the nucleus. In some embodiments, a nuclear localization signal is located on a template RNA. In some embodiments, a gene modifying polypeptide is encoded on a first RNA, and a template RNA is a second, separate, RNA, and a nuclear localization signal is located on the template RNA and not on an RNA encoding the gene modifying polypeptide. While not wishing to be bound by any particular theory, in some embodiments, an RNA encoding a gene modifying polypeptide is targeted primarily to the cytoplasm to promote its translation, while a template RNA is targeted primarily to the nucleus to promote insertion into the genome. In some embodiments a nuclear localization signal is at the 3' end, 5' end, or in an internal region, of a template RNA. In some embodiments, a nuclear localization signal is 3' of the heterologous sequence (e.g., is directly 3' of the heterologous sequence) or is 5' of the heterologous sequence (e.g., is directly 5' of the heterologous sequence). In some embodiments, a nuclear localization signal is placed outside of a 5' UTR or outside of a 3' UTR of a template RNA. In some embodiments, a nuclear localization signal is placed between a 5' UTR and a 3' UTR, wherein optionally the nuclear localization signal is not transcribed with a transgene (e.g., the nuclear localization signal is in an antisense orientation or is downstream of a transcriptional termination signal or polyadenylation signal). In some embodiments, a nuclear localization sequence is situated inside of an intron.Page 117 of 39612985085vlAttorney Docket No.: 2017469-0049In some embodiments a plurality of the same or different nuclear localization signals are in the RNA, e.g., in the template RNA. In some embodiments, a nuclear localization signal is less than 5, less than 10, less than 25, less than 50, less than 75, less than 100, less than 150, less than 200, less than 250, less than 300, less than 350, less than 400, less than 450, less than 500, less than 600, less than 700, less than 800, less than 900, or less than 1000 bp in length. Various RNA nuclear localization sequences can be used. For example, Lubelsky and Ulitsky, Nature 555 (107-111), 2018 describe RNA sequences which drive RNA localization into the nucleus. In some embodiments, a nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, a nuclear localization signal binds a nucl ear-enriched protein. In some embodiments the nuclear localization signal binds the HNRNPK protein. In some embodiments, a nuclear localization signal is rich in pyrimidines, e.g., is a C / T rich, C / U rich, C rich, T rich, or U rich region. In some embodiments, a nuclear localization signal is derived from a long non-coding RNA. In some embodiments, a nuclear localization signal is derived from MALAT1 long non-coding RNA or is the 600 nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738- 751), 2012). In some embodiments, a nuclear localization signal is derived from BORG long non-coding RNA or is a AGCCC (SEQ ID NO: 83) motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014). In some embodiments, a nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, a nuclear localization signal is derived from a retrovirus.
[0437] In some embodiments, a gene modifying polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, for example a nuclear localization sequence (NLS). In some embodiments, the NLS is a bipartite NLS. In some embodiments, an NLS facilitates the import of a protein comprising an NLS into the cell nucleus. In some embodiments, an NLS is fused to the N-terminus of a gene modifying polypeptide as described herein. In some embodiments, an NLS is fused to the C-terminus of a gene modifying polypeptide. In some embodiments, an NLS is fused to the N-terminus or the C-terminus of a Cas domain. In some embodiments, a linker sequence is disposed between an NLS and a neighboring domain of a gene modifying polypeptide.
[0438] In some embodiments, an NLS comprises an amino acid sequence as listed in Table 9. An NLS may be utilized with one or more copies in a polypeptide in one or more locations in a polypeptide, e.g., 1, 2, 3, or more copies of an NLS in an N-terminal domain, between peptide domains, in a C-terminal domain, or in a combination of locations, in orderPage 118 of 39612985085vlAttorney Docket No.: 2017469-0049 to improve subcellular localization to the nucleus. Multiple unique sequences may be used within a single polypeptide. Sequences may be naturally monopartite or bipartite, e.g., having one or two stretches of basic amino acids, or may be used as chimeric bipartite sequences. Sequence references correspond to UniProt accession numbers, except where indicated as SeqNLS for sequences mined using a subcellular localization prediction algorithm (Lin et al BMC Bioinformat 13: 157 (2012), incorporated herein by reference in its entirety).Table 9. Exemplary nuclear localization signals for use in gene modifying systems
[0439] In some embodiments, a gene modifying polypeptide disclosed herein comprises an N-terminal NLS having an amino acid sequence according to SEQ ID NO: 15, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide disclosed herein comprises a C-terminal NLS having an amino acid sequence according to SEQ ID NO: 17, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0440] Additional exemplary NLS sequences are also described in PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences.
[0441] In some embodiments, a gene modifying polypeptide comprises, from N- terminus to C-terminus, one or more (e.g., 1, 2, 3, 4, 5, or all 6) of an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA binding domain, a linker, an RT domain, and / or a second NLS. In some embodiments, a gene modifying polypeptidePage 119 of 39612985085vlAttorney Docket No.: 2017469-0049 comprises, from N-terminus to C-terminus, a first NLS, a DNA binding domain, a linker, an RT domain, and a second NLS.
[0442] In some embodiments, a gene modifying polypeptide comprises: (i) an N- terminal NLS comprising an NLS sequence according to SEQ ID NO: 15, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (ii) a Cas domain comprising an amino acid sequence according to SEQ ID NO: 12, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iii) a linker comprising an amino acid sequence according to SEQ ID NO: 18, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iv) an RT domain comprising an amino acid sequence according to SEQ ID NO: 20, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; and (v) a C-terminal NLS having an amino acid sequence according to SEQ ID NO: 17, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises: (i) an N-terminal NLS comprising an amino acid sequence according to SEQ ID NO: 15, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (ii) a Cas domain comprising an amino acid sequence according to SEQ ID NO: 12, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iii) a linker comprising an amino acid sequence of SEQ ID NO: 19, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iv) an RT domainPage 120 of 39612985085vlAttorney Docket No.: 2017469-0049 comprising an amino acid sequence of SEQ ID NO: 20, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; and (v) a C-terminal NLS comprising an amino acid sequence according to SEQ ID NO: 17, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises the amino acid sequence of SEQ ID NO: 24 or SEQ ID NO: 28 (e.g., as shown in Table 10 below), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0443] In some embodiments, a gene modifying polypeptide comprises: (i) an N- terminal NLS comprising an amino acid sequence according to SEQ ID NO: 15, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (ii) a Cas domain comprising an amino acid sequence according to SEQ ID NO: 12, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iii) a linker comprising an amino acid sequence of SEQ ID NO: 48), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iv) an RT domain comprising an amino acid sequence according to SEQ ID NO: 46, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; and (v) a C-terminal NLS comprising an amino acid sequence according to SEQ ID NO: 17, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises the amino acid sequence of SEQ ID NO: 87 (e.g., as shown in Table 10 below), or an amino acidPage 121 of 39612985085vlAttorney Docket No.: 2017469-0049 sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.
[0444] In some embodiments, a gene modifying polypeptide comprises: (i) an N- terminal NLS comprising an amino acid sequence according to SEQ ID NO: 15, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (ii) a Cas domain comprising an amino acid sequence according to SEQ ID NO: 12, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iii) a linker comprising an amino acid sequence according to SEQ ID NO: 49, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; (iv) an RT domain comprising an amino acid sequence according to SEQ ID NO: 47, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto; and (v) a C-terminal NLS comprising an amino acid sequence according to SEQ ID NO: 17, a functional fragment thereof, or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto. In some embodiments, a gene modifying polypeptide comprises the amino acid sequence of SEQ ID NO: 88 (e.g., as shown in Table 10 below), or an amino acid sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity thereto.Table 10. Exemplary gene modifying polypeptidesPage 122 of 39612985085vlAttorney Docket No.: 2017469-0049Page 123 of 39612985085vlAttorney Docket No.: 2017469-0049Page 124 of 39612985085vlAttorney Docket No.: 2017469-0049Page 125 of 39612985085vlAttorney Docket No.: 2017469-0049
[0445] In some embodiments, a gene modifying polypeptide of the present disclosure comprises an N-terminal methionine residue.
[0446] In some embodiments, a gene modifying polypeptide of the present disclosure comprises (e.g., C-terminal to a second NLS) a T2A sequence and / or a puromycin sequence. In some embodiments, a nucleic acid encoding a gene modifying polypeptide encodes a T2A sequence, e.g., wherein the T2A sequence is situated between a region encoding the gene modifying polypeptide and a second region, wherein the second region optionally encodes a selectable marker, e.g., puromycin.
[0447] In some embodiments, a gene modifying polypeptide of the present disclosure comprises a spacer sequence between a first NLS and a DNA binding domain. In some embodiments, a spacer sequence between a first NLS and a DNA binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, a spacer sequence between a first NLS and a DNA binding domain comprises the amino acid sequence GG.
[0448] In some embodiments, a gene modifying polypeptide of the present disclosure comprises a spacer sequence between a DNA binding domain and a linker. In some embodiments, a spacer sequence between a DNA binding domain and a linker comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, a spacer sequence between a DNA binding domain and a linker comprises the amino acid sequence GG.
[0449] In some embodiments, a gene modifying polypeptide of the present disclosure comprises a spacer sequence between a linker and an RT domain. In some embodiments, a spacer sequence between a linker and an RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, a spacer sequence between a linker and an RT domain comprises the amino acid sequence GG.Page 126 of 39612985085vlAttorney Docket No.: 2017469-0049
[0450] In some embodiments, a gene modifying polypeptide of the present disclosure comprises a spacer sequence between an RT domain and a second NLS. In some embodiments, a spacer sequence between an RT domain and a second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, a spacer sequence between an RT domain and a second NLS comprises the amino acid sequence AG.
[0451] In some embodiments, a gene modifying polypeptide of the present disclosure comprises a spacer sequence between a second NLS and a T2A sequence and / or puromycin sequence. In some embodiments, a spacer sequence between a second NLS and a T2A sequence and / or puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, a spacer sequence between a second NLS and a T2A sequence and / or puromycin sequence comprises the amino acid sequence GSG.7. Inteins
[0452] In some embodiments, an intein-N (intN) domain may be fused to the N- terminal portion of a first domain of a gene modifying polypeptide described herein, and an intein-C (intC) domain may be fused to the C-terminal portion of a second domain of a gene modifying polypeptide described herein for the joining of the N-terminal portion to the C- terminal portion, thereby joining the first and second domains. In some embodiments, the first and second domains are each independently chosen from a DNA binding domain, an RNA binding domain, an RT domain, and an endonuclease domain.
[0453] Inteins can occur as self-splicing protein intron (e.g., peptide), e.g., which ligates flanking N-terminal and C-terminal exteins (e.g., fragments to be joined). An intein may, in some embodiments, comprise a fragment of a protein that is able to excise itself and join the remaining fragments (the exteins) with a peptide bond in a process known as protein splicing. Inteins are also referred to as “protein introns.” The process of an intein excising itself and joining the remaining portions of the protein is herein termed “protein splicing” or “intein-mediated protein splicing.”8. Additional domains
[0454] In some embodiments, a gene modifying polypeptide can bind a target DNA sequence and a template nucleic acid (e.g., template RNA), nick a target site, and write (e.g., reverse transcribe) the template nucleic acid into the target DNA, resulting in a modification of the target site. In some embodiments, additional domains may be added to a gene modifying polypeptide of the present disclosure to enhance the efficiency of the process. InPage 127 of 39612985085vlAttorney Docket No.: 2017469-0049 some embodiments, a gene modifying polypeptide may contain an additional DNA ligation domain to join reverse transcribed DNA to the DNA of at a target site. In some embodiments, a gene modifying polypeptide may comprise a heterologous RNA-binding domain. In some embodiments, a gene modifying polypeptide may comprise a domain having 5' to 3' exonuclease activity (e.g., wherein the 5' to 3' exonuclease activity increases repair of an alteration of a target site, e.g., in favor of an alteration over an original genomic sequence). In some embodiments, a gene modifying polypeptide may comprise a domain having 3' to 5' exonuclease activity, e.g., proof-reading activity. In some embodiments, a writing domain, e.g., an RT domain, has 3' to 5' exonuclease activity, e.g., proof-reading activity.9. Nucleic Acids encoding gene modifying polypeptides
[0455] The present disclosure further provides nucleic acids encoding gene modifying polypeptides described herein for use in gene modifying systems disclosed herein.
[0456] In some embodiments, a nucleic acid encoding a gene modifying polypeptide described herein may be altered from a canonical sequence to have altered codon usage, e.g. improved for expression in human cells. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide described herein comprises one or more silent mutations in a coding region (e.g., in a sequence encoding an RT domain) relative to a nucleic acid molecule as described herein. In some embodiments, a coding region may be followed by 1, 2, or 3 stop codons.
[0457] In some embodiments, a nucleic acid (e.g., RNA) encoding a gene modifying polypeptide described herein comprises a posttranscriptional regulatory element that enhances nuclear export. In some embodiments, a posttranscriptional regulatory element is that of Hepatitis B Virus (HPRE) or Woodchuck Hepatitis Virus (WPRE). Nucleic acid sequences of exemplary WPRE regulatory elements are shown in Table 11, below. In some embodiments, a WPRE regulatory element included in a nucleic acid encoding a gene modifying polypeptide comprises one or more mutations relative to a wild-type WPRE regulatory element, e.g., mutations recited in Zanta-Boussif, M. A. et al. Gene Then 16, 605- 619 (2009).
[0458] In some embodiments, a WPRE regulatory element, e.g., as described herein, comprises a mutation. In some embodiments, a WPRE regulatory element described herein does not comprise a T at position 412 relative to SEQ ID NO: 29. In some embodiments, aPage 128 of 39612985085vlAttorney Docket No.: 2017469-0049WPRE regulatory element described herein comprises a G at position 412 relative to SEQ ID NO: 29. In some embodiments, a WPRE regulatory element described herein comprises a nucleic acid sequence according to SEQ ID NO: 21.
[0459] In some embodiments, a nucleic acid encoding a gene modifying polypeptide is flanked by untranslated regions (UTRs) that modify protein expression levels. Various 5' and 3' UTRs can affect protein expression. For example, in some embodiments, a coding sequence may be preceded by a 5' UTR that modifies RNA stability or protein translation. In some embodiments, a coding sequence may be followed by a 3' UTR that modifies RNA stability or translation. In some embodiments, a coding sequence may be preceded by a 5' UTR and followed by a 3 ' UTR that modify RNA stability or translation. In some embodiments, a 5' and / or 3' UTR may be selected to enhance protein expression. In some embodiments, a 5' and / or 3' UTR may be selected to modify protein expression such that overproduction inhibition is minimized. In some embodiments, UTRs are around a coding sequence, e.g., outside the coding sequence and in other embodiments proximal to the coding sequence.
[0460] In some embodiments, a gene modifying system described herein comprises a DNA encoding a transcript, wherein the DNA comprises the corresponding 5' UTR and 3' UTR sequences, with T substituting for U. In some embodiments, a DNA vector used to produce an RNA component of a gene modifying system described herein further comprises a promoter upstream of a 5' UTR for initiating in vitro transcription, e.g., a T7, T3, or SP6 promoter. In some embodiments, a 5' UTR begins with GGG, which is a suitable start for optimizing transcription using T7 RNA polymerase. For tuning transcription levels and altering the transcription start site nucleotides to fit alternative 5' UTRs, the teachings of Davidson et al. Pac Symp Biocomput 433-443 (2010) describe T7 promoter variants, and the methods of discovery thereof, that fulfill both of these traits.
[0461] Exemplary sequences of 5’ UTRs and 3’ UTRs for use in nucleic acids encoding gene modifying polypeptides described herein are shown in Table 11.
[0462] In some embodiments, a nucleic acid (e.g., an RNA) encoding a gene modifying polypeptide further comprises a poly(A) tail for enhancing expression of the polypeptide.Page 129 of 39612985085vlAttorney Docket No.: 2017469-0049
[0463] In some embodiments, a poly(A) tail is made of exclusively adenosine ribonucleotides. In some embodiments, a poly(A) tail comprises ribonucleotides other than adenosine, e.g., interspersed within / between regions of polyadenosine.
[0464] In some embodiments, a poly(A) tail consists of a nucleic acid sequence comprising a pattern of 12-18 adenosine nucleotides followed by at least two non-adenosine nucleotides, 22-28 adenosine nucleotides followed by at least two non-adenosine nucleotides, 32-38 adenosine nucleotides followed by at least two non-adenosine nucleotides, and 42-48 adenosine nucleotides.
[0465] In some embodiments, a poly(A) tail consists of a nucleic acid sequence comprising a pattern of 13-17 adenosine nucleotides followed by 1-4 non-adenosine nucleotides, 23-27 adenosine nucleotides followed by 1-5 non-adenosine nucleotides, 33-37 adenosine nucleotides followed by 2-6 non-adenosine nucleotides, and 43-47 adenosine nucleotides.
[0466] In some embodiments, a poly(A) tail consists of a nucleic acid sequence comprising a pattern of 14-16 adenosine nucleotides followed by at least two (e.g., 2-6) non- adenosine nucleotides, 24-26 adenosine nucleotides followed by at least two (e.g., 2-6) non- adenosine nucleotides, 34-36 adenosine nucleotides followed by at least two (e.g., 2-6) non- adenosine nucleotides, and 44-46 adenosine nucleotides.
[0467] In some embodiments, a poly(A) tail consists of a nucleic acid sequence comprising a pattern of 15 adenosine nucleotides followed by at least two (e.g., 2-4) non- adenosine nucleotides, 25 adenosine nucleotides followed by at least two (e.g., 2-4) non- adenosine nucleotides, 35 adenosine nucleotides followed by at least two (e.g., 2-4) non- adenosine nucleotides, and 45 adenosine nucleotides.
[0468] In some embodiments, a poly(A) tail consists of a nucleic acid sequence comprising a pattern of 15 adenosine nucleotides followed by 1-4 non-adenosine nucleotides, 25 adenosine nucleotides followed by 1-5 non-adenosine nucleotides, 35 adenosine nucleotides followed by 2-6 non-adenosine nucleotides, and 45 adenosine nucleotides.
[0469] In some embodiments, a poly(A) tail comprises two to ten non-adenosine nucleotides. In some embodiments, from 5’ to 3’, the number of non-adenosine nucleotides in each group of non-adenosine nucleotides increases. In some embodiments, from 5’ to 3’, the number of non-adenosine nucleotides in each group of non-adenosine nucleotides is 2, 3, or 4Page 130 of 39612985085vlAttorney Docket No.: 2017469-0049 non-adenosine nucleotides. In some embodiments, non-adenosine nucleotides are cytosine nucleotides and / or uridine nucleotides.
[0470] In some embodiments, a nucleic acid encoding a gene modifying polypeptide of the present disclosure may comprise uridines that are about 1%- 100%, about 10%-100%, about 20%-100%, about 30%-100%, about 40%-100%, about 50%-100%, about 60%-100%, about 70%-100%, about 80%-100%, about 90%-100%, about l%-90%, about 10%-90%, about 20%-90%, about 30%-90%, about 40%-90%, about 50%-90%, about 60%-90%, about 70%-90%, about 80%-90%, about l%-80%, about 10%-80%, about 20%-80%, about 30%- 80%, about 40%-80%, about 50%-80%, about 60%-80%, about 70%-80%, about l%-70%, about 10%-70%, about 20%-70%, about 30%-70%, about 40%-70%, about 50%-70%, about 60%-70%, about l%-60%, about 10%-60%, about 20%-60%, about 30%-60%, about 40%- 60%, about 50%-60%, about l%-50%, about 10%-50%, about 20%-50%, about 30%-50%, about 40%-50%, about l%-40%, about 10%-40%, about 20%-40%, about 30%-40%, about l%-30%, about 10%-30%, about 20%-30%, about l%-20%, about 10%-20%, or about 1%- 10% substituted with N1 -methyl pseudouridine. In some embodiments, a nucleic acid encoding a gene modifying polypeptide of the present disclosure may comprise uridines that are about 2%, about 5%, about 10%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, or 100% substituted with Nl-methyl-pseudouri dine. In some embodiments, a nucleic acid encoding a gene modifying polypeptide of the present disclosure may comprise uridines that are 100% substituted with Nl-methyl-pseudouri dine.
[0471] Exemplary sequences of poly(A) tails for use in nucleic acids encoding gene modifying polypeptides described herein are shown in Table 11.Table 11. Exemplary regulatory elements, 5’ UTRs, 3’ UTRs, and poly(A) tailsPage 131 of 39612985085vlAttorney Docket No.: 2017469-0049Page 132 of 39612985085vlAttorney Docket No.: 2017469-0049
[0472] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence listed in Table 12, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto.
[0473] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 23, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence of SEQ ID NO: 23.
[0474] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 89, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 89.
[0475] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 90, or a sequence having at leastPage 133 of 39612985085vlAttorney Docket No.: 2017469-004970%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 90.
[0476] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 91, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 91.
[0477] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 92, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 492.
[0478] In some embodiments, a nucleic acid encoding a gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 93, or a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%. at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity thereto. In some embodiments, a nucleic acid encoding the gene modifying polypeptide comprises a nucleic acid sequence according to SEQ ID NO: 93.Table 12. Exemplary Nucleic Acid Sequences Encoding Gene Modifying PolypeptidesPage 134 of 39612985085vlAttorney Docket No.: 2017469-0049Page 135 of 39612985085vlAttorney Docket No.: 2017469-0049Page 136 of 39612985085vlAttorney Docket No.: 2017469-0049Page 137 of 39612985085vlAttorney Docket No.: 2017469-0049Page 138 of 39612985085vlAttorney Docket No.: 2017469-0049Page 139 of 39612985085vlAttorney Docket No.: 2017469-0049Page 140 of 39612985085vlAttorney Docket No.: 2017469-0049Page 141 of 39612985085vlAttorney Docket No.: 2017469-0049Page 142 of 39612985085vlAttorney Docket No.: 2017469-0049Page 143 of 39612985085vlAttorney Docket No.: 2017469-0049Page 144 of 39612985085vlAttorney Docket No.: 2017469-0049Page 145 of 39612985085vlAttorney Docket No.: 2017469-0049Page 146 of 39612985085vlAttorney Docket No.: 2017469-0049Page 147 of 39612985085vlAttorney Docket No.: 2017469-0049Page 148 of 39612985085vlAttorney Docket No.: 2017469-0049Page 149 of 39612985085vlAttorney Docket No.: 2017469-0049Page 150 of 39612985085vlAttorney Docket No.: 2017469-0049Page 151 of 39612985085vlAttorney Docket No.: 2017469-0049Page 152 of 39612985085vlAttorney Docket No.: 2017469-0049B. Template nucleic acids
[0479] Gene modifying systems described herein can modify a host target DNA site using a template nucleic acid sequence. In some embodiments, gene modifying systems described herein transcribe an RNA sequence template into host target DNA sites by target- primed reverse transcription (TPRT). By modifying DNA sequence(s) via reversePage 153 of 39612985085vlAttorney Docket No.: 2017469-0049 transcription of an RNA sequence template directly into the host genome, a gene modifying system can insert an object sequence into a target genome without the need for exogenous DNA sequences to be introduced into the host cell (unlike, for example, CRISPR systems), as well as eliminate an exogenous DNA insertion step. A gene modifying system can also delete a sequence from a target genome or introduce a substitution using an object sequence. Therefore, a gene modifying system provides a platform for the use of customized RNA sequence templates containing object sequences, e.g., sequences comprising heterologous gene coding and / or function information.
[0480] In some embodiments, a template nucleic acid comprises one or more sequence (e.g., 2 sequences) that binds a gene modifying polypeptide.
[0481] In some embodiments, a gene modifying system or method described herein comprises a single template nucleic acid (e.g., template RNA). In some embodiments a system or method described herein comprises a plurality of template nucleic acids (e.g., template RNAs). For example, a gene modifying system described herein comprises a first RNA comprising (e.g., from 5' to 3') a sequence that binds a gene modifying polypeptide (e.g., a DNA-binding domain and / or an endonuclease domain, e.g., a gRNA) and a sequence that binds a target site (e.g., a second strand of a site in a target genome), and a second RNA (e.g., a template RNA) comprising (e.g., from 5' to 3') optionally a sequence that binds the gene modifying polypeptide (e.g., that specifically binds an RT domain), a heterologous object sequence, and a PBS sequence. In some embodiments, when a gene modifying system comprises a plurality of nucleic acids, each nucleic acid comprises a conjugating domain. In some embodiments, a conjugating domain enables association of nucleic acid molecules, e.g., by hybridization of complementary sequences. For example, in some embodiments a first RNA comprises a first conjugating domain and a second RNA comprises a second conjugating domain, and the first and second conjugating domains are capable of hybridizing to one another, e.g., under stringent conditions. In some embodiments, stringent conditions for hybridization include hybridization in 4x sodium chloride / sodium citrate (SSC), at about 65 °C, followed by a wash in IxSSC, at about 65 °C.
[0482] In some embodiments, a template nucleic acid comprises RNA. In some embodiments, a template nucleic acid comprises DNA (e.g., single stranded or double stranded DNA).Page 154 of 39612985085vlAttorney Docket No.: 2017469-0049
[0483] In some embodiments, a template nucleic acid comprises one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more) homology domains that have homology to a target sequence. In some embodiments, a homology domain is about 10-20, about 20-50, or about 50-100 nucleotides in length.
[0484] In some embodiments, a template RNA can comprise a gRNA sequence, e.g., to direct a gene modifying polypeptide to a target site of interest. In some embodiments, a template RNA comprises (e.g., from 5' to 3') (i) optionally a gRNA spacer that binds a target site (e.g., a second strand of a site in a target genome), (ii) optionally a gRNA scaffold that binds a polypeptide described herein (e.g., a gene modifying polypeptide or a Cas polypeptide), (iii) a heterologous object sequence comprising a mutation region (optionally the heterologous object sequence comprises, from 5’ to 3’, a first homology region, a mutation region, and a second homology region), and (iv) a primer binding site (PBS) sequence comprising a 3' target homology domain.
[0485] A template nucleic acid (e.g., template RNA) component of a genome editing system described herein typically is able to bind a gene modifying polypeptide of a gene modifying system described herein. In some embodiments, a template nucleic acid (e.g., template RNA) has a 3 ' region that is capable of binding a gene modifying polypeptide. A binding region, e.g., 3' region, may be a structured RNA region, e.g., having at least 1, at least 2 or at least 3 hairpin loops, capable of binding the gene modifying polypeptide of the system. A binding region may associate with a template nucleic acid (e.g., template RNA) with any of the polypeptide modules. In some embodiments, a binding region of a template nucleic acid (e.g., template RNA) may associate with an RNA-binding domain in a gene modifying polypeptide. In some embodiments, a binding region of a template nucleic acid (e.g., template RNA) may associate with an RT domain of a gene modifying polypeptide (e.g., specifically bind to the RT domain). In some embodiments, a template nucleic acid (e.g., template RNA) may associate with a DNA binding domain of a gene modifying polypeptide, e.g., a gRNA associating with a Cas9-derived DNA binding domain. In some embodiments, a binding region may provide DNA target recognition, e.g., a gRNA hybridizing to a target DNA sequence and binding a gene modifying polypeptide, e.g., a Cas9 domain. In some embodiments, a template nucleic acid (e.g., template RNA) may associate with multiple components of a gene modifying polypeptide, e.g., DNA binding domain and RT domain.Page 155 of 39612985085vlAttorney Docket No.: 2017469-0049
[0486] In some embodiments, a template RNA has a poly-A tail at its 3 ' end. In some embodiments, a template RNA does not have a poly-A tail at its 3 ' end.
[0487] In some embodiments, a template nucleic acid is a template RNA. In some embodiments, a template RNA comprises one or more modified nucleotides. For example, in some embodiments, a template RNA comprises one or more deoxyribonucleotides. In some embodiments, regions of a template RNA are replaced by DNA nucleotides, e.g., to enhance stability of the molecule. For example, the 3' end of a template may comprise DNA nucleotides, while the rest of the template comprises RNA nucleotides that can be reverse transcribed. For instance, in some embodiments, a heterologous object sequence is primarily or wholly made up of RNA nucleotides (e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% RNA nucleotides). In some embodiments, a PBS sequence is primarily or wholly made up of DNA nucleotides (e.g., at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%, or 100% DNA nucleotides). In some embodiments, a heterologous object sequence for writing into the genome may comprise DNA nucleotides. In some embodiments, DNA nucleotides in a template nucleic acid are copied into the genome by a domain capable of DNA-dependent DNA polymerase activity. In some embodiments, DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in a gene modifying polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, e.g., second strand synthesis. In some embodiments, a template nucleic acid is composed of only DNA nucleotides.
[0488] In some embodiments, a gene modifying system described herein comprises two nucleic acids which together comprise the sequences of a template RNA described herein. In some embodiments, two nucleic acids are associated with each other non- covalently, e.g., directly associated with each other (e.g., via base pairing), or indirectly associated as part of a complex comprising one or more additional molecule.
[0489] A template RNA described herein may comprise, from 5’ to 3’ : (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous object sequence (4) a primer binding site (PBS) sequence. Each of these components is now described in more detail.Page 156 of 39612985085vlAttorney Docket No.: 2017469-00491. gRNA spacer and gRNA scaffold
[0490] A template RNA described herein may comprise a gRNA spacer that directs a gene modifying system to a target nucleic acid, and a gRNA scaffold that promotes association of a template RNA with a Cas domain of a gene modifying polypeptide. In some embodiments, a gRNA scaffold has been engineered for improved performance with StlCas9. Gene modifying systems described herein can also comprise a gRNA that is not part of a template nucleic acid. For example, a gRNA that comprises a gRNA spacer and gRNA scaffold, but not a heterologous object sequence or a PBS sequence, can be used, e.g., to induce second strand nicking.
[0491] In some embodiments, the present disclosure provides variant gRNA scaffolds that are compatible with StlCas9. In some embodiments, variant gRNA scaffolds are used in a system comprising a gene modifying polypeptide that comprises an StlCas9 domain.
[0492] The wild-type StlCas9 gRNA scaffold has a hypothesized secondary structure, shown in FIG. 6A. Generally, from 5’ to 3’, the gRNA scaffold comprises: a region comprising a lower stem, an upper stem, and tetraloop (also collectively referred to as Repeat:anti-repeat duplex or RAR); a first single stranded region; a Stem loop 1, a second single stranded region; a Stem loop 2; and a third single stranded region. More specifically, as shown in FIG. 6A, which illustrates an exemplary wild-type StlCas9 gRNA scaffold, the upper stem comprises three paired bases (nucleotides 12-14 pair with nucleotides 19-21) and the 4-nucleotide tetraloop is nucleotides 15-18. At the base of the three paired bases of the upper stem is a region with bulges (nucleotides 22 and nucleotides 25 bulge from the region), and at the base of the region with bulges is a lower stem (nucleotides 1-9 pair with nucleotides 26-34). Moving in a 3’ direction, the next region is the first single stranded region which contains nucleotides 35 and 36. Following the first single stranded region is Stem loop 1, which comprises nucleotides 37-47. Next is the second single stranded region, comprising nucleotides 48-53. Next is Stem loop 2 which comprises nucleotides 54-82. 3’ of Stem loop 2 is a third single stranded region which comprises nucleotides 83-84. The hypothesized structure represents the likely secondary structure of the StlCas9 gRNA scaffold under physiologically relevant conditions. However, even if an StlCas9 gRNA scaffold were to adopt a different structure from the hypothesized structure shown herein, the named regions (such as Stem loop 1, Stem loop 2, RAR upper stem, RAR lower stem, and tetraloop) of variant scaffolds could still be readily identified based at least on sequencePage 157 of 39612985085vlAttorney Docket No.: 2017469-0049 alignments to the wild-type reference sequence (e.g., SEQ ID NO: 45), and optionally using additional tools such as RNA folding algorithms.
[0493] In some embodiments, a spacer is situated at the 5’ end of a gRNA scaffold.
[0494] In some embodiments, a variant gRNA scaffold can comprise mutations in different regions of the gRNA scaffold. For example, in some embodiments, a variant gRNA scaffold comprises a mutation in the upper stem that results in the thermodynamic strengthening of RAR. In some embodiments, the upper stem may be lengthened, e.g., by 1- 8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs).
[0495] In some embodiments, a variant gRNA scaffold comprises a mutation in the tetraloop of RAR, which may optimize performance by improving the thermodynamic stability of RAR. In some embodiments, one or more nucleotides in the loop of the tetraloop may be substituted. In some embodiments, the loop region of the tetraloop may be lengthened, e.g., by 1 nucleotide, resulting in a loop 5 nucleotides in length.
[0496] In some embodiments, a variant gRNA scaffold may comprise a truncation in the stem of Stem loop 2 and / or in one or both single stranded regions at its base (i.e., the second and third single stranded regions). In some embodiments, the stem of Stem loop 2 comprises truncations in 3 ’-5’ direction end ranging from 1- 32 nucleotides, relative to SEQ ID NO: 45.
[0497] In some embodiments, a variant gRNA scaffold may comprise one or more mutations that destabilize the upper RAR stem relative to the wild-type sequence. In some embodiments, a variant gRNA scaffold has a deletion of one or more nucleotides of the upper RAR stem. In some embodiments, a variant gRNA scaffold has a deletion of one or more nucleotides in the region with bulges that is situated between the upper RAR stem and lower RAR stem. In some embodiments, a variant gRNA scaffold has a substitution wherein a G-C base pair in the upper RAR stem is replaced with a base pair other than G-C (e.g., an A-U base pair).
[0498] Different mutations to a variant gRNA scaffold may be combined. For instance, in some embodiments, a variant gRNA scaffold comprises a mutation in the upper stem of the RAR and a mutation in the tetraloop of the RAR. In some embodiments, the upper stem is lengthened, e.g., by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs), and the tetraloop comprises a substitution or an insertion (e.g., a 1 nucleotide insertion).Page 158 of 39612985085vlAttorney Docket No.: 2017469-0049
[0499] In some embodiments, a variant gRNA scaffold comprises a mutation in the upper stem of the RAR and a truncation in the stem of Stem loop 2. In some embodiments, the upper stem is lengthened, e.g., by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs), and the stem of Stem loop 2 comprises a truncation of from 1- 32 nucleotides relative to SEQ ID NO: 45.
[0500] In some embodiments, a variant gRNA scaffold comprises a mutation in the tetraloop of the RAR and a truncation in the stem of Stem loop 2. In some embodiments, the tetraloop comprises a substitution or an insertion (e.g., a 1 nucleotide insertion) and the stem of Stem loop 2 comprises a truncation of from 1- 32 nucleotides, relative to SEQ ID NO: 45.
[0501] In some embodiments, a variant gRNA scaffold comprises: (1) a mutation in the upper stem of the RAR, (2) a mutation in the tetraloop of the RAR, and (3) a truncation in the stem of Stem loop 2. In some embodiments, the upper stem is lengthened, e.g., by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs), the tetraloop comprises a substitution or an insertion (e.g., a 1 nucleotide insertion), and the stem of Stem loop 2 comprises a truncation of from 1- 32 nucleotides, relative to SEQ ID NO: 45.
[0502] In some embodiments, a StlCas9 gRNA scaffold comprises an insertion (e.g., of 10 nucleotides) between positions 15 and 18 of SEQ ID NO: 45, and a deletion of positions 16 and 17, of SEQ ID NO: 45. In some embodiments, the insertion has a sequence according to GACUUCGGUC (SEQ ID NO: 94).
[0503] In some embodiments, a StlCas9 gRNA scaffold comprises an insertion (e.g., of 10 nucleotides) between positions 15 and 18 of SEQ ID NO: 45, and a deletion of positions 16 and 17 of SEQ ID NO: 45. In some embodiments, the insertion has a sequence according to CUAGAAAUAG (SEQ ID NO: 95).
[0504] In some embodiments, a StlCas9 gRNA scaffold comprises an insertion (e.g., of 12 nucleotides) between positions 14 and 19 of SEQ ID NO: 45, and a deletion of positions 15-18 of SEQ ID NO: 45. In some embodiments, the insertion has a sequence according to CGCGGUAACGCG (SEQ ID NO: 96).
[0505] In some embodiments, a StlCas9 gRNA scaffold comprises an insertion. In some embodiments, the insertion results in a sequence, e.g., a tetraloop sequence, of CGCGGUAACGCG (SEQ ID NO: 11).Page 159 of 39612985085vlAttorney Docket No.: 2017469-0049
[0506] In some embodiments, a variant StlCas9 scaffold has a substitution resulting in a G-C base pair in the RAR lower stem. In some embodiments, a substitution resulting in a G-C base pair in the RAR strengthens and / or stabilizes the RAR. In some embodiments, the substitution comprises a substitution of position 4 of SEQ ID NO: 45 with a G and the template further comprises a substitution of position 31 of SEQ ID NO: 45 with a C.
[0507] In some embodiments, a template RNA comprises a substitution in the second single stranded region. In some embodiments, the substitution is a substitution of position 51 of SEQ ID NO: 45 with U or a substitution of position 54 of SEQ ID NO: 45 with C.
[0508] In some embodiments, a template RNA comprises a T-lock (UUCG) tetraloop and an RAR U4G modification (e.g., altering A39 of SEQ ID NO: 45 to C and its paired base to a G). In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified StlCas9-based scaffold sequence (dSL2).
[0509] In some embodiments, a template RNA comprises a T-lock (UUCG) tetraloop, a GAAU linker (e.g., altering A59 of SEQ ID NO: 45 to U), and 3’UCC (altering A62 to C) modifications. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0510] In some embodiments, a template RNA comprises a T-lock (UUCG) tetraloop and RAR U4G and 3’UCC modifications. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0511] In some embodiments, a template RNA comprises a T-lock (UUCG) tetraloop and RAR U4G and GAAU linker modifications. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0512] In some embodiments, a template RNA comprises a T-lock (UUCG) tetraloop and RAR U4G, 3’UCC, and GAAU linker modifications. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.Page 160 of 39612985085vlAttorney Docket No.: 2017469-0049
[0513] In some embodiments, a template RNA comprises a GA A A tetraloop and an RAR U4G modification. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0514] In some embodiments, a template RNA comprises a GA A A tetraloop and a GAAU linker. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0515] In some embodiments, a template RNA comprises a GA A A tetraloop and a 3’UCC modification. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0516] In some embodiments, a template RNA comprises a GA A A tetraloop and RAR ING and 3’UCC modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0517] In some embodiments, a template RNA comprises a GA A A tetraloop and RAR U4G and GAAU linker modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0518] In some embodiments, a template RNA comprises a GA A A tetraloop and RAR U4G, 3’UCC, and GAAU linker modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0519] In some embodiments, a template RNA comprises a GUAA tetraloop and an RAR U4G modification. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0520] In some embodiments, a template RNA comprises a GUAA tetraloop and a GAAU linker. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0521] In some embodiments, a template RNA comprises a GUAA tetraloop and GAAU linker and 3’UCC modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.Page 161 of 39612985085vlAttorney Docket No.: 2017469-0049
[0522] In some embodiments, a template RNA comprises a GUAA tetraloop and RAR U4G and 3’UCC modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0523] In some embodiments, a template RNA comprises a GUAA tetraloop and RAR U4G and GAAU linker modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0524] In some embodiments, a template RNA comprises a GUAA tetraloop and RAR U4G, 3’UCC, and GAAU linker modifications in combination. In some embodiments, this template RNA exhibits increased editing efficiency relative to a similar template RNA with unmodified dSL2 sequence.
[0525] In some embodiments, RAR U4G results in increased editing activity regardless of tetraloop sequence. Without wishing to be bound by any particular theory, it is thought that the RAR U4G modification strengthens and / or stabilizes the RAR domain. In some embodiments, certain 3 ’end structures (e.g., a linker modification (e.g., GAAAto GAAU) or modification of the 3’ end nucleotides (e.g., 3’UCAto 3’UCC)) result in increased editing for some tetraloop sequences. In some embodiments, this editing further increases in combination with RAR strengthening.
[0526] In some embodiments, a template RNA described herein comprises the nucleic acid sequence of UAAGGCUGUGCUGACCAUCGAGUCGUUGUACUCUGCGCGGUAACGCGCAGAAG CUACAACGAUAAGGCUUCAUGCCGAAAUCACCUUUCUCGUCGAUGGUCAG (SEQ ID NO: 5). In some embodiments, a template RNA described herein comprises a gRNA scaffold comprising the nucleic acid sequence of GUCGUUGUACUCUGCGCGGUAACGCGCAGAAGCUACAACGAUAAGGCUUCAUG CCGAAAUCA (SEQ ID NO: 2). Without wishing to be bound by any particular theory, in some embodiments, a gRNA scaffold has hypothesized secondary structure, shown in FIG.6B.
[0527] In some embodiments, a gRNA is a short synthetic RNA composed of a scaffold sequence that participates in CRISPR-associated protein binding and a user-defined ~20 nucleotide targeting sequence for a genomic target. The structure of a complete gRNAPage 162 of 39612985085vlAttorney Docket No.: 2017469-0049 was described by Nishimasu et al. Cell 156, P935-949 (2014). A gRNA (also referred to as sgRNA for single-guide RNA) comprises of crRNA- and tracrRNA-derived sequences connected by an artificial tetraloop. A crRNA sequence can be divided into guide (20 nt) and repeat (12 nt) regions, whereas a tracrRNA sequence can be divided into anti-repeat (14 nt) and three tracrRNA stem loops (Nishimasu et al. Cell 156, P935-949 (2014)). In practice, guide RNA sequences are generally designed to have a length of between 17 - 24 nucleotides (e.g., 19, 20, or 21 nucleotides) and to be complementary to a targeted nucleic acid sequence. Custom gRNA generators and algorithms are available commercially for use in the design of effective guide RNAs. In some embodiments, a gRNA comprises two RNA components from the native CRISPR system, e.g. crRNA and tracrRNA. In some embodiments, a gRNA may also comprise a chimeric, single guide RNA (sgRNA) containing sequence from both a tracrRNA (for binding the nuclease) and at least one crRNA (to guide the nuclease to a sequence targeted for editing / binding). Chemically modified sgRNAs have also been demonstrated to be effective for use with CRISPR-associated proteins; see, for example, Hendel et al. (2015) Nature Biotechnol., 985 - 991. In some embodiments, a gRNA spacer comprises a nucleic acid sequence that is complementary to a DNA sequence associated with a target gene.
[0528] In some embodiments, the region of a template nucleic acid, e.g., template RNA, comprising a gRNA adopts an underwound ribbon-like structure of gRNA bound to target DNA (e.g., as described in Mulepati et al. Science 19 Sep 2014:Vol. 345, Issue 6203, pp. 1479-1484). Without wishing to be bound by any particular theory, this non-canonical structure is thought to be facilitated by rotation of every sixth nucleotide out of the RNA- DNA hybrid. Thus, in some embodiments, the region of a template nucleic acid, e.g., template RNA, comprising a gRNA may tolerate increased mismatching with a target site at some interval, e.g., every sixth base. In some embodiments, the region of a template nucleic acid, e.g., template RNA, comprising a gRNA comprising homology to a target site may possess wobble positions at a regular interval, e.g., every sixth base, that do not need to base pair with the target site.
[0529] In some embodiments, a template nucleic acid (e.g., template RNA) has at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, or at least 24 bases of at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% homology to a target site, e.g., at the 5’ end, e.g., comprising a gRNA spacerPage 163 of 39612985085vlAttorney Docket No.: 2017469-0049 sequence of length appropriate to a Cas9 domain of a gene modifying polypeptide described herein (e.g., as listed in Table 5).
[0530] Table 13 provides parameters to define components for designing gRNA and / or template RNAs to apply Cas variants listed in Table 5 for gene modifying. The cut site indicates the validated or predicted protospacer adjacent motif (PAM) requirements, validated or predicted location of cut site (relative to the most upstream base of the PAM site). A gRNA for a given enzyme can be assembled by concatenating the crRNA, tetraloop, and tracrRNA sequences, and further adding a 5' spacer of a length within Spacer (min) and Spacer (max) that matches a protospacer at a target site. Further, the predicted location of a ssDNA nick at a target site is important for designing a PBS sequence of a template RNAthat can anneal to the sequence immediately 5' of a nick in order to initiate target primed reverse transcription. In some embodiments, a gRNA scaffold described herein comprises a nucleic acid sequence comprising, in the 5’ to 3’ direction, a crRNA of Table 13, a tetraloop from the same row of Table 13, a...
Claims
1. Attorney Docket No.: 2017469-0049LISTING OF THE CLAIMSWHAT IS CLAIMED IS:
1. A gene modifying system comprising: a. a gene modifying polypeptide-encoding polynucleotide comprising, i. a 5’ untranslated region (UTR) sequence, ii. a polynucleotide sequence encoding a first nuclear localization signal (NLS), iii. a polynucleotide sequence encoding an Stl Cas9 domain, iv. a polynucleotide encoding a peptide linker, v. a polynucleotide encoding a reverse transcriptase (RT) domain, vi. a polynucleotide encoding a second NLS, vii. a 3’ UTR sequence, and viii. a poly (A) tail sequence; and b. a template ribonucleotide comprising, i. a gRNA spacer, ii. a gRNA scaffold, iii. a heterologous object sequence comprising a mutation region for correcting a mutation in the human SERPINA1 gene, and iv. a primer binding site (PBS) sequence complementary to a portion of the human SERPINA1 gene.
2. The gene modifying system of claim 1, wherein the RT domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a an amino acid sequence according to SEQ ID NO: 20, SEQ ID NO: 54, SEQ ID NO: 46, SEQ ID NO: 57, SEQ ID NO: 58, SEQ ID NO: 59, SEQ ID NO: 47, SEQ ID NO: 61, or SEQ ID NO: 62.Page 389 of 39612985085vlAttorney Docket No.: 2017469-00493. The gene modifying system of claim 1 or 2, wherein the 5’ UTR sequence comprises a nucleotide sequence at least at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 13.
4. The gene modifying system of any one of claims 1 to 3, wherein the polynucleotide encoding the first NLS encodes a modified c-myc NLS and comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 85.
5. The gene modifying system of any one of claims 1 to 4, wherein the polynucleotide encoding the STI Cas9 domain encodes a modified Stl Cas9 domain comprising an N955A mutation, and wherein the polynucleotide encoding the STI Cas9 domain comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 115.
6. The gene modifying system of any one of claims 1 to 5, wherein the polynucleotide encoding the linker comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 80, SEQ ID NO: 81, or SEQ ID NO: 82.
7. The gene modifying system of any one of claims 1 to 6, wherein the polynucleotide encoding the second NLS comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%,Page 390 of 39612985085vlAttorney Docket No.: 2017469-0049 at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 86.
8. The gene modifying system of any one of claims 1 to 7, wherein the 3’ UTR sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 14.
9. The gene modifying system of any one of claims 1 to 8, wherein the poly (A) tail sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 22.
10. The gene modifying system of any one of claims 1 to 9, wherein the gene modifying polypeptide-encoding polynucleotide further comprises a Woodchuck Hepatitis Virus Posttranscriptional Regulatory Element (WPRE) sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 21.
11. The gene modifying system of any one of claims 1 to 10, wherein the gRNA spacer comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 1 or SEQ ID NO: 6.
12. The gene modifying system of any one of claims 1 to 11, wherein the gRNA scaffold comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at leastPage 391 of 39612985085vlAttorney Docket No.: 2017469-004996%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 2, SEQ ID NO: 7, SEQ ID NO: 100, SEQ ID NO: 110, SEQ ID NO: 101, or SEQ ID NO: 111.
13. The gene modifying system of any one of claims 1 to 12, wherein the heterologous object sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 3, SEQ ID NO: 8, SEQ ID NO: 102, SEQ ID NO: 112, SEQ ID NO: 103, or SEQ ID NO: 113.
14. The gene modifying system of any one of claims 1 to 13, wherein the PBS sequence comprises a nucleotide sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to a nucleotide sequence according to SEQ ID NO: 4, or SEQ ID NO: 9.
15. A gene modifying system comprising: a. a gene modifying polypeptide comprising, i. a first nuclear localization signal (NLS), ii. an Stl Cas9 domain, iii. a peptide linker, iv. a reverse transcriptase (RT) domain, and v. a second NLS; and b. a template ribonucleotide comprising, i. a gRNA spacer, ii. a gRNA scaffold, iii. a heterologous object sequence comprising a mutation region for correcting a mutation in the human SERPINA1 gene, andPage 392 of 39612985085vlAttorney Docket No.: 2017469-0049 iv. a primer binding site (PBS) sequence complementary to a portion of the human SERPINA1 gene.
16. The gene modifying system of claim 15, wherein the first NLS comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 15.
17. The gene modifying system of claim 15 or 16, wherein the Stl Cas9 domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 12.
18. The gene modifying system of any one of claims 15 to 17, wherein the peptide linker comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 19, SEQ ID NO: 48, or SEQ ID NO: 49.
19. The gene modifying system of any one of claims 15 to 18, wherein the RT domain comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 20, SEQ ID NO: 46, or SEQ ID NO: 47.
20. The gene modifying system of any one of claims 15 to 19, wherein the secondNLS comprises an amino acid sequence at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at leastPage 393 of 39612985085vlAttorney Docket No.: 2017469-004996%, at least 97%, at least 98%, at least 99%, or 100% identical to an amino acid sequence according to SEQ ID NO: 17.
21. The gene modifying system of any one of claims 1 to 20, wherein the gene modifying system further comprises a lipid delivery vehicle comprising: i. an ionizable lipid; ii. a sterol; iii. a non-cationic lipid; and iv. a conjugated lipid.
22. The gene modifying system of claim 21, wherein the lipid delivery vehicle comprises: i. 47.5% (w / w) Lipid 1; ii. 40% (w / w) cholesterol; iii. 10% (w / w) DSPC; and iv. 2.5% DMG-PEG2000.
23. The gene modifying system of claim 21 or 22, wherein the lipid delivery vehicle has an N / P ratio of 6.
24. A pharmaceutical composition comprising the gene modifying system of any one of claims 1 to 23 and a pharmaceutically acceptable excipient or carrier.
25. A method for modifying a human SERPINA1 gene in a cell comprising contacting the cell with the gene modifying system of any one of claims 1 to 23, or the pharmaceutical composition of claim 24, thereby modifying the human SERPINA1 gene.Page 394 of 39612985085vlAttorney Docket No.: 2017469-004926. A method of treating a subject having a disease or condition associated with a mutation in the human SERPINA1 gene comprising administering to the subject the gene modifying system of any one of claims 1 to 23, or the pharmaceutical composition of claim 24, thereby treating the subject having the disease or condition.Page 395 of 39612985085vl