SERPINA MODULATORY COMPOSITIONS AND METHODS
Patent Information
- Application Number
- JP2024515068
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-01-27
- Filing Date
- 2022-09-07
- Publication Date
- 2025-09-17
AI Technical Summary
Current treatments for alpha-1 antitrypsin deficiency (AATD) are limited, with augmentation therapy failing to restore normal physiological regulation of alpha-1 antitrypsin (AAT) and not addressing liver disease caused by the toxic gain of function of the Z allele, necessitating the development of more effective genetic modification systems for correcting SERPINA1 gene mutations.
A novel genetic modification system comprising a genetically modified polypeptide with a reverse transcriptase domain and Cas9 nickase, guided by a gRNA scaffold, is used to introduce a heterologous sequence into the SERPINA1 gene, correcting the E342K mutation by inserting, altering, or deleting nucleotides to restore AAT function.
The system effectively corrects the SERPINA1 gene mutation, potentially restoring AAT function and addressing both lung and liver diseases associated with AATD, providing a more effective treatment than existing therapies.
Abstract
Description
[Technical field]
[0001] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML in accordance with WIPO Standard ST.26, and is incorporated herein by reference in its entirety. Said XML copy, created on October 31, 2022, is named V2065-7024WO_SL.XML and is 31,115,675kb in size.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 241,970, filed September 8, 2021, U.S. Provisional Patent Application No. 63 / 253,087, filed October 6, 2021, and U.S. Provisional Patent Application No. 63 / 303,905, filed January 27, 2022. The contents of the foregoing applications are incorporated herein by reference in their entireties. [Background technology]
[0003] The integration of a nucleic acid of interest into a genome occurs at low frequency in the absence of special proteins to facilitate the insertion event, and has little site specificity. Some existing methods, such as CRISPR / Cas9, are more suitable for small edits that rely on host repair pathways, and are less effective at integrating long sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into a genome, and then a second step of inserting a sequence of interest into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, modifying, or deleting a sequence of interest in a genome.
[0004] AATD is characterized by low circulating levels of AAT. AAT is produced primarily in hepatocytes and secreted into the blood, but also in other cell types, including lung epithelial cells and certain leukocytes. AAT inhibits several serine proteases secreted by inflammatory cells (most notably neutrophil elastase [NE], proteinase 3, and cathepsin G), thus protecting organs such as the lungs from protease-induced damage, especially during periods of inflammation.
[0005] The two most common clinical variants of AAT are the E264V (PiS) and E342K (PiZ) alleles. The clinical single nucleotide variant E342K (PiZ) results in a structurally unstable and / or inactive AAT protein, resulting in toxicity in the liver and inactivity in the lungs. Inheritance is autosomal codominant. The majority of AATD patients carry at least one copy of the E342K mutation.
[0006] The mutation most commonly associated with AATD is a glutamic acid to lysine substitution (E342K) in the SERPINA1 gene that encodes the AAT protein. The E342K mutation is located at the hinge between the beta sheet and the reactive center loop (RCL) of the AAT protein, resulting in a loop-sheet dimer that can subsequently elongate to form long chains of loop-sheet polymers that aggregate the AAT-Z protein in the rough endoplasmic reticulum (rER) of hepatocytes during biosynthesis. This mutation, known as the Z mutation or Z allele, causes the translated protein to misfold incorrectly, so that it is not secreted into the bloodstream. As a result, circulating AAT levels in individuals homozygous for the Z allele (PiZZ) are severely reduced, and only approximately 15% of the mutant Z-AAT protein is correctly folded and secreted from cells. A further consequence of the Z mutation is that secreted Z-AAT has reduced activity compared to the wild-type protein, with 40%-80% of the normal antiprotease activity (American thoracic society / European respiratory society, Am J Respir Crit Care Med. 2003;168(7):818-900; and Ogushi et al. J Clin Invest. 1987;80(5):1366-74).
[0007] There are two disease phenotypes associated with the PiZZ genotype. Accumulation of polymerized Z-AAT protein in hepatocytes results in gain-of-function cytotoxicity and neonatal liver disease that can lead to cellular stress, inflammation, fibrosis, cirrhosis, and hepatocellular carcinoma (HCC) in 12% of patients. This accumulation can spontaneously abate, but can be fatal in a minority of children. The loss-of-function phenotype results from reduced systemic levels of AAT, leading to increased protease digestion of connective tissue in the lower airways. Excessive protease digestion of connective tissue and alveolar lining reduces lung elasticity and function, leading to emphysema, a hallmark of chronic obstructive pulmonary disease (COPD). This effect is severe in PiZZ individuals, typically manifesting in middle age, leading to a reduced quality of life and a shortened lifespan (average 68 years) (Tanash et al. Int J Chron Obstruct Pulm Dis. 2016;11:1663-9). This effect is more pronounced in PiZZ individuals who smoke, resulting in an even shorter lifespan (58 years). Piitulainen and Tanash,COPD 2015;12(1):36-41. PiZZ individuals represent the majority of individuals with clinically relevant AATD lung disease. A milder form of AATD is associated with the SZ genotype in which the Z allele is combined with the S allele. The S allele is associated with a slight reduction in circulating AAT levels but does not cause cytotoxicity in liver cells. The result is clinically significant lung disease but not liver disease. Fregonese and Stolk, Orphanet J Rare Dis. 2008;33:16. As with the ZZ genotype, the lack of circulating AAT in subjects with the SZ genotype results in uncontrolled protease activity that can deteriorate lung tissue over time and cause emphysema, especially in smokers. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] American thoracic society / European respiratory society,Am J Respir Crit Care Med.2003;168(7):818-900 [Non-Patent Document 2] Ogushi et al.J Clin Invest.1987;80(5):1366-74 [Non-Patent Document 3] Tanash et al.Int J Chron Obstruct Pulm Dis.2016;11:1663-9 [Non-Patent Document 4] Piitulainen and Tanash, COPD 2015;12(1):36-41 [Non-Patent Document 5] Fregonese and Stolk,Orphanet J Rare Dis.2008;33:16 Summary of the Invention [Problem to be solved by the invention]
[0009] Treatment options for AATD are limited, and there is currently no cure. A small number of neonatal patients and those with advanced liver disease undergo liver transplantation. The current standard of care for AAT-deficient individuals who have or show signs of developing significant lung disease is augmentation therapy or protein replacement therapy. Augmentation therapy involves the administration of human AAT protein concentrate purified from pooled donor plasma to augment the missing AAT. This treatment involves weekly infusions of AAT protein purified from healthy blood donors. Infusions of plasma protein have been shown to improve survival and slow the rate of progression of emphysema, but in challenging situations (e.g., active lung infection), augmentation therapy alone is often insufficient. Augmentation therapy also fails to restore normal physiological regulation of AAT in patients, and its efficacy has been difficult to demonstrate. Furthermore, augmentation therapy cannot address liver disease caused by the toxic gain-of-function of the Z allele. Thus, new, more effective treatments for AATD are needed. [Means for solving the problem]
[0010] The present disclosure relates to novel compositions, systems and methods for modifying the genome at one or more locations in a host cell, tissue or subject in vivo or in vitro. The present disclosure provides genetic modification systems that can modulate alpha 1 antitrypsin (AAT) activity (e.g., insert, modify, or delete a sequence of interest), and methods for treating alpha 1 antitrypsin deficiency (AATD) by administering one or more such systems to modify a genomic sequence by a single nucleotide to correct a SERPINA1 PiZ mutation that causes alpha 1 antitrypsin deficiency.
[0011] In one aspect, the present disclosure relates to a system for modifying DNA to correct a human SERPINA1 gene mutation that causes AATD, the system comprising: (a) a nucleic acid encoding a genetically modified polypeptide capable of target-priming reverse transcription, the polypeptide comprising (i) a reverse transcriptase domain, and (ii) a Cas9 nickase that binds to DNA and has endonuclease activity; and (b) a template RNA comprising (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene, (ii) a gRNA scaffold that binds to the polypeptide, (iii) a heterologous target sequence comprising a mutation region for correcting the mutation, and (iv) a primer binding site (PBS) sequence comprising at least 3, 4, 5, 6, 7 or 8 bases with 100% homology to the target DNA strand at the 3' end of the template RNA. The SERPINA1 gene may comprise an E342K mutation (also referred to as a PiZ mutation). The template RNA sequence can include a sequence described herein, for example, in Tables 1, 3, 4, 5, 6a, 6B, X2, X3, X3a, X5, or XX.
[0012] The gRNA spacer may comprise at least 15 bases of 100% homology to the target DNA at the 5' end of the template RNA. The template RNA may further comprise a PBS sequence comprising at least 5 bases of at least 80% homology to the target DNA strand. The template RNA may comprise one or more chemical modifications.
[0013] The domains of the genetically modified polypeptide can be linked by a peptide linker. The polypeptide can include one or more peptide linkers. The genetically modified polypeptide can further include a nuclear localization signal. The polypeptide can include two or more nuclear localization signals, such as multiple adjacent nuclear localization signals or one or more nuclear localization signals in different regions of the polypeptide, such as one or more nuclear localization signals at the N-terminus of the polypeptide and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding the genetically modified polypeptide can encode one or more intein domains.
[0014] Introduction of the system into a target cell may result in the insertion of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs of exogenous DNA. Introduction of the system into a target cell may result in a deletion, where the deletion is less than 2, 3, 4, 5, 10, 50, or 100 base pairs of genomic DNA upstream or downstream of the insertion. Introduction of the system into a target cell may result in a substitution, for example, of 1, 2, or 3 nucleotides, for example, consecutive nucleotides.
[0015] The heterologous sequence of interest can be at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600 or 700 base pairs.
[0016] In one aspect, the present disclosure relates to a pharmaceutical composition comprising the above-mentioned system and a pharma- ceutically acceptable excipient or carrier, wherein the pharma- ceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle. In one aspect, the present disclosure relates to a pharmaceutical composition comprising the above-mentioned system and a plurality of pharma- ceutically acceptable excipients or carriers, wherein the pharma- ceutically acceptable excipients or carriers are selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle, for example, wherein the above-mentioned system is delivered by two different excipients or carriers, for example, two lipid nanoparticles, two viral vectors, or one lipid nanoparticle and one viral vector. The viral vector can be an adeno-associated virus (AAV).
[0017] In one aspect, the disclosure relates to a host cell (e.g., a mammalian cell, e.g., a human cell) comprising the above-described system.
[0018] In one aspect, the disclosure relates to a method for correcting a mutation in a human SERPINA1 gene in a cell, tissue, or subject, the method comprising administering the above-described system to a cell, tissue, or subject, and optionally, the correction of the mutant SERPINA1 gene comprises an amino acid substitution of K342E (reversal of the pathogenic substitution, which is E342K). The system can be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid of (a) can be integrated into the genome of a host cell. In some embodiments, the nucleic acid of (a) is not integrated into the genome of a host cell. In some embodiments, the heterologous sequence of interest is inserted at only one target site in the genome of the host cell. The heterologous sequence of interest can be inserted at two or more target sites in the genome of the host cell, for example, at the same corresponding sites in two homologous chromosomes or at two different sites on the same or different chromosomes. The heterologous sequence of interest can encode a mammalian polypeptide or a fragment or variant thereof. The components of the system can be delivered on one, two, three, four, or more different nucleic acid molecules. The system may be introduced into the host cell by electroporation or by using at least one vehicle selected from a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle.
[0019] Features of the compositions or methods may include one or more of the embodiments listed below.
[0020] Enumeration of embodiments 1. (i) a gRNA spacer complementary to a first portion of a human SERPINA1 gene, the gRNA spacer having a sequence that includes the core nucleotide of a gRNA spacer sequence of Table 1 or a sequence with one, two or three substitutions therein, and optionally including one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the gRNA spacer (e.g., including one or more flanking nucleotides adjacent to the core nucleotide), or the gRNA spacer having a sequence of a gRNA spacer of Table 6A, 6B, X2, X3, X3a, X5 or XX or a sequence with one, two or three substitutions therein; (ii) a gRNA scaffold that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide); (iii) a heterologous sequence of interest comprising a mutation region for introducing a mutation into a second portion of the human SERPINA1 gene (e.g., for correcting a mutation therein), wherein, optionally, the heterologous sequence of interest comprises, from 5' to 3', a post-edited homology region, a mutation region and a pre-edited homology region; and (iv) a primer binding site (PBS) sequence comprising at least 3, 4, 5, 6, 7, or 8 bases that has 100% identity to the third portion of the human SERPINA1 gene; (e.g., 5' to 3') of the template RNA.
[0021] 2. The template RNA of embodiment 1, wherein the heterologous sequence of interest comprises a core nucleotide of a RT template sequence from Table 3 or a sequence having one, two or three substitutions therein, and optionally comprises one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence, or the heterologous sequence of interest comprises a sequence of a RT template sequence from Table 6A or 6B.
[0022] 3. The template RNA of embodiment 1, wherein the heterologous sequence of interest comprises the core nucleotide of the RT template sequence of Table 3 corresponding to the gRNA spacer sequence or a sequence having one, two or three substitutions therein, and optionally comprises one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence (e.g., comprises one or more flanking nucleotides adjacent to the core nucleotide), or the heterologous sequence of interest comprises a sequence of a RT template sequence from Table 6A or 6B.
[0023] 4. The template RNA of any of the previous embodiments, wherein the heterologous sequence of interest has the sequence of a heterologous sequence of interest from a template RNA listed in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto, or a sequence having one, two or three substitutions thereto.
[0024] 5. The template RNA according to any of the previous embodiments, wherein the heterologous sequence of interest has a length of 6 to 16 nucleotides (e.g., 6, 8, 10, 12, 14, 15 or 16 nucleotides).
[0025] 6. The template RNA of any one of the previous embodiments, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence from the same row of Table 3 as the RT template sequence, or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence (e.g., includes one or more flanking nucleotides adjacent to the core nucleotide).
[0026] 7. The template RNA of any one of embodiments 1-5, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence of Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence, or the PBS sequence has a sequence that includes the PBS sequence of Table 6A or 6B corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence with one, two or three substitutions therein.
[0027] 8. The template RNA of any of the previous embodiments, wherein the PBS sequence has a PBS sequence from a template RNA listed in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto, or a sequence having one, two or three substitutions therein.
[0028] 9. The template RNA of any of the previous embodiments, wherein the PBS sequence has a length of 8 to 12 nucleotides (e.g., 8, 9, 10, 11 or 12 nucleotides).
[0029] 10. The template RNA of any of the preceding embodiments, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold in Table 12 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0030] 11. The template RNA of any of embodiments 1-9, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold of Table 12 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0031] 12. The template RNA of any of the previous embodiments, wherein the gRNA scaffold has a sequence of a gRNA scaffold from a template RNA listed in Table X2, X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto.
[0032] 13. The template RNA of any of the previous embodiments, comprising a sequence of a template RNA in Table X2, X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto.
[0033] 14. (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene; (ii) a gRNA scaffold that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide); (iii) a heterologous target sequence comprising a mutation region for introducing a mutation into a second portion of the human SERPINA1 gene (e.g., for correcting a mutation therein), the heterologous target sequence comprising the core nucleotide of a RT template sequence from Table 3 or a sequence having one, two or three substitutions therein, and optionally comprising one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence, or comprising a RT template sequence from Table 6A or 6B; and (iv) a PBS sequence containing at least 3, 4, 5, 6, 7, or 8 bases that has 100% identity to the third portion of the human SERPINA1 gene. (e.g., 5' to 3') of the template RNA.
[0034] 15. The template RNA of embodiment 14, wherein the gRNA spacer comprises the core nucleotides of the gRNA spacer sequence of Table 1 or a sequence having one, two or three substitutions therein, and optionally one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the gRNA spacer sequence, or the gRNA spacer comprises a gRNA spacer sequence of Table 6A or 6B.
[0035] 16. The template RNA of embodiment 14, wherein the heterologous sequence of interest comprises the core nucleotides of the gRNA spacer sequence of Table 1 corresponding to the RT template sequence or a sequence having one, two or three substitutions therein, and optionally one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the gRNA spacer sequence, or the heterologous sequence of interest comprises the nucleotides of the gRNA spacer sequence of Table 6A or 6B.
[0036] 17. The template RNA of any one of embodiments 14 to 16, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence from the same row of Table 3 as the RT template sequence, or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence.
[0037] 18. The template RNA of any one of embodiments 14-17, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence of Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence, or the PBS sequence has a sequence that includes the PBS sequence of Table 6A or 6B corresponding to the RT template sequence, the gRNA spacer sequence, or both.
[0038] 19. The template RNA of any of embodiments 14 to 18, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold of Table 6A or 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0039] 20. The template RNA of any of embodiments 14-18, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold of Tables 6A or 12 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0040] 21. A genetic modification system for modifying DNA, comprising: (a) a first RNA comprising, from 5' to 3', (i) a guide RNA sequence complementary to a first portion of a human SERPINA1 gene, wherein the guide RNA sequence comprises a core nucleotide of a spacer sequence in Table 1 or a sequence having one, two or three substitutions therein, and optionally comprises one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the guide RNA sequence; and (ii) a sequence (e.g., a scaffold region) that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide), and (b)(iii) a heterologous sequence of interest comprising a nucleotide substitution for introducing a mutation into a second portion of the human SERPINA1 gene, wherein, optionally, the heterologous sequence of interest comprises, from 5' to 3', a post-edited homology region, a mutated region and a pre-edited homology region; (iv) a primer region comprising at least 5, 6, 7 or 8 bases having 100% identity to the third portion of the human SERPINA1 gene; and (v) a second RNA comprising an RRS (RNA binding protein recognition sequence) that binds to the genetically modified protein. A system including:
[0041] 22. The genetic modification system of embodiment 21, wherein the heterologous sequence of interest comprises a core nucleotide of the RT template sequence from Table 3 or a sequence having one, two or three substitutions therein, and optionally comprises one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence.
[0042] 23. The genetic modification system of embodiment 21, wherein the heterologous sequence of interest comprises the core nucleotide of the RT template sequence of Table 3 corresponding to the gRNA spacer sequence or a sequence having one, two or three substitutions therein, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the RT template sequence.
[0043] 24. The genetic modification system according to any one of embodiments 21 to 23, wherein the PBS sequence comprises a core nucleotide of the PBS sequence from the same row as the RT template sequence in Table 3 or a sequence having one, two or three substitutions therein, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0044] 25. The genetic modification system according to any one of embodiments 21 to 23, wherein the PBS sequence comprises a core nucleotide of the PBS sequence of Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having one, two or three substitutions therein, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0045] 26. The genetic modification system according to any one of embodiments 21 to 25, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold in Table 12 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0046] 27. The genetic modification system according to any one of embodiments 21 to 25, wherein the gRNA scaffold comprises a sequence of the gRNA scaffold of Table 12 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0047] 28. A genetic modification system for modifying DNA, comprising: (a) a first RNA comprising, from 5' to 3', (i) a guide RNA sequence complementary to a first portion of a human SERPINA1 gene, and (ii) a sequence (e.g., a scaffold region) that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide); and (b)(iii) a heterologous sequence of interest comprising nucleotide substitutions for introducing a mutation into a second portion of the human SERPINA1 gene, the heterologous sequence of interest comprising a core nucleotide of the RT template sequence from Table 3 or a sequence having one, two or three substitutions therein, and optionally comprising one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence; and (iv) a primer region comprising at least 5, 6, 7 or 8 bases having 100% homology with the third portion of the human SERPINA1 gene; and (v) a second RNA comprising an RRS (RNA binding protein recognition sequence) that binds to the genetically modified protein. A system including:
[0048] 29. The genetic modification system of embodiment 28, wherein the gRNA spacer comprises a core nucleotide of the gRNA spacer sequence of Table 1 or a sequence having one, two or three substitutions therein, and optionally comprises one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the gRNA spacer sequence.
[0049] 30. The genetic modification system of embodiment 28, wherein the heterologous target sequence comprises a core nucleotide of the gRNA spacer sequence of Table 1 corresponding to the RT template sequence or a sequence having one, two or three substitutions therein, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the gRNA spacer sequence.
[0050] 31. The genetic modification system according to any one of embodiments 28 to 30, wherein the PBS sequence has a sequence comprising the core nucleotide of the PBS sequence from the same row as the RT template sequence in Table 3 or a sequence having one, two or three substitutions therein, and optionally comprising one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0051] 32. The genetic modification system according to any one of embodiments 28 to 30, wherein the PBS sequence comprises a core nucleotide of the PBS sequence of Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having one, two or three substitutions therein, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0052] 33. The genetic modification system according to any one of embodiments 28 to 32, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold in Table 12 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0053] 34. The genetic modification system according to any one of embodiments 28 to 32, wherein the gRNA scaffold comprises a sequence of the gRNA scaffold of Table 12 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0054] 35. A gRNA comprising: (i) a gRNA spacer sequence complementary to a first portion of a human SERPINA1 gene, the gRNA spacer sequence comprising a core nucleotide of a gRNA spacer sequence of Table 1, Table 2 or Table 4 or a sequence having one, two or three substitutions therein, and optionally comprising one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the gRNA spacer sequence; and (ii) a gRNA scaffold.
[0055] 36. The gRNA of embodiment 35, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold in Table 12 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0056] 37. The gRNA of embodiment 35, wherein the gRNA scaffold comprises a sequence of the gRNA scaffold of Table 12 corresponding to the gRNA spacer sequence or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0057] 38. (iii) a heterologous sequence of interest comprising a mutation region for introducing a mutation into a second portion of the human SERPINA1 gene, the heterologous sequence of interest comprising the core nucleotide of a RT template sequence from Table 3 or a sequence having one, two or three substitutions therein, and optionally comprising one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the RT template sequence; and (iv) a template RNA comprising a PBS sequence comprising at least 5, 6, 7 or 8 bases having 100% homology with the third portion of the human SERPINA1 gene.
[0058] 39. The template RNA of embodiment 38, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence from the same row of Table 3 as the RT template sequence or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence.
[0059] 40. The template RNA of embodiment 38, wherein the PBS sequence has a sequence that includes the core nucleotide of the PBS sequence of Table 3 corresponding to the RT template sequence or a sequence with one, two or three substitutions therein, and optionally includes one or more contiguous nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0060] 41. A template RNA described in any one of embodiments 1 to 20 or 38 to 40, a genetic modification system described in any one of embodiments 21 to 34, or a gRNA described in any one of embodiments 35 to 37, wherein the mutation introduced by the system is a K342E mutation in the SERPINA1 gene (e.g., to correct a pathogenic E342K mutation).
[0061] 42. The template RNA according to any one of embodiments 1 to 20 or 38 to 41, or the genetic modification system according to any one of embodiments 21 to 34 or 41, wherein the pre-editing sequence comprises a length of about 1 nucleotide to about 35 nucleotides (e.g., about 1 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, or 30 to 35 nucleotides).
[0062] 43. The template RNA according to any one of embodiments 1 to 20 or 38 to 42, or the gene modification system according to any one of embodiments 21 to 34, 41 or 42, wherein the mutated region comprises a single nucleotide.
[0063] 44. The template RNA according to any one of embodiments 1 to 20 or 38 to 42, or the gene modification system according to any one of embodiments 21 to 34, 41 or 42, wherein the mutated region is at least 2 nucleotides in length.
[0064] 45. The template RNA according to any one of embodiments 1 to 20, 38 to 42 or 44, or the genetic modification system according to any one of embodiments 21 to 34, 41, 42 or 44, wherein the mutated region is at most 32 (e.g., at most 5, 10, 15, 20, 25, 30 or 32) nucleotides in length and comprises one, two or three sequence differences relative to the second part of the human SERPINA1 gene.
[0065] 46. The template RNA according to any one of embodiments 1 to 20, 38 to 42, 44 or 45, or the genetic modification system according to any one of embodiments 21 to 34, 41, 42, 44 or 45, wherein the mutated region comprises two sequence differences relative to the second part of the human SERPINA1 gene.
[0066] 47. The template RNA according to any one of embodiments 1 to 20, 38 to 42 or 44 to 46, or the genetic modification system according to any one of embodiments 21 to 34, 41, 42 or 44 to 46, wherein the mutated region comprises a first region (e.g., a first nucleotide) designed to correct a pathogenic mutation in the SERPINA1 gene and a second region (e.g., a second nucleotide) designed to inactivate the PAM sequence (e.g., a "PAM-kill" mutation as described in Table 5).
[0067] 48. A template RNA described in any one of embodiments 1 to 20, 38 to 46 or a genetic modification system described in any one of embodiments 21 to 34 or 41 to 46, wherein the mutated region has less than 80%, 70%, 60%, 50%, 40% or 30% identity with the corresponding portion of the human SERPINA1 gene.
[0068] 49. The template RNA of any one of the previous embodiments, wherein the template RNA comprises one or more silent mutations (e.g., silent substitutions), e.g., as illustrated in Table 7B.
[0069] 50. The template RNA of any of the previous embodiments, wherein the mutated region comprises a first region designed to correct a pathogenic mutation in the SERPINA1 gene and a second region designed to introduce a silent substitution.
[0070] 51. The template RNA of any one of the preceding embodiments, comprising one or more chemically modified nucleotides.
[0071] 52. A genetically modified system comprising: A template RNA according to any one of embodiments 1 to 20, 38 to 42, or a system according to any one of embodiments 21 to 34, or 41 to 46; and A genetically modified polypeptide or a nucleic acid (e.g., RNA) encoding the genetically modified polypeptide. A genetically modified system comprising:
[0072] 53. A genetically modified polypeptide comprising: A reverse transcriptase (RT) domain (e.g., an RT domain from a retrovirus or a polypeptide domain having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto); and A Cas domain (e.g., a Cas9 domain) that binds to a target DNA molecule and is heterologous to the RT domain; and Optionally, a linker disposed between the RT domain and the Cas domain. The genetic modification system of embodiment 52, comprising:
[0073] 54.RT domain is (a) an RT domain of Table 6; or (b) RT domains from murine leukemia virus (MMLV), porcine endogenous retrovirus (PERV); avian reticuloendotheliosis virus (AVIRE), feline leukemia virus (FLV), simian foamy virus (SFV) (e.g., SFV3L), bovine leukemia virus (BLV), Mason-Pfizer monkey virus (MPMV), human foamy virus (HFV) or bovine foamy / syncytial virus (BFV / BSV). The genetic modification system described in embodiment 53, comprising:
[0074] 55. The genetic modification system according to embodiment 53 or 54, wherein the Cas domain comprises a Cas domain of Table X1, XX or X5 or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto.
[0075] 56. A genetic modification system according to any one of embodiments 53 to 55, wherein the spacer comprises a spacer of Table XX or a sequence having one, two or three substitutions therein, and the Cas domain comprises a Cas domain of the same row of Table XX or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto.
[0076] 57. A genetic modification system according to any of embodiments 53 to 56, wherein the spacer comprises a spacer of Table XX and the Cas domain comprises a Cas domain of the same row of Table XX.
[0077] 58. A genetic modification system according to any one of embodiments 53 to 57, wherein the spacer comprises a spacer of Table X5 or a sequence having one, two or three substitutions therein, and the Cas domain comprises a Cas domain of the same row of Table X5 or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto.
[0078] 59. A genetic modification system according to any of embodiments 53 to 58, wherein the spacer comprises a spacer of Table X5 and the Cas domain comprises a Cas domain of the same row of Table X5.
[0079] 60. A genetic modification system according to any one of embodiments 53 to 59, wherein the spacer comprises a spacer of Table 6A or a sequence having one, two or three substitutions therein, and the Cas domain comprises a Cas domain of the same row of Table 6A or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto.
[0080] 61. A genetic modification system according to any of embodiments 53 to 60, wherein the spacer comprises a spacer of Table 6A and the Cas domain comprises a Cas domain of the same row of Table 6A.
[0081] 62. A genetic modification system according to any one of embodiments 53 to 51, wherein the spacer comprises a spacer of Table 6B or a sequence having one, two or three substitutions therein, and the Cas domain comprises a Cas domain of the same row of Table 6B or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto.
[0082] 63. A genetic modification system according to any of embodiments 53 to 62, wherein the spacer comprises a spacer of Table 6B and the Cas domain comprises a Cas domain of the same row of Table 6B.
[0083] 64. A genetic modification system according to any one of embodiments 53 to 63, wherein the Cas domain comprises a Cas domain of Table 7 or Table 8.
[0084] 65. Cas domain, (a) a Cas9 domain; (b) an SpCas9 domain, a BlatCas9 domain, an Nme2Cas9 domain, a PnpCas9 domain, a SauCas9 domain, a SauCas9-KKH domain, a SauriCas9 domain, a SauriCas9-KKH domain, a ScaCas9-Sc++ domain, a SpyCas9 domain, a SpyCas9-NG domain, a SpyCas9-SpRY domain, or a St1Cas9 domain; and / or (c) A Cas9 domain comprising an N670A mutation, an N611A mutation, an N605A mutation, an N580A mutation, an N588A mutation, an N872A mutation, an N863 mutation, an N622A mutation or an H840A mutation.
[0085] 66. The genetic modification system of embodiment 65, wherein the Cas9 domain binds to a PAM sequence listed in Table 7 or Table 12.
[0086] 67. The genetic modification system described in embodiment 66, wherein a second portion of the human SERPINA1 gene overlaps with the PAM recognized by the Cas domain, e.g., the second portion of the human SERPINA1 gene is within the PAM or the PAM is within the second portion of the human SERPINA1 gene.
[0087] 68. A genetic modification system according to any one of embodiments 53 to 67, wherein the gRNA spacer is a gRNA spacer as described in Table 1 and the Cas domain comprises a Cas domain listed in the same row of Table 1 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0088] 69. The genetic modification system of any one of the previous embodiments, wherein the template RNA comprises a sequence of the template RNA sequence of Table 6A or 6B or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0089] 70. (a) the template RNA comprises a sequence of the template RNA sequence in Table 3; (b) the Cas domain comprises a Cas domain of Table 7 or Table 8; (c) the linker comprises a linker sequence of Table 10 (e.g., any of SEQ ID NOs: 5217, 5106, 5190, and 5218); and (d) A genetically modified system according to any one of embodiments 53 to 69, wherein the genetically modified polypeptide comprises one or two NLS sequences from Table 11 (e.g., any of SEQ ID NOs: 5245, 5290, 5323, 5330, 5349, 5350, 5351, and 4001).
[0090] 71. A genetic modification system described in any of embodiments 53 to 70, which generates a first nick in the first strand of the human SERPINA1 gene.
[0091] 72. The genetic modification system described in embodiment 71, further comprising a second strand-targeting gRNA spacer that introduces a second nick into the second strand of the human SERPINA1 gene.
[0092] 73. The genetic modification system of embodiment 72, wherein the gRNA targeting the second strand comprises a sequence comprising the core nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence from Table 2, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence.
[0093] 74. The genetic modification system of embodiment 72, wherein the gRNA targeting the second strand comprises a sequence comprising a core nucleotide of a left gRNA spacer sequence or a right gRNA spacer sequence from Table 2 corresponding to the gRNA spacer sequence of (i), and optionally one or more consecutive nucleotides starting from the 3' end of a flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence.
[0094] 75. The genetic modification system of embodiment 72, wherein the gRNA targeting the second strand comprises a sequence comprising the core nucleotides of the second nicked gRNA sequence from Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, and optionally one or more contiguous nucleotides starting from the 3' end of the flanking nucleotides of the second nicked gRNA sequence.
[0095] 76. The genetic modification system of embodiment 72, wherein the gRNA targeting the second strand comprises a sequence comprising the core nucleotides of the second nicked gRNA sequence from Table 4 corresponding to the gRNA spacer sequence of (i) or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto, and optionally one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the second nicked gRNA sequence.
[0096] 77. The genetic modification system of any one of the preceding embodiments, wherein the gRNA targeting the second strand has a "PAM-in orientation" with the template RNA of the genetic modification system, for example as illustrated in Table 4.
[0097] 78. The genetic modification system of any one of the previous embodiments, wherein the gRNA targeting the second strand targets a sequence that overlaps with the target mutation of the template RNA.
[0098] 79. The gRNA targeting the second strand is (i) a sequence complementary to the SERPINA1 mutation (e.g., a spacer sequence); (ii) a sequence complementary to a wild-type sequence at the target locus (e.g., a spacer sequence); (iii) a sequence (e.g., a spacer sequence) that is complementary to a SNP proximal to the target locus, e.g., a SNP contained in the genomic DNA of a subject (e.g., a patient); (iv) a sequence that is complementary to or contains one or more silent substitutions proximal to the target locus (e.g., a spacer sequence). The genetic modification system of embodiment 78, comprising:
[0099] 80. The template RNA, genetic modification system, or gRNA of any one of the previous embodiments, wherein the gRNA spacer comprises about 1, 2, 3 or more flanking nucleotides of the gRNA spacer.
[0100] 81. The template RNA or genetic modification system according to any one of the previous embodiments, wherein the heterologous sequence of interest comprises about 2, 3, 4, 5, 10, 20, 30, 40 or more flanking nucleotides of the RT template sequence.
[0101] 82. The template RNA or genetic modification system according to any one of the preceding embodiments, wherein the heterologous sequence of interest comprises about 8-30, 9-25, 10-20, 11-16 or 12-15 (e.g., about 11-16) nucleotides.
[0102] 83. The template RNA or gene modification system according to any one of the previous embodiments, wherein the mutated region comprises a sequence difference of one, two or three nucleotide positions relative to the corresponding part of the human SERPINA1 gene.
[0103] 84. The template RNA or gene modification system according to any one of the previous embodiments, wherein the mutated region comprises a sequence difference of at least two nucleotide positions relative to the corresponding portion of the human SERPINA1 gene.
[0104] 85. The template RNA or gene modification system according to any one of the previous embodiments, wherein the post-edited homology region and / or the pre-edited homology region comprises 100% identity with the SERPINA1 gene.
[0105] 86. The template RNA or genetic modification system according to any one of the previous embodiments, wherein the PBS sequence further comprises about 1, 2, 3, 4, 5, 6, 7 or more flanking nucleotides.
[0106] 87. The template RNA or genetic modification system according to any one of the preceding embodiments, wherein the PBS sequence comprises about 5-20, 8-16, 8-14, 8-13, 9-13, 9-12 or 10-12 (e.g., about 9-12) nucleotides.
[0107] 88. The template RNA or gene modification system according to any one of the previous embodiments, wherein the PBS sequence is bound within 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides of the nick site of the SERPINA1 gene.
[0108] 89. A genetically modified system according to any one of the previous embodiments, wherein the domains of the genetically modified polypeptide are linked by a peptide linker.
[0109] 90. The genetic modification system of embodiment 89, wherein the linker comprises a linker sequence of Table 10 (e.g., any of SEQ ID NOs: 5217, 5106, 5190, and 5218).
[0110] 91. A genetically modified system according to any one of the preceding embodiments, wherein the genetically modified polypeptide further comprises one or more nuclear localization sequences (NLS).
[0111] 92. The genetically modified system described in embodiment 91, wherein the genetically modified polypeptide comprises a first NLS and a second NLS.
[0112] 93. A genetic modification system described in embodiment 91 or 92, wherein the NLS comprises an NLS sequence of Table 11 (e.g., any of SEQ ID NOs: 5245, 5290, 5323, 5330, 5349, 5350, 5351, and 4001).
[0113] 94. A template RNA comprising a sequence of a template RNA of Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0114] 95. A template RNA comprising the sequence of a template RNA in Table 4.
[0115] 96. A genetically modified system comprising: (i) a template RNA comprising a sequence of a template RNA in Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto; and (ii) a second nicked gRNA sequence from the same row as (i) of Table 4, having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto. A genetically modified system comprising:
[0116] 97. A genetically modified system comprising: (i) a template RNA comprising a template RNA sequence of Table 4; and (ii) a second nicked gRNA sequence from the same row as (i) of Table 4; A genetically modified system comprising:
[0117] 98. A DNA encoding the template RNA according to any one of embodiments 1 to 20, 38 to 48, 80 to 88, 94 or 95, or the gRNA according to any one of embodiments 35 to 37.
[0118] 99. A pharmaceutical composition comprising the system of any one of embodiments 52-93, 96 or 97 or one or more nucleic acids encoding same and a pharma- ceutically acceptable excipient or carrier.
[0119] 100. The pharmaceutical composition according to embodiment 99, wherein the pharma- ceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle and a lipid nanoparticle.
[0120] 101. The pharmaceutical composition of embodiment 100, wherein the viral vector is an adeno-associated virus.
[0121] 102. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising a template RNA or a genetic modification system according to any one of the previous embodiments.
[0122] 103. A method for producing a template RNA according to any one of embodiments 1-20, 38-48, 80-88, 94, or 953, comprising synthesizing the template RNA by in vitro transcription (e.g., solid-phase synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow for the production of the template RNA.
[0123] 104. A method for modifying a target site in the human SERPINA1 gene in a cell, comprising contacting the cell with a gene modification system of any one of embodiments 52 to 93, 96 or 97 or a DNA encoding the same, thereby modifying the target site in the human SERPINA1 gene in the cell.
[0124] 105. A method for modifying a target site in the human SERPINA1 gene in a cell, comprising contacting a cell with (i) a template RNA or a DNA encoding the same of any one of embodiments 52 to 93, 96 or 97; and (ii) a genetically modified polypeptide or a nucleic acid encoding a genetically modified polypeptide, thereby modifying the target site in the human SERPINA1 gene in the cell.
[0125] 106. A method for treating a subject having a disease or symptom associated with a mutation in the human SERPINA1 gene, comprising administering to the subject a genetic modification system described in any one of embodiments 52 to 93, 96 or 97 or a DNA encoding the same, thereby treating the subject having a disease or symptom associated with a mutation in the human SERPINA1 gene.
[0126] 107. A method for treating a subject having a disease or condition associated with a mutation in the human SERPINA1 gene, comprising administering to the subject a template RNA or a DNA encoding the same of any one of embodiments 52 to 93, 96 or 97; and (ii) a genetically modified polypeptide or a nucleic acid encoding a genetically modified polypeptide, thereby treating the subject having a disease or condition associated with a mutation in the human SERPINA1 gene.
[0127] 108. The method of embodiment 106 or 107, wherein the disease or condition is alpha 1 antitrypsin deficiency (AATD).
[0128] 109. The method of any one of embodiments 106 to 108, wherein the subject has an E342K mutation (i.e., a PiZ mutation).
[0129] 110. A method for treating a subject having AATD, comprising administering to the subject a genetic modification system described in any one of embodiments 52 to 93, 96 or 97, or a DNA encoding the same, thereby treating the subject having AATD.
[0130] 111. A method for treating a subject having AATD, comprising: (i) administering to the subject a template RNA or a DNA encoding the same according to any one of embodiments 52 to 93, 96 or 97; and (ii) a genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, thereby treating the subject having AATD.
[0131] 112. A genetic modification system or method according to any one of the preceding embodiments, wherein introduction of the system into a target cell results in correction of a pathogenic mutation in the SERPINA1 gene.
[0132] 113. A genetic modification system or method according to any one of the preceding embodiments, wherein the pathogenic mutation is an E342K mutation and the correction comprises an amino acid substitution of K342E.
[0133] 114. A genetic modification system or method according to any of the previous embodiments, wherein the correction of the mutation occurs in at least 30% (e.g., 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.
[0134] 115. The genetic modification system or method of any of the previous embodiments, wherein correction of the mutation occurs in at least 30% (e.g., 30%, 40%, 50%, 60%, 70% or more) of the targeted cells.
[0135] 116. The genetic modification system or method of any of the previous embodiments, wherein the genetic modification system comprises a gRNA targeting the second strand, and the correction of mutations in the population of target cells is increased relative to a population of target cells treated with a genetic modification system comprising a template RNA that does not have a gRNA targeting the second strand.
[0136] 117. The genetic modification system or method of any of the previous embodiments, wherein the template RNA contains one or more silent substitutions (e.g., as exemplified in Table 7B), and the correction of mutations in the population of target cells is increased relative to a population of target cells treated with a genetic modification system containing a template RNA that does not contain one or more silent substitutions.
[0137] 118. The method of any of the previous embodiments, wherein the cell is a mammalian cell, such as a human cell.
[0138] 119. The method of any one of the preceding embodiments, wherein the subject is a human.
[0139] 120. The method of any of the previous embodiments, wherein the contacting is performed ex vivo, e.g., the DNA of the cell or subject is modified ex vivo.
[0140] 121. The method of any of the previous embodiments, wherein the contacting is performed in vivo, e.g., the DNA of a cell or a subject is modified in vivo.
[0141] 122. The method of any of the previous embodiments, comprising contacting a cell or a cell within a subject with a nucleic acid (e.g., DNA or RNA) encoding a genetically modified polypeptide under conditions that allow for production of the genetically modified polypeptide, wherein contacting the cell or subject with the system.
[0142] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee. [Brief description of the drawings]
[0143] [Figure 1]1 shows a genetic modification system as described herein. The diagram on the left shows a genetically modified polypeptide, including a Cas nickase domain (e.g., spCas9 N863A) and a reverse transcriptase domain (RT domain) linked by a linker. The diagram on the right shows a template RNA, including, from 5' to 3', a gRNA spacer, a gRNA scaffold, a heterologous sequence of interest, and a primer binding site sequence (PBS sequence). The heterologous sequence of interest may include a mutated region that includes one or more sequence differences with respect to the target site. The heterologous sequence of interest may also include a pre-editing homology region and a post-editing homology region adjacent to the mutated region. Without wishing to be bound by theory, it is believed that the gRNA spacer of the template RNA binds to the second strand of the target site in the genome, and the gRNA scaffold of the template RNA binds to the genetically modified polypeptide, for example, localizing the genetically modified polypeptide to the target site in the genome. The Cas domain of the genetically modified polypeptide is believed to nick the target site (e.g., the first strand of the target site) and, for example, allow the PBS sequence to bind to the sequence adjacent to the site to be modified on the first strand of the target site. The RT domain of the genetically modified polypeptide is believed to use the first strand of the target site bound to the complementary sequence containing the PBS sequence of the template RNA as a primer and the heterologous target sequence of the template RNA as a template, for example to polymerize a sequence complementary to the heterologous target sequence. Without wishing to be bound by theory, it is believed that reverse transcription then proceeds through the pre-editing homology region, then the mutation region, and then the post-editing homology region, thereby generating a DNA strand containing the mutation specified by the heterologous target sequence. [Diagram 2] Graph showing the percentage of rewriting achieved using RNAV209-013 or RNAV214-040 genetically modified polypeptides with the indicated template RNA. [Diagram 3] 1 is a graph showing the amount of Fah mRNA relative to the wild type when template RNA was used together with RNAV209-013 or RNAV214-040 genetically modified polypeptides. [Figure 4]1 is a graph showing the percentage of Cas9-positive hepatocytes 6 hours after administration of LNPs containing various genetically modified polypeptides and template RNA. [Diagram 5] 1 is a graph showing rewrite levels in liver samples 6 days after administration of LNPs containing various genetically modified polypeptides and template RNA. [Figure 6] 1 is a graph showing restoration of wild-type Fah mRNA in liver samples following administration of LNPs containing various genetically modified polypeptides and template RNA, compared to littermate heterozygous mice. [Figure 7] 1 is a graph showing the distribution of Fah protein in liver samples after administration of LNPs containing various genetically modified polypeptides and template RNA. [Figure 8] A series of western blots showing Cas9-RT expression 6 hours after injection of Cas9-RT mRNA+TTR-guided LNP. Each lane represents an individual animal, and 20 μg of tissue homogenate was added per lane. Positive controls are from in vitro cell experiments in which Cas9-RT was expressed (see above). GAPDH was used as a loading control for each sample. n=4 per vehicle or treatment group. [Figure 9] 13A-C are graphs showing gene editing of the TTR locus following treatment with Cas9-RT mRNA+TTR-guided LNP.Levels of indels detected at the TTR locus as measured by TIDE analysis of Sanger sequencing of the protospacer-targeted TTR locus. [Figure 10] Graph showing that TTR serum levels are reduced following treatment with Cas9-RT mRNA+TTR guided LNPs. Measurement of circulating TTR levels 5 days after treatment of mice with LNPs encapsulating Cas9-RT+TTR guide RNA. [Figure 11]Figure 1 shows Cas9-RT expression after injection of Cas9-RT mRNA+TTR guided LNP. Relative expression was quantified by ProteinSimple Jess capillary electrophoresis Western blot. Numbers within symbols are animal numbers within groups. Vehicle n=2, Cas9-RT+TTR guided n=3. [Figure 12] Figure 1 shows gene editing of the TTR locus following injection of Cas9-RT mRNA+TTR-guided LNP. The level of indels detected at the TTR locus was measured by amplicon sequencing of the TTR locus targeted by the protospacer. Eight different biopsies were taken throughout the liver of each animal, and the percentage of reads showing indels was measured by amplicon sequencing. [Figure 13] Graph showing the percentage of indel activity of various gene modification systems including template RNAs containing five SpCas9 spacers in combination with wild-type SpCas9 polypeptides evaluated in HEK293T cells. [Figure 14] 13 is a graph showing the percentage of indels at PiZ mutation sites in HEK293T landing pad cells after treatment with the genetic modification system. [Figure 15] Graph showing ranking of active spacers by indel activity and distance from PiZ mutation after screening evaluation in HEK293T cells. [Figure 16] 1 is a graph showing the percentage of complete rewriting activity of various gene modification systems containing template RNA. [Figure 17A-B] This is a heat map graphing the rewrite rates of a genetically modified system containing various SpRY_ED0 template RNAs (with different PBS and RT lengths) and an exemplary SpRY_Cas9-containing genetically modified polypeptide (Figure 17A) and a genetically modified system containing various St1_ED4 template RNAs (with different PBS and RT lengths) and an exemplary St1Cas9-containing genetically modified polypeptide (Figure 17B). [Figure 18]1 is a graph showing the 17 best performing combinations of template RNA and genetically modified polypeptides containing Cas9 variants (ranked by rewriting activity). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0144] definition The term "expression cassette," as used herein, refers to a nucleic acid construct that contains sufficient nucleic acid elements for expression of a nucleic acid molecule of the present invention.
[0145] "gRNA spacer," as used herein, refers to a portion of a nucleic acid that has complementarity to a target nucleic acid and, together with the gRNA scaffold, can target a Cas protein to the target nucleic acid.
[0146] "gRNA scaffold" as used herein refers to a portion of a nucleic acid that can bind to a Cas protein and, together with a gRNA spacer, target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold comprises a crRNA sequence, a tetraloop, and a tracrRNA sequence.
[0147] "Genetically modified polypeptide" as used herein refers to a polypeptide comprising a retroviral reverse transcriptase or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity to a retroviral reverse transcriptase that is capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule within the host cell). In some embodiments, the genetically modified polypeptide is capable of integrating a sequence substantially independent of the host machinery. In some embodiments, the genetically modified polypeptide integrates a sequence at a random location within the genome, and in some embodiments, the genetically modified polypeptide integrates a sequence at a specific target site. In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively 1) bind to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate integration of at least a portion of the template nucleic acid into the target DNA. Genetically modified polypeptides include both naturally occurring polypeptides and engineered variants thereof, e.g., having one or more amino acid substitutions relative to the naturally occurring sequence. Genetically modified polypeptides also include heterologous constructs, for example, where one or more of the domains listed above are heterologous to each other, whether by heterologous fusion (or other conjugates) of otherwise wild-type domains as well as fusion of modified domains, for example, by substitution or fusion of heterologous subdomains or other replacement domains. Exemplary genetically modified polypeptides that can be used in the methods provided herein, and systems comprising and methods of using them, are described in PCT / US2021 / 020948, which is incorporated herein by reference, for example, for genetically modified polypeptides comprising retroviral reverse transcriptase domains. In some embodiments, the genetically modified polypeptide incorporates a sequence into a gene. In some embodiments, the genetically modified polypeptide incorporates a sequence into a sequence outside of a gene. "Genetically modified system" as used herein refers to a system comprising a genetically modified polypeptide and a template nucleic acid.
[0148] The term "domain" as used herein refers to a structure of a biomolecule that contributes to a specific function of the biomolecule. A domain may include a continuous region (e.g., a continuous sequence) or a discrete non-contiguous region (e.g., a non-contiguous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA binding domains, and reverse transcription domains; examples of nucleic acid domains include regulatory domains, such as transcription factor binding domains. In some embodiments, a domain (e.g., a Cas domain) may include two or more smaller domains (e.g., a DNA binding domain and an endonuclease domain).
[0149] As used herein, the term "exogenous" when used in reference to a biomolecule (such as a nucleic acid sequence or a polypeptide) means that the biomolecule has been introduced into a host genome, cell, or organism by man. For example, a nucleic acid that is added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods is exogenous to the existing nucleic acid sequence, cell, tissue, or subject.
[0150] As used herein, "first strand" and "second strand" used to describe the individual DNA strands of a target DNA distinguish two DNA strands on which the reverse transcriptase domain initiates polymerization, e.g., on which target primed synthesis is initiated. The first strand refers to the strand of the target DNA on which the reverse transcriptase domain initiates polymerization, e.g., target primed synthesis is initiated. The second strand refers to the other strand of the target DNA. The names first strand and second strand do not otherwise describe the target site DNA strands; for example, in some embodiments, the first strand and second strand are nicked by the polypeptides described herein, but the names "first" and "second" strand are independent of the order in which such nicks appear.
[0151] The term "heterologous", when used to describe a first element with respect to a second element, means that the first and second elements do not naturally occur in the arrangement as described. For example, a heterologous polypeptide, nucleic acid molecule, construct or sequence refers to (a) a polypeptide, nucleic acid molecule, or a portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or a portion of a polypeptide or nucleic acid molecule that is modified or mutated relative to its natural state, or (c) a polypeptide or nucleic acid molecule that has an altered expression compared to the native expression level under similar conditions. For example, a heterologous regulatory sequence (e.g., promoter, enhancer) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from that in which the gene or nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide, or a nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide, or a nucleic acid encoding a DNA-binding domain of a polypeptide) can be positioned relative to other domains or can be of a different sequence or from a different source compared to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may be present naturally within the host cell genome, but may have an altered expression level or a different sequence, or both. In other embodiments, a heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but may instead be introduced into the host cell by transformation (e.g., transfection, electroporation), where the added molecule may be integrated into the host genome or may exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids or other self-replicating vectors).
[0152] As used herein, "insertion" of a sequence into a target site refers to the net addition of a DNA sequence at the target site, e.g., where there is a new nucleotide in the heterologous sequence of interest that does not have a cognate position in the non-edited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous sequence of interest to the target nucleic acid sequence will result in an alignment gap in the target nucleic acid sequence.
[0153] As used herein, a "deletion" generated by a heterologous sequence of interest at a target site refers to a net deletion of DNA sequence at the target site, e.g., where there is a nucleotide in the unedited target site that does not have a cognate position in the heterologous sequence of interest. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous sequence of interest to the target nucleic acid sequence will result in an alignment gap in the molecule that contains the PBS sequence and the heterologous sequence of interest.
[0154] The term "inverted terminal repeat" or "ITR" as used herein refers to an AAV viral cis element so named because of its symmetry. The element promotes efficient multiplication of the AAV genome. The minimal elements for ITR function are assumed to be a Rep binding site (RBS; 5'-GCGCGCTCGCTCGCTC-3' for AAV2; SEQ ID NO: 4601) and a terminal separation site (TRS; 5'-AGTTGG-3' for AAV2) and a variable palindrome sequence that allows hairpin formation. According to the present invention, an ITR comprises at least these three elements (RBS, TRS, and a sequence that allows hairpin formation). In addition, in the present invention, the term "ITR" refers to the ITRs of known natural AAV serotypes (e.g., ITRs of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), chimeric ITRs formed by fusion of ITR elements from different serotypes, and functional variants thereof. "Functional variant" refers to a sequence that exhibits at least 80%, 85%, 90%, preferably at least 95% sequence identity with a known ITR and allows for the enrichment of the sequence containing said ITR in the presence of Rep protein.
[0155] The term "mutated region," as used herein, refers to a region in a template RNA that has one or more sequence differences compared to the corresponding sequence in a target nucleic acid. Sequence differences can include, for example, substitutions, insertions, frameshifts, or deletions.
[0156] The term "mutated" when applied to a nucleic acid sequence means that nucleotides within a nucleic acid sequence have been inserted, deleted, or changed relative to a reference (e.g., naturally occurring) nucleic acid sequence. A single alteration may be made at a locus (point mutation), or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.
[0157] "Nucleic acid molecule" refers to both RNA and DNA molecules, including but not limited to complementary DNA ("cDNA"), genomic DNA ("gDNA") and messenger RNA ("mRNA"), including synthetic nucleic acid molecules, such as those chemically synthesized or recombinantly produced, such as RNA templates, as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule can be the sense or antisense strand. Unless otherwise indicated, and by way of example of all sequences described herein under the general format "SEQ ID NO:" or "nucleic acid comprising SEQ ID NO:1", refers to a nucleic acid, at least a portion of which has either (i) the sequence of SEQ ID NO:1, or (ii) a sequence complementary to SEQ ID NO:1. The choice between the two is determined by the context in which SEQ ID NO:1 is used. For example, if the nucleic acid is used as a probe, the choice between the two is determined by the requirement that the probe be complementary to the desired target. The nucleic acid sequences of the present disclosure may be chemically or biochemically modified or may contain non-natural or derivatized nucleotide bases, as will be readily understood by those skilled in the art. Such modifications include, for example, labels, methylation, replacement of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged bonds (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties, (e.g., polypeptides), intercalating agents (e.g., acridines, psoralens, etc.), chelators, alkylators, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Chemically modified bases (e.g., see Table 13), backbones (e.g., see Table 14), and modified caps (e.g., see Table 15) are also included. Synthetic molecules that mimic polynucleotides in their ability to bind to a specified sequence through hydrogen bonds and other chemical interactions are also included. Such molecules are known in the art and include, for example, those that use peptide bonds instead of phosphate bonds in the backbone of the molecule, eg, peptide nucleic acids (PNA).Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structures such as modifications found in "locked" nucleic acids (LNA). In various embodiments, the nucleic acid is in operative association with additional genetic elements, such as tissue-specific expression-control sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences) and additional elements, such as inverted repeats (e.g., inverted terminal repeats, e.g., elements derived from viruses (e.g., AAV ITRs)) and tandem repeats, inverted / direct repeats, regions of homology (segments with differing degrees of homology to the target DNA), untranslated regions (UTRs) (5', 3', or both 5' and 3' UTRs), and various combinations of the foregoing. The nucleic acid elements of the systems provided by the present invention can be provided in various topologies, including single-stranded, double-stranded, circular, linear, open-ended linear, closed-ended linear, and specific versions of these, such as doggybone DNA (dbDNA), closed-ended DNA (ceDNA).
[0158] As used herein, a "gene expression unit" is a nucleic acid sequence that includes at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked to a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For example, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. Operably linked DNA sequences can be contiguous or non-contiguous. When it is necessary to link two protein coding regions, operably linked sequences can be in the same reading frame.
[0159] The term "host genome" or "host cell" as used herein refers to a cell and / or its genome into which proteins and / or genetic material have been introduced. Such terms refer not only to the particular subject cell and / or genome, but also to the progeny of such a cell and / or the genome of the progeny of such a cell. It is understood that such progeny may not in fact be identical to the parent cell, since certain modifications may occur in subsequent generations due to mutations or environmental influences, but are still included within the scope of the term "host cell" as used herein. The host genome or host cell may be an isolated cell or cell line grown in culture or genomic material isolated from such a cell or cell line, or may be a host cell or host genome that constitutes a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, for example, as described herein. In certain examples, the host cell may be a mammalian cell, a human cell, a bird cell, a reptilian cell, a bovine cell, a horse cell, a pig cell, a goat cell, a sheep cell, a chicken cell, or a turkey cell. In particular examples, the host cell can be a corn cell, a soybean cell, a wheat cell, or a rice cell.
[0160] As used herein, "operative association" describes the functional relationship between two nucleic acid sequences, e.g., 1) a promoter and 2) a heterologous sequence of interest, and in such instances means that the promoter and the heterologous sequence of interest (e.g., a gene of interest) are oriented such that, under appropriate conditions, the promoter drives expression of the heterologous sequence of interest. For example, a template nucleic acid having a promoter and a heterologous sequence of interest can be, for example, single-stranded in a (+) or (-) orientation. The "operative association" between the promoter and the heterologous sequence of interest in this template means that, regardless of whether the template nucleic acid is to be transcribed in a particular state, it will be correctly transcribed when in the appropriate state (e.g., in a (+) orientation, in the presence of required catalytic factors and NTPs, etc.). The operative association applies similarly to other pairs of nucleic acids, including sequences encoding other tissue-specific expression control sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homology regions, and heterologous sequences of interest, or retroviral RT domains.
[0161] The term "primer binding site sequence" or "PBS sequence" as used herein refers to a portion of a template RNA that can bind to a region contained in a target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence that contains at least 3, 4, 5, 6, 7, or 8 bases that have 100% identity to a region contained in a target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 bases that have 100% identity to a region contained in a target nucleic acid sequence. Without intending to be bound by a particular theory, in some embodiments, when the template RNA contains a PBS sequence and a heterologous target sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence to enable the reverse transcriptase domain to use that region as a primer for reverse transcription and to use the heterologous target sequence as a template for reverse transcription.
[0162] As used herein, a "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) having a stem containing sufficient self-complementarity to form a stem-loop, e.g., at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs, and a loop having at least 3 (e.g., 4) base pairs. The stem may contain mismatches or bulges.
[0163] As used herein, "tissue-specific expression-control sequence" refers to a nucleic acid element that increases or decreases the level of a transcript that contains a heterologous sequence of interest in a tissue-specific manner in a target tissue, e.g., preferentially in an on-target tissue compared to an off-target tissue. In some embodiments, the tissue-specific expression-control sequence preferentially drives or inhibits the transcription, activity, or half-life of a transcript that contains a heterologous sequence of interest in a target tissue, e.g., preferentially in an on-target tissue compared to an off-target tissue. Exemplary tissue-specific expression-control sequences include tissue-specific promoters, repressors, enhancers, or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to on-target (tissues in which expression or activity of the template nucleic acid is desired or acceptable) and off-target (tissues in which expression or activity of the template nucleic acid is undesirable or unacceptable). For example, a tissue-specific promoter preferentially drives expression in on-target tissue compared to off-target tissue. In contrast, microRNAs that bind to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues compared to on-target tissues, thereby reducing the expression of the template nucleic acid in the off-target tissues. Thus, promoters and microRNA recognition sequences specific to the same tissue, e.g., target tissue, have contrasting functions in terms of transcription, activity, or half-life of the associated sequence in the tissue (matching expression levels, i.e., promoting and suppressing high levels of microRNA in off-target tissues and low levels in on-target tissues, respectively, while the promoter drives high expression in on-target tissues and low expression in off-target tissues).
[0164] List of Headlines 1) Introduction 2) Genetic modification system a) Polypeptide components of the genetic modification system i) Lighting Domain ii) Endonuclease domain and DNA binding domain (1) A genetically engineered polypeptide containing a Cas domain. (2) TAL effectors and zinc finger nucleases iii) Linker iv) Localization sequences for gene modification systems v) Evolved variants of genetically engineered polypeptides and systems vi) Intein vii) Further domains b) Template Nucleic Acid i) gRNA spacer and gRNA scaffold ii) Heterologous sequence of interest iii) PBS sequence iv) Exemplary template sequences c) gRNA with inducible activity d) Circular RNA and ribozymes in genetic engineering systems e) Target nucleic acid site f) Second Strand Nicking 3) Preparation of compositions and systems 4) Therapeutic use 5) Administration and Delivery a) Tissue-specific activity / administration i) Promoter ii) microRNA b) Viral vectors and their components c) AAV administration d) Lipid nanoparticles 6) Kits, Products, and Pharmaceutical Compositions 7) Chemistry, Manufacturing, and Controls (CMC)
[0165] introduction The present disclosure relates to methods for treating Alpha 1 Antitrypsin Deficiency (AATD) and compositions for targeting, editing, modifying or manipulating DNA sequences (e.g., inserting a heterologous sequence of interest at a target site in a mammalian genome), e.g., at one or more locations within a DNA sequence in a cell, tissue or subject in vivo or in vitro. The heterologous DNA sequence of interest may, for example, include a substitution.
[0166] More specifically, the present disclosure provides methods for treating AATD using a reverse transcriptase-based system to alter a subject's genomic DNA sequence, for example, by inserting, deleting, or substituting one or more nucleotides in the sequence of interest.
[0167] The present disclosure provides, in part, a method for treating AATD using a genetic modification system that includes a genetically modified polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the genetic modification system can be used to introduce modifications into a target site in a genome. In some embodiments, the genetically modified polypeptide component includes a writing domain (e.g., reverse transcriptase domain), a DNA binding domain, and an endonuclease domain (e.g., nickase domain). In some embodiments, the template nucleic acid (e.g., template RNA) includes a sequence (e.g., gRNA spacer) that binds to a target site in a genome (e.g., binds to the second strand of the target site), a sequence that binds to the genetically modified polypeptide component (e.g., gRNA scaffold), a heterologous target sequence, and a PBS sequence. Without wishing to be bound by theory, it is believed that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site in a genome and binds to the genetically modified polypeptide component (e.g., localizes the polypeptide component to the target site in a genome). It is believed that the endonuclease (e.g., nickase) of the genetically modified polypeptide component can cleave the target site (e.g., the first strand of the target site) and, for example, the PBS sequence can bind to the sequence adjacent to the site to be modified on the first strand of the target site. It is believed that the writing domain (e.g., reverse transcriptase domain) of the polypeptide component can polymerize, for example, a sequence complementary to the heterologous target sequence, using the first strand of the target site bound to a complementary sequence including the PBS sequence of the template nucleic acid as a primer and the heterologous target sequence of the template nucleic acid as a template. Without wishing to be bound by theory, it is believed that the selection of an appropriate heterologous target sequence can result in the substitution, deletion and / or insertion of one or more nucleotides at the target site.
[0168] Genetic modification system In some embodiments, the genetically modified system described herein comprises (A) a genetically modified polypeptide or a nucleic acid encoding a genetically modified polypeptide, wherein the genetically modified polypeptide comprises (i) a reverse transcriptase domain, and either (x) an endonuclease domain with DNA binding functionality or (y) an endonuclease domain and a separate DNA binding domain; and (B) a template RNA. In some embodiments, the genetically modified polypeptide acts as a substantially autonomous protein machinery capable of incorporating a template nucleic acid sequence into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule within the host cell) without substantial reliance on the host machinery. For example, the genetically modified protein may comprise a DNA binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, the DNA binding function may comprise an RNA component, e.g., a gRNA spacer, that directs the protein to the DNA sequence. In other embodiments, the genetically modified polypeptide may comprise a reverse transcriptase domain and an endonuclease domain. The RNA template element of the genetically modified system is typically heterologous to the genetically modified polypeptide element and provides a sequence of interest to be inserted (reverse transcribed) into the host genome. In some embodiments, the genetically modified polypeptide is capable of target priming reverse transcription. In some embodiments, the genetically modified polypeptide is capable of second strand synthesis.
[0169] In some embodiments, the gene modification system is combined with a second polypeptide. In some embodiments, the second polypeptide may comprise an endonuclease domain. In some embodiments, the second polypeptide may comprise a polymerase domain, for example, a reverse transcriptase domain. In some embodiments, the second polypeptide may comprise a DNA-dependent DNA polymerase domain. In some embodiments, the second polypeptide assists in completing genome editing, for example, by contributing to second strand synthesis or DNA repair recovery.
[0170] A functional genetically modified polypeptide can be composed of unrelated DNA binding, reverse transcription, and endonuclease domains. This modular structure allows the combination of functional domains, such as dCas9 (DNA binding), MMLV reverse transcriptase (reverse transcription), FokI (endonuclease). In some embodiments, multiple functional domains can occur from a single protein, such as Cas9 or Cas9 nickase (DNA binding, endonuclease).
[0171] In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively 1) bind to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the genetically modified polypeptide is a modified polypeptide that comprises one or more amino acid substitutions with respect to the corresponding native sequence. In some embodiments, the genetically modified polypeptide comprises two or more domains that are heterologous to each other, for example, by heterologous fusion (or other conjugate) of an otherwise wild-type domain as well as fusion of a modified domain, for example, by substitution or fusion of a heterologous subdomain or other replacement domain. For example, in some embodiments, one or more of the following are present: the RT domain is heterologous to the DBD; the DBD is heterologous to the endonuclease domain; or the RT domain is heterologous to the endonuclease domain.
[0172] In some embodiments, a template RNA molecule for use in the system comprises, from 5' to 3', (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous sequence of interest; and (4) a primer binding site (PBS) sequence. (1) a gRNA spacer of about 18 to 22 nt, for example, 20 nt; (2) A gRNA scaffold that includes one or more hairpin loops, e.g., one, two, or three loops for associating a template with a Cas domain, e.g., a nickase Cas9 domain. In some embodiments, the gRNA scaffold includes, from 5' to 3', the sequence: GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 5008). (3) In some embodiments, the heterologous sequence of interest is, for example, 7 to 74, for example, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70, or 70 to 80 or 80 to 90 nt in length. In some embodiments, the first (usually 5') base of the sequence is not C. (4) In some embodiments, the PBS sequence that binds to the target priming sequence after nicking has occurred is, for example, 3 to 20 nt, for example, 7 to 15 nt, for example, 12 to 14 nt, In some embodiments, the PBS sequence has a GC content of 40 to 60%.
[0173] In some embodiments, a second gRNA associated with the system may help drive complete integration. In some embodiments, the second gRNA may target a position 0-200 nt away from the first strand nick, e.g., 0-50, 50-100, 100-200 nt away from the first strand nick. In some embodiments, the second gRNA may only bind to its target sequence after editing has been made, e.g., the gRNA binds to a sequence present in the heterologous sequence of interest but not in the initial target sequence.
[0174] In some embodiments, the genetic modification systems described herein are used to perform editing in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the genetic modification systems are used to perform editing in primary cells, such as primary cortical neurons from E18.5 mice.
[0175] In some embodiments, the genetically modified polypeptides described herein comprise a reverse transcriptase or RT domain (e.g., as described herein) comprising a MoMLV RT sequence or a variant thereof. In some embodiments, the MoMLV RT sequence comprises one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In some embodiments, the MoMLV RT sequence comprises a combination of mutations such as D200N, L603W, and T330P, optionally further comprising T306K and / or W313F.
[0176] In some embodiments, the endonuclease domain (e.g., as described herein) of nCas9 includes, for example, an N863A mutation (e.g., in spCas9) or an H840A mutation.
[0177] In some embodiments, a heterologous sequence of interest (e.g., of a system described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, or more nucleotides in length.
[0178] In some embodiments, the RT and endonuclease domains are linked by a flexible linker comprising, for example, the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 5006).
[0179] In some embodiments, the endonuclease domain is N-terminal to the RT domain. In some embodiments, the endonuclease domain is C-terminal to the RT domain.
[0180] In some embodiments, the system incorporates a heterologous sequence of interest into a target site by TPRT, for example, as described herein.
[0181] In some embodiments, the genetically modified polypeptide comprises a DNA binding domain. In some embodiments, the genetically modified polypeptide comprises an RNA binding domain. In some embodiments, the RNA binding domain comprises the RNA binding domain of B-box protein, MS2 coat protein, dCas, or an element of the sequence in the table herein. In some embodiments, the RNA binding domain can bind to template RNA with higher affinity than standard RNA binding domains.
[0182] In some embodiments, the genetic modification system can generate an insertion at a target site of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally up to 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetic modification system can generate an insertion at a target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally up to 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the genetic modification system can generate insertions at target sites of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally up to 1, 5, 10, or 20 kilobases). In some embodiments, the genetic modification system can generate deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally up to 500, 400, 300, or 200 nucleotides). In some embodiments, the genetic modification system is capable of generating deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or 200 nucleotides (and optionally up to 500, 400, 300 or 200 nucleotides). In some embodiments, the genetic modification system can generate deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190 or 200 nucleotides (and optionally up to 500, 400, 300 or 200 nucleotides).In some embodiments, the genetic modification system can generate deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally up to 1, 5, 10, or 20 kilobases). In some embodiments, the genetic modification system can generate substitutions at the target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides. In some embodiments, the genetic modification system can generate substitutions at a target site of 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotides.
[0183] In some embodiments, the substitution is a transition mutation. In some embodiments, the substitution is a transversion mutation. In some embodiments, the substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.
[0184] In some embodiments, the insertion, deletion, substitution, or combination thereof increases or decreases expression (e.g., transcription or translation) of a gene. In some embodiments, the insertion, deletion, substitution, or combination thereof increases or decreases expression (e.g., transcription or translation) of a gene by modifying, adding, or deleting sequences in a promoter or enhancer, such as sequences that bind transcription factors. In some embodiments, the insertion, deletion, substitution, or combination thereof modifies the translation of a gene (e.g., modifies the amino acid sequence), inserts or deletes a start or stop codon, modifies or restores the translation frame of a gene. In some embodiments, the insertion, deletion, substitution, or combination thereof modifies the splicing of a gene, for example, by inserting, deleting, or modifying a splice acceptor or donor site. In some embodiments, the insertion, deletion, substitution, or combination thereof modifies the half-life of a transcript or protein. In some embodiments, the insertion, deletion, substitution, or combination thereof alters protein localization in a cell (e.g., from the cytoplasm to mitochondria, from the cytoplasm to the extracellular space (e.g., adding a secretion tag)). In some embodiments, the insertion, deletion, substitution, or combination thereof alters (e.g., improves) protein folding (e.g., to prevent accumulation of misfolded proteins). In some embodiments, the insertion, deletion, substitution, or combination thereof alters, increases, decreases the activity of a gene, e.g., a protein encoded by the gene.
[0185] Exemplary genetically modified polypeptides, and systems comprising and methods of using them, are described in PCT / US2021 / 020948, which is incorporated by reference herein, for retroviral RT domains, including, for example, amino acid and nucleic acid sequences therein.
[0186] Exemplary genetically modified polypeptide and retroviral RT domain sequences are also described, for example, in International Patent Application No. PCT / US21 / 20948, filed March 4, 2021, including Tables 30, 31, and 44 therein; the entire application is incorporated herein by reference, for example, with respect to the sequences and retroviral RTs in the tables. Thus, the genetically modified polypeptides described herein may include an amino acid sequence according to any of the tables described in this paragraph or a domain thereof (e.g., a retroviral RT domain), or any of the above functional fragments or variants thereof, or an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0187] In some embodiments, a polypeptide for use in any of the systems described herein may be a molecular reconstitution or an ancestral reconstitution based on aligned polypeptide sequences of multiple homologous proteins. In some embodiments, a reverse transcriptase domain for use in any of the systems described herein may be a molecular reconstitution or an ancestral reconstitution, or may be modified at specific residues based on the alignment of reverse transcriptase domains from the same or different sources. Those skilled in the art can align polypeptide or nucleic acid sequences based on the accession numbers provided herein, for example, by using routine sequence analysis tools such as Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis. Molecular reconstitutions can be made based on sequence consensus, for example, using the techniques described in Ivics et al., Cell 1997, 501-510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99.
[0188] Polypeptide components of genetic modification systems In some embodiments, the genetically modified polypeptide has the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing, e.g., reverse transcription. In some embodiments, each function is contained within a different domain. In some embodiments, a function can be attributed to two or more domains (e.g., two or more domains exhibit functionality together). In some embodiments, two or more domains can have the same or similar functions (e.g., two or more domains each have DNA binding functionality independently, e.g., in two different DNA sequences). In other embodiments, one or more domains can perform one or more functions, e.g., a Cas9 domain can perform both DNA binding and target site cleavage. In some embodiments, the domains are all located within a single polypeptide. In some embodiments, the first domain is present in a first polypeptide and the second domain is present in a second polypeptide. For example, in some embodiments, the sequence may be split between a first polypeptide and a second polypeptide, e.g., the first polypeptide includes a reverse transcriptase (RT) domain and the second polypeptide includes a DNA binding domain and an endonuclease domain, e.g., a nickase domain. By way of further example, in some embodiments, each of the first and second polypeptides includes a DNA binding domain (e.g., a first DNA binding domain and a second DNA binding domain). In some embodiments, the first and second polypeptides may be post-translationally linked via a split intein to form a single genetically modified polypeptide.
[0189] In some embodiments, the genetically modified polypeptides described herein (e.g., the systems described herein include genetically modified polypeptides that include: 1) a Cas domain (e.g., a Cas nickase domain, e.g., a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain of Table D or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto, wherein the RT domain is C-terminal to the Cas domain; and a linker disposed between the RT domain and the Cas domain, the linker having a sequence from the same row as the RT domain of Table D or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto.
[0190] In some embodiments, the RT domain has a sequence with 100% identity to the RT domain of Table D, and the linker has a sequence with 100% identity to the linker sequence from the same row as the RT domain of Table D. In some embodiments, the Cas domain comprises a sequence of Table 8 or a sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence according to any of SEQ ID NOs: 1-3332 in the sequence listing, or a sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto.
[0191] In some embodiments, the genetically modified polypeptide comprises a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain. In some embodiments, the genetically modified polypeptide comprises a sequence of SEQ ID NO: 4000 comprising a first NLS and a Cas domain, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises a sequence of SEQ ID NO: 4001 comprising a second NLS, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98% or 99% identity thereto.
[0192] Exemplary N-Terminal NLS-Cas9 Domains [ka]
[0193] Exemplary C-Terminal Sequences Containing an NLS AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001)
[0194] Lighting Domain (RT Domain) In certain aspects of the present invention, the writing domain of the genetic modification system has reverse transcriptase activity and is also referred to as reverse transcriptase domain (RT domain). In some embodiments, the RT domain comprises an RT catalytic portion and an RNA binding region (e.g., a region that binds to template RNA).
[0195] In some embodiments, the nucleic acid encoding the reverse transcriptase is modified from its native sequence, e.g., to have modified codon usage, improved for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a retrovirus. In some embodiments, the RT domain comprising the genetically modified polypeptide is mutated from its original amino acid sequence, e.g., has at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions. In some embodiments, the RT domain is derived from a retroviral RT, e.g., HIV-1 RT, Moloney Murine Leukemia Virus (MMLV) RT, avian myeloblastosis virus (AMV) RT, or Rous Sarcoma Virus (RSV) RT.
[0196] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits increased stringency of target primed reverse transcription (TPRT) initiation, e.g., compared to the endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3 nt within the target site immediately upstream of the first strand nick, e.g., the genomic DNA priming the RNA template, has at least 66% or 100% complementarity to the 3 nt of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when there is less than a 5 nt mismatch (e.g., less than 1, 2, 3, 4, or 5 nt mismatch) between the template RNA homology and the target DNA primed reverse transcription. In some embodiments, the RT domain is modified to increase stringency in mismatches in priming the TPRT reaction, e.g., the RT domain tolerates no mismatches or tolerates less mismatches within the priming region compared to a wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain comprises an HIV-1 RT domain. In embodiments, the HIV-1 RT domain initiates synthesis at a lower level, even with three nucleotide mismatches, compared to alternative RT domains (e.g., as described by Jamburuthugoda and Eickbush J Mol Biol 407(5):661-672 (2011), which is incorporated by reference in its entirety).
[0197] In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain is a monomer. In some embodiments, the RT domain naturally functions as a monomer or a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain naturally functions as a monomer, e.g., is derived from a virus that functions as a monomer. In embodiments, the RT domain is an antibody specific for any of the following viruses: murine leukemia viruses (MLV; also referred to as MoMLV) (e.g., P03355), porcine endogenous retroviruses (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), avian reticuloendotheliosis virus (AVIRE) (e.g., UniProtKB Accession No.: P03360), feline leukemia virus (FLV or FeLV) (e.g., UniProtKB Accession No.: P10273); Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foamy virus (HFV) (e.g., UniProt P14350), simian foamy virus (SFV) (e.g., SFV3L) (e.g., UniProt P23074 or P27401), or bovine foamy / syncytial virus (BFV / BSV) (e.g., UniProt O41894), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity). In some embodiments, the RT domain is a dimer in its native functionality. In some embodiments, the RT domain is derived from a virus that functions as a dimer.In embodiments, the RT domain is selected from the group consisting of avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastosis virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci 67(16):2717-2747(2010)), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95% or 99% identity thereto). Naturally, a heterodimeric RT domain may in some embodiments also be functional as a homodimer. In some embodiments, the dimeric RT domain is expressed as a fusion protein, e.g., as a homodimeric or heterodimeric fusion protein. In some embodiments, the RT function of the system is fulfilled by multiple RT domains (e.g., as described herein).In further embodiments, the multiple RT domains can be fused or separated, for example, on the same polypeptide or on different polypeptides.
[0198] In some embodiments, the genetic modification system described herein comprises an integrase domain, e.g., the integrase domain can be part of the RT domain. In some embodiments, the RT domain (e.g., as described herein) comprises an integrase domain. In some embodiments, the RT domain (e.g., as described herein) lacks an integrase domain or comprises an integrase domain that is inactivated by mutation or deletion. In some embodiments, the genetic modification system described herein comprises an RNase H domain, e.g., the RNase H domain can be part of the RT domain. In some embodiments, the RNase H domain is not part of the RT domain and is covalently linked via a flexible linker. In some embodiments, the RT domain (e.g., as described herein) comprises an RNase H domain, e.g., an endogenous RNase H domain or a heterologous RNase H domain. In some embodiments, the RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, the RT domain (e.g., as described herein) comprises a RNase H domain that has been added, deleted, mutated, or exchanged for a heterologous RNase H domain. In some embodiments, the polypeptide comprises an inactivated endogenous RNase H domain. In some embodiments, an endogenous RNase H domain from one of the other domains of the polypeptide is genetically removed such that it is not included in the polypeptide, e.g., the endogenous RNase H domain is partially or completely truncated from the polypeptide that includes the domain. In some embodiments, the mutation of the RNase H domain produces a polypeptide that exhibits lower RNase activity, e.g., by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% lower, compared to an otherwise similar domain that does not have the mutation, e.g., as measured by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988), which is incorporated by reference in its entirety.In some embodiments, RNase H activity is abolished.
[0199] In some embodiments, the RT domain undergoes mutations to increase fidelity compared to other similar domains that do not have mutations. For example, in some embodiments, the YADD (SEQ ID NO: 25690) or YMDD (SEQ ID NO: 25691) motif in the RT domain (e.g., in reverse transcriptase) is replaced with YVDD (SEQ ID NO: 25692). In some embodiments, the replacement of YADD (SEQ ID NO: 25690), or YMDD (SEQ ID NO: 25691), or YVDD (SEQ ID NO: 25692) results in higher fidelity in retroviral reverse transcriptase activity (e.g., as described in Jamburuthugoda and Eickbush J Mol Biol 2011; the entirety of which is incorporated herein by reference).
[0200] In some embodiments, the genetically modified polypeptides described herein comprise an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto. In some embodiments, the nucleic acids described herein encode an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98% or 99% identity thereto.
[0201] [Table 1]
[0202] [Table 2]
[0203] [Table 3]
[0204] [Table 4]
[0205]
Table 5
[0206]
Table 6
[0207]
Table 7
[0208]
Table 8
[0209]
Table 9
[0210]
Table 10
[0211]
Table 11
[0212]
Table 12
[0213]
Table 13
[0214]
Table 14
[0215] [Table 15]
[0216] In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutagenesis. In some embodiments, the reverse transcriptase domain is modified to have improved properties, such as SuperScript IV (SSIV) reverse transcriptase from MMLV RT. In some embodiments, the reverse transcriptase domain can be modified to have a lower error rate, for example, as described in WO2001068895, which is incorporated herein by reference. In some embodiments, the reverse transcriptase domain can be modified to have increased thermostability. In some embodiments, the reverse transcriptase domain can be modified to have increased processivity. In some embodiments, the reverse transcriptase domain can be modified to have resistance to inhibitors. In some embodiments, the reverse transcriptase domain can be modified to be faster. In some embodiments, the reverse transcriptase domain can be modified to have increased tolerance to modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain can be modified to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is modified to bind to the template RNA. In some embodiments, the one or more mutations are selected from D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, H8Y, T306K or D653N in the RT domain of murine leukemia virus reverse transcriptase, or a corresponding mutation at a corresponding position in another RT domain.
[0217] In some embodiments, the genetically modified polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as, for example, wild-type M-MLV RT comprising the following sequence: M-MLV(WT): [ka]
[0218] In some embodiments, the genetically modified polypeptide comprises an RT domain from a retroviral reverse transcriptase, such as, for example, M-MLV RT, which comprises the following sequence: [ka]
[0219] In some embodiments, the genetically modified polypeptide comprises an RT domain from a retroviral reverse transcriptase comprising the sequence of amino acids 659 to 1329 of NP_057933. In some embodiments, the genetically modified polypeptide further comprises one additional amino acid at the N-terminus of the sequence of amino acids 659 to 1329 of NP_057933, e.g., as shown below. [ka] [ka]
[0220] In embodiments, the genetically modified polypeptide further comprises one additional amino acid at the C-terminus of the sequence of amino acids 659 to 1329 of NP_057933. In some embodiments, the genetically modified polypeptide comprises an RNaseH1 domain (e.g., amino acids 1178 to 1318 of NP_057933).
[0221] In some embodiments, a retroviral reverse transcriptase domain, e.g., M-MLV RT, can contain one or more mutations from the wild-type sequence that can improve the characteristics of RT, e.g., thermostability, processivity, and / or template binding. In some embodiments, the M-MLV RT domain comprises a combination of mutations, such as, for example, D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, K103L, relative to the M-MLV(WT) sequence above, such as D200N, L603W, and T330P, optionally further comprising T306K and W313F. In some embodiments, the M-MLV RT as used herein comprises the mutations D200N, L603W, T330P, T306K, and W313F. In embodiments, the mutant M-MLV RT comprises the following amino acid sequence: M-MLV(PE2): [ka]
[0222] In some embodiments, the writing domain (e.g., the RT domain) comprises an RNA binding domain that specifically binds, for example, to an RNA sequence. In some embodiments, the template RNA comprises an RNA sequence that is specifically bound by the RNA binding domain of the writing domain.
[0223] In some embodiments, the reverse transcription domain simply recognizes and reverse transcribes a particular template of the system, for example, a template RNA. In some embodiments, the template comprises a sequence or structure that allows recognition and reverse transcription by the reverse transcription domain. In some embodiments, the template comprises a sequence or structure that allows association with the RNA binding domain of a polypeptide component of the genome modification system described herein. In some embodiments, the genome modification system preferably reverse transcribes a template that includes an association sequence over a template that lacks the association sequence.
[0224] The writing domain may also include a DNA-dependent DNA polymerase activity, for example an enzyme activity capable of writing DNA into a genome from a template DNA sequence. In some embodiments, DNA-dependent DNA polymerization is used to complete the second strand synthesis of the target site editing. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in the polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, for example, second strand synthesis. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a second polypeptide of the system. In some embodiments, the DNA-dependent DNA polymerase activity is provided by an endogenous host cell polymerase that is optionally recruited to the target site by a component of the genome modification system.
[0225] In some embodiments, the reverse transcriptase domain exhibits a lower probability of poor termination (P) in vitro compared to a reference reverse transcriptase domain. off In some embodiments, the reference reverse transcriptase domain is a viral reverse transcriptase domain, for example, the RT domain from M-MLV.
[0226] In some embodiments, the reverse transcriptase domain has an in vitro transcription factor of about 5×10, e.g., as measured on a 1094 nt RNA. -3 / nt, 5×10 -4 / nt, or 5×10 -6 A lower probability of insufficient termination (P off In some embodiments, insufficient termination rates in vitro are determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845, which is incorporated herein by reference in its entirety.
[0227] In some embodiments, the reverse transcriptase domain can complete at least about 30% or 50% of integration in the cell. The percentage of complete integration can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites that contain at least 98% of the expected integration sequence) by the number of total integration events (including substantially full-length and partial) in the by cell population. In some embodiments, integration in the cell is determined (e.g., through the integration site) using long-read amplicon sequencing, for example, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903, which is incorporated herein by reference in its entirety.
[0228] In embodiments, quantifying integration in a cell comprises counting the percentage of integrations that comprise at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% or 100% of the DNA sequence corresponding to the template RNA (e.g., a template RNA having a length of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 3, 4 or 5 kb, e.g., 0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, 1.0-1.2, 1.2-1.4, 1.4-1.6, 1.6-1.8, 1.8-2.0, 2-3, 3-4 or 4-5 kb).
[0229] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at a rate of 0.1-50 nt / sec (e.g., 0.1-1, 1-10, or 10-50 nt / sec). In some embodiments, polymerization of dNTPs by the reverse transcriptase domain is measured by a single molecule assay, e.g., as described in Schwartz and Quake (2009) PNAS 106(48):20294-20299, which is incorporated by reference in its entirety.
[0230] In some embodiments, the reverse transcriptase domain is at least 1×10 nucleotides, e.g., as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153, which is incorporated by reference in its entirety. -3 ~1×10 -4 Or 1×10 -4 ~1×10 -5 In some embodiments, the reverse transcriptase domain has an in vitro error rate (e.g., nucleotide misincorporation) of 1×10 substitutions / nt. In some embodiments, the reverse transcriptase domain is sequenced at 1×10 in cells (e.g., HEK293T cells) by, e.g., long-read amplicon sequencing, e.g., as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903, which is incorporated by reference in its entirety. -3 ~1×10 -4 Or 1×10 -4 ~1×10 -5 It has an error rate (e.g., nucleotide misincorporation) of .
[0231] In some embodiments, the reverse transcriptase domain is capable of performing reverse transcription of target RNA in vitro. In some embodiments, the reverse transcriptase requires a primer of at least 3 nucleotides to initiate reverse transcription of the template. In some embodiments, reverse transcription of target RNA is determined by detection of cDNA from target RNA (e.g., when a ssDNA primer is provided that anneals to the target with at least 3, 4, 5, 6, 7, 8, 9 or 10 nt at the 3' end), for example, as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845, which is incorporated herein by reference in its entirety.
[0232] In some embodiments, the reverse transcriptase domain performs reverse transcription (e.g., by generating cDNA) at least 5-fold or 10-fold more efficiently, e.g., when converting the RNA template to cDNA, as compared to an RNA template lacking a protein binding motif (e.g., the 3'UTR). In embodiments, the efficiency of reverse transcription is measured as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153, which is incorporated herein by reference in its entirety.
[0233] In some embodiments, the reverse transcriptase domain specifically binds to a particular RNA template, e.g., at a higher frequency (e.g., about 5-fold or 10-fold higher frequency) than any endogenous cellular RNA when expressed in a cell (e.g., HEK293T cell). In some embodiments, the frequency of specific binding between the reverse transcriptase domain and the template RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501, which is incorporated by reference in its entirety.
[0234] Template nucleic acid binding domain Genetically modified polypeptides typically have a region that can associate with template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain that can associate with RNA molecules that contain specific signatures, such as structural motifs. In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within a reverse transcription domain, for example, a component derived from reverse transcriptase has a known signature for RNA preference.
[0235] In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within the target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-associated protein that recognizes the structure of the template nucleic acid (e.g., template RNA) including the gRNA. In some embodiments, the genetically modified polypeptide comprises a DNA binding domain that comprises a CRISPR-associated protein that associates with a gRNA scaffold, allowing the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the gRNA scaffold and gRNA spacer are contained within the template nucleic acid (e.g., template RNA), such that the DNA binding domain is also a template nucleic acid binding domain. In some embodiments, the polypeptide has RNA binding functions within multiple domains, for example, it may bind to a gRNA structure within the CRISPR-associated DNA binding domain and additional sequences or structures within the reverse transcriptase domain.
[0236] In some embodiments, the RNA binding domain can bind to the template RNA with higher affinity than a standard RNA binding domain. In some embodiments, the standard RNA binding domain is an RNA binding domain from Cas9 of S. pyogenes. In some embodiments, the RNA binding domain can bind to the template RNA with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the RNA binding domain to its template RNA is measured in vitro, e.g., by thermophoresis, e.g., as described in Asmari et al. Methods 146:107-119 (2018), which is incorporated herein by reference in its entirety. In some embodiments, the affinity of the RNA binding domain to its template RNA is measured in a cell (e.g., by FRET or CLIP-Seq).
[0237] In some embodiments, the RNA binding domain associates with the template RNA at least about 5-fold or 10-fold more frequently than scrambled RNA in vitro. In some embodiments, the frequency of association between the RNA binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501, which is incorporated herein by reference in its entirety. In some embodiments, the RNA binding domain associates with the template RNA in a cell (e.g., in a HEK293T cell) at least about 5-fold or 10-fold more frequently than scrambled RNA. In some embodiments, the frequency of association between the RNA binding domain and the template RNA or scrambled RNA is measured by CLIP-seq, e.g., as described in Lin and Miles (2019) supra.
[0238] In some embodiments, an RT domain (e.g., one listed in Table 6) comprises one or more mutations listed in Table 2A below. In some embodiments, an RT domain listed in Table 6 comprises one, two, three, four, five, or six of the mutations listed in the corresponding row of Table 2A below.
[0239] [Table 16]
[0240] [Table 17]
[0241] [Table 18]
[0242] [Table 19]
[0243] [Table 20]
[0244] Endonuclease domain and DNA binding domain In some embodiments, the genetically modified polypeptide has the function of DNA target site cleavage via an endonuclease domain. In some embodiments, the genetically modified polypeptide comprises, for example, a DNA binding domain for binding to a target nucleic acid. In some embodiments, the domain (e.g., Cas domain) of the genetically modified polypeptide comprises two or more smaller domains, for example, a DNA binding domain and an endonuclease domain. When a DNA binding domain (e.g., Cas domain) is described as binding to a target nucleic acid sequence, it is understood that in some embodiments, the binding is mediated by a gRNA.
[0245] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA) binding domain. For example, in some embodiments, the polypeptide comprises a CRISPR-associated endonuclease domain that binds to template RNA, including gRNA, binds to a target DNA sequence (e.g., has complementarity to a portion of gRNA), and cleaves the target DNA sequence. In some embodiments, the endonuclease domain or endonuclease / DNA binding domain from a heterologous source can be used or modified (e.g., by inserting, deleting, or substituting one or more residues) in the genetic modification system described herein.
[0246] In some embodiments, the nucleic acid encoding the endonuclease domain or endonuclease / DNA binding domain is modified from its native sequence to have altered codon usage, e.g., improved for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9), a type II restriction endonuclease (e.g., Fok1), a meganuclease (e.g., I-SceI), or other endonuclease domain.
[0247] In certain aspects, the DNA binding domain of the genetically modified polypeptide described herein is selected, designed, or constructed for binding to the desired host DNA target sequence. In certain embodiments, the DNA binding domain of the polypeptide is a heterologous DNA binding factor. In some embodiments, the heterologous DNA binding factor is a zinc finger factor or a TAL effector factor, such as a zinc finger or a TAL polypeptide or a functional fragment thereof. In some embodiments, the heterologous DNA binding factor is a sequence-induced DNA binding factor, such as Cas9, Cpf1, or other CRISPR-associated proteins, that has been modified to have no endonuclease activity. In some embodiments, the heterologous DNA binding factor retains endonuclease activity. In some embodiments, the heterologous DNA binding factor retains partial endonuclease activity, such as cleaving ssDNA, e.g., has nickase activity. In certain embodiments, the heterologous DNA binding domain can be any one or more of Cas9, TAL domain, ZF domain, Myb domain, combinations thereof, or complexes thereof.
[0248] In some embodiments, the DNA-binding domain is modified, e.g., by site-directed mutagenesis, to increase or decrease DNA-binding factors (e.g., the number and / or specificity of zinc fingers), etc., to alter DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is modified from its native sequence to have altered codon usage, e.g., improved for human cells. In several embodiments, the DNA-binding domain includes one or more modifications relative to the wild-type DNA-binding domain, e.g., modifications by directed evolution, e.g., phage-assisted continuous evolution (PACE).
[0249] In some embodiments, the DNA binding domain comprises a meganuclease domain (e.g., in an endonuclease domain portion, e.g., as described herein), or a functional fragment thereof. In some embodiments, the meganuclease domain has endonuclease activity, e.g., double-strand cleavage and / or nickase activity. In other embodiments, the meganuclease domain has reduced activity, e.g., lacks endonuclease activity, e.g., the meganuclease is catalytically inactive. In some embodiments, a catalytically inactive meganuclease is used as a DNA binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012), which is incorporated herein by reference in its entirety.
[0250] In some embodiments, the genetically modified polypeptide comprises modifications to the DNA binding domain, for example, compared to the wild-type polypeptide. In some embodiments, the DNA binding domain comprises additions, deletions, substitutions, or modifications to the amino acid sequence of the original DNA binding domain. In some embodiments, the DNA binding domain is modified to comprise a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain replaces at least a portion (e.g., the entirety) of a previous DNA binding domain of the polypeptide. In some embodiments, the functional domain comprises a zinc finger (e.g., a zinc finger that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain comprises a Cas domain (e.g., a Cas domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the Cas domain comprises Cas9 or a mutant or variant thereof (e.g., as described herein). In some embodiments, the Cas domain is associated with a guide RNA (gRNA), e.g., as described herein. In some embodiments, the Cas domain is guided to the target nucleic acid (e.g., DNA) sequence of interest by the gRNA. In some embodiments, the Cas domain is encoded in the same nucleic acid (e.g., RNA) molecule as the gRNA. In some embodiments, the Cas domain is encoded in a different nucleic acid (e.g., RNA) molecule than the gRNA.
[0251] In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with higher affinity than a standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain from Cas9 of S. pyogenes. In some embodiments, the DNA-binding domain can bind to a target sequence (e.g., a dsDNA target sequence) with an affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).
[0252] In some embodiments, the affinity of a DNA binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, e.g., by thermophoresis, e.g., as described in Asmari et al. Methods 146:107-119 (2018), which is incorporated by reference in its entirety.
[0253] In embodiments, the DNA binding domain can bind to its target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM) in the presence of a molar excess, e.g., about a 100-fold molar excess, of scrambled sequence competitor dsDNA.
[0254] In some embodiments, the DNA binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured, e.g., by ChIP-seq (e.g., in HEK293T cells), e.g., as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, which is incorporated by reference in its entirety. In some embodiments, the DNA binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at least about 5-fold or 10-fold higher frequency than any other sequence in the genome of the target cell, as measured, e.g., by ChIP-seq (e.g., in HEK293T), e.g., as described in He and Pu (2010) supra.
[0255] In some embodiments, the endonuclease domain has a nickase activity and cleaves one strand of the target DNA. In some embodiments, the nickase activity reduces the formation of double-stranded breaks at the target site. In some embodiments, the endonuclease domain generates a staggered nick structure in the first and second strands of the target DNA. In some embodiments, the staggered nick structure generates a free 3' overhang at the target site. In some embodiments, the free 3' overhang at the target site improves editing efficiency, for example, by enhancing access and annealing of the 3' homology region of the template nucleic acid. In some embodiments, the staggered nick structure reduces the formation of double-stranded breaks at the target site.
[0256] In some embodiments, the endonuclease domain cleaves both strands of the target DNA, for example resulting in a blunt end cleavage of the target with no ssDNA overhangs on either side of the cleavage site. The amino acid sequence of the endonuclease domain of the genetic modification system described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of the endonuclease domain described herein, for example, the endonuclease domain from Table 8.
[0257] In certain embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonuclease is a Holliday junction resolvase or a homolog thereof, such as the Holliday junction cleavage enzyme (Ssol Hje) from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonuclease is a large fragment endonuclease of a spliceosomal protein, such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). In certain embodiments, the heterologous endonuclease is derived from a CRISPR-associated protein, such as Cas9. In certain embodiments, the heterologous endonuclease is modified to have only ssDNA cleavage activity, e.g., only nickase activity, e.g., a Cas9 nickase, e.g., SpCas9 with D10A, H840A, or N863A mutations. Table 8 shows exemplary Cas proteins and mutations associated with nickase activity. In yet other embodiments, the homologous endonuclease domain is modified, e.g., by site-directed mutagenesis, to modify DNA endonuclease activity. In yet other embodiments, the endonuclease domain is modified to reduce DNA sequence specificity, e.g., by truncation to remove a domain that confers DNA sequence specificity or a mutation to inactivate a region that confers DNA sequence specificity.
[0258] In some embodiments, the endonuclease domain has nickase activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks more frequently than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98% or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2% or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease does not substantially form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.
[0259] In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the first strand; for example, in some embodiments, the endonuclease domain cuts the genomic DNA of the target site near the modification site on the strand that will be extended by the writing domain. In some embodiments, the endonuclease domain has a nickase activity that nicks the target site DNA of the first strand and does not nicks the target site DNA of the second strand. For example, when the polypeptide comprises a CRISPR-associated endonuclease domain with nickase activity, in some embodiments, the CRISPR-associated endonuclease domain nicks the target site DNA strand that contains the PAM site (e.g., does not nicks the target site DNA strand that does not contain the PAM site). As a further example, when a polypeptide comprises a CRISPR-associated endonuclease domain with nickase activity, in some embodiments, the CRISPR-associated endonuclease domain nicks the target site DNA strand that does not contain a PAM site (e.g., and does not nick the target site DNA strand that includes a PAM site).
[0260] In some other embodiments, the endonuclease domain has a nickase activity that makes a nick in the target site DNA of the first strand and the second strand. Without intending to be bound by a particular theory, after the writing domain (e.g., RT domain) of the polypeptide described herein polymerizes (e.g., reverse transcribes) from a heterologous target sequence of a template nucleic acid (e.g., a template RNA), the cellular DNA repair machinery must repair the nick on the first DNA strand. The target site DNA now comprises two different sequences for the first DNA strand: the first corresponds to the original genomic DNA (e.g., with a free 5' end) and the second corresponds to that polymerized from the heterologous target sequence (e.g., with a free 3' end). It is believed that the two different sequences equilibrate with each other, first one and then the other hybridize with the second strand, where the order of incorporation of the cellular DNA repair machinery into its repair target site is a stochastic process. Without intending to be bound by any particular theory, it is believed that the introduction of an additional nick into the second strand may bias the cellular DNA repair machinery to use sequences based on the heterologous target sequence more frequently than the original genomic sequence (Anzalone et al. Nature 576:149-157 (2019)). In some embodiments, the additional nick is positioned at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution) or to the nick on the first strand.
[0261] Alternatively or additionally, without intending to be bound by any particular theory, it is believed that an additional nick to the second strand may facilitate second strand synthesis. In some embodiments, if the genetic modification system inserts or replaces a portion of the first strand, synthesis of a new sequence corresponding to the insertion / substitution in the second strand is required.
[0262] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain), which nicks both the first strand and the second strand. For example, in such embodiments, the endonuclease domain can be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs nicking of the first strand and a further gRNA spacer that directs nicking of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, where a first endonuclease domain nicks the first strand and a second endonuclease domain nicks the second strand (optionally, the first endonuclease domain does not (e.g., cannot) nick the second strand and the second endonuclease domain does not (e.g., cannot) nick the first strand).
[0263] In some embodiments, the endonuclease domain can nick the first and second strands. In some embodiments, the first and second strand nicks occur at the same position in the target site, but not on opposite strands. In some embodiments, the second strand nick occurs at a staggered position, e.g., upstream or downstream from the first nick. In some embodiments, the endonuclease domain generates a deletion of the target site when the second strand nick is upstream of the first strand nick. In some embodiments, the endonuclease domain generates a duplication of the target site when the second strand nick is downstream of the first strand nick. In some embodiments, the endonuclease domain does not generate a duplication and / or deletion when the first and second strand nicks occur at the same position in the target site. In some embodiments, the endonuclease domain has altered activity depending on the protein conformation or RNA binding state, for example, to promote first strand or second strand nicking (e.g., as described in Christensen et al. PNAS 2006; incorporated herein by reference in its entirety).
[0264] In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a homing endonuclease, or a functional fragment thereof. In some embodiments, the endonuclease domain comprises a meganuclease from the LAGLIDADG (SEQ ID NO: 25693), GIY-YIG, HNH, His-Cys Box, or PD-(D / E)XK family, e.g., having conserved amino acid motifs, e.g., as indicated in the family name, or a functional fragment or variant thereof. In some embodiments, the endonuclease domain comprises a meganuclease selected from, for example, I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI (Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6), or a fragment thereof. In some embodiments, the meganuclease is naturally monomeric, e.g., I-SceI, I-TevI, or dimeric, e.g., I-CreI, in its functional form. For example, LAGLIDADG (SEQ ID NO: 25693) meganucleases with a single copy of the LAGLIDADG motif (SEQ ID NO: 25693) generally form homodimers, while members with two copies of the LAGLIDADG (SEQ ID NO: 25693) motif are generally found as monomers. In some embodiments, meganucleases that normally form as dimers are expressed as fusions, e.g., two subunits are expressed as a single ORF, optionally linked by a linker, e.g., an I-CreI dimer fusion (Rodriguez-Fornes et al. Gene Therapy 2020; incorporated herein by reference in its entirety).In some embodiments, the meganuclease, or a functional fragment thereof, is engineered to preferentially have nickase activity in one strand of a double-stranded DNA molecule, e.g., I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, the meganuclease or a functional fragment thereof with this preference for single-strand cleavage is used, e.g., as an endonuclease domain with nickase activity. In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof, that naturally targets or is engineered to target a safe harbor site, e.g., an SH6 site that targets I-CreI (Rodriguez-Fornes et al., supra). In some embodiments, the endonuclease domain comprises a meganuclease, or a functional fragment thereof, that has a sequence-tolerant catalytic domain, e.g., I-TevI, which recognizes the minimal motif CNNNG (Kleinstiver et al. PNAS 2012). In some embodiments, the target sequence-resistant catalytic domain is fused to a DNA binding domain, e.g., activity is induced by fusing I-TevI to (i) Zn fingers to create Tev-ZFEs (Kleinstiver et al. PNAS 2012), (ii) other meganucleases to create MegaTev (Wolfs et al. Nucleic Acids Res 2014), and / or (iii) Cas9 to create TevCas9 (Wolfs et al. PNAS 2016).
[0265] In some embodiments, the endonuclease domain comprises a restriction enzyme, e.g., a Type IIS or Type IIP restriction enzyme. In some embodiments, the endonuclease domain comprises a Type IIS restriction enzyme, e.g., FokI, or a fragment or variant thereof. In some embodiments, the endonuclease domain comprises a Type IIP restriction enzyme, e.g., PvuII, or a fragment or variant thereof. In some embodiments, the dimeric restriction enzyme is expressed as a fusion, e.g., a FokI dimer fusion, such that it functions as a single strand (Minczuk et al. Nucleic Acids Res 36(12):3926-3938(2008)).
[0266] The use of additional endonuclease domains is described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565 (2017), which is incorporated by reference in its entirety.
[0267] In some embodiments, the genetically modified polypeptide comprises a modification to the endonuclease domain, e.g., compared to a wild-type Cas protein. In some embodiments, the endonuclease domain comprises an addition, deletion, substitution, or modification to the amino acid sequence of a wild-type Cas protein. In some embodiments, the endonuclease domain is modified to comprise a heterologous functional domain that specifically binds to and / or directs endonucleolytic cleavage of a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the endonuclease domain comprises a zinc finger. In some embodiments, the endonuclease domain, including a Cas domain, associates with a guide RNA (gRNA), e.g., as described herein. In some embodiments, the endonuclease domain is modified to comprise a functional domain that does not target a specific target nucleic acid (e.g., DNA) sequence. In some embodiments, the endonuclease domain comprises a Fok1 domain.
[0268] In some embodiments, the endonuclease domain associates with the target dsDNA at least about 5-fold or 10-fold more frequently than scrambled dsDNA in vitro. In some embodiments, the endonuclease domain associates with the target dsDNA at least about 5-fold or 10-fold more frequently than scrambled dsDNA in vitro, for example, in a cell (e.g., HEK293T cell). In some embodiments, the frequency of association between the endonuclease domain and the target DNA or scrambled DNA is measured by ChIP-seq, for example, as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21, which is incorporated herein by reference in its entirety.
[0269] In some embodiments, the endonuclease domain may catalyze the formation of nicks at the target sequence, e.g., by at least about a 5-fold or 10-fold increase, relative to non-target sequences (e.g., relative to any other genomic sequences in the genome of the target cell). In some embodiments, the level of nicking is measured using NickSeq, e.g., as described in Elacqua et al. (2019) bioRxiv doi.org / 10.1101 / 867937, which is incorporated by reference in its entirety.
[0270] In some embodiments, the endonuclease domain is capable of nicking DNA in vitro. In some embodiments, the nick results in an exposed base. In some embodiments, the exposed base can be detected using a nuclease sensitivity assay, for example, as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res 23(19):3805-3809, which is incorporated by reference in its entirety. In some embodiments, the level of exposed base (e.g., as detected by a nuclease sensitivity assay) is increased by at least 10%, 50% or more compared to the reference endonuclease domain. In some embodiments, the reference endonuclease domain is an endonuclease domain from Cas9 of S. pyogenes.
[0271] In some embodiments, the endonuclease domain is capable of nicking DNA in a cell. In some embodiments, the endonuclease domain is capable of nicking DNA in a HEK293T cell. In some embodiments, an unrepaired nick that undergoes replication in the absence of Rad51 results in an increased rate of NHEJ at the site of the nick, detectable, for example, by using a Rad51 inhibition assay, e.g., as described in Bothmer et al. (2017) Nat Commun 8:13905, which is incorporated by reference in its entirety. In some embodiments, the rate of NHEJ is increased by more than 0-5%. In some embodiments, the rate of NHEJ is increased, for example, by 20-70% (e.g., 30%-60% or 40-50%) upon Rad51 inhibition.
[0272] In some embodiments, the endonuclease domain releases the target after cleavage. In some embodiments, release of the target is indicated indirectly by assessing multiple turnover by the enzyme, for example, as described in Yourik at al. RNA 25(1):35-44 (2019), which is incorporated by reference in its entirety, and as shown in FIG. 2. In some embodiments, the k of the endonuclease domain is exp is measured by this method and is 1×10 -3 ~1×10 -5 The time is minus 1.
[0273] In some embodiments, the endonuclease domain is capable of binding to about 1×10 8 s -1 M -1 Catalytic efficiency (k cat / K m In some embodiments, the endonuclease domain has an in vitro affinity of about 1×10 5 , 1×10 6 , 1×10 7 , or 1 × 10 8 s -1 M -1 In some embodiments, the catalytic efficiency is determined as described in Chen et al. (2018) Science 360(6387):436-439, which is incorporated by reference in its entirety. In some embodiments, the endonuclease domain has a catalytic efficiency of greater than about 1×10 8 s -1 M -1 Catalytic efficiency (k cat / K m In some embodiments, the endonuclease domain has a concentration of about 1×10 5 , 1×10 6 , 1×10 7 , or 1 × 10 8 s -1 M -1 The catalytic efficiency is greater than 1.
[0274] Genetically engineered polypeptides containing a Cas domain In some embodiments, the genetically modified polypeptide described herein comprises a Cas domain. In some embodiments, the Cas domain can guide the genetically modified polypeptide to the target site specified by the gRNA spacer, thereby modifying the target nucleic acid sequence in "cis". In some embodiments, the genetically modified polypeptide is fused to a Cas domain. In some embodiments, the genetically modified polypeptide comprises a CRISPR / Cas domain (also referred to herein as CRISPR-associated protein). In some embodiments, the CRISPR / Cas domain comprises a protein involved in clustered regularly interspaced short palindromic repeats (CRISPR) system, such as a Cas protein, and optionally binds to a guide RNA, such as a single guide RNA (sgRNA).
[0275] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases, termed CRISPR-associated or "Cas" endonucleases (e.g., Cas9 or Cpf1), to cleave foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is guided to a target nucleotide sequence (e.g., a site in a genome to be sequence-edited) by a sequence-specific non-coding "guide RNA" that targets a single-stranded or double-stranded DNA sequence. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One Class II CRISPR system includes a type II Cas endonuclease, such as Cas9, a CRISPR RNA ("crRNA"), and a trans-activating crRNA ("tracrRNA"). The crRNA typically contains a "spacer" sequence (protospacer), which is an RNA sequence of about 20 nucleotides that corresponds to the target DNA sequence. In wild-type systems, and in some modified systems, the crRNA binds to the tracrRNA and forms a partially double-stranded structure that is cleaved by RNase III.cIt also contains a region that results in a rRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then guides the Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is generally adjacent to a "protospacer adjacent motif" ("PAM") that is specific for a given Cas endonuclease and required for cleavage activity at the target site that matches the spacer of the crRNA. CRISPR endonucleases identified from various prokaryotic species have unique PAM sequence requirements, for example, as listed for the exemplary Cas enzymes in Table 7; example PAM sequences include 5'-NGG (Streptococcus pyogenes), 5'-NNAGAA (Streptococcus thermophilus CRISPR1), 5'-NGGNG (Streptococcus thermophilus CRISPR3), and 5'-NNNGATT (Neisseria meningiditis). Some endonucleases, for example, the Cas9 endonuclease, associate with a G-rich PAM site, e.g., 5'-NGG, and perform a blunt-end cut of the target DNA three nucleotides upstream (5') from the PAM site. Another class II CRISPR system includes a V-type endonuclease Cpf1, which is smaller than Cas9; examples include AsCpf1 (from Acidaminococcus sp.) and LbCpf1 (from Lachnospiraceae sp.). Cpf1-associated CRISPR arrays do not require tracrRNA and are processed into mature crRNA; in other words, the Cpf1 system, in some embodiments, exclusively includes Cpf1 nuclease and crRNA to cleave the target DNA sequence. Cpf1 endonuclease typically associates with T-rich PAM sites, such as 5'-TTN. Cpf1 can also recognize the 5'-CTA PAM motif.Cpf1 typically cleaves target DNA by introducing offset or staggered double-stranded breaks into 4- or 5-nucleotide 5' overhangs, e.g., cleaving the target DNA with a 5-nucleotide offset or staggered break located 18 nucleotides downstream (3') from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand; the 5-nucleotide overhangs resulting from such offset breaks allow for more precise genome editing by DNA insertion via homologous recombination rather than insertion with blunt-end cut DNA. See, e.g., Zetsche et al. (2015) Cell, 163:759-771.
[0276] A variety of CRISPR-associated (Cas) genes or proteins can be used in the technology provided by the present disclosure, and the choice of Cas protein will depend on the specific conditions of the method. Specific examples of Cas proteins include class II systems including Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein, e.g., Cas9 protein, can be derived from any of a variety of prokaryotic species. In some embodiments, a particular Cas protein, e.g., a particular Cas9 protein, is selected to recognize a particular protospacer adjacent motif (PAM) sequence. In some embodiments, the DNA binding domain or endonuclease domain comprises a sequence-targeting polypeptide, such as a Cas protein, e.g., Cas9. In certain embodiments, the Cas protein, e.g., Cas9 protein, can be obtained from bacteria or archaea, or can be synthesized using known methods. In certain embodiments, the Cas protein can be derived from gram-positive or gram-negative bacteria. In certain embodiments, the Cas protein is selected from the group consisting of Streptococcus (e.g., S. pyogenes, or S. thermophilus), Francisella (e.g., F. novicida), Staphylococcus (e.g., S. aureus), Acidaminococcus (e.g., Acidaminococcus sp. BV3L6), sp. BV3L6), Neisseria (e.g., N. meningitidis), Cryptococcus, Corynebacterium, Haemophilus, Eubacterium, Pasteurella, Prevotella, Veillonella, or Marinobacter.
[0277] In some embodiments, the genetically modified polypeptide may comprise the amino acid sequence of SEQ ID NO: 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto. In some embodiments, the amino acid sequence of SEQ ID NO: 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto is located at the N-terminus of the genetically modified polypeptide. In several embodiments, the amino acid sequence of SEQ ID NO: 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids of the N-terminus of the genetically modified polypeptide.
[0278] Exemplary N-Terminal NLS-Cas9 Domains [ka]
[0279] In some embodiments, the genetically modified polypeptide may comprise the amino acid sequence of SEQ ID NO: 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto. In some embodiments, the amino acid sequence of SEQ ID NO: 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto is located at the C-terminus of the genetically modified polypeptide. In several embodiments, the amino acid sequence of SEQ ID NO: 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids of the C-terminus of the genetically modified polypeptide.
[0280] Exemplary C-Terminal Sequences Containing an NLS AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001)
[0281] Example benchmark sequence [ka] [ka] [ka]
[0282] In some embodiments, the genetically modified polypeptide may comprise a Cas domain listed in Tables 7 or 8, or a functional fragment thereof, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99% identity thereto.
[0283] [Table 21]
[0284] [Table 22]
[0285] [Table 23]
[0286] [Table 24]
[0287] [Table 25]
[0288] [Table 26]
[0289]
Table 27
[0290]
Table 28
[0291]
Table 29
[0292]
Table 30
[0293]
Table 31
[0294]
Table 32
[0295]
Table 33
[0296]
Table 34
[0297]
Table 35
[0298]
Table 36
[0299] [Table 37]
[0300] [Table 38]
[0301] In some embodiments, the Cas protein requires that a protospacer adjacent motif (PAM) is present in or adjacent to the target DNA sequence to which the Cas protein binds and / or functions. In some embodiments, the PAM is or includes, from 5' to 3', NGG, YG, NNGRRT, NNNRRT, NGA, TYCV, TATV, NTTN, or NNNGATT, where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A or C or G. In some embodiments, the Cas protein is a protein listed in Table 7 or 8. In some embodiments, the Cas protein comprises one or more mutations that modify its PAM. In some embodiments, the Cas protein comprises the mutations E1369R, E1449H, and R1556A or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises E782K, N968K, and R1015H mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises D1135V, R1335Q, and T1337R mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises S542R and K607R mutations or analogous substitutions for the amino acids corresponding to said positions. In some embodiments, the Cas protein comprises S542R, K548V, and N552R mutations or analogous substitutions for the amino acids corresponding to said positions. Exemplary advances in the modification of Cas enzymes to recognize modified PAM sequences are reviewed in Collias et al Nature Communications 12:555 (2021), the entirety of which is incorporated herein by reference.
[0302] In some embodiments, the Cas protein is catalytically active and cleaves one or both strands of the target DNA site, in some embodiments, following cleavage of the target DNA site, an alteration, e.g., an insertion or deletion, is formed, e.g., by a cellular repair mechanism.
[0303] In some embodiments, the Cas protein is modified to inactivate or partially inactivate the nuclease, e.g., nuclease-deficient Cas9. While wild-type Cas9 generates double-strand breaks (DSBs) at the specific DNA sequence targeted by the gRNA, several CRISPR endonucleases with modified functionality are available, e.g., a partially inactivated "nickase" version of Cas9 generates only single-strand breaks; catalytically inactive Cas9 ("dCas9") does not cleave the target DNA. In some embodiments, binding of dCas9 to a DNA sequence may interfere with transcription at that site by steric hindrance. In some embodiments, binding of dCas9 to an anchor sequence may interfere with (e.g., reduce or prevent) the formation and / or maintenance of a genomic complex (e.g., ASMC). In some embodiments, the DNA binding domain comprises a catalytically inactive Cas9, e.g., dCas9. Many catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 comprises a mutation, e.g., a D10A and an H840A or an N863A mutation, within each endonuclease domain of the Cas protein. In some embodiments, the catalytically inactive or partially inactive CRISPR / Cas domain comprises a Cas protein comprising one or more mutations, e.g., one or more of the mutations listed in Table 7. In some embodiments, a Cas protein listed in a given row of Table 7 comprises one, two, three, or all of the mutations listed in the same row of Table 7. In some embodiments, for example, a Cas protein not listed in Table 7 comprises one, two, three, or all of the mutations listed in a row of Table 7, or a corresponding mutation at a corresponding site in the Cas protein.
[0304] In some embodiments, the catalytically inactive, e.g., dCas9, or partially inactivated Cas9 protein comprises a D11 mutation (e.g., a D11A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive, e.g., dCas9, or partially inactivated Cas9 protein comprises a H969 mutation (e.g., a H969A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive, e.g., dCas9, or partially inactivated Cas9 protein comprises a N995 mutation (e.g., a N995A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive, e.g., dCas9, protein comprises a mutation at one, two, or three of positions D11, H969, and N995 (e.g., a D11A, H969A, and N995A mutation) or an analogous substitution for the amino acid corresponding to said position.
[0305] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D10 mutation (e.g., a D10A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a H557 mutation (e.g., a H557A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, comprises a D10 mutation (e.g., a D10A mutation) and a H557 mutation (e.g., a H557A mutation) or an analogous substitution for the amino acid corresponding to said position.
[0306] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D839 mutation (e.g., a D839A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a H840 mutation (e.g., a H840A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a N863 mutation (e.g., a N863A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), an H840 mutation (e.g., H840A), and an N863 mutation (e.g., N863A) or an analogous substitution for the amino acid corresponding to said positions.
[0307] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially deactivated Cas9 protein, comprises an E993 mutation (e.g., an E993A mutation) or an analogous substitution for the amino acid corresponding to said position.
[0308] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D917 mutation (e.g., a D917A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises an E1006 mutation (e.g., an E1006A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D1255 mutation (e.g., a D1255A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, comprises a D917 mutation (e.g., D917A), an E1006 mutation (e.g., E1006A), and a D1255 mutation (e.g., D1255A) or an analogous substitution for the amino acid corresponding to said position.
[0309] In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D16 mutation (e.g., a D16A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a D587 mutation (e.g., a D587A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, the partially inactivated Cas domain has nickase activity. In some embodiments, the partially inactivated Cas9 domain is a Cas9 nickase domain. In some embodiments, the catalytically inactive Cas domain or inactive Cas domain does not form a detectable double-stranded break. In some embodiments, the catalytically inactive Cas9 protein, e.g., dCas9, or partially inactivated Cas9 protein, comprises a H588 mutation (e.g., a H588A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein, comprises an N611 mutation (e.g., an N611A mutation) or an analogous substitution for the amino acid corresponding to said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, comprises a D16 mutation (e.g., D16A), a D587 mutation (e.g., D587A), an H588 mutation (e.g., H588A), and an N611 mutation (e.g., N611A) or an analogous substitution for the amino acid corresponding to said position.
[0310] In some embodiments, the DNA binding domain or endonuclease domain may comprise a Cas molecule that includes or is linked (e.g., covalently) to a gRNA (e.g., a template nucleic acid that includes a gRNA, e.g., a template RNA).
[0311] In some embodiments, the endonuclease domain or DNA binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a modified SpCas9. In some embodiments, the modified SpCas9 comprises a modification that alters the protospacer adjacent motif (PAM) specificity. In some embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In some embodiments, the modified SpCas9 comprises one or more amino acid substitutions, e.g., at one or more of the following positions: L1111, D1135, G1218, E1219, A1322, or R1335, e.g., selected from the following: L1111R, D1135V, G1218R, E1219F, A1322R, R1335V. In some embodiments, the modified SpCas9 comprises the amino acid substitution T1337R and one or more additional amino acid substitutions selected from the following: L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof. In some embodiments, the modified SpCas9 comprises: (i) one or more amino acid substitutions selected from the following: D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more additional amino acid substitutions selected from the following: L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or a corresponding amino acid substitution thereof.
[0312] In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas domain, such as a Cas9 domain. In some embodiments, the endonuclease domain or DNA binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In some embodiments, the endonuclease domain or DNA binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA binding domain comprises S. pyogenes or S. thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas9 sequence, e.g., as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5,726-737, which is incorporated herein by reference. In some embodiments, the endonuclease domain or DNA binding domain comprises the HNH nuclease subdomain and / or the RuvC1 subdomain of a Cas, e.g., Cas9, or a mutant thereof, as described herein.In some embodiments, the endonuclease domain or DNA binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA binding domain comprises a Cas polypeptide (e.g., an enzyme), or a functional fragment thereof. In some embodiments, the Cas polypeptide (e.g., an enzyme) is selected from the following: Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpf, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. l, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Cs y3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1 , Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, C sa5, a type II Cas effector protein, a type V Cas effector protein, a type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, hyper accurate Cas9 mutant (HypaCas9), homologues thereof, modified or engineered versions thereof, and / or functional fragments thereof.In some embodiments, the Cas9 comprises one or more substitutions selected from the following, e.g., H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A. In some embodiments, the Cas9 comprises one or more mutations at a position selected from the following, e.g., D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, e.g., one or more substitutions selected from the following, D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the endonuclease domain or DNA binding domain is selected from the group consisting of Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Staphylococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, meningitidis, Streptococcus pyogenes or Staphylococcus aureus, or a functional fragment or variant thereof.
[0313] In some embodiments, the endonuclease domain or DNA binding domain comprises a Cpf1 domain that includes one or more substitutions, e.g., at positions D917, E1006A, D1255, or any combination thereof, selected from the following: D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A.
[0314] In some embodiments, the endonuclease domain or DNA binding domain comprises spCas9, spCas9-VRQR, spCas9-VRER, xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER, spCas9-LRKIQK, or spCas9-LRVSQL.
[0315] In some embodiments, the genetically modified polypeptide has an endonuclease domain that includes a Cas9 nickase, such as Cas9 H840A. In some embodiments, Cas9 H840A has the following amino acid sequence: Cas9 Nickase (H840A): [ka]
[0316] In some embodiments, the genetically modified polypeptide comprises a dCas9 sequence that includes a D10A and / or H840A mutation, for example, the following sequence: [ka]
[0317] TAL effectors and zinc finger nucleases In some embodiments, the endonuclease domain or DNA binding domain comprises a TAL effector molecule. A TAL effector molecule, for example, a TAL effector molecule that specifically binds to a DNA sequence, typically comprises multiple TAL effector domains or fragments thereof, and optionally one or more additional portions of a naturally occurring TAL effector (for example, the N-terminus and / or C-terminus of multiple TAL effector domains). Many TAL effectors are known to those skilled in the art and are commercially available, for example, from Thermo Fisher Scientific.
[0318] Naturally occurring TALEs are natural effector proteins secreted by numerous species of bacterial pathogens, including the plant pathogen Xanthomonas, that regulate gene expression in the host plant and promote bacterial colonization and survival. The specific binding of TAL effectors is typically based on a central repeat domain (repeated variable dinucleotide, RVD domain) of nearly identical tandem repeats of 33 or 34 amino acids.
[0319] Members of the TAL effector family differ primarily in the number and order of their repeats. The number of repeats typically ranges from 1.5 to 33.5 repeats, with the C-terminal repeats usually being shorter in length (e.g., about 20 amino acids) and commonly referred to as "half-repeats." Each repeat of a TAL effector is generally characterized by a one-repeat-to-one base-pair correlation (one repeat recognizes one base pair on the target gene sequence) with different repeat types exhibiting different base-pair specificities. In general, a reduced number of repeats leads to a weakening of the protein-DNA interaction. It has been shown that several 6.5 repeats are sufficient to activate transcription of a reporter gene (Scholze et al., 2010).
[0320] The variation between repeats occurs primarily at amino acid positions 12 and 13, which are therefore termed "hypervariable" and are responsible for the specificity of the interaction with the target DNA promoter sequence, as shown in Table 9, which lists exemplary repeat variable diresidues (RVDs) and their correspondence to nucleobase targets.
[0321] [Table 39]
[0322] Thus, it is possible to modify the repeats of TAL effectors to target specific DNA sequences. Furthermore, studies have demonstrated that RVD NK can target G. Furthermore, the target sites of TAL effectors tend to contain a T adjacent to the 5' base targeted by the first repeat, although the exact mechanism of this recognition is unknown. More than 113 TAL effector sequences are known to date. Non-limiting examples of TAL effectors from Xanthomonas include Hax2, Hax3, Hax4, AvrXa7, AvrXa10, and AvrBs3.
[0323] Thus, the TAL effector domain of the TAL effector molecules described herein can be derived from TAL effectors from any bacterial species (e.g., Xanthomonas species, such as African strains of Xanthomonas oryzae pv. Oryzae (Yu et al. 2011), Xanthomonas campestris pv. raphani strain 756C, and Xanthomonas oryzae pv. Oryzicola strain BLS256 (Bogdanove et al. 2011)). In some embodiments, the TAL effector domain also comprises an RVD domain and flanking sequences (sequences N-terminal and / or C-terminal to the RVD domain) from a naturally occurring TAL effector. It may contain more or fewer repeats of the RVD of the naturally occurring TAL effector. TAL effector molecules can be designed to target a given DNA sequence based on the above codes or others known in the art. The number of TAL effector domains (e.g., repeats (monomers or modules)) and their particular sequences can be selected based on the desired DNA target sequence. For example, TAL effector domains, e.g., repeats, can be removed or added as appropriate for a particular target sequence. In certain embodiments, the TAL effector molecules of the invention comprise between 6.5 and 33.5 TAL effector domains, e.g., repeats. In certain embodiments, the TAL effector molecules of the invention comprise between 8 and 33.5 TAL effector domains, e.g., repeats, for example, between 10 and 25 TAL effector domains, e.g., repeats, for example, between 10 and 14 TAL effector domains, e.g., repeats.
[0324] In some embodiments, the TAL effector molecule comprises a TAL effector domain that corresponds to a perfect match with the DNA target sequence. In some embodiments, mismatches between the repeats and the target base pairs on the DNA target sequence are tolerated as long as they allow the function of the polypeptide comprising the TAL effector molecule. In general, TALE binding is inversely correlated with the number of mismatches. In some embodiments, the TAL effector molecule of the polypeptide of the present invention comprises at most 7 mismatches, 6 mismatches, 5 mismatches, 4 mismatches, 3 mismatches, 2 mismatches, or 1 mismatch, and optionally no mismatches, with the target DNA sequence. Without intending to be bound by a particular theory, in general, as the number of TAL effector domains of the TAL effector molecule decreases, a reduced number of mismatches will not only be tolerated, but will allow the function of the polypeptide comprising the TAL effector molecule. It is believed that the binding affinity depends on the sum of the matching repeat-DNA combinations. For example, a TAL effector molecule with 25 or more TAL effector domains may be able to tolerate up to 7 mismatches.
[0325] In addition to the TAL effector domain, the TAL effector molecule of the present invention may contain additional sequences derived from naturally occurring TAL effectors. The length of the C-terminal and / or N-terminal sequences included on both sides of the TAL effector domain portion of the TAL effector molecule may vary and may be selected by those skilled in the art, for example, based on the study of Zhang et al. (2011). Zhang et al. have characterized several C-terminal and N-terminal truncation mutants in proteins based on Hax3-derived TAL effectors and identified key elements that contribute to optimal binding to target sequences and thus transcription activation. In general, transcription activity was found to be inversely correlated with the length of the N-terminus. Regarding the C-terminus, a key element in DNA binding residues in the first 68 amino acids of the Hax3 sequence was identified. Thus, in some embodiments, the first 68 amino acids on the C-terminal side of the TAL effector domain of a naturally occurring TAL effector are included in the TAL effector molecule. Thus, in one embodiment, a TAL effector molecule comprises: 1) one or more TAL effector domains derived from a naturally occurring TAL effector; 2) at least 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260, 270, 280 or more amino acids from a naturally occurring TAL effector N-terminal to the TAL effector domain; and / or 3) at least 68, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260 or more amino acids from a naturally occurring TAL effector C-terminal to the TAL effector domain.
[0326] In some embodiments, the endonuclease domain or DNA binding domain is or comprises a Zn finger molecule. The Zn finger molecule comprises a Zn finger protein, such as a naturally occurring or modified Zn finger protein, or a fragment thereof. Many Zn finger proteins are known to those skilled in the art and are commercially available, for example, from Sigma-Aldrich.
[0327] In some embodiments, the zinc finger molecule comprises a non-naturally occurring zinc finger protein engineered to bind to a selected target DNA sequence. See, e.g., Beerli, et al. (2002) Nature Biotechnol. 20:135-141; Pabo, et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan, et al. (2001) Nature Biotechnol. 19:656-660; Segal, et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo, et al. al. (2000) Curr. Opin. Struct. Biol. 10:411-416; U.S. Pat. No. 6,453,242; U.S. Pat. No. 6,534,261; U.S. Pat. No. 6,599,692; U.S. Pat. No. 6,503,717; U.S. Pat. No. 6,689,558; U.S. Pat. No. 7,030,215; U.S. Pat. No. 6,794,136; U.S. Pat. No. 7,067,317 Nos. 7,070,934, 7,361,635, 7,253,273, and U.S. Patent Application Publication Nos. 2005 / 0064474, 2007 / 0218528, and 2005 / 0267061 (all of which are incorporated by reference in their entireties).
[0328] The modified Zn finger protein may have a novel binding specificity compared to naturally occurring Zn finger protein. The modification method includes, but is not limited to, rational design and various types of selection. Rational design includes, for example, the use of a database that includes triplet (or quadruplet) nucleotide sequences and individual Zn finger amino acid sequences, where each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of Zn fingers that bind to a particular triplet or quadruplet sequence. See, for example, U.S. Patent No. 6,453,242 and U.S. Patent No. 6,534,261 (their entireties are incorporated herein by reference).
[0329] Exemplary selection methods, including phage display and two-hybrid systems, are disclosed in U.S. Patent No. 5,789,538; U.S. Patent No. 5,925,523; U.S. Patent No. 6,007,988; U.S. Patent No. 6,013,453; U.S. Patent No. 6,410,248; U.S. Patent No. 6,140,466; U.S. Patent No. 6,200,759; and U.S. Patent No. 6,242,568; as well as International Patent Publications WO 98 / 37186; WO 98 / 53057; WO 00 / 27878; and WO 01 / 88197 and GB Patent No. 2,338,237. Additionally, increased binding specificity in zinc finger proteins is described, for example, in International Patent Publication WO 02 / 077227.
[0330] Furthermore, as disclosed in these and other references, zinc finger domains and / or multi-fingered zinc finger proteins can be linked together using any suitable linker sequence, including, for example, linkers of 5 or more amino acids in length. See also U.S. Pat. Nos. 6,479,626; 6,903,185; and 7,153,949 for exemplary linker sequences of 6 or more amino acids in length. The proteins described herein can include any combination of suitable linkers between the individual zinc fingers of the protein. Furthermore, increased binding specificity in zinc finger binding domains is described, for example, in co-owned International Patent Publication WO 02 / 077227.
[0331] Zinc finger proteins and methods for the design and construction of fusion proteins (and polynucleotides encoding same) are known to those of skill in the art and include those disclosed in U.S. Pat. Nos. 6,140,0815; 789,538; 6,453,242; 6,534,261; 5,925,523; 6,007,988; 6,013,453; and 6,200,759; International Patent Publication Nos. WO 95 / 19431; WO 96 / 19432; WO 97 / 19436; WO 98 / 19438; WO 99 / 19439 ... These are described in detail in WO 98 / 53057; WO 98 / 54311; WO 00 / 27878; WO 01 / 60970; WO 01 / 88197; WO 02 / 099084; WO 98 / 53058; WO 98 / 53059; WO 98 / 53060; WO 02 / 016536; and WO 03 / 016496.
[0332] Furthermore, as disclosed in these and other references, the zinc finger proteins and / or multi-fingered zinc finger proteins can be linked together, e.g., as a fusion protein, using any suitable linker sequence, including, e.g., linkers of 5 or more amino acids in length. See also U.S. Pat. Nos. 6,479,626; 6,903,185; and 7,153,949 for exemplary linker sequences of 6 or more amino acids in length. The zinc finger molecules described herein can include any combination of suitable linkers between the individual zinc finger proteins and / or multi-fingered zinc finger proteins of the zinc finger molecule.
[0333] In certain embodiments, the DNA binding domain or endonuclease domain comprises a Zn finger molecule comprising an engineered Zn finger protein that binds (in a sequence-specific manner) to a target DNA sequence. In some embodiments, the Zn finger molecule comprises one Zn finger protein or a fragment thereof. In other embodiments, the Zn finger molecule comprises multiple Zn finger proteins (or fragments thereof), for example, 2, 3, 4, 5, 6 or more Zn finger proteins (and optionally at most 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2 Zn finger proteins). In some embodiments, the Zn finger molecule comprises at least three Zn finger proteins. In some embodiments, the Zn finger molecule comprises 4, 5, or 6 fingers. In some embodiments, the Zn finger molecule comprises 8, 9, 10, 11, or 12 fingers. In some embodiments, a Zn finger molecule comprising three Zn finger proteins recognizes a target DNA sequence comprising 9 or 10 nucleotides. In some embodiments, a zinc finger molecule comprising four zinc finger proteins recognizes a target DNA sequence comprising 12-14 nucleotides, while in some embodiments, a zinc finger molecule comprising six zinc finger proteins recognizes a target DNA sequence comprising 18-21 nucleotides.
[0334] In some embodiments, the Zn finger molecule comprises a double-handed Zn finger protein. A double-handed Zn finger protein is a protein in which two clusters of Zn finger proteins are separated by an intervening amino acid, such that the two Zn finger domains bind to two discontinuous target DNA sequences. An example of a double-handed Zn finger binding protein is SIP1, where a cluster of four Zn finger proteins is located at the amino terminus of the protein and a cluster of three Zn finger proteins is located at the carboxyl terminus (see Remade, et al. (1999) EMBO Journal 18(18):5073-5084). Each cluster of Zn fingers in these proteins can bind to a unique target sequence, and the interval between the two target sequences can include a number of nucleotides.
[0335] Linker In some embodiments, the genetically modified polypeptide may include a linker, e.g., a peptide linker, e.g., a linker described in Table 10. In some embodiments, the genetically modified polypeptide includes, from N-terminal to C-terminal, a Cas domain (e.g., a Cas domain in Table 8), a linker in Table 10 (or a sequence having at least 70%, 80%, 85%, 90%, 95% or 99% identity thereto), and an RT domain (e.g., an RT domain in Table 6). In some embodiments, the genetically modified polypeptide includes a flexible linker between the endonuclease and the RT domain, e.g., a linker comprising the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 11,002). In some embodiments, the RT domain of the genetically modified polypeptide may be located C-terminal to the endonuclease domain. In some embodiments, the RT domain of the genetically modified polypeptide may be located N-terminal to the endonuclease domain.
[0336] [Table 40]
[0337] [Table 41]
[0338] [Table 42]
[0339] [Table 43]
[0340] In some embodiments, the linker of the genetically modified polypeptide is (SGGS) n (SEQ ID NO:5025), (GGGS) n (SEQ ID NO:5026), (GGGGS) n (SEQ ID NO:5027), (G) n , (EAAAK) n (SEQ ID NO:5028), (GGS) n , or (XP) n The motif is selected from:
[0341] Selection of genetically modified polypeptides by pooled screening Candidate gene modification polypeptide can be screened to evaluate the gene editing ability of the candidate.For example, RNA gene modification system designed for targeted editing of coding sequence in human genome can be used.In certain embodiments, such gene modification system can be used with pool screening method.
[0342] For example, a library of genetically modified polypeptide candidates and template guide RNA (tgRNA) can be introduced into a mammalian cell to test the gene editing ability of the candidates by pool screening method. In certain embodiments, the library of genetically modified polypeptide candidates is introduced into a mammalian cell, and then tgRNA is introduced into the cell.
[0343] Representative, non-limiting examples of mammalian cells that can be used in the screen include HEK293T cells, U2OS cells, HeLa cells, HepG2 cells, Huh7 cells, K562 cells, or iPS cells.
[0344] The genetically modified polypeptide candidates may include: 1) a Cas-nuclease, such as a wild-type Cas nuclease, such as a wild-type Cas9 nuclease, a mutant Cas nuclease, such as a Cas nickase, such as a Cas9 nickase, such as a Cas9 N863A nickase, or a Cas nuclease selected from Table 7 or Table 8; 2) a peptide linker, such as a sequence from Table D or Table 10, which may exhibit various degrees of length, flexibility, hydrophobicity, and / or secondary structure; and 3) a reverse transcriptase (RT), such as an RT domain from Table D or Table 6. The genetically modified polypeptide candidate library includes a plurality of different genetically modified polypeptide candidates that differ from each other with respect to one, two, or all three of the Cas nuclease, peptide linker, or RT domain components, or a plurality of nucleic acid expression vectors encoding such genetically modified polypeptide candidates.
[0345] For screening of genetically modified polypeptide candidates, a two-component system can be used that includes a genetically modified polypeptide component and a tgRNA component. The genetically modified component can include, for example, an expression vector, such as an expression plasmid or a lentiviral vector that encodes a genetically modified polypeptide candidate, and includes, for example, a human codon-optimized nucleic acid that encodes a genetically modified polypeptide candidate, such as the above-mentioned Cas-linker-RT fusion. In certain embodiments, a lentiviral cassette is used that includes (i) a promoter for expression in mammalian cells, such as a CMV promoter; (ii) a genetically modified library candidate, such as a Cas-linker-RT fusion that includes a Cas nuclease in Table 7 or Table 8, a peptide linker in Table 10, and a RT in Table 6, such as a Cas-linker-RT fusion as in Table D; (iii) a self-cleaving polypeptide, such as a T2A peptide; (iv) a marker that allows selection in mammalian cells, such as a puromycin resistance gene; and (v) a termination signal, such as a polyA tail.
[0346] The tgRNA component can include a tgRNA or an expression vector, e.g., an expression plasmid that generates the tgRNA and drives expression of the tgRNA, e.g., using a U6 promoter, where the tgRNA is a non-coding RNA sequence that is recognized by Cas, localizing it to the genomic locus of interest, and that templates the reverse transcription of the desired edit into the genome via the RT domain.
[0347] To prepare a pool of cells expressing genetically modified polypeptide library candidates, mammalian cells, such as HEK293T or U2OS cells, can be transduced with a pooled genetically modified polypeptide candidate expression vector preparation, such as a lentiviral preparation of the genetically modified candidate polypeptide library. In certain embodiments, lentiviral plasmids are used and HEK293 Lenti-X cells are seeded in 15 cm plates (approximately 12×10 cells) prior to lentiviral plasmid transfection. 6In such an embodiment, lentiviral plasmid transfection can be performed using Lentiviral Packaging Mix (Biosettia), and transfection of plasmid DNA for gene modification candidate library can be performed using Lipofectamine 2000 and Opti-MEM medium according to the manufacturer's protocol. In such an embodiment, extracellular DNA can be removed by complete medium change the next day, and virus-containing medium can be collected after 48 hours. Lentiviral medium can be concentrated using Lenti-X Concentrator (TaKaRa Biosciences), and 5 mL lentiviral aliquots can be made and stored at -80°C. Lentiviral titration is performed after selection, for example by counting colony forming units after puromycin selection.
[0348] To monitor gene editing of the target DNA, mammalian cells, such as HEK293T or U2OS cells carrying the target DNA, may be used. In other embodiments for monitoring gene editing of the target DNA, mammalian cells, such as HEK293T or U2OS cells carrying a target DNA genomic landing pad, may be used. In certain embodiments, the target DNA genomic landing pad may contain a gene to be edited for the treatment of a disease or disorder of interest. In other specific embodiments, the target DNA is a genetic sequence that expresses a protein that exhibits a detectable characteristic that can be monitored to determine whether gene editing has occurred. For example, in certain embodiments, a blue fluorescent protein (BFP)-expressing or green fluorescent protein (GFP)-expressing genomic landing pad is used. In certain embodiments, mammalian cells, such as HEK293T or U2OS cells containing the target DNA, such as a target DNA genomic landing pad, are seeded in culture plates at 500x-3000x cells per genetically modified library candidate and transduced at 0.2-0.3 multiplicity of infection (MOI) to minimize multiple infections per cell. Puromycin (2.5ug / mL) can be added 48 hours after infection to allow for selection of infected cells. In such an embodiment, the cells are placed under puromycin selection for at least 7 days and then scaled up for tgRNA introduction, e.g., tgRNA electroporation.
[0349] To confirm whether gene editing occurs, mammalian cells containing the target DNA to be edited can be infected with the candidate genetically modified polypeptide library and then transfected with a tgRNA designed for use in editing the target DNA. The cells can then be analyzed, for example, by using cell sorting and sequence analysis, to determine whether editing of the target locus occurred according to the designed results, or whether no editing or incomplete editing occurred.
[0350] In certain embodiments, to determine whether genome editing occurs, BFP- or GFP-expressing mammalian cells, e.g., HEK293T or U2OS cells, may be infected with genetically modified library candidates and then transfected or electroporated by electroporation of 250,000 cells / well with tgRNA plasmids or RNA, e.g., 200 ng of tgRNA plasmids designed to convert BFP to GFP or GFP to BFP, at a cell number that ensures >250×-1000× coverage per library candidate. In such embodiments, the genome editing capabilities of the various constructs in this assay may be assessed by sorting the cells by Fluorescence-Activated Cell Sorting (FACS) for expression of color-converted fluorescent proteins (FPs) 4-10 days after electroporation. Cells are sorted and collected as distinct populations of non-edited cells (exhibiting original fluorescent protein signal), edited cells (exhibiting converted fluorescent protein signal), and incompletely edited (exhibiting no fluorescent protein signal) cells. A sample of unsorted cells may also be collected as an input population to determine candidate enrichment during analysis.
[0351] To determine whether the genetically modified library candidates exhibit genome editing capability in the assay, genomic DNA (gDNA) is collected from the sorted cell population and analyzed by sequencing the genetically modified library candidates in each population. Briefly, the genetically modified candidates are amplified from the genome using primers specific to the genetically modified polypeptide expression vector, e.g., lentiviral cassette, amplified in a second round of PCR to dilute the genomic DNA, and then sequenced, e.g., by a next-generation sequencing platform. After quality control of the sequencing reads, reads of at least about 1500 nucleotides and generally no more than about 3200 nucleotides are mapped to the genetically modified polypeptide library sequence, and those that contain a minimum of about 80% match with the library sequence are considered to be successfully aligned with a given candidate for this pooled screen. To identify candidates that can perform genetic editing in the assay, e.g., BFP to GFP or GFP to BFP editing, the read count of each library candidate in the edited population is compared to its read count in the initial non-sorted population.
[0352] For pool screening, genetic modification candidates with genome editing capabilities are identified based on the enrichment of the edited (converted FP) population compared to non-sorted (input) cells. In some embodiments, an enrichment of at least 1.0, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or at least 100 times the input indicates potentially useful gene editing activity, for example, at least 2-fold enrichment. In some embodiments, enrichment is converted to a log value by taking the log base 2 of the enrichment ratio. In some embodiments, a log2 enrichment score of at least 0, 1, 2, 3, 4, 5, 5.5, 6.0, 6.2, 6.3, 6.4, 6.5, or at least 6.6 indicates potentially useful gene editing activity, for example, a log2 enrichment score of at least 1.0. In certain embodiments, the enrichment value observed for a candidate genetic modification can be compared to the enrichment value observed under similar conditions using a reference, for example, element ID number: 17380.
[0353] In some embodiments, multiple tgRNAs can be used to screen gene modification candidate library.In certain embodiments, multiple tgRNAs can be used to optimize template / Cas-linker-RT fusion pairs, for example, for gene editing of specific target genes, for example, gene targets for disease treatment.In certain embodiments, pooling method for screening gene modification candidates can be performed with many different tgRNAs in array format.
[0354] In some embodiments, multiple types of edits, e.g., insertions, substitutions, and / or deletions of different lengths, can be used to screen a library of genetic modification candidates.
[0355] In some embodiments, multiple target sequences, such as different fluorescent proteins, can be used to screen the genetically modified candidate library. In some embodiments, multiple target sequences, such as different fluorescent proteins, can be used to screen the genetically modified candidate library. In some embodiments, multiple cell types, such as HEK293T or U2OS, can be used to screen the genetically modified candidate library. Those skilled in the art will understand that a given candidate may show an increased or decreased altered editing ability or observable or useful activity across different conditions, including tgRNA sequence (e.g., nucleotide modification, PBS length, RT template length), target sequence, target position, type of editing, position of mutation relative to the first strand nick of the genetically modified polypeptide, or cell type. Thus, in some embodiments, genetically modified library candidates are screened across multiple parameters, for example, with at least two different tgRNAs in at least two cell types, and genetic editing activity is identified by enrichment in any single condition. In other embodiments, candidates with more robust activity across different tgRNAs and cell types are identified by enrichment in at least two conditions, for example, all conditions screened. To be clear, candidates found to show little to no enrichment under any given condition are not presumed to be inactive across all conditions and may be screened using different parameters or reconstituted at the polypeptide level, for example by exchanging, swapping or altering domains (e.g., RT domains), linkers or other signals (e.g., NLS).
[0356] Exemplary Cas9-Linker-RT Fusion Sequences In some embodiments, the genetically modified polypeptide comprises a linker sequence and a RT sequence. In some embodiments, the genetically modified polypeptide comprises a linker sequence listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of an RT domain listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises a linker sequence listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto; and an amino acid sequence of an RT domain listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises: (i) a linker sequence listed in a row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto; and (ii) an amino acid sequence of an RT domain listed in the same row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0357] Exemplary Genetically Modified Polypeptides In some embodiments, the genetically modified polypeptide (e.g., a genetically modified polypeptide that is part of a system described herein) comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 80% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 90% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 95% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of any one of SEQ ID NOs: 1-7743. In some embodiments, the genetically modified polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0358] In some embodiments, the genetically modified polypeptide comprises an amino acid sequence listed in Table A1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0359] In some embodiments, the genetically modified polypeptide comprises an amino acid sequence listed in Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises a linker comprising a linker sequence listed in Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an RT domain comprising an RT domain sequence listed in Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises: (i) a linker comprising a linker sequence listed in a row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto; and (ii) an RT domain comprising an RT domain sequence listed in the same row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0360] [Table 44]
[0361] In some embodiments, the genetically modified polypeptide comprises an amino acid sequence listed in Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises a linker comprising a linker sequence listed in Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises an RT domain comprising an RT domain sequence listed in Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide comprises: (i) a linker comprising a linker sequence listed in a row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto; and (ii) an RT domain comprising an RT domain sequence listed in the same row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0362] [Table 45]
[0363] [Table 46]
[0364] [Table 47]
[0365] Exemplary Genetically Modified Polypeptide Subsequences In some embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, one or more (e.g., one, two, three, four, five, or all six) of an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA binding domain, a linker, an RT domain, and / or a second NLS. In some embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, an NLS (e.g., a first NLS), a DNA binding domain, a linker, and an RT domain, wherein the linker and RT domain are the linker and RT domain of any one of the genetically modified polypeptides of SEQ ID NOs: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to said linker and RT domain. In some embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, a DNA binding domain, a linker, an RT domain, and an NLS (e.g., a second NLS), where the linker and RT domain are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to the linker and RT domain of any one of the genetically modified polypeptides of SEQ ID NOs: 1-7743, or said linker and RT domain. In some embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, a first NLS, a DNA binding domain, a linker, an RT domain, and a second NLS, where the linker and RT domain are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to the linker and RT domain of any one of the genetically modified polypeptides of SEQ ID NOs: 1-7743, or said linker and RT domain. In some embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.
[0366] In some embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, an N-terminal methionine residue, a first nuclear localization signal (NLS) (e.g., a genetically modified polypeptide listed in any one of SEQ ID NOs: 1-7743 and / or in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto), a DNA binding domain (e.g., a Cas domain, e.g., a SpyCas9 domain, e.g., a Cas domain listed in Table 8, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto; or a DNA binding domain of a genetically modified polypeptide listed in any one of SEQ ID NOs: 1-7743 and / or in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto). , a linker (e.g., a genetically modified polypeptide listed in any one of SEQ ID NOs: 1 to 7743 and / or any one of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto), an RT domain (e.g., a genetically modified polypeptide listed in any one of SEQ ID NOs: 1 to 7743 and / or any one of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto), 5%, 80%, 85%, 90%, 95% or 99% identity thereto), and one or more (e.g., one, two, three, four, five, or all six) of a first NLS (e.g., a genetically modified polypeptide listed in any one of SEQ ID NOs: 1-7743 and / or any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto).In some embodiments, the genetically modified polypeptide further comprises (e.g., C-terminal to the second NLS) a T2A sequence and / or a puromycin sequence (e.g., any one of SEQ ID NOs: 1-7743 and / or any of Tables A1, T1 or T2 listed in or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto). In some embodiments, a nucleic acid encoding a genetically modified polypeptide (e.g., as described herein) encodes a T2A sequence, e.g., the T2A sequence is located between a region encoding the genetically modified polypeptide and a second region, which optionally encodes a selection marker, e.g., puromycin.
[0367] In certain embodiments, the first NLS comprises a first NLS sequence of a genetically modified polypeptide having an amino acid sequence of any of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the first NLS comprises a first NLS sequence of a genetically modified polypeptide listed in any of Tables A1, T1 or T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the first NLS sequence comprises a C-myc NLS. In certain embodiments, the first NLS comprises the amino acid sequence PAAKRVKLD (SEQ ID NO: 11,095) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0368] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the first NLS and the DNA binding domain. In certain embodiments, the spacer sequence between the first NLS and the DNA binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In certain embodiments, the spacer sequence between the first NLS and the DNA binding domain comprises the amino acid sequence GG.
[0369] In certain embodiments, the DNA binding domain comprises a DNA binding domain of a genetically modified polypeptide of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the DNA binding domain comprises a DNA binding domain of a genetically modified polypeptide listed in any of Tables A1, T1 or T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the DNA binding domain comprises a Cas domain (e.g., those listed in Table 8). In certain embodiments, the DNA binding domain comprises an amino acid sequence of a SpyCas9 polypeptide (e.g., those listed in Table 8, e.g., Cas9 N863A polypeptide) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the DNA binding domain has the amino acid sequence: [ka] or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0370] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the DNA binding domain and the linker. In certain embodiments, the spacer sequence between the DNA binding domain and the linker comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In certain embodiments, the spacer sequence between the DNA binding domain and the linker comprises the amino acid sequence GG.
[0371] In certain embodiments, the linker comprises a linker sequence of a genetically modified polypeptide of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises a linker sequence of a genetically modified polypeptide listed in any of Tables A1, T1 or T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises an amino acid sequence listed in Tables D or 10 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0372] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the linker and the RT domain. In certain embodiments, the spacer sequence between the linker and the RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In certain embodiments, the spacer sequence between the linker and the RT domain comprises the amino acid sequence GG.
[0373] In certain embodiments, the RT domain comprises an RT domain sequence of a genetically modified polypeptide of any one of SEQ ID NOs: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the RT domain comprises an RT domain sequence of a genetically modified polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the RT domain comprises an amino acid sequence listed in Tables D or 6, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain has a length of about 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 amino acids.
[0374] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the RT domain and the second NLS. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises the amino acid sequence AG.
[0375] In certain embodiments, the second NLS comprises a second NLS sequence of a genetically modified polypeptide of any one of SEQ ID NOs: 1-7743. In certain embodiments, the second NLS comprises a second NLS sequence of a genetically modified polypeptide listed in any of Tables A1, T1, or T2. In certain embodiments, the second NLS sequence comprises a plurality of partial NLS sequences. In embodiments, the NLS sequence, e.g., the second NLS sequence, comprises a first partial NLS sequence (e.g., comprising the amino acid sequence KRTADGSEFE (SEQ ID NO: 11,097)), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the NLS sequence, e.g., the second NLS sequence, comprises a second partial NLS sequence. In embodiments, the NLS sequence, e.g., the second NLS sequence, comprises an SV40A5 NLS, e.g., a bisecting SV40A5 NLS (e.g., comprising the amino acid sequence KRTADGSEFESPKKKAKVE (SEQ ID NO: 11,098)), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the NLS sequence, e.g., the second NLS sequence, comprises the amino acid sequence KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 11,099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0376] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acids. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises the amino acid sequence GSG.
[0377] Linker and RT domains In some embodiments, the genetically modified polypeptide comprises a linker (e.g., as described herein) and an RT domain (e.g., as described herein). In certain embodiments, the genetically modified polypeptide comprises, in order from N-terminus to C-terminus, a linker (e.g., as described herein) and an RT domain (e.g., as described herein).
[0378] In certain embodiments, the linker comprises a linker sequence listed in Table 10 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises a linker sequence of any one of SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises a linker sequence of any one of SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises a linker sequence of any one of SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the linker comprises a linker sequence of an exemplary genetically modified polypeptide listed in any of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain comprises an RT domain sequence listed in Table 6, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain comprises an RT domain sequence of an exemplary genetically modified polypeptide listed in any of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0379] In some embodiments, the genetically modified polypeptide comprises a portion of any one of SEQ ID NOs: 1-7743, the portion comprising the linker and RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said portion.
[0380] In some embodiments, the genetically modified polypeptide comprises a linker of any one of SEQ ID NOs: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said linker. In some embodiments, the genetically modified polypeptide comprises a linker of any one of SEQ ID NOs: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said linker. In some embodiments, the genetically modified polypeptide comprises a linker of any one of SEQ ID NOs: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said linker. In some embodiments, the genetically modified polypeptide comprises a linker comprising a genetically modified polypeptide linker listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0381] In some embodiments, the genetically modified polypeptide comprises an RT domain of any one of SEQ ID NOs: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said RT domain. In some embodiments, the genetically modified polypeptide comprises an RT domain of any one of SEQ ID NOs: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said RT domain. In some embodiments, the genetically modified polypeptide comprises an RT domain of any one of SEQ ID NOs: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said RT domain. In some embodiments, the genetically modified polypeptide comprises an RT domain of a genetically modified polypeptide listed in any of Tables A1, T1 or T2, or an RT domain comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0382] In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence of the linker and RT domain of the genetically modified polypeptide having the amino acid sequence of any one of SEQ ID NOs: 1-7743 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto). In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence of the linker and RT domain having at least 80% identity thereto. In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence of the linker and RT domain having at least 90% identity thereto. In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence of the linker and RT domain having at least 95% identity thereto. In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises an amino acid sequence of the linker and RT domain having at least 99% identity to any one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises an amino acid sequence of the linker and RT domain of the genetically modified polypeptide having an amino acid sequence of any one of SEQ ID NOs: 6001 to 7743 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto). In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises an amino acid sequence of the linker and RT domain of the genetically modified polypeptide having an amino acid sequence of any one of SEQ ID NOs: 4501 to 4541 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto).In certain embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence of the linker and RT domain from a single row of any of Tables A1, T1, or T2 (e.g., a single exemplary genetically modified polypeptide listed in any of Tables A1, T1, or T2) (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).
[0383] In certain embodiments, the linker and RT domains of the genetically modified polypeptides comprise linker and RT domain amino acid sequences from two different amino acid sequences selected from SEQ ID NOs: 1-7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto). In certain embodiments, the linker and RT domains of the genetically modified polypeptides comprise linker and RT domain amino acid sequences from different rows of any of Tables A1, T1, or T2 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto).
[0384] In certain embodiments, the genetically modified polypeptide further comprises a first NLS (e.g., a 5'NLS), e.g., as described herein. In certain embodiments, the genetically modified polypeptide further comprises a second NLS (e.g., a 3'NLS), e.g., as described herein. In certain embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.
[0385] RT family and mutants In certain embodiments, the genetically modified polypeptide comprises the amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, XMRV6, BLVAU, BLVJ, HTL1A, HTL1C, HTL1L, HTL32, HTL3P, HTLV2, JSRV, MLVF5, MLVRD, MMTVB, MPMV, SFVCP, SMRVH, SRV1, SRV2, and WDSV. In certain embodiments, the genetically modified polypeptide comprises the amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.
[0386] In certain embodiments, the genetically modified polypeptide comprises the amino acid sequence of an RT domain sequence from the MLVMS RT domain. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations as listed in column 1 of Table M1, or corresponding point mutations thereto. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations as listed in column 3 of Table M1 (Gen1 MLVMS), or corresponding point mutations thereto. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations at amino acid positions of the RT domain as listed in columns 1 and 2 of Table M2, or corresponding amino acid positions thereto.
[0387] In certain embodiments, the genetically modified polypeptide comprises the amino acid sequence of an RT domain sequence from the AVIRE RT domain. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations as listed in column 2 of Table M1, or corresponding point mutations thereto. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations as listed in column 4 of Table M1 (Gen2 AVIRE), or corresponding point mutations thereto. In embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations at amino acid positions of the RT domain as listed in columns 3 and 4 of Table M2, or corresponding amino acid positions thereto. In certain embodiments, the RT domain comprises an IENSSP (e.g., at the C-terminus).
[0388] [Table 48]
[0389] [Table 49]
[0390] In certain embodiments, the genetically modified polypeptide comprises an RT domain from a gammaretrovirus. In certain embodiments, the RT domain of the genetically modified polypeptide from a gammaretrovirus comprises the amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In some embodiments, the RT domain of the genetically modified polypeptide from a gammaretrovirus is not derived from PERV. In some embodiments, the RT comprises one, two, three, four, five, six or more mutations corresponding to the mutations D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N in the RT domain of murine leukemia virus reverse transcriptase as shown in Table 2A. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% identity to the linker domain of any one of SEQ ID NOs: 1-7743. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity to SEQ ID NO: 5217 or SEQ ID NO: 11,041.
[0391] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of AVIRE RT (e.g., the AVIRE_P03360 sequence, e.g., SEQ ID NO: 8001), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of AVIRE RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D200N, G330P, L605W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of AVIRE RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D200N, G330P, and L605W.
[0392] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of BAEVM RT (e.g., the BAEVM_P10272 sequence, e.g., SEQ ID NO: 8004), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of BAEVM RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D198N, E328P, L602W, T304K, and W311F. In some embodiments, the RT domain comprises the amino acid sequence of BAEVM RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D198N, E328P, and L602W.
[0393] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of FFV RT (e.g., FFV_O93209 sequence, e.g., SEQ ID NO: 8012), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of FFV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three or four mutations selected from the group consisting of D21N, T293N, T419P and L393K. In some embodiments, the RT domain comprises the amino acid sequence of FFV RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D21N, T293N and T419P. In some embodiments, the RT domain comprises the amino acid sequence of FFV RT, further comprising the mutation D21N. In some embodiments, the RT domain comprises the amino acid sequence of FFV RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of T207N, T333P, and L307K, In some embodiments, the RT domain comprises the amino acid sequence of FFV RT, or a corresponding position in a homologous RT domain, further comprising one or two mutations selected from the group consisting of T207N and T333P.
[0394] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of FLV RT (e.g., the FLV_P10273 sequence, e.g., SEQ ID NO: 8019), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of FLV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, or four mutations selected from the group consisting of D199N, L602W, T305K, and W312F. In some embodiments, the RT domain comprises the amino acid sequence of FLV RT, or a corresponding position in a homologous RT domain, further comprising one or two mutations selected from the group consisting of D199N and L602W.
[0395] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of FOAMV RT (e.g., the FOAMV_P14350 sequence, e.g., SEQ ID NO: 8021), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of FOAMV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three or four mutations selected from the group consisting of D24N, T296N, S420P, and L396K. In some embodiments, the RT domain comprises the amino acid sequence of FOAMV RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D24N, T296N, and S420P. In some embodiments, the RT domain comprises the amino acid sequence of FOAMV RT, or a corresponding position in a homologous RT domain, further comprising the mutation D24N. In some embodiments, the RT domain comprises the amino acid sequence of FOAMV RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of T207N, S331P, and L307K. In some embodiments, the RT domain comprises the amino acid sequence of FOAMV RT, or a corresponding position in a homologous RT domain, further comprising one or two mutations selected from the group consisting of T207N and S331P.
[0396] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of GALV RT (e.g., the GALV_P21414 sequence, e.g., SEQ ID NO: 8027), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of GALV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain comprises the amino acid sequence of GALV RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D198N, E328P, and L600W.
[0397] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of KORV RT (e.g., the KORV_Q9TTC1 sequence, e.g., SEQ ID NO: 8047), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of GALV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four, five, or six mutations selected from the group consisting of D32N, D322N, E452P, L274W, T428K, and W435F. In some embodiments, the RT domain comprises the amino acid sequence of GALV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, or four mutations selected from the group consisting of D32N, D322N, E452P, and L274W. In some embodiments, the RT domain comprises the amino acid sequence of GALV RT further comprising the mutation D32N. In some embodiments, the RT domain comprises the amino acid sequence of KORV RT further comprising one, two, three, four or five mutations selected from the group consisting of D231N, E361P, L633W, T337K, and W344F, or at the corresponding position in a homologous RT domain. In some embodiments, the RT domain comprises the amino acid sequence of KORV RT further comprising one, two or three mutations selected from the group consisting of D231N, E361P, and L633W, or at the corresponding position in a homologous RT domain.
[0398] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVAV RT (e.g., the MLVAV_P03356 sequence, e.g., SEQ ID NO: 8053), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of MLVAV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of MLVAV RT, or a corresponding position in a homologous RT domain, further comprising one, two, or three mutations selected from the group consisting of D200N, T330P, and L603W.
[0399] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVBM RT (e.g., the MLVBM_Q7SVK7 sequence, e.g., SEQ ID NO: 8056), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of MLVBM RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four, or five mutations selected from the group consisting of D199N, T329P, L602W, T305K, and W312F. In some embodiments, the RT domain comprises the amino acid sequence of MLVBM RT, or a corresponding position in a homologous RT domain, further comprising one, two, and three mutations selected from the group consisting of D200N, T330P, and L603W.
[0400] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVCB RT (e.g., MLVCB_P08361 sequence, e.g., SEQ ID NO: 8062), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of MLVCB RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of MLVCB RT, or a corresponding position in a homologous RT domain, further comprising one, two and three mutations selected from the group consisting of D200N, T330P, and L603W.
[0401] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVFF RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of MLVFF RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of MLVFF RT, or a corresponding position in a homologous RT domain, further comprising one, two and three mutations selected from the group consisting of D200N, T330P, and L603W.
[0402] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVMS RT (e.g., MLVMS_refSeq, e.g., SEQ ID NO: 8137; or MLVMS_P03355 sequence, e.g., SEQ ID NO: 8070), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of MLVMS RT, or the corresponding positions in a homologous RT domain, further comprising one, two, three, four, five, or six mutations selected from the group consisting of D200N, T330P, L603W, T306K, W313F, and H8Y. In some embodiments, the RT domain comprises the amino acid sequence of MLVMS RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of MLVMS RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D200N, T330P, and L603W.
[0403] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of a PERV RT (e.g., the PERV_Q4VFZ2 sequence, e.g., SEQ ID NO: 8099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of a PERV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four, or five mutations selected from the group consisting of D196N, E326P, L599W, T302K, and W309F. In some embodiments, the RT domain comprises the amino acid sequence of a PERV RT, or a corresponding position in a homologous RT domain, further comprising one, two, or three mutations selected from the group consisting of D196N, E326P, and L599W.
[0404] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of SFV1 RT (e.g., SFV1_P23074 sequence, e.g., SEQ ID NO: 8105), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of SFV1 RT, or a corresponding position in a homologous RT domain, further comprising one, two, three or four mutations selected from the group consisting of D24N, T296N, N420P, and L396K. In some embodiments, the RT domain comprises the amino acid sequence of SFV1 RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D24N, T296N, and N420P. In some embodiments, the RT domain comprises the amino acid sequence of SFV1 RT, or a corresponding position in a homologous RT domain, further comprising D24N.
[0405] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of SFV3L RT (e.g., the SFV3L_P27401 sequence, e.g., SEQ ID NO: 8111), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of SFV3L RT, or a corresponding position in a homologous RT domain, further comprising one, two, three or four mutations selected from the group consisting of D24N, T296N, N422P, and L396K. In some embodiments, the RT domain comprises the amino acid sequence of SFV3L RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D24N, T296N, and N422P. In some embodiments, the RT domain comprises the amino acid sequence of SFV3L RT, or the corresponding position in a homologous RT domain, further comprising the mutation D24N. In some embodiments, the RT domain comprises the amino acid sequence of SFV3L RT, or the corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of T307N, N333P, and L307K. In some embodiments, the RT domain comprises the amino acid sequence of SFV3L RT, or the corresponding position in a homologous RT domain, further comprising one or two mutations selected from the group consisting of T307N and N333P.
[0406] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of WMSV RT (e.g., the WMSV_P03359 sequence, e.g., SEQ ID NO: 8131), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of WMSV RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four, or five mutations selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain comprises the amino acid sequence of WMSV RT, or a corresponding position in a homologous RT domain, further comprising one, two, or three mutations selected from the group consisting of D198N, E328P, and L600W.
[0407] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of XMRV6 RT (e.g., the XMRV6_A1Z651 sequence, e.g., SEQ ID NO: 8134), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the RT domain comprises the amino acid sequence of XMRV6 RT, or a corresponding position in a homologous RT domain, further comprising one, two, three, four or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises the amino acid sequence of XMRV6 RT, or a corresponding position in a homologous RT domain, further comprising one, two or three mutations selected from the group consisting of D200N, T330P, and L603W.
[0408] In certain embodiments, the RT domain of the genetically modified polypeptide comprises the amino acid sequence of the RT domain of AVIRE RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In embodiments, the RT domain comprises the amino acid sequence of the RT domain contained in the sequence listed in column 1 of Table A5 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity thereto with SEQ ID NO:5217 or SEQ ID NO:11,041.
[0409] In certain embodiments, the RT domain of the genetically modified polypeptide comprises the amino acid sequence of the RT domain of MLVMS RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In embodiments, the RT domain comprises the amino acid sequence of the RT domain contained in any of the sequences listed in columns 2-6 of Table A5 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity thereto, SEQ ID NO:5217 or SEQ ID NO:11,041.
[0410] [Table 50]
[0411] [Table 51]
[0412] [Table 52]
[0413] [Table 53]
[0414] [Table 54]
[0415] [Table 55]
[0416] [Table 56]
[0417] [Table 57]
[0418] system In one aspect, the disclosure relates to a system comprising a nucleic acid molecule (e.g., as described herein) encoding a genetically modified polypeptide and a template nucleic acid (e.g., a template RNA, e.g., as described herein). In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises one or more silent mutations in the coding region (e.g., a sequence encoding the RT domain) relative to the nucleic acid molecule described herein. In certain embodiments, the system further comprises a gRNA (e.g., a gRNA that binds to a polypeptide that induces a nick, e.g., on the opposite strand of the target DNA to which the genetically modified polypeptide is bound).
[0419] In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, a nucleic acid molecule encoding a genetically modified polypeptide encodes a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0420] In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NOs: 1-7743, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NOs: 6001-7743, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NOs: 4501-4541, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, a nucleic acid molecule encoding a genetically modified polypeptide comprises a sequence encoding a portion of a polypeptide listed in any of Tables A1, T1, or T2, the portion including the linker and RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to said portion.
[0421] In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a linker of an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding a linker of a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0422] In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding the RT domain of an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding the RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding the RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide comprises a sequence encoding the RT domain of a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0423] In one aspect, the disclosure relates to a system comprising a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., a template RNA, e.g., as described herein).
[0424] In certain embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0425] In certain embodiments, the genetically modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NOs: 1-7743, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, the genetically modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NOs: 6001-7743, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, the genetically modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NOs: 4501-4541, the portion comprising an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to the linker and RT domain, or the portion. In certain embodiments, the genetically modified polypeptide comprises a portion of a polypeptide listed in any of Tables A1, T1 or T2, the portion comprising the linker and RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity to said portion.
[0426] In certain embodiments, the genetically modified polypeptide comprises a linker of an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a sequence encoding a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a sequence encoding a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a linker of a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0427] In certain embodiments, the genetically modified polypeptide comprises an RT domain of an amino acid sequence selected from SEQ ID NOs: 1-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001-7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501-4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto. In certain embodiments, the genetically modified polypeptide comprises an RT domain of a polypeptide listed in any of Tables A1, T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identity thereto.
[0428] [Table 58]
[0429] [Table 59]
[0430] [Table 60]
[0431] [Table 61]
[0432] [Table 62]
[0433]
Table 63
[0434]
Table 64
[0435]
Table 65
[0436]
Table 66
[0437]
Table 67
[0438]
Table 68
[0439]
Table 69
[0440]
Table 70
[0441]
Table 71
[0442]
Table 72
[0443]
Table 73
[0444]
Table 74
[0445]
Table 75
[0446]
Table 76
[0447]
Table 77
[0448]
Table 78
[0449]
Table 79
[0450]
Table 80
[0451]
Table 81
[0452]
Table 82
[0453]
Table 83
[0454]
Table 84
[0455]
Table 85
[0456]
Table 86
[0457]
Table 87
[0458]
Table 88
[0459]
Table 89
[0460]
Table 90
[0461]
Table 91
[0462]
Table 92
[0463]
Table 93
[0464]
Table 94
[0465]
Table 95
[0466]
Table 96
[0467]
Table 97
[0468]
Table 98
[0469]
Table 99
[0470]
Table 100
[0471]
Table 101
[0472]
Table 102
[0473]
Table 103
[0474]
Table 104
[0475]
Table 105
[0476]
Table 106
[0477]
Table 107
[0478]
Table 108
[0479]
Table 109
[0480]
Table 110
[0481]
Table 111
[0482]
Table 112
[0483]
Table 113
[0484]
Table 114
[0485]
Table 115
[0486]
Table 116
[0487]
Table 117
[0488]
Table 118
[0489]
Table 119
[0490]
Table 120
[0491]
Table 121
[0492]
Table 122
[0493]
Table 123
[0494]
Table 124
[0495]
Table 125
[0496]
Table 126
[0497]
Table 127
[0498]
Table 128
[0499]
Table 129
[0500]
Table 130
[0501]
Table 131
[0502]
Table 132
[0503]
Table 133
[0504]
Table 134
[0505]
Table 135
[0506]
Table 136
[0507]
Table 137
[0508]
Table 138
[0509]
Table 139
[0510]
Table 140
[0511]
Table 141
[0512]
Table 142
[0513]
Table 143
[0514]
Table 144
[0515]
Table 145
[0516]
Table 146
[0517]
Table 147
[0518]
Table 148
[0519]
Table 149
[0520]
Table 150
[0521]
Table 151
[0522]
Table 152
[0523]
Table 153
[0524]
Table 154
[0525]
Table 155
[0526]
Table 156
[0527]
Table 157
[0528]
Table 158
[0529]
Table 159
[0530]
Table 160
[0531]
Table 161
[0532]
Table 162
[0533]
Table 163
[0534]
Table 164
[0535]
Table 165
[0536]
Table 166
[0537]
Table 167
[0538]
Table 168
[0539]
Table 169
[0540]
Table 170
[0541]
Table 171
[0542]
Table 172
[0543]
Table 173
[0544]
Table 174
[0545]
Table 175
[0546]
Table 176
[0547]
Table 177
[0548]
Table 178
[0549]
Table 179
[0550]
Table 180
[0551]
Table 181
[0552]
Table 182
[0553]
Table 183
[0554]
Table 184
[0555]
Table 185
[0556]
Table 186
[0557]
Table 187
[0558]
Table 188
[0559]
Table 189
[0560]
Table 190
[0561]
Table 191
[0562]
Table 192
[0563]
Table 193
[0564]
Table 194
[0565]
Table 195
[0566]
Table 196
[0567]
Table 197
[0568]
Table 198
[0569]
Table 199
[0570]
Table 200
[0571]
Table 201
[0572]
Table 202
[0573]
Table 203
[0574]
Table 204
[0575]
Table 205
[0576]
Table 206
[0577]
Table 207
[0578]
Table 208
[0579]
Table 209
[0580]
Table 210
[0581]
Table 211
[0582]
Table 212
[0583]
Table 213
[0584]
Table 214
[0585]
Table 215
[0586]
Table 216
[0587]
Table 217
[0588]
Table 218
[0589]
Table 219
[0590]
Table 220
[0591]
Table 221
[0592]
Table 222
[0593]
Table 223
[0594]
Table 224
[0595]
Table 225
[0596]
Table 226
[0597]
Table 227
[0598]
Table 228
[0599]
Table 229
[0600]
Table 230
[0601]
Table 231
[0602]
Table 232
[0603]
Table 233
[0604]
Table 234
[0605]
Table 235
[0606]
Table 236
[0607]
Table 237
[0608]
Table 238
[0609]
Table 239
[0610]
Table 240
[0611]
Table 241
[0612]
Table 242
[0613]
Table 243
[0614]
Table 244
[0615]
Table 245
[0616]
Table 246
[0617]
Table 247
[0618]
Table 248
[0619]
Table 249
[0620]
Table 250
[0621]
Table 251
[0622]
Table 252
[0623]
Table 253
[0624]
Table 254
[0625]
Table 255
[0626]
Table 256
[0627]
Table 257
[0628]
Table 258
[0629]
Table 259
[0630]
Table 260
[0631]
Table 261
[0632]
Table 262
[0633]
Table 263
[0634]
Table 264
[0635]
Table 265
[0636]
Table 266
[0637]
Table 267
[0638]
Table 268
[0639]
Table 269
[0640]
Table 270
[0641]
Table 271
[0642]
Table 272
[0643]
Table 273
[0644]
Table 274
[0645]
Table 275
[0646]
Table 276
[0647]
Table 277
[0648]
Table 278
[0649]
Table 279
[0650]
Table 280
[0651]
Table 281
[0652]
Table 282
[0653]
Table 283
[0654]
Table 284
[0655]
Table 285
[0656]
Table 286
[0657]
Table 287
[0658]
Table 288
[0659]
Table 289
[0660]
Table 290
[0661]
Table 291
[0662]
Table 292
[0663]
Table 293
[0664]
Table 294
[0665]
Table 295
[0666]
Table 296
[0667]
Table 297
[0668]
Table 298
[0669]
Table 299
[0670]
Table 300
[0671]
Table 301
[0672]
Table 302
[0673]
Table 303
[0674]
Table 304
[0675]
Table 305
[0676]
Table 306
[0677]
Table 307
[0678]
Table 308
[0679]
Table 309
[0680]
Table 310
[0681]
Table 311
[0682]
Table 312
[0683]
Table 313
[0684]
Table 314
[0685]
Table 315
[0686]
Table 316
[0687]
Table 317
[0688]
Table 318
[0689]
Table 319
[0690]
Table 320
[0691]
Table 321
[0692]
Table 322
[0693]
Table 323
[0694]
Table 324
[0695]
Table 325
[0696]
Table 326
[0697]
Table 327
[0698]
Table 328
[0699]
Table 329
[0700]
Table 330
[0701]
Table 331
[0702]
Table 332
[0703]
Table 333
[0704]
Table 334
[0705]
Table 335
[0706]
Table 336
[0707]
Table 337
[0708]
Table 338
[0709]
Table 339
[0710]
Table 340
[0711]
Table 341
[0712]
Table 342
[0713]
Table 343
[0714]
Table 344
[0715]
Table 345
[0716]
Table 346
[0717]
Table 347
[0718]
Table 348
[0719]
Table 349
[0720]
Table 350
[0721]
Table 351
[0722]
Table 352
[0723]
Table 353
[0724]
Table 354
[0725]
Table 355
[0726]
Table 356
[0727]
Table 357
[0728]
Table 358
[0729]
Table 359
[0730] Table 360
[0731]
Table 361
[0732]
Table 362
[0733]
Table 363
[0734]
Table 364
[0735]
Table 365
[0736]
Table 366
[0737]
Table 367
[0738]
Table 368
[0739]
Table 369
[0740]
Table 370
[0741]
Table 371
[0742]
Table 372
[0743]
Table 373
[0744]
Table 374
[0745]
Table 375
[0746]
Table 376
[0747]
Table 377
[0748]
Table 378
[0749]
Table 379
[0750]
Table 380
[0751]
Table 381
[0752]
Table 382
[0753]
Table 383
[0754]
Table 384
[0755]
Table 385
[0756]
Table 386
[0757]
Table 387
[0758]
Table 388
[0759]
Table 389
[0760]
Table 390
[0761]
Table 391
[0762]
Table 392
[0763]
Table 393
[0764]
Table 394
[0765]
Table 395
[0766]
Table 396
[0767]
Table 397
[0768]
Table 398
[0769]
Table 399
[0770]
Table 400
[0771]
Table 401
[0772]
Table 402
[0773]
Table 403
[0774]
Table 404
[0775]
Table 405
[0776]
Table 406
[0777]
Table 407
[0778]
Table 408
[0779]
Table 409
[0780]
Table 410
[0781]
Table 411
[0782]
Table 412
[0783]
Table 413
[0784]
Table 414
[0785]
Table 415
[0786]
Table 416
[0787]
Table 417
[0788]
Table 418
[0789]
Table 419
[0790] Table 420
[0791]
Table 421
[0792]
Table 422
[0793]
Table 423
[0794]
Table 424
[0795]
Table 425
[0796]
Table 426
[0797]
Table 427
[0798]
Table 428
[0799]
Table 429
[0800]
Table 430
[0801]
Table 431
[0802]
Table 432
[0803]
Table 433
[0804]
Table 434
[0805]
Table 435
[0806]
Table 436
[0807]
Table 437
[0808]
Table 438
[0809]
Table 439
[0810]
Table 440
[0811]
Table 441
[0812]
Table 442
[0813]
Table 443
[0814]
Table 444
[0815]
Table 445
[0816]
Table 446
[0817]
Table 447
[0818]
Table 448
[0819]
Table 449
[0820] Table 450
[0821]
Table 451
[0822]
Table 452
[0823]
Table 453
[0824]
Table 454
[0825]
Table 455
[0826]
Table 456
[0827]
Table 457
[0828]
Table 458
[0829]
Table 459
[0830]
Table 460
[0831]
Table 461
[0832]
Table 462
[0833]
Table 463
[0834]
Table 464
[0835]
Table 465
[0836]
Table 466
[0837]
Table 467
[0838]
Table 468
[0839]
Table 469
[0840]
Table 470
[0841]
Table 471
[0842]
Table 472
[0843]
Table 473
[0844]
Table 474
[0845]
Table 475
[0846]
Table 476
[0847]
Table 477
[0848]
Table 478
[0849]
Table 479
[0850]
Table 480
[0851]
Table 481
[0852]
Table 482
[0853]
Table 483
[0854]
Table 484
[0855]
Table 485
[0856]
Table 486
[0857]
Table 487
[0858]
Table 488
[0859]
Table 489
[0860]
Table 490
[0861]
Table 491
[0862]
Table 492
[0863]
Table 493
[0864]
Table 494
[0865]
Table 495
[0866] [Table 496]
[0867] [Table 497]
[0868] [Table 498]
[0869] [Table 499]
[0870] [Table 500]
[0871] Localization sequences for gene modification systems In certain embodiments, the gene editor system RNA further comprises a subcellular localization sequence, e.g., a nuclear localization sequence (NLS). In some embodiments, the genetically modified polypeptide comprises an NLS contained in SEQ ID NO: 4000 and / or SEQ ID NO: 4001, or an NLS having an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
[0872] The nuclear localization sequence can be an RNA sequence that promotes the entry of the RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the genetically modified polypeptide is encoded on a first RNA, the template RNA is a second, separate RNA, and the nuclear localization signal is located on the template RNA and not on the RNA that encodes the genetically modified polypeptide. Without intending to be bound by a particular theory, in some embodiments, the RNA that encodes the genetically modified polypeptide is targeted primarily to the cytoplasm to promote its translation, while the template RNA is targeted primarily to the nucleus to promote insertion into the genome. In some embodiments, the nuclear localization signal is present at the 3' end, 5' end, or within an internal region of the template RNA. In some embodiments, the nuclear localization signal is 3' of the heterologous sequence (e.g., directly 3' of the heterologous sequence) or 5' of the heterologous sequence (e.g., directly 5' of the heterologous sequence). In some embodiments, the nuclear localization signal is located outside the 5'UTR or outside the 3'UTR of the template RNA. In some embodiments, the nuclear localization signal is located between the 5'UTR and the 3'UTR, and optionally the nuclear localization signal is not transcribed by the transgene (e.g., the nuclear localization signal is in an antisense orientation or downstream of a transcription termination signal or a polyadenylation signal). In some embodiments, the nuclear localization sequence is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present in the RNA, e.g., in the template RNA. In some embodiments, the nuclear localization signal is less than 5 bp, 10 bp, 25 bp, 50 bp, 75 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1000 bp in length. Various RNA nuclear localization sequences can be used.For example, Lubelsky and Ulitsky, Nature 555(107-111), 2018, describe the RNA sequence that drives RNA to localize in the nucleus.In some embodiments, the nuclear localization signal is SINE-derived nuclear RNA localization (SIRLOIN) signal.In some embodiments, the nuclear localization signal binds to a nuclear enriched protein. In some embodiments, the nuclear localization signal binds to a HNRNPK protein. In some embodiments, the nuclear localization signal is rich in pyrimidines, such as a C / T-rich, C / U-rich, C-rich, T-rich, or U-rich region. In some embodiments, the nuclear localization signal is derived from a long untranslated RNA. In some embodiments, the nuclear localization signal is derived from the MALAT1 long untranslated RNA or is the 600 nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738-751), 2012). In some embodiments, the nuclear localization signal is derived from the BORG long untranslated RNA or is an AGCCC motif (Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.
[0873] In some embodiments, the polypeptides described herein comprise one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, e.g., nuclear localization sequences (NLSs). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS promotes the entry of a protein comprising the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of the genetically modified polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the genetically modified polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of the Cas domain. In some embodiments, a linker sequence is disposed between the NLS and an adjacent domain of the genetically modified polypeptide.
[0874] In some embodiments, the NLS is selected from the group consisting of the amino acid sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 5009), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5010), RKSGKIAAIWKRPRKPKKKRKV (SEQ ID NO: 5011) KRTADGSEFESPKKKRKV (SEQ ID NO: 5012), KKTELQTTNAENKTKKL (SEQ ID NO: 5013), or KRGINDRNFWRGENGRKTR (SEQ ID NO:5014), KRPAATKKAGQAKKKK (SEQ ID NO:5015), PAAKRVKLD (SEQ ID NO:4644), KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO:4649), KRTADGSEFE (SEQ ID NO:4650), KRTADGSEFESPKKKAKVE (SEQ ID NO:4651), AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO:4001), or a functional fragment or variant thereof. Exemplary NLS sequences are also described in PCT / EP2000 / 011690, the contents of which are incorporated by reference herein for their disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises an amino acid sequence as disclosed in Table 11. The NLSs in this table may be utilized in one or more copies in a polypeptide at one or more positions in the polypeptide, for example, one, two, three or more copies of the NLS in the N-terminal domain, between peptide domains, in the C-terminal domain, or in a combination of positions, to improve subcellular localization to the nucleus. Multiple unique sequences may be used in a single polypeptide. The sequences may be monopartite or bipartite in nature, for example, with one or two stretches of basic amino acids, or may be used as chimeric bipartite sequences. Sequence references correspond to UniProt accession numbers, except where indicated as SeqNLS, for sequences retrieved using subcellular localization prediction algorithms (Lin et al BMC Bioinformat 13:157 (2012), incorporated herein by reference in its entirety).
[0875] [Table 501]
[0876] [Table 502]
[0877] [Table 503]
[0878] [Table 504]
[0879] [Table 505]
[0880] [Table 506]
[0881] In some embodiments, the NLS is a bi-knot NLS. A bi-knot NLS typically comprises two basic amino acid clusters (e.g., may be about 10 amino acids long) separated by a spacer sequence. A mono-knot NLS typically lacks a spacer. An example of a bi-knot NLS is the nucleoplasmin NLS with the sequence KR[PAATKKAGQA]KKKK (SEQ ID NO: 5015) (spacer in brackets). Another exemplary bi-knot NLS has the sequence PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5016). Exemplary NLSs are described in WO2020051561 (which is incorporated herein by reference in its entirety, including the disclosure regarding nuclear localization sequences).
[0882] In certain embodiments, the gene editor system polypeptide (e.g., the genetically modified polypeptide described herein) further comprises a subcellular localization sequence, e.g., a nuclear localization sequence and / or a nucleolar localization sequence. The nuclear localization sequence and / or the nucleolar localization sequence may be an amino acid sequence that promotes the entry of the protein into the nucleus and / or the nucleolus, where it may promote the integration of the heterologous sequence into the genome. In certain embodiments, the gene editor system polypeptide (e.g., the genetically modified polypeptide described herein, for example) further comprises a nucleolar localization sequence. In certain embodiments, the genetically modified polypeptide is encoded on a first RNA, the template RNA is a second separate RNA, and the nucleolar localization signal is encoded on the RNA encoding the genetically modified polypeptide and is not encoded on the template RNA. In some embodiments, the nucleolar localization signal is located at the N-terminus, C-terminus, or within an internal region of the polypeptide. In some embodiments, multiple identical or different nucleolar localization signals are used. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, or 100 amino acids in length. Nucleolar localization signals of various polypeptides can be used. For example, Yang et al., Journal of Biomedical Science 22, 33 (2015) describes a nuclear localization signal that also functions as a nucleolar localization signal. In some embodiments, the nucleolar localization signal can also be a nuclear localization signal. In some embodiments, the nucleolar localization signal can overlap with a nuclear localization signal. In some embodiments, the nucleolar localization signal can include a stretch of basic residues. In some embodiments, the nucleolar localization signal can be enriched in arginine and lysine residues. In some embodiments, the nucleolar localization signal can be derived from a protein that is enriched in the nucleolus. In some embodiments, the nucleolar localization signal can be derived from a protein that is enriched in the ribosomal RNA locus. In some embodiments, the nucleolar localization signal may be derived from a protein that binds to rRNA. In some embodiments, the nucleolar localization signal may be derived from MSP58. In some embodiments, the nucleolar localization signal may be a monoknot motif.In some embodiments, the nucleolar localization signal may be a bi-knot motif. In some embodiments, the nucleolar localization signal may consist of multiple mono- or bi-knot motifs. In some embodiments, the nucleolar localization signal may consist of a mixture of mono- and bi-knot motifs. In some embodiments, the nucleolar localization signal may be a double bi-knot motif. In some embodiments, the nucleolar localization motif may be KRASSQALGTIPKRRSSSRFIKRKK (SEQ ID NO: 5017). In some embodiments, the nucleolar localization signal may be derived from nuclear factor kappa B-inducing kinase. In some embodiments, the nucleolar localization signal may be a RKKRKKK motif (SEQ ID NO: 5018) (described in Birbach et al., Journal of Cell Science, 117(3615-3624), 2004).
[0883] Evolutionary variants of genetically engineered polypeptides and systems In some embodiments, the present invention provides evolved variants of the genetically modified polypeptides described herein. The evolved variants can be produced in some embodiments by mutagenizing the reference genetically modified polypeptide or one of the fragments or domains contained therein. In some embodiments, one or more of the domains (e.g., reverse transcriptase domain) are evolved. One or more of these evolved variant domains can evolve alone or together with other domains in some embodiments. One or more evolved variant domains can be combined in some embodiments with non-evolved cognate components or evolved variants of cognate components (e.g., those that may have evolved in a parallel or sequential manner).
[0884] In some embodiments, the process of mutagenizing the reference genetically modified polypeptide or a fragment or domain thereof comprises mutagenizing the reference genetically modified polypeptide or a fragment or domain thereof. In some embodiments, the mutagenization comprises, for example, a progressive evolution method (e.g., PACE) or a non-progressive evolution method (e.g., PANCE), as described herein. In some embodiments, the evolved genetically modified polypeptide or a fragment or domain thereof comprises one or more amino acid mutations introduced into its amino acid sequence compared to the amino acid sequence of the reference genetically modified polypeptide or a fragment or domain thereof. In some embodiments, the amino acid sequence mutation may comprise one or more mutated residues (e.g., conservative substitutions, non-conservative substitutions, or a combination thereof) in the amino acid sequence of the reference genetically modified polypeptide, for example, as a result of a change in the nucleotide sequence encoding the genetically modified polypeptide resulting in a change in the codon at any particular position in the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination thereof. An evolved variant genetically modified polypeptide may include variants in one or more components or domains of the genetically modified polypeptide (eg, variants introduced into the reverse transcriptase domain).
[0885] In some aspects, the disclosure provides genetically modified polypeptides, systems, kits, and methods that use or include evolved variants of genetically modified polypeptides, e.g., using evolved variants of genetically modified polypeptides or genetically modified polypeptides produced or producible by PACE or PANCE. In some embodiments, the non-evolved reference genetically modified polypeptide is a genetically modified polypeptide disclosed herein.
[0886] The term "phage-assisted continuous evolution (PACE)" as used herein generally refers to incremental evolution using phages as viral vectors. Examples of PACE technology are described, for example, in International PCT Application No. PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010 as WO 2010 / 028347; International PCT Application, PCT / US2011 / 066747, filed December 22, 2011, and published June 28, 2012 as WO 2012 / 088381; U.S. Patent No. 9,023,594, issued May 5, 2015; U.S. Patent No. 9,771,574, issued September 26, 2017; U.S. Patent No. 9,771,574, issued July 19, 2016; and U.S. Patent No. 9,821,575, issued July 19, 2017. No. 9,394,537, filed on January 20, 2015, and published on September 11, 2015 as WO 2015 / 134121; U.S. Pat. No. 10,179,911, issued on January 15, 2019; and International PCT Application PCT / US2016 / 027795, filed on April 15, 2016, and published on October 20, 2016 as WO 2016 / 168631, each of which is incorporated by reference in its entirety.
[0887] The term "phage-assisted non-gradual evolution (PANCE)" as used herein generally refers to non-gradual evolution using phages as viral vectors. Examples of PANCE technology are described, for example, in Suzuki T. et al, Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase, Nat Chem Biol. 13(12):1261-1266 (2017), which is incorporated herein by reference in its entirety. Briefly, PANCE is a technique for rapid in vivo directed evolution using sequential flask transfers of evolutionary selection phages (SPs) containing the genes of interest to be evolved into whole fresh host cells (e.g., E. coli cells). Genes inside the host cells can be held constant while genes contained in the SPs are evolved incrementally. Following phage propagation, an aliquot of the infected cells can be used to transfect the next flask containing the host E. coli. This process can be repeated and / or continued until the desired phenotype has been evolved, for example, for as many transfers as desired.
[0888] Methods for applying PACE and PANCE to genetically modified polypeptides will be readily understood by those skilled in the art by reference, inter alia, to the aforementioned references. Further exemplary methods for directing the progressive evolution of a genomically modified protein or system, for example, in a population of host cells, using, for example, phage particles, can be applied to generate evolved variants of a genetically modified polypeptide or a fragment or subdomain thereof. Non-limiting examples of such methods are described in International PCT Application PCT / US2009 / 056194, filed September 8, 2009, and published March 11, 2010 as WO 2010 / 028347; International PCT Application PCT / US2009 / 056194, filed December 22, 2011, and published June 28, 2012 as WO 2012 / 088381; No. PCT / US2011 / 066747; U.S. Patent No. 9,023,594 issued May 5, 2015; U.S. Patent No. 9,771,574 issued September 26, 2017; U.S. Patent No. 9,394,537 issued July 19, 2016; International Publication No. WO 2015 / 134121 filed January 20, 2015 and published September 11, 2015. No. PCT / US2015 / 012022, published as International PCT Publication No. WO 2019 / 023680, filed on June 14, 2019, and published on January 31, 2019 as International PCT Publication No. PCT / US2019 / 37216, filed on April 15, 2016, and published on October 20, 2016 as International PCT Publication No. PCT / US2016 / 027795, filed on April 15, 2016, and published on October 20, 2016 as International PCT Publication No. WO 2016 / 168631, and International Application PCT / US2019 / 47996, filed on August 23, 2019, each of which is incorporated herein by reference in its entirety.
[0889] In some non-limiting exemplary embodiments, the method of evolving an evolved mutant genetically modified polypeptide, fragment or domain thereof includes (a) contacting a population of host cells with a population of viral vectors (starting genetically modified polypeptide or fragment or domain thereof) containing a gene of interest, where (1) the host cells are suitable for infection with the viral vector; (2) the host cells express viral genes necessary for the production of viral particles; (3) the expression of at least one viral gene necessary for the production of infectious viral particles is dependent on the function of the gene of interest; and / or (4) the viral vector allows expression of a protein in the host cells and can be replicated by the host cells and packaged into viral particles. In some embodiments, the method includes (b) contacting the host cells with a mutagen using a host cell containing a mutation that enhances the mutation rate (e.g., by delivering a mutant plasmid or any genomic modification such as a damaged DNA proofreading polymerase, SOS genes, e.g., UmuC, UmuD', and / or RecA (these mutations can be under the control of an inducible promoter when bound to the plasmid), or a combination thereof). In some embodiments, the method includes (c) incubating the population of host cells under conditions that allow for viral replication and viral particle production, where the host cells are removed from the population of host cells and fresh, uninfected host cells are introduced into the population of host cells, thus replenishing the population of host cells and forming a stream of host cells. In some embodiments, the cells are incubated under conditions that allow the gene of interest to acquire mutations. In some embodiments, the method further includes (d) isolating from the population of host cells a mutated version of the viral vector that encodes an evolved gene product (e.g., an evolved mutant genetically modified polypeptide or a fragment or domain thereof).
[0890] Those skilled in the art will appreciate the various features that can be used within the above framework. For example, in some embodiments, the viral vector or phage is a filamentous phage, such as an M13 phage, e.g., an M13 selection phage. In certain embodiments, the gene required for the production of infectious viral particles is the M13 gene III (gIII). In some embodiments, the phage may lack a functional gIII, but instead contains gI, gII, gIV, gV, gVI, gVII, gVIII, gIX, and gX. In some embodiments, the generation of infectious VSV particles contains the envelope protein VSV-G. In various embodiments, different retroviral vectors can be used, e.g., Murine Leukemia Virus vectors, or Lentiviral vectors. In some embodiments, the retroviral vector can be efficiently packaged, e.g., using the VSV-G envelope protein as a replacement for the native envelope protein of the virus.
[0891] In some embodiments, the host cells are incubated for a suitable number of viral life cycles, e.g., at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10000, or more consecutive viral life cycles, where an illustrative and non-limiting example for M13 phage is 10-20 minutes per viral life cycle. Similarly, conditions can be adjusted to adjust the residence time of the host cells in the population of host cells (e.g., about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 70, about 80, about 90, about 100, about 120, about 150, or about 180 minutes). The host cell population can be adjusted to adjust the density of host cells, or in some embodiments, the host cell density in the inflow, e.g., about 10 3 cells / ml, approximately 10 4 cells / ml, approximately 10 5 cells / ml, approximately 5-10 5 cells / ml, approximately 10 6 cells / ml, approximately 5-10 6 cells / ml, approximately 10 7 cells / ml, approximately 5-10 7 cells / ml, approximately 10 8 cells / ml, approximately 5-10 8 cells / ml, approximately 10 9 cells / ml, approximately 5-10 9 cells / ml, approximately 10 10 cells / ml, or approximately 5-10 10 The cells / ml can be partly controlled.
[0892] Intein In some embodiments, as described in more detail below, an intein-N (intN) domain may be fused to the N-terminal portion of a first domain of a genetically modified polypeptide described herein, and an intein-C (intC) domain may be fused to the C-terminal portion of a second domain of a genetically modified polypeptide described herein, linking the N-terminal portion to the C-terminal portion, thereby linking the first and second domains. In some embodiments, the first and second domains are each independently selected from a DNA binding domain, an RNA binding domain, an RT domain, and an endonuclease domain.
[0893] Inteins can exist, for example, as self-splicing protein introns (e.g., peptides) that link flanking N-terminal and C-terminal exteins (e.g., fragments to be linked). Inteins can, in some cases, comprise a fragment of a protein that can be automatically excised to link the remaining fragment (extein) with a peptide bond in a process known as protein splicing. Inteins are also referred to as "protein introns." The process of inteins being automatically excised to link the remaining portion of a protein is referred to herein as "protein splicing" or "intein-mediated protein splicing."
[0894] In some embodiments, the inteins of a precursor protein (an intein-containing protein before intein-mediated protein splicing) are derived from two genes. Such inteins are referred to herein as split inteins (e.g., split intein-N and split intein-C). Thus, an intein-based approach can be used to link a first polypeptide sequence and a second polypeptide sequence together. For example, in cyanobacteria, DnaE, the catalytic subunit of DNA polymerase III, is encoded by two separate genes, dnaE-n and dnaE-c. When an intein-N domain, such as that encoded by the dnaE-n gene, is located as part of a first polypeptide sequence, it can link the first polypeptide sequence with a second polypeptide sequence, where the second polypeptide sequence includes an intein-C domain, such as that encoded by the dnaE-c gene. Thus, in some embodiments, a protein may be made by providing nucleic acids encoding a first and a second polypeptide sequence (e.g., a first nucleic acid molecule encoding the first polypeptide sequence and a second nucleic acid molecule encoding the second polypeptide sequence), which are introduced into a cell under conditions that allow for the production of the first and second polypeptide sequences, and ligation of the first polypeptide sequence to the second polypeptide sequence via an intein-based mechanism.
[0895] The use of inteins to link heterologous protein fragments has been described, for example, in Wood et al., J. Biol. Chem. 289(21);14512-9(2014), which is incorporated herein by reference in its entirety. For example, when fused to separate protein fragments, the inteins IntN and IntC can recognize each other and splice from themselves and / or simultaneously ligate the flanking N- and C-terminal exteins of the protein fragments to which they are fused, thereby reconstituting a full-length protein from the two protein fragments.
[0896] In some embodiments, synthetic inteins based on dnaE intein, Cfa-N (e.g., split intein-N) and Cfa-C (e.g., split intein-C) intein pairs are used. Examples of such inteins are described, for example, in Stevens et al., J Am Chem Soc. 2016 Feb.24;138(7):2162-5, which is incorporated herein by reference in its entirety. Non-limiting examples of intein pairs that can be used according to the present disclosure include Cfa DnaE intein, Ssp GyrB intein, Ssp DnaX intein, Ter DnaE3 intein, Ter ThyX intein, Rma DnaB intein and Cne Prp8 intein (e.g., as described in U.S. Pat. No. 8,394,604, which is incorporated herein by reference).
[0897] In some embodiments involving split Cas9, the intein-N and intein-C domains may be fused to the N-terminal portion of split Cas9 and the C-terminal portion of split Cas9, respectively, for linking the N-terminal portion of split Cas9 with the C-terminal portion of split Cas9. For example, in some embodiments, intein-N is fused to the C-terminus of the N-terminal portion of split Cas9, i.e., forming the structure N-[N-terminal portion of split Cas9]-[intein-N]-C. In some embodiments, intein-C is fused to the N-terminus of the C-terminal portion of split Cas9, i.e., forming the structure N-[intein-C]-[C-terminal portion of split Cas9]-C. The mechanism of intein-mediated protein splicing for linking intein-linked proteins (e.g., split Cas9) is described in Shah et al., Chem Sci. 2014;5(1):446-461, which is incorporated herein by reference. Methods for designing and using inteins are known in the art and are described, for example, in WO2020051561, WO2014004336, WO2017132580, US20150344549, and US20180127780, each of which is incorporated herein by reference in its entirety.
[0898] In some embodiments, split refers to division into two or more fragments. In some embodiments, split Cas9 protein or split Cas9 comprises a Cas9 protein provided as an N-terminal fragment and a C-terminal fragment encoded by two separate nucleotide sequences. Polypeptides corresponding to the N-terminal and C-terminal portions of the Cas9 protein can be spliced to form a reconstituted Cas9 protein. In embodiments, the Cas9 protein is split into two fragments within a denatured region of the protein, for example, as described in Nishimasu et al., Cell, Volume 156, Issue 5, pp. 935-949, 2014, or as described in Jiang et al. (2016) Science 351:867-871 and PDB file: 5F9R, each of which is incorporated herein by reference in its entirety. The denatured region can be determined by one or more protein structure determination techniques known in the art, including, but not limited to, X-ray crystallography, NMR spectrosc...
Claims
1. A template RNA comprising, from 5' to 3': (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene, said gRNA spacer having a sequence comprising the core nucleotides of the nucleic acid sequence of SEQ ID NO: 20816 or a gRNA spacer sequence of Table 1, and optionally comprising one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of said gRNA spacer, or having a sequence of a spacer selected from Table 6A, 6B, X2, X3, X3a, X5, or XX; (ii) a gRNA scaffold that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide); (iii) a heterologous sequence of interest comprising a mutation region for introducing a mutation into a second portion of the human SERPINA1 gene (e.g., for correcting a mutation therein), wherein optionally, the heterologous sequence of interest comprises, from 5' to 3', a post-edited homology region, a mutation region, and a pre-edited homology region; and (iv) a primer binding site (PBS) sequence comprising at least 5, 6, 7, or 8 bases having 100% identity with the third portion of the human SERPINA1 gene; A template RNA comprising: (i) the heterologous target sequence comprises the core nucleotides of an RT template sequence from Table 3, and optionally comprises one or more contiguous nucleotides beginning at the 3' end of a flanking nucleotide of the RT template sequence, or the heterologous target sequence comprises the sequence of an RT template sequence from Table 6A or 6B; or (ii) The template RNA of claim 1, wherein the heterologous sequence of interest comprises the sequence of a heterologous sequence of interest from a template RNA listed in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto, or a sequence having one, two, or three substitutions thereto.
3. A template RNA comprising, from 5' to 3': (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene; (ii) a gRNA scaffold that binds to a genetically modified polypeptide (e.g., binds to a Cas domain of the genetically modified polypeptide); (iii) a heterologous target sequence comprising a mutation region for introducing a mutation into a second portion of the human SERPINA1 gene (e.g., for correcting a mutation therein), the heterologous target sequence comprising the core nucleotides of an RT template sequence of Table 3 and, optionally, one or more contiguous nucleotides starting from the 3′ end of a flanking nucleotide of the RT template sequence, or comprising an RT template sequence of Table 6A, 6B, X2, X3, X3a, X5, or XX; and (iv) a PBS sequence comprising at least 5, 6, 7, or 8 bases of 100% identity to a third portion of the human SERPINA1 gene. A template RNA comprising: (i) the PBS sequence comprises a sequence comprising the core nucleotide of a PBS sequence of Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, and optionally comprises one or more contiguous nucleotides starting from the 5' end of a flanking nucleotide of the PBS sequence, or the PBS sequence has a sequence comprising a PBS sequence of Table 6A or 6B corresponding to the RT template sequence, the gRNA spacer sequence, or both; or (ii) The template RNA of claim 2, wherein the PBS sequence comprises a PBS sequence from a template RNA listed in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto, or a sequence having one, two, or three substitutions thereto.
5. 2. The template RNA of Claim 1, wherein the gRNA scaffold comprises a sequence of a gRNA scaffold of Table 6A or 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
6. The template RNA of claim 1 , wherein the mutation is a K342E mutation in the SERPINA1 gene (e.g., to correct a pathogenic E342K mutation). (a) the mutated region comprises a first region (e.g., a first nucleotide) designed to correct a pathogenic mutation in the SERPINA1 gene and a second region (e.g., a second nucleotide) designed to inactivate a PAM sequence (e.g., a "PAM-kill" mutation as described in Table 5); (b) the template RNA includes one or more silent mutations (e.g., silent substitutions), e.g., as exemplified in Table 7B; (c) the mutation region comprises a first region designed to correct a pathogenic mutation in the SERPINA1 gene and a second region designed to introduce a silent substitution.
8. The template RNA of claim 1, comprising one or more chemically modified nucleotides.
9. 1. A genetic modification system comprising: (i) the template RNA of claim 1, and (ii) a genetically modified polypeptide or a nucleic acid (e.g., RNA) encoding the genetically modified polypeptide A genetically modified system comprising:
10. The genetically modified polypeptide is (a) a reverse transcriptase (RT) domain (e.g., an RT domain from a retrovirus or a polypeptide domain having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity thereto); (b) a Cas domain (e.g., a Cas9 domain) that binds to a target DNA molecule and is heterologous to the RT domain; and (c) optionally, a linker disposed between the RT domain and the Cas domain; and optionally (i) the RT domain is (1) an RT domain of Table 6; or (2) RT domains from murine leukemia virus (MMLV), porcine endogenous retrovirus (PERV); avian reticuloendotheliosis virus (AVIRE), feline leukemia virus (FLV), simian foamy virus (SFV) (e.g., SFV3L), bovine leukemia virus (BLV), Mason-Pfizer monkey virus (MPMV), human foamy virus (HFV), or bovine foamy / syncytial virus (BFV / BSV). Includes; (ii) the spacer comprises a spacer of Table XX, X5, 1, 6A or 6B, or a sequence having one, two or three substitutions therein, and the Cas domain comprises a Cas domain of the same row of Table XX, X5, 1, 6A or 6B, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity thereto; (iii) the Cas domain is (1) Contains a Cas domain of Table 7 or Table 8; (2) a Cas9 domain; and / or (3) SpyCas9 domain, SpCas9 domain, BlatCas9 domain, Nme2Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain; (iv) the linker comprises the sequence of a linker in Table 10 (e.g., any of SEQ ID NOs: 5217, 5106, 5190, and 5218); (v) the genetically modified polypeptide comprises one or two NLS sequences from Table 11 (e.g., any of SEQ ID NOs: 5245, 5290, 5323, 5330, 5349, 5350, 5351, and 4001); and / or (vi) The genetic modification system of claim 9, wherein the genetic modification system generates a first nick in a first strand of the human SERPINA1 gene, and optionally, the genetic modification system further comprises a second-strand-targeting gRNA spacer that directs a second nick in a second strand of the human SERPINA1 gene.
11. The template RNA is a template RNA sequence of Table 6A or 6B or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, or a template RNA sequence of Table X2, X3, or X3a or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. The genetic modification system of claim 9.
12. (a) (i) the second strand-targeting gRNA comprises a sequence comprising the core nucleotides of a left gRNA spacer sequence or a right gRNA spacer sequence from Table 2, and optionally one or more contiguous nucleotides starting from the 3′ end of a flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence; (ii) the gRNA targeting the second strand comprises a sequence comprising the core nucleotides of a left gRNA spacer sequence or a right gRNA spacer sequence from Table 2 corresponding to the gRNA spacer sequence of (i), and optionally one or more contiguous nucleotides starting from the 3′ end of a flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence; (iii) the gRNA targeting the second strand comprises a sequence comprising the core nucleotides of a second nicked gRNA sequence from Table 4 or a sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and optionally one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the second nicked gRNA sequence; (iv) the second strand-targeting gRNA comprises a sequence comprising the core nucleotides of a second nicked gRNA sequence from Table 4 corresponding to the gRNA spacer sequence of (i), or a sequence with at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and optionally one or more contiguous nucleotides starting from the 3' end of a flanking nucleotide of the second nicked gRNA sequence; (b) the gRNA targeting the second strand has a "PAM-in orientation" with the template RNA of the genetic modification system, e.g., as illustrated in Table 4; and / or (c) the second strand-targeting gRNA targets a sequence that overlaps with a target mutation in the template RNA, and optionally, the second strand-targeting gRNA (1) a sequence complementary to the SERPINA1 mutation (e.g., a spacer sequence); (2) a sequence complementary to the wild-type sequence at the target locus (e.g., a spacer sequence); (3) a sequence (e.g., a spacer sequence) complementary to a SNP proximal to the target locus, e.g., a SNP contained in the genomic DNA of a subject (e.g., a patient); (4) a sequence (e.g., a spacer sequence) that is complementary to or contains one or more silent substitutions proximal to the target locus. The genetic modification system of claim 10, comprising:
13. The heterologous sequence of interest comprises about 8 to 30, 9 to 25, 10 to 20, 11 to 16, or 12 to 15 (e.g., about 11 to 16) nucleotides. The template RNA according to any one of claims 1 to 8 or the genetic modification system according to any one of claims 9 to 12.
14. The template RNA according to any one of claims 1 to 8 or the genetic modification system according to any one of claims 9 to 12, wherein the PBS sequence comprises about 5 to 20, 8 to 16, 8 to 14, 8 to 13, 9 to 13, 9 to 12, or 10 to 12 (e.g., about 9 to 12) nucleotides.
15. A template RNA comprising the sequence of a template RNA of Table 4, 6A or 6B, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto.
16. 1. A genetic modification system comprising: A template RNA comprising a template RNA sequence of Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto; and A second nicked gRNA sequence from the same row as (i) in Table 4, a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity thereto. A genetically modified system comprising:
17. A DNA encoding the template RNA according to any one of claims 1 to 8 or 15, or the gene modification system according to any one of claims 9 to 12 or 16.
18. A pharmaceutical composition comprising the genetic modification system of any one of claims 9 to 12 or 16, or one or more nucleic acids encoding the same, and a pharmaceutically acceptable excipient or carrier, optionally wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle, and optionally wherein the viral vector is an adeno-associated virus.
19. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising the template RNA of any one of claims 1 to 8 or 15 or the genetic modification system of any one of claims 9 to 12 or 16.
20. A lipid nanoparticle (LNP) component comprising a template RNA described in any one of claims 1 to 8 or 15, or a genetic modification system described in any one of claims 9 to 12 or 16, or DNA encoding the same.
21. 10. A method for producing a template RNA according to claim 1, the method comprising synthesizing the template RNA by in vitro transcription or by introducing DNA encoding the template RNA into a host cell under conditions that allow for the production of the template RNA.
22. 10. An in vitro or ex vivo method for modifying a target site in a human SERPINA1 gene in a cell, the method comprising contacting the cell with the gene modification system of claim 9 or a DNA encoding the same, thereby modifying the target site in the human SERPINA1 gene in the cell, and optionally (a) the cells are derived from a subject with alpha-1 antitrypsin deficiency (AATD); and / or (b) the cell is a mammalian cell, such as a human cell.
23. 10. A pharmaceutical composition for use in treating a subject having a disease or condition associated with a mutation in the human SERPINA1 gene, the pharmaceutical composition comprising the genetic modification system of claim 9 or a DNA encoding the same, and optionally (a) the disease or condition is alpha-1 antitrypsin deficiency (AATD); (b) the subject has an E342K mutation; and / or (c) The pharmaceutical composition, wherein the subject is a human.
24. Use of the genetic modification system of claim 9 or DNA encoding same in the manufacture of a pharmaceutical for a subject having a disease or condition associated with a mutation in the human SERPINA1 gene, optionally comprising: (a) the disease or condition is alpha-1 antitrypsin deficiency (AATD); (b) the subject has an E342K mutation; and / or (c) The use, wherein the subject is a human.
25. 10. A pharmaceutical composition for use in treating a subject having AATD, the pharmaceutical composition comprising the genetic modification system of claim 9 or DNA encoding same, and optionally the subject is a human.
26. Use of the genetic modification system of claim 9 or DNA encoding it in the manufacture of a pharmaceutical for treating a subject having AATD, optionally wherein the subject is a human.
27. (a) introduction of the system into a target cell results in correction of a pathogenic mutation in the SERPINA1 gene, optionally wherein the pathogenic mutation is an E342K mutation and the correction comprises an amino acid substitution of K342E; (b) introducing the system into a target cell results in a mutation that restores function of the SERPINA1 gene, and optionally (i) the mutation correction occurs in at least 30% (e.g., 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acids; and / or (ii) correction of the mutation occurs in at least 30% (e.g., 30%, 40%, 50%, 60%, 70% or more) of the target cells; (c) the genetic modification system comprises a gRNA targeting the second strand, and correction of the mutation in the population of target cells is increased relative to a population of target cells treated with a genetic modification system comprising a template RNA without a gRNA targeting the second strand; and / or (d) the template RNA comprises one or more silent substitutions (e.g., as exemplified in Table 7B), and correction of the mutation in a population of target cells is increased relative to a population of target cells treated with the gene modification system comprising a template RNA that does not comprise one or more silent substitutions.