SERPINA modifiable composition and method
A gene modification system using a reverse transcriptase and St1Cas9 domain with a mutant gRNA scaffold corrects SERPINA1 gene mutations, effectively treating alpha-1 antitrypsin deficiency by restoring AAT levels and activity, improving outcomes for both liver and lung diseases.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2026-03-19
AI Technical Summary
Current methods for treating alpha-1 antitrypsin deficiency (AATD) are inadequate, particularly in addressing liver disease and lung disease manifestations, with augmentation therapy being insufficient and unable to restore normal physiological regulation of AAT levels, and there is a need for more effective genome editing techniques to correct the SERPINA1 gene mutations causing the deficiency.
A gene modification system comprising a reverse transcriptase (RT) domain and an St1Cas9 domain, along with a template RNA containing a mutant gRNA scaffold, is used to target and correct the SERPINA1 gene mutations, enabling insertion, deletion, or substitution of sequences to modulate AAT activity, delivered via plasmid vectors, viral vectors, or lipid nanoparticles.
This system effectively corrects SERPINA1 gene mutations, potentially restoring normal AAT levels and activity, thereby addressing both liver and lung disease associated with AATD, providing a more effective treatment than existing therapies.
Smart Images

Figure 2026509461000342 
Figure 2026509461000343 
Figure 2026509461000344
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the benefits of U.S. Provisional Patent Application No. 63 / 490,462, filed on 15 March 2023. The contents of the aforementioned application are incorporated herein by reference in whole.
[0002] Sequence List This application is submitted electronically in XML format in accordance with WIPO standard ST.26, and includes a sequence listing incorporated herein by reference in whole. The XML copy prepared on 12 December 2023 is named V2065-704901_SL.XML and has a size of 15,726,626 bytes. [Background technology]
[0003] The insertion of target nucleic acids into the genome occurs infrequently and with little site specificity, in the absence of specialized proteins to facilitate insertion events. Some existing methods, such as CRISPR / Cas9, are better suited to small edits dependent on host repair pathways and are less effective for inserting longer sequences. Other existing methods, such as Cre / loxP, require a first step of inserting a loxP site into the genome, followed by a second step of inserting the target sequence into the loxP site. In this art, there is a need for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, modifying, or deleting target sequences within the genome.
[0004] AATD is characterized by low circulating AAT levels. AAT is primarily produced and secreted into the bloodstream by hepatocytes, but it is also produced by other types of cells, including lung epithelial cells and certain types of leukocytes. AAT inhibits several serine proteases secreted by inflammatory cells (most notably neutrophil elastase [NE], proteinase 3, and cathepsin G), thus protecting organs such as the lungs from protease-induced damage, especially during the inflammatory phase.
[0005] The two most common clinical variants of AAT are the E264V(PiS) and E342K(PiZ) alleles. The clinical single nucleotide variant E342K(PiZ) leads to a structurally unstable and / or inactive AAT protein, resulting in hepatotoxicity and inactive lung disease. The inheritance is autosomal codominant. More than half of AATD patients harbor at least one copy of the E342K mutation.
[0006] The most frequently associated mutation with AATD is the substitution of glutamate to lysine (E342K) in the SERPINA1 gene encoding the AAT protein. This E342K mutation is located at the hinge between the β-sheet and the reaction center loop (RCL) of the AAT protein, resulting in the formation of a loop-sheet dimer, which can later elongate to form a long-chain loop-sheet polymer, causing the AAT-Z protein to aggregate inside the rough endoplasmic reticulum (rER) of hepatocytes during biosynthesis. This mutation, known as the Z mutation or Z allele, leads to incorrect folding of the translated protein, and therefore it is not secreted into the bloodstream, resulting in significantly reduced circulating AAT levels in individuals homozygous for the Z allele (PiZZ). Only about 15% of mutant Z-AAT proteins fold correctly and are secreted by cells. A further consequence of the Z mutation is that secreted Z-AAT exhibits reduced activity compared to wild-type protein, reaching 40% to 80% of normal antiprotease activity (American thoracic society / European respiratory society, Am J Respir Crit Care Med. 2003;168(7):818-900 and Ogushi et al. J Clin Invest. 1987;80(5):1366-74).
[0007] There are two disease phenotypes associated with the PiZZ genotype. Accumulation of polymerized Z-AAT protein in hepatocytes leads to gain-of-function cytotoxicity, resulting in cellular stress, inflammation, fibrosis, cirrhosis, and hepatocellular carcinoma (HCC) and neonatal liver disease in 12% of patients. This accumulation may resolve spontaneously, but can be fatal in a small number of children. The loss-of-function phenotype arises from decreased systemic AAT levels, leading to increased protease digestion of connective tissue in the lower respiratory tract. Excessive protease digestion of connective tissue and alveolar lining impairs lung elasticity and function, leading to emphysema, a prominent feature of chronic obstructive pulmonary disease (COPD). This effect is severe in individuals with PiZZ and typically manifests in middle age, resulting in a reduced quality of life and shortened lifespan (average 68 years) (Tanash et al. Int J Chron Obstruct Pulm Dis. 2016;11:1663-9). This effect is more pronounced in PiZZ individuals who smoke, resulting in an even shorter lifespan (58 years). Piitulainen and Tanash, COPD 2015;12(1):36-41. PiZZ individuals constitute the majority of people with clinically relevant AATD lung disease. A milder form of AATD is associated with the SZ genotype, in which the Z-allele is combined with the S-allele. The S-allele is associated with a somewhat reduced circulating AAT level but does not cause cytotoxicity in hepatocytes. The result is a clinically significant lung disease, but not a liver disease. Fregonese and Stolk, Orphanet J Rare Dis. 2008;33:16. Similar to the ZZ genotype, in subjects with the SZ genotype, the deficiency of circulating AAT leads to upregulation of protease activity, causing lung tissue to weaken over time, and emphysema may develop, especially in smokers. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] American thoracic society / European respiratory society,Am J Respir Crit Care Med.2003;168(7):818-900 [Non-Patent Document 2] Ogushi et al.J Clin Invest.1987;80(5):1366-74 [Non-Patent Document 3] Tanash et al.Int J Chron Obstruct Pulm Dis.2016;11:1663-9 [Non-Patent Document 4] Piitulainen and Tanash, COPD 2015;12(1):36-41 [Non-Patent Document 5] Fregonese and Stolk,Orphanet J Rare Dis.2008;33:16 [Overview of the project] [Problems that the invention aims to solve]
[0009] While limited treatment options exist for AATD, there is currently no cure. A small number of neonatal patients and patients in the advanced stages of liver disease receive liver transplants. The current standard treatment for individuals with AAT deficiency who have or show signs of significant lung disease is augmentation therapy or protein replacement therapy. Augmentation therapy involves the administration of human AAT protein concentrate purified from pooled donor plasma to augment the deficient AAT. This treatment involves weekly infusions of AAT protein purified from healthy donor blood. While plasma protein infusions have been shown to improve survival rates or slow the rate of emphysema progression, augmentation therapy is often insufficient under challenging conditions (e.g., active lung infection). Augmentation therapy also fails to restore the normal physiological regulation of the patient's AAT, making it difficult to demonstrate its effectiveness. In addition, augmentation therapy cannot address the liver disease resulting from the acquisition of Z allele toxicity. Therefore, novel and more effective treatments for AATD are needed. [Means for solving the problem]
[0010] This disclosure relates to novel compositions, systems, and methods for modifying the genome at one or more locations in a host cell, tissue, or subject in vivo or in vitro. This disclosure provides a gene modification system comprising, for example, a gene modification polypeptide comprising a reverse transcriptase (RT) domain and an St1Cas9 domain, and a template RNA comprising a mutant gRNA scaffold that has been engineered to improve performance when used in cooperation with, for example, the St1Cas9 domain. This disclosure also provides gene modification systems having the ability to modulate α-1 antitrypsin (AAT) activity (e.g., by inserting, modifying, or deleting a sequence of interest), and a method for treating α-1 antitrypsin deficiency (AATD) by modifying a SERPINA1 PiZ mutation that causes α-1 antitrypsin deficiency by administering one or more such systems to modify the genomic sequence at a single nucleotide.
[0011] In one embodiment, the disclosure relates to a system for modifying DNA to correct a human SERPINA1 gene mutation that causes AATD, comprising: (a) a nucleic acid encoding a gene modification polypeptide having the ability of targeted priming reverse transcription, wherein the polypeptide comprises (i) a reverse transcriptase domain and (ii) an endonuclease-active St1Cas9 niccasse that binds to DNA; and (b) a template RNA comprising (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene, (ii) a gRNA scaffold that binds to the polypeptide, (iii) a heterologous target sequence containing a mutation region for correcting the mutation, and (iv) a primer-binding site (PBS) sequence containing at least 3, 4, 5, 6, 7, or 8 bases that are 100% homologous to the target DNA strand at the 3' end of the template RNA. The SERPINA1 gene may contain the E342K mutation (also known as the PiZ mutation). The template RNA sequence may include sequences listed in Table 1, Table 3, Table 4, Table 5, Table 6a, Table 6B, Table X3, or Table X3a of this specification.
[0012] The gRNA spacer may contain at least 15 bases with 100% homology to the target DNA at the 5' end of the template RNA. The template RNA may further contain a PBS sequence containing at least 5 bases with at least 80% homology to the target DNA strand. The template RNA may contain one or more chemical modifications.
[0013] The domains of a genetically modified polypeptide may be linked by a peptide linker. A polypeptide may contain one or more peptide linkers. A genetically modified polypeptide may further contain nuclear localization signals. A polypeptide may contain two or more nuclear localization signals, for example, multiple adjacent nuclear localization signals, or one or more nuclear localization signals in different regions of the polypeptide, for example, one or more nuclear localization signals at the N-terminus of the polypeptide and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding a genetically modified polypeptide may encode one or more intein domains.
[0014] The introduction of the system into target cells may result in insertions of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs into exogenous DNA. The introduction of the system into target cells may result in deletions of less than 2, 3, 4, 5, 10, 50, or 100 base pairs into the genomic DNA upstream or downstream of the insertion. The introduction of the system into target cells may result in substitutions, such as substitutions of 1, 2, or 3 nucleotides, such as substitutions of consecutive nucleotides.
[0015] The heterologous target sequence may consist of at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600, or 700 base pairs.
[0016] In one embodiment, the disclosure relates to a pharmaceutical composition comprising the above system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles. In one embodiment, the disclosure relates to a pharmaceutical composition comprising the above system and a plurality of pharmaceutically acceptable excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles, for example, the above system is delivered by two different excipients or carriers, for example, two lipid nanoparticles and two viral vectors or one lipid nanoparticle and one viral vector. The viral vector may be an adeno-associated virus (AAV).
[0017] In one embodiment, the present disclosure relates to a host cell (e.g., a mammalian cell, e.g., a human cell) including the above-described system.
[0018] In one embodiment, the disclosure relates to a method for correcting mutations in the human SERPINA1 gene of a cell, tissue, or subject, comprising administering the system described above to the cell, tissue, or subject, wherein, optionally, the correction of the mutant SERPINA1 gene includes an amino acid substitution of K342E (reverting to the pathogenic substitution K342K). The system may be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid of (a) may be incorporated into the genome of a host cell. In some embodiments, the nucleic acid of (a) is not incorporated into the genome of a host cell. In some embodiments, the heterologous target sequence is inserted into only one target site in the host cell genome. The heterologous target sequence may be inserted into two or more target sites in the host cell genome, for example, the same corresponding site on two homologous chromosomes or two different sites on the same or different chromosomes. The heterologous target sequence may encode a mammalian polypeptide or a fragment or variant thereof. The components of the system may be delivered on one, two, three, four, or more different nucleic acid molecules. The system can be introduced into host cells by electroporation or by using at least one vehicle selected from plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.
[0019] The features of this composition or method may include one or more of the following listed embodiments.
[0020] Enumerated embodiments 1. A template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having a partial or complete deletion of stem loop 2, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0021] 2. The deletion is the template RNA according to Embodiment 1, having a deletion length of 1 to 32 nucleotides (e.g., 2 to 29, 2 to 20, 2 to 10, or 10 to 20).
[0022] 3. The deletion is the entirety of stem-loop 2, the template RNA as described in Embodiment 1.
[0023] 4. The deletion is at positions 55 to 84, the template RNA described in Embodiment 1.
[0024] 5. The St1Cas9 scaffold is the template RNA according to Embodiment 1, comprising a deletion of a portion of the second single-stranded region (e.g., nucleotides 1, 2, 3, or 4 at the 3' end of the single-stranded region).
[0025] 6. The mutant St1Cas9 scaffold is the template RNA described in any one of the preceding embodiments, having either or both an elongated RAR upper stem or a substitution that generates a GC base pair in the RAR upper stem.
[0026] 7. The mutant St1Cas9 scaffold is the template RNA described in any one of the preceding embodiments, having a mutation in the tetraloop.
[0027] 8. The St1Cas9 scaffold comprises an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion has a sequence relating to GACUUCGGUC, as described in any one of the preceding embodiments of the template RNA.
[0028] 9. The St1Cas9 scaffold comprises an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion has a sequence relating to CUAGAAAUAG, as described in any one of the preceding embodiments of the template RNA.
[0029] 10. The St1Cas9 scaffold comprises an insertion (e.g., 12 nucleotides) at positions 14–19 and a deletion at positions 15–18, and optionally the insertion has a sequence relating to CGCGGUAACGCG, as described in any one of the preceding embodiments of the template RNA.
[0030] 11. The mutant St1Cas9 scaffold has a substitution that generates a GC base pair in the lower RAR stem, optionally comprising a substitution by G at position 4, and the template RNA as described in any one of the prior embodiments, further comprising a substitution by C at position 31.
[0031] 12. A template RNA according to any one of the preceding embodiments, comprising a substitution in the second single-stranded region, wherein the substitution is optionally a U substitution at position 51 or a C substitution at position 54.
[0032] 13. Template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having an elongated RAR upper stem or one or both of the substitutions that produce a GC base pair on the RAR upper stem, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0033] 14. The St1Cas9 scaffold is the template RNA according to Embodiment 13, comprising an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion having a sequence relating to GACUUCGGUC.
[0034] 15. The St1Cas9 scaffold comprises an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion has a sequence relating to CUAGAAAUAG, as described in Embodiment 13.
[0035] 16. The St1Cas9 scaffold is the template RNA according to Embodiment 13, comprising an insertion (e.g., 12 nucleotides) at positions 14-19 and a deletion at positions 15-18, wherein optionally the insertion has a sequence relating to CGCGGUAACGCG.
[0036] 17. The template RNA described in any one of the preceding embodiments, wherein the RAR upper stem is extended by 1 to 8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs) compared to the wild-type sequence of Sequence ID No. 25999.
[0037] 18. The template RNA according to Embodiment 17, wherein at least 50%, 60%, 70%, 80%, or 90% of the base pairs that are novel compared to SEQ ID NO: 25999 are GC base pairs.
[0038] 19. Template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having a substitution that generates a GC base pair in the lower RAR stem, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0039] The template RNA according to Embodiment 19, comprising a substitution with G at position 20.4 and a substitution with C at position 31.
[0040] 21. Template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold with a mutation in the tetraloop, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0041] 22. A template RNA described in any one of the preceding embodiments, in which one or more nucleotides in the tetraloop are substituted.
[0042] 23. A tetraloop template RNA according to any one of the prior embodiments, comprising a sequence selected from AACA, AAUA, ACCA, ACUA, AGUA, AGCA, AUCA, AUUA, CAAC, CUCG, CUUG, GAAA, GAGA, GCAA, GCGA, GGAA, GGAG, GGGA, GUAA, GUGA, UAAC, UACG, UCAC, UCCG, UGAA, UGAC, UGCG, UUAC, or UUCG.
[0043] 24. The tetraloop is, for example, the template RNA described in any one of the preceding embodiments, which extends up to 5 nucleotides.
[0044] 25. The template RNA according to Embodiment 24, wherein the elongated tetraloop contains a sequence selected from GAAGA or GACAA.
[0045] 26. The St1Cas9 scaffold comprises an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion has a sequence relating to GACUUCGGUC, as described in any one of embodiments 21-25.
[0046] 27. The St1Cas9 scaffold comprises an insertion (e.g., 10 nucleotides) at positions 15-18 and deletions at positions 16 and 17, and optionally the insertion has a sequence relating to CUAGAAAUAG, as described in any one of embodiments 21-25.
[0047] 28. The St1Cas9 scaffold comprises an insertion (e.g., 12 nucleotides) at positions 14-19 and a deletion at positions 15-18, and optionally the insertion has a sequence relating to CGCGGUAACGCG, as described in any one of embodiments 21-25.
[0048] 29. Template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having a substitution in the second single-stranded region, (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0049] The template RNA according to Embodiment 29, including a substitution with U at position 30.51.
[0050] The template RNA according to Embodiment 29, including a substitution with C at position 31.54.
[0051] 32. The mutant gRNA scaffold is a template RNA according to any one of the preceding embodiments, comprising the sequence shown in Table 23 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0052] 33. A template RNA according to any one of the preceding embodiments, comprising a sequence relating to any of Tables 20, 21, 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15, or E16, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0053] 34. A template RNA according to any one of the prior embodiments, comprising the sequence relating to Sequence ID No. 27131 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0054] 35. A template RNA according to any one of the prior embodiments, comprising the sequence relating to Sequence ID No. 27132 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0055] 36. A template RNA according to any one of the prior embodiments, comprising the sequence relating to Sequence ID No. 27133 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0056] 37. A template RNA according to any one of the prior embodiments, comprising the sequence relating to Sequence ID No. 27134 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0057] 38. The mutant St1Cas9 scaffold is a template RNA described in any one of the preceding embodiments, having a length of 50-60, 60-70, 70-80, or 80-84 nucleotides.
[0058] 39. Template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant gRNA scaffold containing the sequence shown in Table 23 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto. (3) heterogeneous sequences, and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing [the specified RNA].
[0059] 40. A template RNA, for example, from 5' to 3', (i) gRNA spacers complementary to the first portion of the human SERPINA1 gene, having a sequence containing the core nucleotide of the gRNA spacer sequence in Table 1 or a sequence having one, two, or three substitutions thereto, and optionally containing one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the gRNA spacer (e.g., containing one or more flanking nucleotides adjacent to the core nucleotide), or gRNA spacers having the sequence of the gRNA spacer in Table 6A, Table 6B, Table X3, or Table X3a or a sequence having one, two, or three substitutions thereto. (ii) gRNA scaffolds that bind to gene-modified polypeptides (for example, by binding to the Cas domain of gene-modified polypeptides), (iii) a heterologous target sequence containing a mutation region for introducing a mutation into the second portion of the human SERPINA1 gene (for example, for correcting an existing mutation therein) (optionally, the heterologous target sequence includes a post-edited homology region, a mutation region and a pre-edited homology region from 5' to 3'), and (iv) A primer binding site (PBS) sequence containing at least 3, 4, 5, 6, 7, or 8 bases that is 100% identical to the third portion of the human SERPINA1 gene. Template RNA containing this RNA.
[0060] 41. The heterologous target sequence comprises a core nucleotide of the RT template sequence from Table 3 or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the RT template sequence, or the heterologous target sequence comprises a sequence of the RT template sequence from Table 6A or Table 6B, as described in any one of the preceding embodiments of the template RNA.
[0061] 42. The heterologous target sequence comprises a core nucleotide of the RT template sequence in Table 3 corresponding to the gRNA spacer sequence, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the RT template sequence (e.g., one or more flanking nucleotides adjacent to the core nucleotide), or the heterologous target sequence comprises a template RNA as described in any one of the preceding embodiments, wherein the heterologous target sequence comprises a sequence of the RT template sequence from Table 6A or Table 6B.
[0062] 43. The heterologous target sequence is a template RNA according to any one of the preceding embodiments, having a heterologous target sequence sequence from the template RNA shown in Table X3 or Table X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto, or a sequence having one, two, or three substitutions thereto.
[0063] 44. The heterologous target sequence is a template RNA described in any one of the preceding embodiments, having a length of 6 to 16 nucleotides (e.g., 6, 8, 10, 12, 14, 15, or 16 nucleotides).
[0064] 45. The PBS sequence is a template RNA according to any one of the preceding embodiments, having a sequence containing the core nucleotide of the PBS sequence in the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally containing one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence (e.g., one or more flanking nucleotides adjacent to the core nucleotide).
[0065] 46. The PBS sequence is a template RNA according to any one of Embodiments 1 to 44, comprising a sequence containing the core nucleotide of the PBS sequence corresponding to the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, a gRNA spacer sequence, or both, and optionally comprising one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence, or the PBS sequence is a sequence containing the PBS sequence corresponding to the RT template sequence in Table 6A or Table 6B, or a sequence having one, two, or three substitutions thereto, a gRNA spacer sequence, or both.
[0066] 47. The PBS sequence is the template RNA according to any one of the preceding embodiments, having a PBS sequence from the template RNA shown in Table X3 or Table X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto, or a sequence having one, two, or three substitutions thereto.
[0067] 48. The PBS sequence is a template RNA described in any one of the preceding embodiments, having a length of 8 to 12 nucleotides (e.g., 8, 9, 10, 11, or 12 nucleotides).
[0068] 49. The gRNA scaffold is a template RNA according to any one of the prior embodiments, comprising the sequence of the gRNA scaffold in Table 12 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0069] 50. The gRNA scaffold is a template RNA according to any one of Embodiments 1 to 48, comprising a sequence of a gRNA scaffold in Table 12 corresponding to an RT template sequence, a gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0070] 51. The gRNA scaffold is a template RNA according to any one of the prior embodiments, having a sequence of a gRNA scaffold from the template RNA shown in Table X3 or Table X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0071] 52. A template RNA according to any one of the prior embodiments, comprising a sequence of template RNA shown in Table X3 or Table X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0072] 53. A template RNA, for example, from 5' to 3', (i) A gRNA spacer complementary to the first part of the human SERPINA1 gene, (ii) gRNA scaffolds that bind to gene-modified polypeptides (for example, by binding to the Cas domain of gene-modified polypeptides), (iii) heterologous target sequences comprising a mutant region for introducing a mutation into the second portion of the human SERPINA1 gene (for example, for correcting an existing mutation), comprising a core nucleotide of the RT template sequence in Table 3 or a sequence having one, two, or three substitutions thereto, and optionally comprising one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the RT template sequence, or heterologous target sequences comprising the RT template sequence in Table 6A or Table 6B, and (iv) A PBS sequence containing at least 3, 4, 5, 6, 7, or 8 bases that is 100% identical to the third portion of the human SERPINA1 gene. Template RNA containing this RNA.
[0073] 54. The gRNA spacer comprises the core nucleotide of the gRNA spacer sequence in Table 1 or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the gRNA spacer sequence, or the gRNA spacer comprises the template RNA described in any one of the preceding embodiments, wherein the gRNA spacer comprises the gRNA spacer sequence in Table 6A or Table 6B.
[0074] 55. The heterologous target sequence comprises a core nucleotide of a gRNA spacer sequence corresponding to the RT template sequence in Table 1, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the gRNA spacer sequence, or the heterologous target sequence comprises a nucleotide of the gRNA spacer sequence in Table 6A or Table 6B, as described in any one of the preceding embodiments of the template RNA.
[0075] 56. The PBS sequence is a template RNA according to any one of the preceding embodiments, having a sequence containing the core nucleotide of the PBS sequence from the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally containing one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0076] 57. The template RNA according to any one of Embodiments 1 to 55, wherein the PBS sequence has a sequence containing the core nucleotide of the PBS sequence in Table 3 corresponding to the RT template sequence, the gRNA spacer sequence, or both, or a sequence having one, two, or three substitutions thereto, and optionally contains one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence, or the PBS sequence has a sequence containing the PBS sequence in Table 6A or 6B corresponding to the RT template sequence, the gRNA spacer sequence, or both.
[0077] 58. The gRNA scaffold is a template RNA according to any one of Embodiments 1 to 57, comprising a sequence of the gRNA scaffold in Table 6A or Table 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0078] 59. The gRNA scaffold is a template RNA according to any one of Embodiments 1 to 57, comprising a sequence of a gRNA scaffold from Table 6A or Table 12 corresponding to an RT template sequence, a gRNA spacer sequence, or both, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0079] 60. (iii) Heteropurpose sequences comprising a mutant region for introducing a mutation into the second portion of the human SERPINA1 gene, comprising a core nucleotide of the RT template sequence in Table 3 or a sequence having one, two, or three substitutions thereto, and optionally comprising one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the RT template sequence, and (iv) template RNA comprising a PBS sequence containing at least 5, 6, 7, or 8 bases that are 100% homologous to the third portion of the human SERPINA1 gene.
[0080] 61. The PBS sequence is a template RNA according to any one of the preceding embodiments, having a sequence containing the core nucleotide of the PBS sequence in the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally containing one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0081] 62. The PBS sequence is a template RNA according to any one of Embodiments 1 to 60, having a sequence containing the core nucleotide of the PBS sequence corresponding to the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally containing one or more consecutive nucleotides starting from the 5' end of the flanking nucleotide of the PBS sequence.
[0082] 63. The mutation introduced by the system is a K342E mutation in the SERPINA1 gene (for example, to correct a pathogenic E342K mutation) as described in any one of the preceding embodiments of the template RNA.
[0083] 64. The pre-edited sequence is a template RNA described in any one of the preceding embodiments, comprising a length of approximately 1 to 35 nucleotides (for example, a length of approximately 1 to 5, 5 to 10, 10 to 15, 15 to 20, 20 to 25, 25 to 30, or 30 to 35 nucleotides).
[0084] 65. The mutated region is a template RNA described in any one of the preceding embodiments, comprising a single nucleotide.
[0085] 66. The mutated region comprises at least 2 nucleotides in length, the template RNA according to any one of embodiments 1 to 64.
[0086] 67. The mutated region is up to 32 nucleotides long (e.g., up to 5, 10, 15, 20, 25, 30, or 32) and comprises one, two, or three sequence differences compared to the second portion of the human SERPINA1 gene, as described in any one of the preceding embodiments of the template RNA.
[0087] 68. The mutated region comprises two sequence differences compared to the second portion of the human SERPINA1 gene, as described in any one of the preceding embodiments of the template RNA.
[0088] 69. The template RNA described in any one of the preceding embodiments, wherein the mutant region comprises a first region (e.g., a first nucleotide) designed to correct a pathogenic mutation in the SERPINA1 gene and a second region (e.g., a second nucleotide) designed to inactivate the PAM sequence (e.g., a “PAM-kill” mutation as described in Table 5).
[0089] 70. The mutated region contains 80%, 70%, 60%, 50%, 40%, or less than 30% identity with the corresponding portion of the human SERPINA1 gene, as described in any one of the prior embodiments.
[0090] 71. The template RNA is the template RNA described in any one of the preceding embodiments, comprising, for example, one or more silent mutations (e.g., silent substitutions) as illustrated in Table 7B.
[0091] 72. The mutated region comprises a template RNA according to any one of the prior embodiments, comprising a first region designed to correct a pathogenic mutation in the SERPINA1 gene and a second region designed to introduce a silent substitution.
[0092] 73. A template RNA according to any one of the prior embodiments, comprising one or more chemically modified nucleotides.
[0093] 74. A gene modification system, The template RNA described in any one of the prior embodiments, and Genetically modified polypeptides or nucleic acids (e.g., RNA) that encode genetically modified polypeptides. A gene modification system that includes this.
[0094] 75. Genetically modified polypeptides are Reverse transcriptase (RT) domains (for example, retrovirus-derived RT domains or polypeptide domains having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity thereto), and A Cas domain (e.g., Cas9 domain) that binds to the target DNA molecule and is heterogeneous to the RT domain, and Optional linker placed between the RT domain and the Cas domain. A gene modification system according to Embodiment 74, including the above.
[0095] 76.RT domain is, (a) RT domains in Table 6, or (b) RT domains derived from mouse leukemia virus (MMLV), porcine endogenous retrovirus (PERV), reticuloendotheliosis virus of birds (AVIRE), feline leukemia virus (FLV), monkey foam virus (SFV) (e.g., SFV3L), bovine leukemia virus (BLV), Mason-Pfizer monkey virus (MPMV), human foam virus (HFV), or bovine foam virus / syntitis virus (BFV / BSV). A gene modification system according to Embodiment 75, including the above.
[0096] 77. The gene modification system according to Embodiment 75 or 76, wherein the Cas domain includes the Cas domain of Table X1 or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity thereto.
[0097] 78. The gene modification system according to any one of embodiments 75 to 77, comprising a Cas9 nickase domain and an St1Cas9 nickase domain.
[0098] 79. A gene modification system according to any one of embodiments 75 to 78, wherein the gene-modified polypeptide comprises a first NLS, an St1Cas9 niccasse domain, a linker, an RT domain, and a second NLS, oriented from the N-terminal to the C-terminal.
[0099] 80. The gene modification system according to Embodiment 79, wherein the first NLS contains the sequence of SEQ ID NO: 11,095, and the second NLS contains the sequence of SEQ ID NO: 11,099, or both.
[0100] 81. A gene modification system according to any one of embodiments 75 to 80, wherein the linker includes the sequence relating to sequence number 5006.
[0101] 82. A gene modification system according to any one of Embodiments 75 to 81, wherein the spacer comprises a spacer from Table 6A or a sequence having one, two, or three substitutions therefor, and the Cas domain comprises a Cas domain from the same row in Table 6A or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity therefor.
[0102] 83. A gene modification system according to any one of Embodiments 75 to 82, wherein the spacer includes the spacer shown in Table 6A, and the Cas domain includes the Cas domain in the same row of Table 6A.
[0103] 84. A gene modification system according to any one of embodiments 75 to 83, wherein the spacer comprises a spacer from Table 6B or a sequence having one, two, or three substitutions therefor, and the Cas domain comprises a Cas domain from the same row in Table 6B or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity therefor.
[0104] 85. A gene modification system according to any one of Embodiments 75 to 84, wherein the spacer includes the spacer shown in Table 6B, and the Cas domain includes the Cas domain in the same row of Table 6B.
[0105] 86. A gene modification system according to any one of Embodiments 75 to 85, comprising a Cas domain as shown in Table 7 or Table 8.
[0106] The 87.Cas domain is, (a) Cas9 domain, (b) SpCas9 domain, BlatCas9 domain, Nme2Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain or St1Cas9 domain, and / or (c) A gene modification system according to any one of embodiments 75 to 86, wherein the Cas9 domain includes an N670A mutation, an N611A mutation, an N605A mutation, an N580A mutation, an N588A mutation, an N872A mutation, an N863 mutation, an N622A mutation, or an H840A mutation.
[0107] 88. The gene modification system according to Embodiment 87, wherein the Cas9 domain binds to a PAM sequence listed in Table 7 or Table 12.
[0108] 89. The gene modification system according to Embodiment 88, wherein the second portion of the human SERPINA1 gene overlaps with the PAM recognized by the Cas domain, for example, the second portion of the human SERPINA1 gene is within the range of the PAM, or the PAM is within the range of the second portion of the human SERPINA1 gene.
[0109] 90. A gene modification system according to any one of Embodiments 75 to 89, wherein the gRNA spacer is a gRNA spacer according to Table 1, and the Cas domain includes a Cas domain listed in the same row of Table 1 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0110] 91. A gene modification system according to any one of Embodiments 75 to 90, wherein the template RNA includes a sequence of the template RNA sequence in Table 6A or Table 6B, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0111] 92. (a) The template RNA contains the sequence of the template RNA sequence shown in Table 3. (b) The Cas domain includes the Cas domains in Table 7 or Table 8, (c) The linker includes a linker sequence from Table 10 (for example, any of sequence numbers 5217, 5106, 5190 and 5218), and (d) The gene modification system according to any one of Embodiments 75 to 91, wherein the gene modification polypeptide comprises one or two NLS sequences from Table 11 (for example, any of SEQ ID NOs. 5245, 5290, 5323, 5330, 5349, 5350, 5351 and 4001).
[0112] 93. A gene modification system according to any one of embodiments 75 to 92, which creates a first nick in the first strand of the human SERPINA1 gene.
[0113] 94. The gene modification system according to Embodiment 93, further comprising a second strand targeting gRNA spacer that directs a second nick to the second strand of the human SERPINA1 gene.
[0114] 95. The gene modification system according to Embodiment 94, wherein the second strand targeting gRNA comprises a sequence containing the core nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence from Table 2, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence.
[0115] 96. The gene modification system according to Embodiment 94, wherein the second strand targeting gRNA comprises a sequence containing the core nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence from Table 2 corresponding to the gRNA spacer sequence of (i), and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the left gRNA spacer sequence or the right gRNA spacer sequence.
[0116] 97. The gene modification system according to Embodiment 94, wherein the second strand targeting gRNA comprises a sequence containing the core nucleotide of the second nick gRNA sequence from Table 4 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the second nick gRNA sequence.
[0117] 98. The gene modification system according to Embodiment 94, wherein the second strand targeting gRNA comprises a sequence containing the core nucleotide of the second nick gRNA sequence from Table 4 corresponding to the gRNA spacer sequence of (i), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotide of the second nick gRNA sequence.
[0118] 99. The gene modification system according to any one of the preceding embodiments, wherein the second strand-targeting gRNA has a "PAM-in orientation" with respect to the template RNA of the gene modification system, as illustrated in Table 4, for example.
[0119] 100. A gene modification system according to any one of the prior embodiments, wherein the second strand-targeted gRNA targets a sequence that overlaps with a target mutation in the template RNA.
[0120] 101. The second strand-targeting gRNA is, (i) A sequence complementary to the SERPINA1 mutation (e.g., a spacer sequence), (ii) A sequence complementary to the wild-type sequence at the target locus (e.g., a spacer sequence), (iii) A SNP proximal to the target gene locus, for example, a sequence complementary to the SNP contained in the genomic DNA of the subject (e.g., a patient) (e.g., a spacer sequence), (iv) A sequence that is complementary to or contains one or more silent substitutions proximal to the target gene locus (e.g., a spacer sequence) A gene modification system according to Embodiment 100, including the above.
[0121] 102. A template RNA or gene modification system according to any one of the prior embodiments, wherein the gRNA spacer comprises about one, two, three or more flanking nucleotides of the gRNA spacer.
[0122] 103. A template RNA or gene modification system according to any one of the prior embodiments, wherein the heterologous target sequence comprises approximately 2, 3, 4, 5, 10, 20, 30, 40 or more flanking nucleotides of the RT template sequence.
[0123] 104. A template RNA or gene modification system according to any one of the preceding embodiments, wherein the heterologous target sequence comprises approximately 8 to 30, 9 to 25, 10 to 20, 11 to 16, or 12 to 15 (e.g., approximately 11 to 16) nucleotides.
[0124] 105. A template RNA or gene modification system according to any one of the preceding embodiments, wherein the mutated region comprises one, two, or three nucleotide positions that are sequence-differentiated compared to the corresponding portion of the human SERPINA1 gene.
[0125] 106. A template RNA or gene modification system according to any one of the preceding embodiments, wherein the mutated region includes at least two nucleotide positions that are sequence-differentiated compared to the corresponding portion of the human SERPINA1 gene.
[0126] 107. A template RNA or gene modification system according to any one of the prior embodiments, wherein the post-edited homology region and / or pre-edited homology region have 100% identity with respect to the SERPINA1 gene.
[0127] 108. A template RNA or gene modification system according to any one of the prior embodiments, wherein the PBS sequence additionally comprises about 1, 2, 3, 4, 5, 6, 7 or more flanking nucleotides.
[0128] 109. A template RNA or gene modification system according to any one of the preceding embodiments, wherein the PBS sequence comprises approximately 5–20, 8–16, 8–14, 8–13, 9–13, 9–12, or 10–12 (e.g., approximately 9–12) nucleotides.
[0129] 110. A template RNA or gene modification system according to any one of the preceding embodiments, wherein the PBS sequence is bound to within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides from a nicking site in the SERPINA1 gene.
[0130] 111. A gene modification system according to any one of the prior embodiments, wherein the domains of a gene-modified polypeptide are linked together by a peptide linker.
[0131] 112. The gene modification system according to Embodiment 111, wherein the linker includes a linker sequence from Table 10 (for example, any of sequence numbers 5217, 5106, 5190, and 5218).
[0132] 113. The gene modification system according to any one of the prior embodiments, wherein the gene-modified polypeptide further comprises one or more nuclear localization sequences (NLS).
[0133] 114. The gene modification system according to Embodiment 113, wherein the gene modification polypeptide comprises a first NLS and a second NLS.
[0134] 115. The gene modification system according to Embodiment 113 or 114, comprising an NLS sequence from Table 11 (for example, any of SEQ ID NOs. 5245, 5290, 5323, 5330, 5349, 5350, 5351, and 4001).
[0135] 116. DNA encoding the template RNA described in any one of the prior embodiments.
[0136] 117. A pharmaceutical composition comprising a gene modification system according to any one of embodiments 74 to 115 or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier.
[0137] 118. The pharmaceutical composition according to Embodiment 117, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.
[0138] 119. The pharmaceutical composition according to Embodiment 118, wherein the viral vector is an adeno-associated virus.
[0139] 120. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising a template RNA or gene modification system as described in any one of the prior embodiments.
[0140] 121. A method for preparing a template RNA according to any one of Embodiments 1 to 110, comprising synthesizing the template RNA by introducing DNA encoding the template RNA into a host cell under conditions that enable in vitro transcription (e.g., solid-phase synthesis) or the production of the template RNA.
[0141] 122. A method for modifying a target site in the human SERPINA1 gene of a cell, comprising contacting the cell with a gene modification system described in any one of Embodiments 74 to 115 or DNA encoding the same, thereby modifying a target site in the human SERPINA1 gene of the cell.
[0142] 123. A method for modifying a target site in the human SERPINA1 gene of a cell, comprising contacting the cell with (i) a template RNA or DNA encoding it as described in any one of Embodiments 1 to 73, and (ii) a gene-modified polypeptide or nucleic acid encoding a gene-modified polypeptide, thereby modifying a target site in the human SERPINA1 gene of the cell.
[0143] 124. A method for treating a subject having a disease or condition related to a mutation in the human SERPINA1 gene, comprising administering to the subject a gene modification system described in any one of embodiments 74 to 115 or DNA encoding the same, thereby treating the subject having a disease or condition related to a mutation in the human SERPINA1 gene.
[0144] 125. A method for treating a subject having a disease or condition related to a mutation in the human SERPINA1 gene, comprising administering to the subject a template RNA or DNA encoding it as described in any one of Embodiments 1 to 73, and (ii) a gene-modified polypeptide or nucleic acid encoding a gene-modified polypeptide, thereby treating the subject having a disease or condition related to a mutation in the human SERPINA1 gene.
[0145] 126. The method according to Embodiment 124 or 125, wherein the disease or condition is alpha-1 antitrypsin deficiency (AATD).
[0146] 127. The subject is the method according to any one of embodiments 124 to 126, having the E342K mutation (i.e., the PiZ mutation).
[0147] 128. A method for treating a subject having AATD, comprising administering to the subject a gene modification system described in any one of embodiments 74 to 115 or DNA encoding the same, thereby treating the subject having AATD.
[0148] 129. A method for treating a subject having AATD, comprising administering to the subject (i) a template RNA or DNA encoding it as described in any one of embodiments 74 to 115, and (ii) a genetically modified polypeptide or nucleic acid encoding a genetically modified polypeptide, thereby treating the subject having AATD.
[0149] 130. A gene modification system or method according to any one of the prior embodiments, wherein the introduction of the system into target cells results in the modification of a pathogenic mutation in the SERPINA1 gene.
[0150] 131. A gene modification system or method according to any one of the prior embodiments, wherein the pathogenic mutation is an E342K mutation, and the modification comprises an amino acid substitution of K342E.
[0151] 132. A gene modification system or method according to any one of the prior embodiments, wherein the correction of mutations occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target nucleic acid.
[0152] 133. A gene modification system or method according to any one of the prior embodiments, wherein the correction of mutations occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target cells.
[0153] 134. A gene modification system or method according to any one of the preceding embodiments, wherein the gene modification system comprises a second-strand targeted gRNA, and the correction of mutations in the target cell population is increased compared to a target cell population treated with a gene modification system comprising a template RNA without a second-strand targeted gRNA.
[0154] 135. A gene modification system or method according to any one of the prior embodiments, wherein the template RNA contains one or more silent substitutions (e.g., as illustrated in Table 7B), and the correction of mutations in the target cell population is increased compared to a target cell population treated with a gene modification system containing template RNA that does not contain one or more silent substitutions.
[0155] 136. The method according to any one of the prior embodiments, wherein the cells are mammalian cells such as human cells.
[0156] 137. The method according to any one of the prior embodiments, wherein the subject is a human.
[0157] 138. The method according to any one of the preceding embodiments, wherein contact is performed ex vivo, and for example, the DNA of a cell or subject is modified ex vivo.
[0158] 139. The method according to any one of the preceding embodiments, wherein the contact is performed in vivo, and for example, the cell or DNA of the subject is modified in vivo.
[0159] 140. The method according to any one of the prior embodiments, wherein contacting cells or objects with a system involves contacting cells or objects with nucleic acids (e.g., DNA or RNA) encoding gene-modified polypeptides under conditions that enable the production of gene-modified polypeptides.
[0160] 141. The template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments, wherein the gRNA scaffold is a mutant gRNA scaffold comprising a sequence relating to Table 23 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0161] 142. The variant gRNA scaffold is a template RNA, system, pharmaceutical composition, cell, or method described in Embodiment 141, comprising the sequence shown in Table 23.
[0162] 143. A variant gRNA scaffold comprising the sequence GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCA (Sequence ID 26000), the template RNA, system, pharmaceutical composition, cell, or method according to Embodiment 141 or 142.
[0163] 144. A template RNA, system, pharmaceutical composition, cell, or method according to any one of the preceding embodiments, wherein the gRNA spacer includes a sequence relating to Table 22 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0164] 145. The gRNA spacer is a template RNA, system, pharmaceutical composition, cell, or method described in any one of the prior embodiments, comprising the sequence shown in Table 22.
[0165] 146. The gRNA spacer is a template RNA, system, pharmaceutical composition, cell, or method described in any one of the prior embodiments, comprising the sequence related to AAGGCUGUGCUGACCAUCGA (SEQ ID NO: 26001).
[0166] 147. A heterologous target sequence includes a sequence relating to Table 24 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto, as described in any one of the preceding embodiments, which is a template RNA, system, pharmaceutical composition, cell, or method.
[0167] 148. The heterologous target sequence is a template RNA, system, pharmaceutical composition, cell, or method described in any one of the prior embodiments, including the sequence shown in Table 24.
[0168] 149. A template RNA, system, pharmaceutical composition, cell, or method according to any one of the preceding embodiments, wherein the PBS sequence includes the sequence shown in Table 25 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0169] 150. The PBS sequence includes the sequence described in Table 25, and is a template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments.
[0170] 151. A template RNA, system, pharmaceutical composition, cell, or method according to any one of the prior embodiments, comprising a sequence relating to Table 20 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
[0171] 152. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the prior embodiments, comprising the sequence relating to Table 20.
[0172] 153. A template RNA, system, pharmaceutical composition, cell, or method according to any one of the prior embodiments, comprising a sequence relating to Table 21 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
[0173] 154. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the prior embodiments, comprising the sequence relating to Table 21.
[0174] 155. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments, comprising a sequence relating to any one of Table 27, Table E3, Table E3A, Table E7, Table E8, Table E9, Table El1A, Table E11B, Table E12, Table E12A, Table E14, Table E14A, Table E15, or Table E16, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
[0175] 156. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments, comprising a sequence relating to any one of Table 27, Table E3, Table E3A, Table E7, Table E8, Table E9, Table El1A, Table E11B, Table E12, Table E12A, Table E14, Table E14A, Table E15, or Table E16.
[0176] 157. A template RNA, system, pharmaceutical composition, cell, or method according to any one of the prior embodiments, comprising one or more chemical modifications.
[0177] 158. A template RNA, system, pharmaceutical composition, cell, or method according to Embodiment 153, comprising one or more phosphorothioate bonds.
[0178] 159. A template RNA, system, pharmaceutical composition, cell, or method according to Embodiment 157 or 158, comprising one or more 2'-O-methylnucleotides.
[0179] 160. A template RNA, system, pharmaceutical composition, cell, or method according to any one of Embodiments 153 to 157, comprising a sequence relating to the third column of Table 20 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0180] 161. A template RNA, system, pharmaceutical composition, cell, or method according to any one of Embodiments 153 to 157, comprising a sequence relating to the third column of Table 21 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
[0181] 162. A gene modification system, A template RNA described in any one of Embodiments 141 to 161, A genetically modified polypeptide or a nucleic acid encoding a genetically modified polypeptide, wherein the genetically modified polypeptide is (1) St1Cas9 domain, (2) Linker, and (3) Reverse transcriptase (RT) domain Genetically modified polypeptides or nucleic acids, including A gene modification system that includes this.
[0182] The system according to embodiment 162, wherein the 163.St1Cas9 domain is a nickasase.
[0183] 164. The system according to Embodiment 162 or 163, wherein the St1Cas9 domain includes the sequence relating to Sequence ID No. 23818 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0184] 165. The system according to any one of embodiments 162 to 164, wherein the linker includes the sequence relating to sequence number 5006 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0185] 166. The system according to any one of embodiments 162 to 165, wherein the linker includes the sequence relating to sequence number 5217 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0186] 167. The system according to any one of Embodiments 162 to 166, wherein the linker includes the sequence of Table 10 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0187] 168. The system according to any one of Embodiments 162 to 167, wherein the RT domain includes the sequence relating to Sequence ID No. 26006 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0188] 169. The system according to any one of Embodiments 162 to 167, wherein the RT domain includes a sequence relating to any of Sequence IDs 8,001 to 8,003 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0189] 170. The system according to any one of Embodiments 162 to 167, wherein the RT domain includes the sequence in Table 6 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0190] 171. The system according to any one of Embodiments 162 to 167, wherein the gene-modified polypeptide includes the sequence relating to Sequence ID No. 26002 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0191] 172. The system according to any one of embodiments 162 to 171, further comprising a second nicked gRNA (ngRNA), wherein the second nicked gRNA optionally directs the second nicked gRNA to the second strand of the human SERPINA1 gene.
[0192] 173. The system according to Embodiment 172, wherein the second nick gRNA includes a sequence relating to Table 26 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
[0193] 174. The system according to any one of embodiments 162 to 173, wherein the nucleic acid encoding the gene-modified polypeptide includes RNA, for example, mRNA.
[0194] 175. A template RNA or system according to any one of the prior embodiments, wherein the nucleic acid molecule is formulated into lipid nanoparticles (LNPs).
[0195] 176. A system according to any one of the prior embodiments, wherein a template RNA, a nucleic acid molecule encoding a gene-modified polypeptide, and / or gRNA are formulated into an LNP.
[0196] 177. A pharmaceutical composition comprising a system or one or more nucleic acids encoding the system described in any one of embodiments 160 to 174, and a pharmaceutically acceptable excipient or carrier.
[0197] 178. The pharmaceutical composition according to Embodiment 177, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).
[0198] 179. The pharmaceutical composition according to Embodiment 178, wherein the viral vector is an adeno-associated virus.
[0199] 180. A host cell (e.g., a mammalian cell, e.g., a human cell) containing a gene modification system or template RNA as described in any one of the prior embodiments.
[0200] 181. A method for preparing a template RNA according to any one of Embodiments 141 to 161, comprising synthesizing the template RNA in vitro (for example, by in vitro transcription or solid-phase synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that enable the production of the template RNA.
[0201] 182. A method for modifying a target site in a cell (for example, a target site in the human SERPINA1 gene), comprising contacting the cell with a gene modification system described in any one of the preceding embodiments, DNA encoding it, or a pharmaceutical composition, thereby modifying the target site.
[0202] 183. A method for treating a subject having a disease or condition related to a mutation of a gene (e.g., the human SERPINA1 gene), comprising administering to the subject a gene modification system, DNA encoding it, or a pharmaceutical composition described in any one of the preceding embodiments, thereby treating the subject having the disease or condition.
[0203] 184. The method according to Embodiment 183, wherein the disease or condition is alpha-1 antitrypsin deficiency (AATD).
[0204] 185. The subject is the method according to Embodiment 183 or 184, having the E342K mutation.
[0205] 186. A method for treating a subject having AATD, comprising administering to the subject a gene modification system, DNA encoding it, or a pharmaceutical composition described in any one of the prior embodiments, thereby treating the subject having AATD.
[0206] 187. A gene modification system or method according to any one of the prior embodiments, wherein the introduction of the system into target cells results in the modification of a pathogenic mutation in a gene, for example, the SERPINA1 gene.
[0207] 188. A gene modification system or method according to any one of the prior embodiments, wherein the pathogenic mutation is an E342K mutation, and the modification comprises an amino acid substitution of K342E.
[0208] 189. A gene modification system or method according to any one of the prior embodiments, wherein the introduction of the system into target cells results in a mutation that causes the restoration of function of a gene, for example, the SERPINA1 gene.
[0209] 190. A gene modification system or method according to any one of the prior embodiments, wherein the correction of mutations occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target nucleic acid.
[0210] 191. The correction of the mutation occurs in at least 10% of the target cells (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more), and the gene modification system or method according to any one of the preceding embodiments.
[0211] 192. The gene modification system includes a second nick gRNA, and the correction of the mutation in the population of target cells is increased compared to the population of target cells treated with a gene modification system containing a template RNA without the second nick gRNA, and the gene modification system or method according to any one of the preceding embodiments.
[0212] 193. The template RNA includes one or more silent substitutions, and the correction of the mutation in the population of target cells is increased compared to the population of target cells treated with a gene modification system containing a template RNA without one or more silent substitutions, and the gene modification system or method according to any one of the preceding embodiments.
[0213] 194. The cell is a mammalian cell such as a human cell, and the method according to any one of the preceding embodiments.
[0214] 195. The subject is a human, and the method according to any one of the preceding embodiments.
[0215] 196. The contact is performed ex vivo. For example, the DNA of the cell or the subject is modified ex vivo, and the method according to any one of the preceding embodiments.
[0216] 197. The contact is performed in vivo. For example, the DNA of the cell or the subject is modified in vivo, and the method according to any one of the preceding embodiments.
[0217] 198. A method according to any one of the prior embodiments, comprising contacting cells or objects with a system, under conditions that enable the production of genetically modified polypeptides, by contacting cells within a cell or object with nucleic acids (e.g., DNA or RNA) that encode genetically modified polypeptides.
[0218] 199. A method according to any one of the prior embodiments, comprising administering a gene modification system or DNA encoding it twice.
[0219] 200. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments, comprising a sequence relating to any one of Tables 20, 21, 27, E3, E3A, E7, E8, E9, El1A, E11B, E12, E12A, E14, E14A, E15, or E16, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
[0220] 201. A template RNA, system, pharmaceutical composition, cell, or method described in any one of the preceding embodiments, comprising a sequence relating to any one of Tables 20, 21, 27, E3, E3A, E7, E8, E9, El1A, E11B, E12, E12A, E14, E14A, E15, or E16.
[0221] 202.gRNA, (for example, from 5' to 3') (1) gRNA spacer, and (2) Mutant St1Cas9 scaffold having partial or complete deletion of stem-loop 2 gRNA containing this substance.
[0222] 203. The gRNA according to Embodiment 202, wherein the deletion is 1 to 32 nucleotides long (e.g., 2 to 29, 2 to 20, 2 to 10, or 10 to 20).
[0223] 204. The deletion is the entirety of stem-loop 2, as described in Embodiment 202 of the gRNA.
[0224] 205. The deletion is at positions 55-84, as described in Embodiment 202.
[0225] 206. The St1Cas9 scaffold is the gRNA according to Embodiment 202, comprising a deletion of a portion of the second single-stranded region (e.g., nucleotides 1, 2, 3, or 4 at the 3' end of the single-stranded region).
[0226] 207. The mutant St1Cas9 scaffold is a gRNA according to any one of embodiments 202 to 206, having either or both an elongated RAR upper stem or a substitution that generates a GC base pair in the RAR upper stem.
[0227] 208. The mutant St1Cas9 scaffold is a gRNA described in any one of embodiments 202 to 207, having a mutation in the tetraloop.
[0228] 209. It is a gRNA, (for example, from 5' to 3') (1) gRNA spacer, and (2) A mutant St1Cas9 scaffold having either or both of the following: an elongated RAR upper stem or a substitution that generates a GC base pair on the RAR upper stem. gRNA containing this substance.
[0229] 210. The gRNA according to any one of embodiments 202 to 209, wherein the upper stem of the RAR is extended by 1 to 8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs) compared to the wild-type sequence of sequence number 25999.
[0230] 211. The gRNA according to Embodiment 210, wherein at least 50%, 60%, 70%, 80%, or 90% of the base pairs that are novel compared to SEQ ID NO: 25999 are GC base pairs.
[0231] 212.gRNA, (for example, from 5' to 3') (1) gRNA spacer, and (2) Mutant St1Cas9 scaffold with mutation in the tetraloop A gRNA comprising
[0232] The gRNA according to any one of Embodiments 202 to 212, wherein one or more nucleotides in the tetraloop are substituted.
[0233] The gRNA according to any one of Embodiments 202 to 213, wherein the tetraloop comprises a sequence selected from AACA, AAUA, ACCA, ACUA, AGUA, AGCA, AUCA, AUUA, CAAC, CUCG, CUUG, GAAA, GAGA, GCAA, GCGA, GGAA, GGAG, GGGA, GUAA, GUGA, UAAC, UACG, UCAC, UCCG, UGAA, UGAC, UGCG, UUAC or UUCG.
[0234] The gRNA according to any one of Embodiments 202 to 214, wherein the tetraloop extends, for example, up to 5 nucleotides.
[0235] The gRNA according to Embodiment 215, wherein the extended tetraloop comprises a sequence selected from GAAGA or GACAA.
[0236] The gRNA according to any one of Embodiments 202 to 216, wherein the mutant gRNA scaffold comprises a sequence according to Table 23 or a sequence having 1, 2 or 3 or fewer sequence modifications (e.g., substitutions) compared thereto.
[0237] The gRNA according to any one of Embodiments 202 to 217, wherein the mutant St1Cas9 scaffold has a length of 50 to 60, 60 to 70, 70 to 80 or 80 to 84 nucleotides.
[0238] A gRNA comprising, for example, in the 5' to 3' direction (1) a gRNA spacer, and (2) a mutant gRNA scaffold comprising a sequence according to Table 23 or a sequence having 1, 2 or 3 or fewer sequence modifications (e.g., substitutions) compared thereto A gRNA comprising
[0239] 220. The mutant gRNA scaffold is a gRNA described in any one of Embodiments 202 to 217, comprising the sequence shown in Table 23.
[0240] 221. The mutant gRNA scaffold is the gRNA described in any one of Embodiments 202 to 220, which includes the sequence related to GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCA (Sequence ID 26000).
[0241] A copy of this patent publication or patent application publication having color drawings, including at least one color-finished drawing, will be provided by the Office immediately upon request and payment of the required fees. [Brief explanation of the drawing]
[0242] [Figure 1]The gene modification system as described herein is illustrated. The left-hand figure shows the gene modification polypeptide, which includes a Cas nickase domain (e.g., spCas9 N863A) and a reverse transcriptase domain (RT domain) linked by a linker. The right-hand figure shows the template RNA, including a gRNA spacer, gRNA scaffold, heterologous target sequence, and primer-binding site sequence (PBS sequence) from 5' to 3'. The heterologous target sequence may include a mutant region containing one or more sequence differences compared to the target site. The heterologous target sequence may also include pre-edited homology regions and post-edited homology regions adjacent to the mutant region. Although we do not wish to be constrained by theory, it is thought that the gRNA spacer of the template RNA binds to the second strand of the target site in the genome, and the gRNA scaffold of the template RNA binds to the gene modification polypeptide, thereby positioning the gene modification polypeptide at the target site in the genome, for example. It is thought that the Cas domain of the gene-modified polypeptide introduces a nick at the target site (for example, the first strand of the target site), allowing, for example, the PBS sequence to bind to a sequence adjacent to the site to be modified on the first strand of the target site. The RT domain of the gene-modified polypeptide is thought to polymerize, for example, a sequence complementary to the heterologous target sequence, by using the first strand of the target site, which binds to the complementary sequence containing the PBS sequence of the template RNA, as a primer, and by using the heterologous target sequence of the template RNA as a template. Although we do not wish to be constrained by theory, it is thought that reverse transcription then proceeds through the pre-edited homology region, then through the mutation region, and then through the post-edited homology region, thereby creating a DNA strand containing the mutation specified by the heterologous target sequence. [Figure 2] The hypothetical secondary structure of the wild-type St1Cas9 gRNA scaffold is shown, with the descriptions of the variants described herein superimposed on it. [Figure 3-1] (Figure 3A) A graph showing the rewriting performance of an exemplary template RNA containing various scaffolds truncated in the SL2 region of a St1Cas9-based gene modification system. [Figure 3-2](Figure 3B) A graph shows rewriting by an exemplary template RNA containing scaffolds in which the TL and RAR regions have been further manipulated using various modified tetraloops, using an St1Cas9-based gene modification system. [Figure 3-3] (Figure 3C) A graph showing rewriting by an exemplary template RNA containing spacers of various lengths using an St1Cas9-based gene modification system. [Figure 4-1] (Figure 4A) Graphs showing the rewriting efficiency of gene modification systems containing different St1Cas9 compatible template RNAs, including modified scaffold sequences. (Figure 4B) Graphs showing the % indel levels of the same gene modification system evaluated in Figure 4A. [Figure 4-2] (Figure 4C) This graph shows the rewriting efficiency of gene modification systems containing different St1Cas9 compatible template RNAs, including modified scaffold sequences. [Figure 5] This graph shows the rewriting efficiency of gene modification systems, including St1Cas9-based gene modification polypeptides. [Figure 6] The graphs on the left and right show the rewriting efficiency (% indel level) of gene modification systems, including St1Cas9-based gene modification polypeptides, with and without ngRNA. [Figure 7] This graph shows the rewriting activity of a St1Cas9-based gene modification system in primary hepatocytes, including a dSL2 mutant gRNA scaffold, PBS sequences of various lengths, and exemplary template RNAs containing heterologous target sequences. [Figure 8-1] (Figure 8A) A series of graphs showing the percentage rewriting rates achieved using gene modification systems containing different St1Cas9 compatible template RNAs, including mutant scaffolds with various exemplary mutant tetraloop structures, in primary hepatocytes (left panel), high-dose treated HEK293T cells (center panel), or low-dose treated HEK293T cells (right panel). [Figure 8-2] (Figure 8B) A hypothetical secondary structure of the dSL2 truncated St1Cas9 gRNA scaffold is shown, with the description of the variant described herein superimposed on it. [Figure 9-1] (Figure 9A) This graph shows the rewriting activity of an exemplary St1Cas9-based gene modification system, including a mutant template RNA containing the nucleotide sequence of RNACS9201 (which has a dSL2 mutant gRNA scaffold), which has various 2-O'-methyl chemical modifications in the gRNA scaffold region. [Figure 9-2] (Figure 9B) This graph shows the results on rewriting activity when the three nucleotides of the scaffold are modified simultaneously with 2'-O-methyl chemical modification. [Figure 9-3] (Figures 9C-9G) The pattern of 2'-O-methyl-modified nucleotides in the dSL2 St1Cas9 scaffold sequence in Figure 9A is illustrated. Gray bases represent unmodified nucleotide positions. Black and bold bases represent modified nucleotide positions. In this example, the modified nucleotide is a 2'-O-methyl nucleotide. [Figure 9-4] Same as above. [Figure 9-5] Same as above. [Figure 9-6] Same as above. [Figure 9-7] Same as above. [Figure 10] This graph shows the rewriting efficiency of gene modification systems containing different St1Cas9-compatible template RNAs with different patterns of 2'-O-methyl chemical modifications on the dSL2 St1Cas9 scaffold at the primary hepatocyte. [Figure 11] (Figures 11A-11C) are a series of graphs showing the rewriting activity in the liver (Figure 11A), indel activity in the liver (Figure 11B), and hA1AT in serum (Figure 11C) of gene modification systems formulated into LNPs containing different St1Cas9 compatible template RNAs with dSL2 mutant gRNA scaffolds. [Figure 12-1](Figures 12A-12C) A series of graphs showing Amp-Seq results in the liver for complete rewriting rate (Figure 12A), indel rate (Figure 12B), and serum A1AT levels (Figure 12C) using a gene modification system formulated in LNPs containing different St1Cas9 compatible template RNAs with different patterns of 2'-O-methyl chemical modifications in the dSL2 St1Cas9 scaffold. [Figure 12-2] Same as above. [Figure 13] (Figure 13A) This figure shows the location of the St1Cas9 scaffold sequence in reference dSL2. (Figure 13B) This figure shows the location of the wild-type St1Cas9 scaffold sequence. (Figure 13C) This figure shows the hypothetical structure of RNACS13597 with the RAR+4_UUCG mutation compared to dSL2. (Figure 13D) This figure shows the hypothetical structure of RNACS17210 with the RAR+4_AGCA mutation compared to dSL2. [Figure 14] (Figures 14A-14B) Graphs showing the % edit rate (Figure 14A) or indel (Figure 14B) in the liver after administration of one or two doses of gene-modified polypeptide and template RNA, as evaluated by Amp-Seq. [Figure 15] (Figures 15A-15B) Graphs showing the % editing rate (Figure 15A) or indels (Figure 15B) in the liver after administration of gene-modified polypeptides and template RNA, as evaluated by Amp-Seq. [Figure 16] This shows the percentage rewriting rate achieved in the primary hepatocyte after administration of gene-modified polypeptides and template RNA. [Figure 17] The hypothetical secondary structure of the dSL2 truncated St1Cas9 gRNA scaffold is shown, with the description of the variant described herein superimposed on it. [Figure 18] This shows the percentage rewriting rate achieved at the primary hepatocyte after administration of gene-modified polypeptides and template RNA. [Modes for carrying out the invention]
[0243] definition As used herein, the term "expression cassette" refers to a nucleic acid construct comprising sufficient nucleic acid elements for the expression of the nucleic acid molecule of the present invention.
[0244] As used herein, "gRNA spacer" refers to a portion of a nucleic acid that is complementary to the target nucleic acid and, together with the gRNA scaffold, can target the Cas protein to the target nucleic acid.
[0245] As used herein, "gRNA scaffold" refers to a portion of a nucleic acid that can bind to a Cas protein and, together with a gRNA spacer, can target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold includes a crRNA sequence, a tetraloop, and a tracrRNA sequence.
[0246] As used herein, "mutant gRNA scaffold" refers to a gRNA scaffold having a sequence that does not exist in nature. In some embodiments, the mutant sequence includes one or more substitutions compared to the nearest naturally occurring sequence. In some embodiments, the mutant sequence includes one or more insertions compared to the nearest naturally occurring sequence. In some embodiments, the mutant sequence includes one or more deletions compared to the nearest naturally occurring sequence.
[0247] The term “St1Cas9 scaffold” as used herein refers to a gRNA scaffold that can bind to the St1Cas9 protein and, together with a gRNA spacer, can target the St1Cas9 protein to a target nucleic acid. In some embodiments, the St1Cas9 scaffold includes a crRNA sequence, a tetraloop, and a tracerRNA sequence. Exemplary locations of the St1Cas9 scaffold within exemplary template RNA are shown in Figures 13A and 13B.
[0248] In some embodiments, the St1Cas9 scaffold contains a full-length wild-type sequence. In some embodiments, the St1Cas9 scaffold contains a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the sequence GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUGUUUU (Sequence ID 25999). In some embodiments, the St1Cas9 scaffold contains a sequence identical to Sequence ID 25999. In some embodiments, the St1Cas9 scaffold is a truncation mutant. In some embodiments, the St1Cas9 scaffold has at least 80%, 85% identity with the sequence GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCA (Sequence ID 26000). The St1Cas9 scaffold contains sequences having 90%, 95%, 96%, 97%, 98%, or 99% identity. In some embodiments, the St1Cas9 scaffold contains a sequence identical to sequence number 26000. In some embodiments, the St1Cas9 scaffold contains insertions, deletions, or substitutions compared to the reference sequence of sequence number 25999 or 26000. In some embodiments, the St1Cas9 scaffold contains chemically modified nucleotides.
[0249] When used herein, "genetically modified polypeptide" refers to a polypeptide comprising a retroviral reverse transcriptase or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to a retroviral reverse transcriptase capable of incorporating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule in a host cell). In some embodiments, the genetically modified polypeptide is capable of incorporating the sequence substantially independently of the host mechanism. In some embodiments, the genetically modified polypeptide incorporates the sequence at a random location in the genome, and in some embodiments, the genetically modified polypeptide incorporates the sequence at a specific target site. In some embodiments, the genetically modified polypeptide collectively comprises one or more domains that 1) facilitate binding to the template nucleic acid, 2) facilitate binding to the target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. The genetically modified polypeptide includes both the natural polypeptide and its engineered variants having, for example, one or more amino acid substitutions to the natural sequence. Genetically modified polypeptides include heterogeneous constructs, for example, one or more of the domains listed above are heterogeneous with each other, whether otherwise by heterogeneous fusion (or other conjugate) of wild-type domains and fusion of modified domains, for example by substitution or fusion of heterogeneous subdomains or other substituted domains. Exemplary genetically modified polypeptides that can be used in the manner provided herein, as well as systems comprising them and methods of using them, are described, for example, in PCT / US2021 / 020948, which is incorporated herein by reference with respect to genetically modified polypeptides comprising retroviral reverse transcriptase domains. In some embodiments, the genetically modified polypeptide incorporates a sequence into a gene. In some embodiments, the genetically modified polypeptide incorporates a sequence into a sequence outside of a gene. When used herein, “genetically modified system” refers to a system comprising a genetically modified polypeptide and a template nucleic acid.
[0250] As used herein, the term “domain” refers to a structure of a biomolecule that contributes to a specific function of that biomolecule. A domain may include a continuous region (e.g., a continuous sequence) or a separate discontinuous region (e.g., a discontinuous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcription domains, while examples of nucleic acid domains include regulatory domains, such as transcription factor-binding domains. In some embodiments, a domain (e.g., a Cas domain) may include two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).
[0251] As used herein, the term “exomorphic,” when used in relation to biomolecules (e.g., nucleic acid sequences or polypeptides), means that the biomolecule has been artificially introduced into a host genome, cell, or organism. For example, nucleic acids added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods are exomorphic to the existing nucleic acid sequence, cell, tissue, or subject.
[0252] As used herein, the terms “first strand” and “second strand,” used to describe individual DNA strands of target DNA, identify two DNA strands on which the reverse transcriptase domain initiates polymerization, for example, based on the location where targeted primed synthesis begins. “First strand” refers to the strand of target DNA on which the reverse transcriptase domain initiates polymerization, for example, initiating targeted primed synthesis. “Second strand” refers to the other strand of target DNA. The designations “first strand” and “second strand” do not otherwise describe the target site DNA strand. For example, in some embodiments, the first and second strands are nicked by the polypeptides described herein, but the designations “first” and “second” strands are independent of the order in which such nicks appear.
[0253] The term “heterogeneous,” when used to describe a first element in relation to a second element, means that the first and second elements do not naturally exist in the configuration described. For example, heterogeneous polypeptides, nucleic acid molecules, constructs, or sequences refer to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule, that has been modified or mutated from its native state, or (c) a polypeptide or nucleic acid molecule having modified expression compared to its native expression level under similar conditions. For example, heterogeneous regulatory sequences (e.g., promoters, enhancers) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from that which is normally expressed in nature. In another example, a heterogeneous domain or nucleic acid sequence of a polypeptide (e.g., the DNA-binding domain of a polypeptide or the nucleic acid encoding the DNA-binding domain of a polypeptide) may be configured in relation to other domains, or may have a different sequence or origin from a different source compared to other domains or portions of the polypeptide or its encoding nucleic acid. In certain embodiments, heterologous nucleic acid molecules may be present in the native host cell genome, but may have modified expression levels, different sequences, or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous in the host cell or host genome, but may instead be introduced into the host cell by transformation (e.g., transfection, electroporation), and the added molecules may be integrated into the host genome or exist as extrachromosomal genetic material, either transiently (e.g., mRNA) or semi-stable for two or more generations (e.g., episomal viral vectors, plasmids, or other self-replicating vectors).
[0254] As used herein, “insertion” of a sequence into a target site refers to the net addition of a DNA sequence at the target site, such as a new nucleotide in a heterologous target sequence that does not have a cognitive position at the non-edited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous target sequence with respect to the target nucleic acid sequence will result in an alignment gap in the target nucleic acid sequence.
[0255] As used herein, the “deletion” generated by the heterologous target sequence at the target site refers to a net deletion of the DNA sequence at the target site, such as a nucleotide at a non-edited target site that does not have a cognitive position in the heterologous target sequence. In some embodiments, the nucleotide alignment of the PBS sequence and the heterologous target sequence to the target nucleic acid sequence will result in an alignment gap in the molecule containing the PBS sequence and the heterologous target sequence.
[0256] As used herein, the term “inverted terminal repeat” or “ITR” refers to an AAV viral cis-element named for its symmetry. This element promotes effective augmentation of the AAV genome. The minimum elements for ITR function are assumed to be a Rep-binding site (RBS, 5'-GCGCGCTCGCTCGCTC-3' for AAV2, SEQ ID NO: 4601) and a terminal separation site (TRS, 5'-AGTTGG-3' for AAV2, SEQ ID NO: 4602), as well as a variable palindromic sequence that enables hairpin formation. According to the present invention, an ITR comprises at least three of these elements (RBS, TRS, and the sequence that enables hairpin formation). In addition, in the present invention, the term “ITR” refers to ITRs of known native AAV serotypes (e.g., ITRs of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), chimeric ITRs formed by the fusion of ITR elements from different serotypes, and their functional variants. A "functional variant" refers to a sequence that exhibits at least 80%, 85%, or 90%, preferably at least 95%, sequence identity with a known ITR, and that allows for an increase in the sequence containing the ITR in the presence of the Rep protein.
[0257] As used herein, the term "mutant region" refers to a region in template RNA that has one or more sequence differences compared to the corresponding sequence in the target nucleic acid. Sequence differences may include, for example, substitutions, insertions, frameshifts, or deletions.
[0258] The term "(sudden) mutation," when applied to nucleic acid sequences, means that nucleotides within the nucleic acid sequence are inserted, deleted, or altered compared to a reference (e.g., natural) nucleic acid sequence. A single alteration may occur at one locus (point mutation), or multiple nucleotides may be inserted, deleted, or altered at a single locus. In addition, one or more alterations may occur at any number of loci within a single nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.
[0259] "Nucleic acid molecules" refer to, but are not limited to, RNA and DNA molecules, including complementary DNA ("cDNA"), genomic DNA ("gDNA"), and messenger RNA ("mRNA"), and also include synthetic nucleic acid molecules, such as those chemically synthesized or produced by recombination, as described herein, like RNA templates. Nucleic acid molecules may be double-stranded or single-stranded, cyclic, or linear. In the case of single-stranded nucleic acid molecules, the nucleic acid molecule may be a sense strand or an antisense strand. Unless otherwise stated and as described herein in the general format "Sequence ID," an example of a nucleic acid containing any sequence or "Sequence ID 1" is a nucleic acid in which at least a portion has either (i) the sequence of Sequence ID 1, or (ii) a sequence complementary to Sequence ID 1. The choice between the two depends on the context in which Sequence ID 1 is used. For example, when a nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target. The nucleic acid sequences of this disclosure may be chemically or biochemically modified, or may contain unnatural or derivatized nucleotide bases, as will be readily apparent to those skilled in the art. Such modifications include, for example, labeling, methylation, substitution of one or more spontaneously occurring nucleotides by analogs, internucleotide modifications such as uncharged bonds (e.g., methylphosphonic acid, triestriates, phosphoramidates, carbamates, etc.), charged bonds (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), insertants (e.g., acridine, psoralens, etc.), chelating agents, alkylating agents, and modifying bonds (e.g., α-anomeric nucleic acids, etc.). Chemically modified bases (e.g., see Table 13), backbones (e.g., see Table 14), and modified caps (e.g., see Table 15). Synthetic molecules that mimic polynucleotides in terms of their ability to bind to specified sequences by hydrogen bonding and other chemical interactions are also included. Such molecules are known in the art, and examples include those that use peptide bonds instead of phosphate bonds in the molecular backbone, such as peptide nucleic acids (PNAs).Other modifications include analogues that include other structures such as modifications where the ribose ring is found in a crosslinking portion or in a "locked" nucleic acid (LNA). In various embodiments, nucleic acids are associated with additional genetic elements, such as tissue-specific expression-regulatory sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences) and additional elements, such as inverted repeats (e.g., inverted terminal repeats, e.g., viral elements (e.g., AAV ITR)) and tandem repeats, inverted repeats / direct repeats, homologous regions (segments with different degrees of homology to target DNA), untranslated regions (UTRs) (5', 3', or both 5' and 3' UTRs) and various combinations of the above. The nucleic acid elements of the systems provided by the present invention may be provided in various topologies, including single-stranded, double-stranded, circular, linear, open-ended linear, closed-ended linear and specific versions thereof, e.g., doggybone DNA (dbDNA), closed-ended DNA (ceDNA).
[0260] As used herein, “gene expression unit” is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably ligated to at least one effector sequence. The first nucleic acid sequence is operably ligated to the second nucleic acid sequence when the first nucleic acid sequence is positioned functionally in relation to the second nucleic acid sequence. For example, a promoter or enhancer is operably ligated to a coding sequence if the promoter or enhancer affects the transcription or expression of the coding sequence. The operably ligated DNA sequences may be continuous or discontinuous. If it is necessary to ligate two protein coding regions, the operably ligated sequences may reside within the same reading frame.
[0261] The terms “host genome” or “host cell,” as used herein, refer to the cell into which the protein and / or genetic material has been introduced and / or its genome. These terms refer not only to a specific target cell and / or genome, but also to the offspring of such cells and / or the genomes of such offspring. It should be understood that such offspring may not be identical to the parent cell in fact, as certain modifications may occur in later generations due to mutation or environmental influences, but are still included in the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell or a cell line grown in culture or genomic material isolated from such cells or cell lines, or it may be a host cell or host genome constituting a living tissue or organism. In some cases, the host cell may be an animal cell or a plant cell, as described herein, for example. In certain examples, the host cell may be a mammalian cell, a human cell, a bird cell, a reptile cell, a bovine cell, a horse cell, a pig cell, a goat cell, a sheep cell, a chicken cell, or a turkey cell. In certain cases, the host cells may be maize cells, soybean cells, wheat cells, or rice cells.
[0262] As used herein, “action-related” describes a functional relationship between two nucleic acid sequences, for example, 1) a promoter and 2) a heterologous target sequence, where the promoter and the heterologous target sequence (e.g., the gene of interest) are oriented such that, under appropriate conditions, the promoter drives the expression of the heterologous target sequence. For example, a template nucleic acid having a promoter and a heterologous target sequence may be, for example, a single strand with a (+) or (-) orientation. The “action-related” between the promoter and the heterologous target sequence in this template means that, regardless of whether the template nucleic acid is to be transcribed in a particular state, it will be transcribed accurately if it is in an appropriate state (e.g., with a (+) orientation in the presence of the required catalyst and NTPs, etc.). Action-related applies similarly to other pairs of nucleic acids, including other tissue-specific expression regulatory sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homologous regions and sequences encoding heterologous target sequences or retroviral RT domains.
[0263] As used herein, the term “position” in relation to the St1Cas9 scaffold refers to the nucleotide of the St1Cas9 scaffold that aligns with the corresponding nucleotide of the reference sequence of SEQ ID NO: 25999. The position of the reference sequence is shown in Figure 13B. Alignment of nucleic acid or polypeptide sequences can be performed using a sequence analysis tool such as the Basic Local Alignment Search Tool (BLAST), for example, NIH megablast with default parameters.
[0264] In some embodiments, the position of the St1Ca9 scaffold can be identified by providing an alignment of the St1Cas9 scaffold (query sequence) with the reference sequence of SEQ ID NO: 25999 (full-length wild-type sequence, see, for example, Figure 13B) or SEQ ID NO: 26000 (truncation mutant, see, for example, Figure 13A), and identifying the position in the query sequence that corresponds to the position in the reference sequence. For example, in the St1Cas9 scaffold consisting of the sequence of SEQ ID NO: 25999 except that the 5'most G is replaced with a single nucleotide other than G, the substituted position is position 1.
[0265] As another example, in the St1Cas9 scaffold consisting of sequence number 25999, except that a single novel nucleotide is inserted immediately 5' to the 5' side of the G on the 5' side, the G is still at position 1.
[0266] As yet another example, in the St1Cas9 scaffold consisting of sequence number 25999, except that a sequence of n nucleotides is inserted between the G at position 1 and the U at position 2, the nucleotides on the 3' side of the insertion retain their original position numbers. For example, the U at position 2 is still at position 2, rather than at position n+2. There is no need to assign position numbers to the inserted nucleotides by comparing them to the reference sequence.
[0267] The terms “primer-binding site sequence” or “PBS sequence,” as used herein, refer to a portion of template RNA capable of binding to a region contained in the target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence containing at least 3, 4, 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. While not intended to be bound by any particular theory, in some embodiments, if the template RNA contains both a PBS sequence and a heterologous target sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence, enabling the reverse transcriptase domain to use that region as a primer for reverse transcription and the heterologous target sequence as a template for reverse transcription.
[0268] As used herein, “stem-loop sequence” refers to a nucleic acid sequence (e.g., RNA sequence) having a stem containing at least 2 (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop containing at least 3 (e.g., 4) base pairs, where self-complementarity is sufficient to form a stem-loop. The stem may contain mismatches or bulges.
[0269] As used herein, “tissue-specific expression-regulatory sequence” means a nucleic acid element that increases or decreases the level of a transcript containing a heterologous target sequence in a tissue-specific manner within a target tissue, for example, preferentially in an on-target tissue compared to an off-target tissue. In some embodiments, the tissue-specific expression-regulatory sequence preferentially drives or represses the transcription, activity, or half-life of a transcript containing the heterologous target sequence in a tissue-specific manner within a target tissue, for example, preferentially in an on-target tissue compared to an off-target tissue. Exemplary tissue-specific expression-regulatory sequences include tissue-specific promoters, repressors, enhancers or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to on-target (tissues in which expression or activity of the template nucleic acid is desired or acceptable) and off-target (tissues in which expression or activity of the template nucleic acid is undesirable or unacceptable). For example, a tissue-specific promoter preferentially drives expression in an on-target tissue compared to an off-target tissue. In contrast, microRNAs that bind to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues compared to on-target tissues, thereby reducing the expression of template nucleic acids in off-target tissues. Therefore, promoters and microRNA recognition sequences specific to the same tissue, such as a target tissue, have contrasting functions with respect to the transcription, activity, or half-life of related sequences within the tissue (matching expression levels, i.e., promoting and suppressing high levels of microRNA in off-target tissues and low levels in on-target tissues, respectively, while promoters drive high expression in on-target tissues and low expression in off-target tissues).
[0270] List of Headlines 1) Introduction 2) Gene modification systems a) polypeptide components of gene modification systems i) Writing Domain ii) Endonuclease domain and DNA binding domain (1) Genetically modified polypeptide containing a Cas domain (2) TAL effector and zinc finger nuclease iii) Linker iv) Localization sequences for gene modification systems v) Evolutionary variants of gene-modified polypeptides and systems vi) Intein vii) Further domains b) Template nucleic acid i) gRNA spacers and gRNA scaffolds ii) Heterogeneous sequencing iii) PBS sequence iv) Exemplary template arrangement c) gRNAs with inducible activity d) Circular RNAs and ribozymes in gene modification systems e) Target nucleic acid site f) Second chain nicking 3) Preparation of compositions and systems 4) Therapeutic use 5) Administration and Delivery a) Tissue-specific activity / administration i) Promoto ii) microRNA b) Viral vectors and their components c) AAV administration d) Lipid nanoparticles 6) Kits, products and pharmaceutical compositions 7) Chemicals, Manufacturing and Control (CMC)
[0271] introduction This disclosure relates to methods for treating alpha-1 antitrypsin deficiency (AATD) and compositions for targeting, editing, modifying, or manipulating DNA sequences at one or more locations within a DNA sequence in a cell, tissue, or subject, for example, in vivo or in vitro (e.g., inserting a heterologous target sequence into a target site in a mammalian genome). The heterologous target DNA sequence may include, for example, substitutions.
[0272] More specifically, the present disclosure provides a method for treating AATD using a reverse transcriptase-based system for modifying a target genomic DNA sequence, for example, by inserting, deleting, or substituting one or more nucleotides into / from a target sequence.
[0273] This disclosure provides a method for treating AATD using a gene modification system comprising a gene modification polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the gene modification system can be used to introduce modifications into target sites within the genome. In some embodiments, the gene modification polypeptide component includes a writing domain (e.g., reverse transcriptase domain), a DNA-binding domain, and an endonuclease domain (e.g., nickase domain). In some embodiments, the template nucleic acid (e.g., template RNA) includes a sequence that binds to a target site within the genome (e.g., to the second strand of the target site) (e.g., a gRNA spacer), a sequence that binds to the gene modification polypeptide component (e.g., a gRNA scaffold), a heterologous target sequence, and a PBS sequence. While not intended to be constrained by theory, it is assumed that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site within the genome and then binds to the gene modification polypeptide component (e.g., localizes the polypeptide component to the target site within the genome). The endonuclease (e.g., nickase) of the gene-modified polypeptide component is thought to cleave the target site (e.g., the first strand of the target site) and bind to a sequence adjacent to the site where the PBS sequence will be modified on the first strand of the target site. The writing domain (e.g., reverse transcriptase domain) of the polypeptide component is thought to polymerize a sequence complementary to the heterologous target sequence using the first strand of the target site bound to a complementary sequence containing the PBS sequence of the template nucleic acid as a primer and the heterologous target sequence of the template nucleic acid as a template. While we do not wish to be constrained by theory, it is thought that the selection of an appropriate heterologous target sequence may result in the substitution, deletion, and / or insertion of one or more nucleotides at the target site.
[0274] Gene modification system In some embodiments, the gene modification systems described herein include (A) a gene modification polypeptide or a nucleic acid encoding a gene modification polypeptide, the gene modification polypeptide including (i) a reverse transcriptase domain and (x) an endonuclease domain with DNA-binding functionality or (y) an endonuclease domain and a separate DNA-binding domain, and (B) a template RNA. In some embodiments, the gene modification polypeptide acts as a substantially autonomous protein mechanism capable of incorporating a template nucleic acid sequence into a target DNA molecule (e.g., within a mammalian host cell, such as a genomic DNA molecule in a host cell) substantially independently of host mechanisms. For example, the gene modification protein may include a DNA-binding domain, a reverse transcriptase domain and an endonuclease domain. In some embodiments, the DNA-binding functionality may include an RNA component that guides the protein to the DNA sequence, such as a gRNA spacer. In other embodiments, the gene modification polypeptide may include a reverse transcriptase domain and an endonuclease domain. The RNA template element of the gene modification system is typically heterogeneous to the gene modification polypeptide element and provides a target sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the gene-modified polypeptide is capable of targeted priming reverse transcription. In some embodiments, the gene-modified polypeptide is capable of second chain synthesis.
[0275] In some embodiments, the gene modification system is combined with a second polypeptide. In some embodiments, the second polypeptide may include an endonuclease domain. In some embodiments, the second polypeptide may include a polymerase domain, such as a reverse transcriptase domain. In some embodiments, the second polypeptide may include a DNA-dependent DNA polymerase domain. In some embodiments, the second polypeptide assists in the completion of genome editing, for example, by contributing to second strand synthesis or DNA repair recovery.
[0276] Functionally modified polypeptides can consist of unrelated DNA-binding, reverse transcription, and endonuclease domains. This modular structure allows for combinations of functional domains, e.g., dCas9 (DNA binding), MMLV reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains may arise from a single protein, e.g., Cas9 or Cas9 nickase (DNA binding, endonuclease).
[0277] In some embodiments, the genetically modified polypeptide collectively comprises one or more domains that 1) facilitate binding to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the genetically modified polypeptide is a modified polypeptide comprising one or more amino acid substitutions relative to the corresponding native sequence. In some embodiments, the genetically modified polypeptide comprises two or more domains that are heterogeneous to each other, for example, by heterogeneous fusion (or other conjugate) of the otherwise wild-type domain and fusion of modified domains, for example by substitution or fusion of heterogeneous subdomains or other substitution domains. For example, in some embodiments, one or more of the following are the case: the RT domain is heterogeneous to the DBD, the DBD is heterogeneous to the endonuclease domain, or the RT domain is heterogeneous to the endonuclease domain.
[0278] In some embodiments, the template RNA molecule for use in the system includes, from 5' to 3', (1) a gRNA spacer, (2) a gRNA scaffold, (3) a heterologous target sequence, and (4) a primer binding site (PBS) sequence. In some embodiments, (1) The gRNA spacer is approximately 18-22 nt long, for example, 20 nt. (2) A gRNA scaffold comprising one or more hairpin loops, e.g., one, two, or three loops for associating a template with a Cas domain, e.g., a nickase Cas9 domain. In some embodiments, the gRNA scaffold comprises the sequence GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 5008) from 5' to 3'. (3) In some embodiments, the heterologous target sequence has a length of, for example, 7 to 74, for example, 10 to 20, 20 to 30, 30 to 40, 40 to 50, 50 to 60, 60 to 70 or 70 to 80 nt or 80 to 90 nt. In some embodiments, the first (usually 5') base of the sequence is not C. (4) In some embodiments, the PBS sequence that binds to the target priming sequence after nicking is, for example, 3-20 nt, 7-15 nt, or 12-14 nt. In some embodiments, the PBS sequence has a GC content of 40-60%.
[0279] In some embodiments, a second gRNA associated with the system may assist in driving complete integration. In some embodiments, the second gRNA may target positions 0–200 nt away from the first strand nick, for example, 0–50, 50–100, or 100–200 nt away from the first strand nick. In some embodiments, the second gRNA may bind only to its target sequence after editing, for example, the gRNA binds to a sequence that is present in the heterologous target sequence but not in the initial target sequence.
[0280] In some embodiments, the gene modification system described herein is used to perform editing on HEK293, K562, U2OS, or HeLa cells. In some embodiments, the gene modification system is used to perform editing on primary cells, for example, primary cortical neurons from E18.5 mice.
[0281] In some embodiments, the gene-modified polypeptides described herein include a reverse transcriptase or RT domain (e.g., as described herein) containing a MoMLV RT sequence or a variant thereof. In some embodiments, the MoMLV RT sequence includes one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In some embodiments, the MoMLV RT sequence optionally includes a combination of mutations such as D200N, L603W, and T330P, further comprising T306K and / or W313F.
[0282] In some embodiments, the endonuclease domain (e.g., as described herein) is an nCAS9 containing, for example, an N863A mutation (e.g., in spCas9) or an H840A mutation.
[0283] In some embodiments, the heterologous target sequences (for example, in the systems described herein) have nucleotide lengths of approximately 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000 or more.
[0284] In some embodiments, the RT and endonuclease domains are linked by a flexible linker, for example, containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 5006).
[0285] In some embodiments, the endonuclease domain is the N-terminus of the RT domain. In some embodiments, the endonuclease domain is the C-terminus of the RT domain.
[0286] In some embodiments, the system incorporates a heterogeneous target sequence into the target site by TPRT, for example, as described herein.
[0287] In some embodiments, the genetically modified polypeptide includes a DNA-binding domain. In some embodiments, the genetically modified polypeptide includes an RNA-binding domain. In some embodiments, the RNA-binding domain includes an RNA-binding domain of a B-box protein, an MS2 coat protein, a dCas, or an element of the sequence in the table herein. In some embodiments, the RNA-binding domain can bind to template RNA with higher affinity than a standard RNA-binding domain.
[0288] In some embodiments, the gene modification system can generate insertions at target sites of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally 500 or fewer, 400, 300, 200, or 100 nucleotides). In some embodiments, the gene modification system can generate insertions at target sites of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally 500, 400, 300, 200, or 100 nucleotides). In some embodiments, the gene modification system can generate insertions at target sites of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally 1, 5, 10, or 20 kilobases or less). In some embodiments, the gene modification system can generate deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less). In some embodiments, the gene modification system can generate deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less). In some embodiments, the gene modification system can generate deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally 500, 400, 300, or 200 nucleotides or less).In some embodiments, the gene modification system can generate deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally 1, 5, 10, or 20 kilobases or less). In some embodiments, the gene modification system can generate substitutions at target sites of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides or more. In some embodiments, the gene modification system can generate substitutions at target sites of 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotides.
[0289] In some embodiments, the substitution is a transposition mutation. In some embodiments, the substitution is a transversion mutation. In some embodiments, the substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.
[0290] In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by modifying, adding, or deleting sequences within promoters or enhancers, such as sequences that bind to transcription factors. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify gene translation (e.g., by modifying amino acid sequences), insert or delete start or stop codons, or modify or repair gene translation frames. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify gene splicing, for example, by inserting, deleting, or modifying splice acceptors or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof modify the half-life of transcripts or proteins. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in cells (e.g., from cytoplasm to mitochondria, from cytoplasm to extracellular space (e.g., adding secretory tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein folding (e.g., improve it) (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease the activity of a gene, for example, a protein encoded by that gene.
[0291] Exemplary gene-modified polypeptides and systems comprising them and methods of using them are described in PCT / US2021 / 020948, incorporated herein by reference with respect to retroviral RT domains, for example, including amino acids and nucleic acid sequences therein.
[0292] Exemplary gene-modified polypeptides and retroviral RT domain sequences are also described, for example, in International Patent Application No. PCT / US21 / 20948, filed on March 4, 2021, for example in Tables 30, 31, and 44 therein, and the entire application is incorporated herein by reference, for example, with respect to the aforementioned sequences and retroviral RTs in the tables. Accordingly, the gene-modified polypeptides described herein may include an amino acid sequence or its domain (e.g., a retroviral RT domain) according to any of the tables described in this paragraph, or a functional fragment or variant thereof of any of the above, or an amino acid sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0293] In some embodiments, polypeptides for use in any of the systems described herein may be molecular reconstitutes or ancestral reconstitutes based on aligned polypeptide sequences of multiple homologous proteins. In some embodiments, reverse transcriptase domains for use in any of the systems described herein may be molecular reconstitutes or ancestral reconstitutes, or may be modified with specific residues based on the alignment of reverse transcriptase domains from the same or different sources. Those skilled in the art can align polypeptide or nucleic acid sequences by using common sequence analysis tools, such as the Basic Local Alignment Search Tool (BLAST) or CD-Search for conserved domain analysis, based on the acceptance numbers provided herein. Molecular reconstitutes may be prepared based on sequence consensus using methods, for example, those described in Ivics et al., Cell 1997, 501-510; Wagstaff et al., Molecular Biology and Evolution 2013, 88-99.
[0294] Polypeptide components of gene modification systems In some embodiments, the gene-modified polypeptide has the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing, such as reverse transcription. In some embodiments, each function is contained within a different domain. In some embodiments, the function may be attributed to two or more domains (e.g., two or more domains exhibit functionality together). In some embodiments, two or more domains may have the same or similar functions (e.g., two or more domains independently have DNA binding functionality, for example, with two different DNA sequences). In other embodiments, one or more domains may perform one or more functions; for example, the Cas9 domain may be capable of both DNA binding and target site cleavage. In some embodiments, all domains are located within a single polypeptide. In some embodiments, the first domain is present in the first polypeptide, and the second domain is present in the second polypeptide. For example, in some embodiments, the sequence may be split between a first polypeptide and a second polypeptide, for instance, the first polypeptide comprising a reverse transcriptase (RT) domain and the second polypeptide comprising a DNA-binding domain and an endonuclease domain, such as a nickasase domain. In further examples, in some embodiments, each of the first and second polypeptides comprises a DNA-binding domain (e.g., a first DNA-binding domain and a second DNA-binding domain). In some embodiments, the first and second polypeptides may be joined post-translation via a split intein to form a single gene-modified polypeptide.
[0295] In some embodiments, the gene-modified polypeptides described herein include an St1Cas9 domain. The St1Cas9 domain may include a naturally occurring St1Cas9 amino acid sequence or a variant thereof. In some embodiments, the St1Cas9 domain is a nickase. In some embodiments, the St1Cas9 domain includes the sequence relating to SEQ ID NO: 23818 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide containing the St1Cas9 domain is used together with a compatible template RNA, including a variant gRNA scaffold described herein.
[0296] In some embodiments, the gene-modified polypeptides described herein (for example, the systems described herein include gene-modified polypeptides comprising): 1) a Cas domain (e.g., a Cas nickase domain, e.g., a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain from Table D or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto, wherein the RT domain is the C-terminus of the Cas domain; and a linker positioned between the RT domain and the Cas domain, comprising a sequence from the same row as the RT domain from Table D or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0297] In some embodiments, the RT domain has a sequence that is 100% identical to the RT domain in Table D, and the linker has a sequence that is 100% identical to the linker sequence from the same row as the RT domain in Table D. In some embodiments, the Cas domain includes the sequence in Table 8 or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identical thereto. In some embodiments, the gene-modified polypeptide includes an amino acid sequence that follows any of sequence numbers 1 to 3332 in the sequence listing or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.
[0298] In some embodiments, the gene-modified polypeptide includes a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain. In some embodiments, the gene-modified polypeptide includes the sequence of SEQ ID NO: 4000, which includes the first NLS and the Cas domain, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes the sequence of SEQ ID NO: 4001, which includes the second NLS, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
[0299] Exemplary N-terminal NLS-Cas9 domain [ka]
[0300] Exemplary C-terminal sequence including NLS AGKRTADGSEFEKRTADGSEFESPKKKAKVE (Sequence ID 4001)
[0301] Writing domain (RT domain) In certain embodiments of the present invention, the writing domain of the gene modification system has reverse transcriptase activity and is also referred to as the reverse transcriptase domain (RT domain). In some embodiments, the RT domain includes an RT catalytic moiety and an RNA-binding region (for example, a region that binds to template RNA).
[0302] In some embodiments, the nucleic acid encoding the reverse transcriptase is modified from its native sequence to have improved, modified codon usage, for example, for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a retrovirus. In some embodiments, the RT domain, which includes a genetically modified polypeptide, is mutated from its original amino acid sequence, for example, having substitutions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In some embodiments, the RT domain is derived from a retrovirus RT, for example, HIV-1 RT, Moloney Murine Leukemia Virus (MMLV) RT, avian myeloblastosis virus (AMV) RT, or Rous Sarcoma Virus (RSV) RT.
[0303] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits increased stringency for target-primed reverse transcription (TPRT) initiation compared to, for example, an endogenous RT domain. In some embodiments, the RT domain initiates TPRT when 3nts within the target site immediately upstream of the first strand nick, for example, the genomic DNA priming the RNA template, have at least 66% or 100% complementarity to 3nts of homology in the RNA template. In some embodiments, the RT domain initiates TPRT when less than 5nts of mismatch exist between the homology of the template RNA and the target DNA priming reverse transcription (e.g., less than 1, 2, 3, 4, or 5nts of mismatch). In some embodiments, the RT domain is modified to increase stringency in priming mismatches in the TPRT reaction, for example, the RT domain either does not tolerate any mismatches within the priming region or tolerates fewer mismatches compared to a wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain includes an HIV-1 RT domain. In several embodiments, the HIV-1 RT domain initiates synthesis at a lower level, even with three nucleotide mismatches, compared to alternative RT domains (e.g., as described by Jamburthugoda and Eickbush J Mol Biol 407(5):661-672 (2011) (the entire text of which is incorporated herein by reference)). In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or homodimer). In some embodiments, the RT domain is a monomer. In some embodiments, the RT domain functions naturally as a monomer or dimer (e.g., a heterodimer or homodimer). In some embodiments, the RT domain is derived from a virus that functions naturally as a monomer, for example, a monomer.In the embodiment, the RT domain is mouse leukemia virus (MLV, also known as MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary cancer virus (MMTV) (e.g., UniProt P03365), avian reticuloendotheliopathy virus (AVIRE) (e.g., UniProtKB accession: P03360), feline leukemia virus (FLV or FeLV) (e.g., UniProtKB accession: P10273), Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foam virus (HFV) (e.g., UniProt The RT domain or its functional fragment or variant (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto) is selected from P14350), monkey foam virus (SFV) (e.g., SFV3L) (e.g., UniProt P23074 or P27401), or bovine foam virus / syntiocytotic virus (BFV / BSV) (e.g., UniProt O41894). In some embodiments, the RT domain is dimerized in its innate functionality. In some embodiments, the RT domain is derived from a virus that functions as a dimer.In several embodiments, the RT domain is derived from avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastosis virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt The RT domain is selected from P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci 67(16):2717-2747(2010)), or a functional fragment or variant thereof (e.g., an amino acid sequence having at least 70%, 80%, 90%, 95%, or 99% identity thereto). In nature, heterodimeric RT domains may also function as homodimers in some embodiments. In some embodiments, the dimeric RT domain is expressed as a fusion protein, e.g., a homodimeric fusion protein or a heterodimeric fusion protein. In some embodiments, the RT function of the system is satisfied by multiple RT domains (e.g., as described herein).In further embodiments, multiple RT domains may be fused or separated and may reside, for example, on the same polypeptide or on different polypeptides.
[0304] In some embodiments, the gene modification system described herein includes an integrase domain, for example, the integrase domain may be part of the RT domain. In some embodiments, the RT domain (for example, as described herein) includes an integrase domain. In some embodiments, the RT domain (for example, as described herein) includes an integrase domain that lacks an integrase domain or is inactivated by mutation or deletion. In some embodiments, the gene modification system described herein includes a ribonuclease H domain, for example, the ribonuclease H domain may be part of the RT domain. In some embodiments, the RNase H domain is not part of the RT domain but is covalently linked via a flexible linker. In some embodiments, the RT domain (for example, as described herein) includes a ribonuclease H domain, for example, an endogenous ribonuclease H domain or a heterologous ribonuclease H domain. In some embodiments, the RT domain (for example, as described herein) lacks a ribonuclease H domain. In some embodiments, the RT domain (e.g., as described herein) comprises a ribonuclease H domain that has been added to, deleted from, mutated, or replaced with a heterologous ribonuclease H domain. In some embodiments, the polypeptide comprises an inactivated endogenous RNase H domain. In some embodiments, the endogenous RNase H domain from one of the other domains of the polypeptide is genetically removed so that it is not included in the polypeptide, for example, by partially or completely cleaving the endogenous RNase H domain from the polypeptide containing the domain. In some embodiments, a mutation in the ribonuclease H domain produces a polypeptide exhibiting lower ribonuclease activity, for example, by measurement by the method in Kotewicz et al. Nucleic Acids Res 16(1):265-277 (1988) (which is incorporated herein by reference in its entirety), compared to other similar domains without the mutation, by at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%.In some embodiments, ribonuclease H activity is lost.
[0305] In some embodiments, the RT domain undergoes mutation, resulting in increased fidelity compared to other similar domains that do not mutate. For example, in some embodiments, the YADD or YMDD motif within the RT domain (e.g., in the reverse transcriptase) is substituted with YVDD. In several embodiments, the substitution of YADD, YMDD, or YVDD results in higher fidelity in retroviral reverse transcriptase activity (e.g., described in Jamburthugoda and Eickbush J Mol Biol 2011, which is incorporated herein by reference in its entirety).
[0306] In some embodiments, the gene-modified polypeptides described herein include an RT domain having an amino acid sequence according to Table 6 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto. In some embodiments, the nucleic acids described herein encode an RT domain having an amino acid sequence according to Table 6 or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.
[0307] [Table 6-1]
[0308] [Table 6-2]
[0309] [Table 6-3]
[0310] [Table 6-4]
[0311] Table 6-5
[0312] Table 6-6
[0313] Table 6-7
[0314] Table 6-8
[0315] Table 6-9
[0316] Table 6-10
[0317] Table 6-11
[0318] Table 6-12
[0319] Table 6-13
[0320] Table 6-14
[0321] Table 6-15
[0322] Table 6-16
[0323] Table 6-17
[0324] Table 6-18
[0325] Table 6-19
[0326] In some embodiments, the reverse transcriptase domain is modified, for example, by site-directed mutation. In some embodiments, the reverse transcriptase domain is modified to have improved properties, such as SuperScript IV (SSIV) reverse transcriptase derived from MMLV RT. In some embodiments, the reverse transcriptase domain may be modified to have a lower error rate, for example, as described in International Publication No. 2001068895, which is incorporated herein by reference. In some embodiments, the reverse transcriptase domain may be modified to have increased thermal stability. In some embodiments, the reverse transcriptase domain may be modified to have increased processability. In some embodiments, the reverse transcriptase domain may be modified to have resistance to inhibitors. In some embodiments, the reverse transcriptase domain may be modified to be faster. In some embodiments, the reverse transcriptase domain may be modified to have increased resistance to modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain may be modified to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is modified to bind to template RNA. In some embodiments, one or more mutations are selected from D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, H8Y, T306K, or D653N in the RT domain of mouse leukemia virus reverse transcriptase, or from corresponding mutations at corresponding positions in other RT domains.
[0327] In some embodiments, the gene-modified polypeptide includes a RT domain from retroviral reverse transcriptase, for example, wild-type M-MLV RT containing the following sequence: M-MLV(WT): [ka]
[0328] In some embodiments, the gene-modified polypeptide includes an RT domain from retroviral reverse transcriptase, for example, M-MLV RT containing the following sequence: [ka]
[0329] In some embodiments, the genetically modified polypeptide comprises an RT domain from a retroviral reverse transcriptase containing the sequence of amino acids 659-1329 of NP_057933. In several embodiments, the genetically modified polypeptide further comprises one additional amino acid at the N-terminus of the sequence of amino acids 659-1329 of NP_057933, for example, as shown below. [ka] Core RT (bold), above annotation Ribonuclease H (underlined), see above note RNAseH (underlined), see above annotation
[0330] In some embodiments, the gene-modified polypeptide further comprises one additional amino acid at the C-terminus of the sequence of amino acids 659-1329 of NP_057933. In some embodiments, the gene-modified polypeptide comprises an RNaseH1 domain (e.g., amino acids 1178-1318 of NP_057933).
[0331] In some embodiments, the retroviral reverse transcriptase domain, e.g., M-MLV RT, may contain one or more mutations from the wild-type sequence that can improve the characteristics of the RT, e.g., thermal stability, processing capacity, and / or template binding. In some embodiments, the M-MLV RT domain may contain a combination of mutations such as D200N, L603W, and T330P, e.g., one or more mutations selected from the above M-MLV(WT) sequence, e.g., D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, K103L, e.g., D200N, L603W, and T330P, further optionally including T306K and W313F. In some embodiments, the M-MLV RT used herein includes the mutants D200N, L603W, T330P, T306K, and W313F. In several embodiments, the mutant M-MLV RT includes the following amino acid sequence. M-MLV(PE2): [ka]
[0332] In some embodiments, the writing domain (e.g., the RT domain) includes, for example, an RNA-binding domain that specifically binds to an RNA sequence. In some embodiments, the template RNA includes an RNA sequence that is specifically bound by the RNA-binding domain of the writing domain.
[0333] In some embodiments, the reverse transcription domain simply recognizes and reverse transcribes a specific template of the system, e.g., template RNA. In some embodiments, the template includes a sequence or structure that enables recognition and reverse transcription by the reverse transcription domain. In some embodiments, the template includes a sequence or structure that enables association of the polypeptide component of the genome modification system described herein with the RNA-binding domain. In some embodiments, the genome modification system preferentially reverse transcribes a template containing an association sequence over a template lacking the association sequence.
[0334] The writing domain may also include DNA-dependent DNA polymerase activity, such as enzymatic activity capable of writing DNA to the genome from a template DNA sequence. In some embodiments, DNA-dependent DNA polymerization is used to complete the second strand synthesis of the targeted site edit. In some embodiments, DNA-dependent DNA polymerase activity is provided by the DNA polymerase domain in the polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain capable of DNA-dependent DNA polymerization, such as the second strand synthesis. In some embodiments, DNA-dependent DNA polymerase activity is provided by the second polypeptide of the system. In some embodiments, DNA-dependent DNA polymerase activity is optionally provided by an endogenous host cell polymerase recruited to the target site by a component of the genome modification system.
[0335] In some embodiments, the reverse transcriptase domain exhibits a lower probability of insufficient termination (P) compared to the reference reverse transcriptase domain in vitro. off ) has. In some embodiments, the reference reverse transcriptase domain is the RT domain from a viral reverse transcriptase domain, for example, M-MLV.
[0336] In some embodiments, the reverse transcriptase domain, for example, measured in 1094 nt RNA, is approximately 5 × 10¹⁴ in vitro. -3 / nt, 5×10 -4 / nt or 5×10 -6 A lower probability of insufficient termination (P) is less than / nt. off ) has. In some embodiments, the insufficient termination rate in vitro is determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (which is incorporated herein by reference in its entirety).
[0337] In some embodiments, the reverse transcriptase domain can complete at least about 30% or 50% of the integration in cells. The percentage of complete integration can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites containing at least 98% of the expected integration sequence) by the total number of integration events (including substantially full-length and partial) in the cell population. In several embodiments, integration in cells is determined using long-read amplicon sequencing (e.g., through integration sites), as described, for example, in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).
[0338] In several embodiments, the quantification of integration in cells involves counting the percentage of integration that includes at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the DNA sequence corresponding to the template RNA (e.g., template RNA having lengths of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9, 1.0-1.2, 1.2-1.4, 1.4-1.6, 1.6-1.8, 1.8-2.0, 2-3, 3-4, or 4-5 kb).
[0339] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In several embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at rates of 0.1 to 50 nt / second (e.g., 0.1 to 1, 1 to 10, or 10 to 50 nt / second). In several embodiments, polymerization of dNTPs by the reverse transcriptase domain is measured by a single-molecule assay, for example, as described in Schwartz and Quake (2009) PNAS 106(48):20294-20299 (which is incorporated in its entirety by reference).
[0340] In some embodiments, the reverse transcriptase domain has, for example, an in vitro error rate of 1×10 -3 to 1×10 -4 or 1×10 -4 to 1×10 -5 substitutions / nt (e.g., nucleotide misincorporation). In some embodiments, the reverse transcriptase domain has, for example, an error rate of 1×10 -3 to 1×10 -4 or 1×10 -4 to 1×10 -5 substitutions / nt (e.g., nucleotide misincorporation) in cells (e.g., HEK293T cells), for example, by long-read amplicon sequencing, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903, which is incorporated herein by reference in its entirety.
[0341] In some embodiments, the reverse transcriptase domain is capable of performing reverse transcription of target RNA in vitro. In some embodiments, the reverse transcriptase requires at least a 3-nucleotide primer to initiate reverse transcription of the template. In some embodiments, reverse transcription of the target RNA is determined by detection of cDNA from the target RNA, as described in, for example, Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845, which is incorporated herein by reference in its entirety (e.g., when an ssDNA primer that anneals to the target having at least 3, 4, 5, 6, 7, 8, 9 or 10 nt at the 3' end is provided).
[0342] In some embodiments, the reverse transcriptase domain, for example, when converting its RNA template to cDNA, carries out reverse transcription at least 5 or 10 times more efficiently (e.g., by cDNA generation) compared to an RNA template lacking a protein-binding motif (e.g., 3'UTR). In several embodiments, the efficiency of reverse transcription is measured as described in Yasukawa et al. (2017) Biochem Biophys Res Commun 492(2):147-153 (which is incorporated herein by reference in its entirety).
[0343] In some embodiments, the reverse transcriptase domain, when expressed in cells (e.g., HEK293T cells), specifically binds to a particular RNA template at a higher frequency (e.g., about 5 or 10 times higher) than any endogenous cellular RNA. In several embodiments, the frequency of specific binding between the reverse transcriptase domain and template RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501 (which is incorporated herein by reference in its entirety).
[0344] Template nucleic acid binding domain Genetically modified polypeptides typically have a region capable of associating with a template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid-binding domain is an RNA-binding domain. In some embodiments, the RNA-binding domain is a modular domain capable of associating with RNA molecules containing a specific signature, such as a structural motif. In other embodiments, the template nucleic acid-binding domain (e.g., RNA-binding domain) is contained within a reverse transcription domain, where, for example, a component derived from reverse transcriptase has a known signature regarding RNA preference.
[0345] In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) is contained within the target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-related protein that recognizes the structure of a template nucleic acid (e.g., template RNA) containing gRNA. In some embodiments, the genetically modified polypeptide includes a DNA binding domain containing a CRISPR-related protein that associates with a gRNA scaffold, enabling the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the DNA binding domain is also the template nucleic acid binding domain, with the gRNA scaffold and gRNA spacer contained within the template nucleic acid (e.g., template RNA). In some embodiments, the polypeptide has RNA binding function within multiple domains, which may bind to gRNA structures within the CRISPR-related DNA binding domain and additional sequences or structures within the reverse transcriptase domain.
[0346] In some embodiments, the RNA-binding domain can bind to template RNA with higher affinity than the standard RNA-binding domain. In some embodiments, the standard RNA-binding domain is the RNA-binding domain derived from Cas9 of Streptococcus pyogenes. In some embodiments, the RNA-binding domain can bind to template RNA with affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM). In some embodiments, the affinity of the RNA-binding domain to the template RNA is measured in vitro, for example, by thermophoresis, as described in Asmari et al. Methods 146:107-119 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the affinity of the RNA-binding domain to the template RNA is measured in cells (e.g., by FRET or CLIP-Seq).
[0347] In some embodiments, the RNA-binding domain associates with template RNA at a frequency at least 5 to 10 times higher than that of scrambled RNA in vitro. In some embodiments, the frequency of association between the RNA-binding domain and template RNA or scrambled RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) Nucleic Acids Res 47(11):5490-5501 (which is incorporated herein by reference in its entirety). In some embodiments, the RNA-binding domain associates with intracellular template RNA (e.g., HEK293T cells) at a frequency at least 5 to 10 times higher than that of scrambled RNA. In some embodiments, the frequency of association between the RNA-binding domain and template RNA or scrambled RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) (cited above).
[0348] In some embodiments, the RT domain (for example, as listed in Table 6) contains one or more mutations listed in Table 2A below. In some embodiments, the RT domain as listed in Table 6 contains one, two, three, four, five, or six mutations listed in the corresponding rows of Table 2A below.
[0349] [Table 2A-1]
[0350] [Table 2A-2]
[0351] [Table 2A-3]
[0352] [Table 2A-4]
[0353] Endonuclease domain and DNA-binding domain In some embodiments, the genetically modified polypeptide has the function of cleaving a target DNA site via an endonuclease domain. In some embodiments, the genetically modified polypeptide includes, for example, a DNA-binding domain for binding to a target nucleic acid. In some embodiments, the domain of the genetically modified polypeptide (e.g., the Cas domain) includes two or more smaller domains, such as a DNA-binding domain and an endonuclease domain. Where it is stated that the DNA-binding domain (e.g., the Cas domain) binds to a target nucleic acid sequence, in some embodiments, it is understood that the binding is mediated by gRNA.
[0354] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-associated endonuclease domain that binds to a template RNA containing gRNA, binds to a target DNA sequence (e.g., having complementarity to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, an endonuclease domain or endonuclease / DNA-binding domain derived from a heterologous source may be used in or modified (e.g., by insertion, deletion, or substitution of one or more residues) in the gene modification systems described herein.
[0355] In some embodiments, the nucleic acid encoding the endonuclease domain or endonuclease / DNA binding domain is modified from its native sequence to have improved, modified codon usage, for example, for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9), a type II restriction endonuclease (e.g., Fok1), a meganuclease (e.g., I-SceI), or another endonuclease domain.
[0356] In certain embodiments, the DNA-binding domain of the gene-modified polypeptide described herein is selected, designed, or constructed for binding to a desired host DNA target sequence. In certain embodiments, the DNA-binding domain of the polypeptide is a heterogeneous DNA-binding factor. In some embodiments, the heterogeneous DNA-binding factor is a zinc finger factor or a TAL effector factor, e.g., a zinc finger or TAL polypeptide or a functional fragment thereof. In some embodiments, the heterogeneous DNA-binding factor is a sequence-inducible DNA-binding factor such as Cas9, Cpf1, or other CRISPR-related protein that has been modified to lack endonuclease activity. In some embodiments, the heterogeneous DNA-binding factor retains endonuclease activity. In some embodiments, the heterogeneous DNA-binding factor retains partial endonuclease activity such as cleaving ssDNA, e.g., having nickase activity. In certain embodiments, the heterogeneous DNA-binding domain may be one or more of Cas9, a TAL domain, a ZF domain, a Myb domain, a combination thereof, or a complex thereof.
[0357] In some embodiments, the DNA-binding domain is modified, for example, by site-directed mutation, to increase or decrease DNA-binding factors (e.g., the number and / or specificity of zinc fingers), thereby altering DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is modified from its native sequence to have improved, modified codon use, for example, in human cells. In several embodiments, the DNA-binding domain includes one or more modifications to the wild-type DNA-binding domain, such as modifications by directed evolution, such as phage-assisted continuous evolution (PACE).
[0358] In some embodiments, the DNA-binding domain comprises a meganuclease domain (e.g., an endonuclease domain portion, as described herein) or a functional fragment thereof. In some embodiments, the meganuclease domain has endonuclease activity, e.g., double-strand break and / or nickase activity. In other embodiments, the meganuclease domain has reduced activity, e.g., lack of endonuclease activity, e.g., the meganuclease is catalytically inactive. In some embodiments, a catalytically inactive meganuclease is used as the DNA-binding domain, e.g., as described in Fonfara et al. Nucleic Acids Res 40(2):847-860 (2012) (the whole of which is incorporated herein by reference).
[0359] In some embodiments, the genetically modified polypeptide includes modifications to the DNA-binding domain compared to, for example, the wild-type polypeptide. In some embodiments, the DNA-binding domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain is substituted for at least a portion (e.g., all) of the DNA-binding domain of the polypeptide. In some embodiments, the functional domain includes a zinc finger (e.g., a zinc finger that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest). In some embodiments, the functional domain includes a Cas domain (e.g., a Cas domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest). In some embodiments, the Cas domain includes Cas9 or its mutants or variants (e.g., as described herein). In several embodiments, the Cas domain associates with a guide RNA (gRNA), as described herein, for example. In several embodiments, the Cas domain is guided by the gRNA to a target nucleic acid (e.g., DNA) sequence of interest. In several embodiments, the Cas domain is encoded by the same nucleic acid (e.g., RNA) molecule as the gRNA. In several embodiments, the Cas domain is encoded by a different nucleic acid (e.g., RNA) molecule than the gRNA.
[0360] In some embodiments, the DNA-binding domain can bind to the target sequence (e.g., a dsDNA target sequence) with higher affinity than the standard DNA-binding domain. In some embodiments, the standard DNA-binding domain is a DNA-binding domain derived from Cas9 of Streptococcus pyogenes. In some embodiments, the DNA-binding domain can bind to the target sequence (e.g., a dsDNA target sequence) with affinity of 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM).
[0361] In some embodiments, the affinity of the DNA-binding domain for its target sequence (e.g., a dsDNA target sequence) is measured in vitro, for example, by thermophoresis, as described in Asmari et al. Methods 146:107-119 (2018) (the entire text of which is incorporated herein by reference).
[0362] In the embodiment, the DNA-binding domain can bind to its target sequence (e.g., a dsDNA target sequence) with an affinity of, for example, 100 pM to 10 nM (e.g., 100 pM to 1 nM or 1 nM to 10 nM) in the presence of a molar excess, for example, about 100-fold molar excess, of scrambled sequence-competitive dsDNA.
[0363] In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a higher frequency than any other sequence in the genome of the target cell, e.g., a human target cell, as measured, for example, by ChIP-seq (e.g., in HEK293T cells), as described, for example, in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21 (the entirety of which is incorporated herein by reference). In some embodiments, the DNA-binding domain is found to bind to its target sequence (e.g., a dsDNA target sequence) at a frequency at least about 5 or 10 times higher than any other sequence in the genome of the target cell, as measured, for example, by ChIP-seq (e.g., in HEK293T), as described, for example, in He and Pu (2010) (cited above).
[0364] In some embodiments, the endonuclease domain has nickase activity and cleaves one strand of target DNA. In some embodiments, the nickase activity reduces the formation of double-strand breaks at the target site. In some embodiments, the endonuclease domain generates a twisted nick structure in the first and second strands of target DNA. In some embodiments, the twisted nick structure generates a free 3' overhang at the target site. In some embodiments, the free 3' overhang at the target site improves editing efficiency, for example, by enhancing access to and annealing of the 3' homology region of the template nucleic acid. In some embodiments, the twisted nick structure reduces the formation of double-strand breaks at the target site.
[0365] In some embodiments, the endonuclease domain cleaves both strands of target DNA, resulting in a blunt-end cleavage of the target that, for example, does not have an ssDNA overhang on either side of the cleavage site. The amino acid sequences of the endonuclease domains of the gene modification systems described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, and at least about 99% identical to the amino acid sequences of the endonuclease domains described herein, for example, from Table 8.
[0366] In certain embodiments, the heterologous endonucleases are Fok1 or a functional fragment thereof. In certain embodiments, the heterologous endonucleases are Holliday junction resolvers or their homologs, e.g., Holliday junction cleavage enzyme (Ssol Hje) from Sulfolobus solfataricus (Govindaraju et al., Nucleic Acids Research 44:7, 2016). In certain embodiments, the heterologous endonucleases are endonucleases of large fragments of spliceosome proteins such as Prp8 (Mahbub et al., Mobile DNA 8:16, 2017). In certain embodiments, the heterologous endonucleases are derived from CRISPR-related proteins, e.g., Cas9. In certain embodiments, heterologous endonucleases are modified to possess only ssDNA cleavage activity, for example, only nickase activity, such as Cas9 nickase, or SpCas9 with D10A, H840A, or N863A mutations. Table 8 shows exemplary Cas proteins and mutations associated with nickase activity. In yet another embodiment, homologous endonuclease domains are modified to alter DNA endonuclease activity, for example, by site-directed mutation. In yet another embodiment, endonuclease domains are modified to reduce DNA sequence specificity, for example, by cleavage to remove domains that confer DNA sequence specificity or mutations, in order to inactivate the region that confers DNA sequence specificity.
[0367] In some embodiments, the endonuclease domain has nickase activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks more frequently than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease substantially does not form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.
[0368] In some embodiments, the endonuclease domain has nickase activity that cleaves target site DNA on the first strand; for example, in some embodiments, the endonuclease domain cleaves genomic DNA at a target site near a modification site on the strand that will be extended by the writing domain. In some embodiments, the endonuclease domain has nickase activity that cleaves target site DNA on the first strand but not on target site DNA on the second strand. For example, when the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain cleaves target site DNA strands containing PAM sites (and, for example, does not cleave target site DNA strands that do not contain PAM sites). As a further example, when the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain cleaves the target site DNA strand that does not contain a PAM site (but does not cleave the target site DNA strand that does contain a PAM site, for example).
[0369] In some other embodiments, the endonuclease domain has nickase activity that cleaves the target site DNA of the first and second strands. Although not intended to be bound by any particular theory, after the writing domain (e.g., RT domain) of the polypeptide described herein is polymerized (e.g., reverse transcribed) from a heterologous target sequence of the template nucleic acid (e.g., template RNA), the cellular DNA repair mechanism must repair the nick on the first DNA strand. The target site DNA here comprises two distinct sequences relative to the first DNA strand: the first corresponding to the original genomic DNA (e.g., with a free 5' end) and the second corresponding to the one polymerized from the heterologous target sequence (e.g., with a free 3' end). It is assumed that the two distinct sequences equilibrate with each other, and that one hybridizes with the second strand first, then the other, and the order of their incorporation into its repair target site by the cellular DNA repair mechanism is considered to be a stochastic process. While not intended to be bound by any particular theory, it is thought that the introduction of further nicks into the second strand may bias cellular DNA repair mechanisms to use sequences based on heterologous target sequences more frequently than the original genome sequence (Anzalone et al. Nature 576:149-157 (2019)). In some embodiments, the further nicks are located at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides relative to the 5' or 3' of the target site modification (e.g., insertion, deletion, or substitution) or the nicks on the first strand.
[0370] Alternatively or additionally, and without intending to be bound by any particular theory, it is thought that further nicks to the second strand may facilitate the synthesis of the second strand. In some embodiments, if the gene modification system inserts or substitutes a portion of the first strand, the synthesis of a new sequence corresponding to the insertion / substitution in the second strand is required.
[0371] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain), the domain cleaving both a first and a second strand. For example, in such embodiments, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that leads to nicking of the first strand and a further gRNA spacer that leads to nicking of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, the first endonuclease domain cleaving the first strand and the second endonuclease domain cleaving the second strand (optionally, the first endonuclease domain may not cleaving the second strand (e.g., is unable to do so) and the second endonuclease domain may not cleaving the first strand (e.g., is unable to do so)).
[0372] In some embodiments, the endonuclease domain can create nicks in the first and second chains. In some embodiments, the first and second chain nicks occur at the same location in the target site but not on the reverse chain. In some embodiments, the second chain nick occurs at a twisted location, for example, upstream or downstream of the first nick. In some embodiments, the endonuclease domain generates a deletion in the target site if the second chain nick is upstream of the first chain nick. In some embodiments, the endonuclease domain generates a duplication of the target site if the second chain nick is downstream of the first chain nick. In some embodiments, the endonuclease domain does not generate duplication and / or deletion if the first and second chain nicks occur at the same location in the target site. In some embodiments, the endonuclease domain has modified activity depending on the protein structure or RNA binding state, for example (as described in Christensen et al. PNAS 2006, the whole of which is incorporated herein by reference), which promotes nicking of the first or second strand.
[0373] In some embodiments, the endonuclease domain includes a meganuclease or a functional fragment thereof. In some embodiments, the endonuclease domain includes a homing endonuclease or a functional fragment thereof. In some embodiments, the endonuclease domain includes a meganuclease or a functional fragment or variant thereof from the LAGLIDADG, GIY-YIG, HNH, His-Cys Box, or PD-(D / E)XK families, for example, having a conserved amino acid motif as indicated by the family name. In some embodiments, the endonuclease domain comprises a meganuclease or fragment thereof selected from, for example, I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI (Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6). In some embodiments, the meganuclease is naturally present in its functional form as a monomer, e.g., I-SceI, I-TevI, or a dimer, e.g., I-CreI. For example, LAGLIDADG meganucleases having a single copy of the LAGLIDADG motif generally form homodimers, while members having two copies of the LAGLIDADG motif are generally found as monomers. In some embodiments, meganucleases that normally form as dimers are expressed as fusions, for example, as an I-CreI dimer fusion, where the two subunits are optionally linked together by a linker as a single ORF (Rodriguez-Fornes et al. Gene Therapy 2020, the whole of which is incorporated herein by reference).In some embodiments, the meganuclease or functional fragment thereof is modified to preferentially target nickase activity on one strand of a double-stranded DNA molecule, such as I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), or I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, the meganuclease or functional fragment thereof having this preference for single-strand breaks is used, for example, as an endonuclease domain having nickase activity. In some embodiments, the endonuclease domain includes a meganuclease or functional fragment thereof that naturally targets or is modified to target a safe port site, such as an SH6 site targeting I-CreI (Rodriguez-Fornes et al., above). In some embodiments, the endonuclease domain includes I-TevI, which recognizes a meganuclease or functional fragment thereof having a sequence-resistant catalytic domain, such as the minimal motif CNNNG (Kleinstiver et al. PNAS 2012). In some embodiments, the target sequence-resistant catalytic domain is fused to a DNA-binding domain, and activity is induced, for example, by fusing I-TevI to (i) a Zn finger for producing Tev-ZFE (Kleinstiver et al. PNAS 2012), (ii) another meganuclease for producing MegaTev (Wolfs et al. Nucleic Acids Res 2014), and / or (iii) Cas9 for producing TevCas9 (Wolfs et al. PNAS 2016).
[0374] In some embodiments, the endonuclease domain includes a restriction enzyme, such as a type IIS or type IIP restriction enzyme. In some embodiments, the endonuclease domain includes a type IIS restriction enzyme, such as FokI or a fragment or variant thereof. In some embodiments, the endonuclease domain includes a type IIP restriction enzyme, such as PvuII or a fragment or variant thereof. In some embodiments, the dimeric restriction enzyme is expressed as a fusion, such as a FokI dimer fusion, so that it functions as a single chain (Minczuk et al. Nucleic Acids Res 36(12):3926-3938(2008)).
[0375] Further uses of endonuclease domains are described, for example, in Guha and Edgell Int J Mol Sci 18(22):2565(2017) (the entire text of which is incorporated herein by reference).
[0376] In some embodiments, the genetically modified polypeptide includes modifications to the endonuclease domain compared to, for example, the wild-type Cas protein. In some embodiments, the endonuclease domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the wild-type Cas protein. In some embodiments, the endonuclease domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest and / or induces its endonuclease cleavage. In some embodiments, the endonuclease domain includes a zinc finger. In several embodiments, the endonuclease domain, including the Cas domain, associates with a guide RNA (gRNA), for example, as described herein. In some embodiments, the endonuclease domain is modified to include a functional domain that does not target a specific target nucleic acid (e.g., DNA) sequence. In several embodiments, the endonuclease domain includes a Fok1 domain.
[0377] In some embodiments, the endonuclease domain associates with target dsDNA at a frequency at least approximately 5 or 10 times higher than that of scrambled dsDNA in vitro. In some embodiments, the endonuclease domain associates with target dsDNA at a frequency at least approximately 5 or 10 times higher than that of scrambled dsDNA in vitro, for example, within cells (e.g., HEK293T cells). In some embodiments, the frequency of association between the endonuclease domain and target DNA or scrambled DNA is measured by ChIP-seq, for example, as described in He and Pu (2010) Curr. Protoc Mol Biol Chapter 21 (the entirety of which is incorporated herein by reference).
[0378] In some embodiments, the endonuclease domain can catalyze nick formation at target sequences by at least about 5-fold or 10-fold compared to, for example, non-target sequences (compared to, for example, any other genomic sequences in the genome of the target cell). In some embodiments, the level of nick formation is measured using nickSeq, for example, as described in Elacqua et al. (2019) bioRxiv doi.org / 10.1101 / 867937 (which is incorporated herein by reference in its entirety).
[0379] In some embodiments, the endonuclease domain allows for in vitro nicking of DNA. In several embodiments, the nicking results in exposed bases. In several embodiments, the exposed bases can be detected using a nuclease sensitivity assay, for example, as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res 23(19):3805-3809 (which is incorporated herein by reference in its entirety). In several embodiments, the level of exposed bases (detected, for example, by a nuclease sensitivity assay) is increased by at least 10%, 50%, or more compared to a reference endonuclease domain. In some embodiments, the reference endonuclease domain is an endonuclease domain derived from Cas9 of Streptococcus pyogenes.
[0380] In some embodiments, the endonuclease domain is capable of nicking intracellular DNA. In several embodiments, the endonuclease domain is capable of nicking DNA in HEK293T cells. In several embodiments, unrepaired nicks that replicate in the absence of Rad51 result in an increased NHEJ rate at the nick site, which can be detected, for example, by using a Rad51 inhibition assay as described in Bothmer et al. (2017) Nat Commun 8:13905 (which is incorporated herein by reference in its entirety). In several embodiments, the NHEJ rate increases by 0 to over 5%. In several embodiments, the NHEJ rate increases to, for example, 20 to 70% (e.g., 30% to 60% or 40 to 50%) upon Rad51 inhibition.
[0381] In some embodiments, the endonuclease domain releases the target after cleavage. In some embodiments, the release of the target is indirectly indicated by evaluating multiple turnovers by the enzyme, for example, as described in Yourik at al. RNA 25(1):35-44(2019) (the whole of which is incorporated herein by reference) and as shown in Figure 2. In some embodiments, the k of the endonuclease domain exp This can be measured using these methods and yield 1 × 10⁻⁶ -3 ~1 × 10 -5 It is min-1.
[0382] In some embodiments, the endonuclease domain is approximately 1 × 10⁶ in vitro. 8 s -1 M -1 Catalytic efficiency exceeding (k cat / K m ) has. In some embodiments, the endonuclease domain is approximately 1 × 10 in vitro. 5 , 1 x 10 6 , 1 x 10 7 or 1 × 10 8 s -1 M -1 It has a catalytic efficiency exceeding 1 × 10⁻¹⁶. In several embodiments, the catalytic efficiency is determined as described in Chen et al. (2018) Science 360(6387):436-439 (which is incorporated herein by reference in its entirety). In some embodiments, the endonuclease domain has a catalytic efficiency of approximately 1 × 10⁻¹⁶ in the cell. 8 s -1 M -1 Catalytic efficiency exceeding (k cat / K m ) has. In some embodiments, the endonuclease domain is located in the cell at approximately 1 × 10⁻¹⁶ 5 , 1 x 10 6 , 1 x 10 7 or 1 × 10 8 s -1 M -1 It has a catalytic efficiency exceeding [a certain level].
[0383] Genetically modified polypeptides containing the Cas domain In some embodiments, the genetically modified polypeptide described herein includes a Cas domain. In some embodiments, the Cas domain can guide the genetically modified polypeptide to a target site identified by a gRNA spacer, thereby modifying the target nucleic acid sequence at "cis". In some embodiments, the genetically modified polypeptide is fused to the Cas domain. In some embodiments, the genetically modified polypeptide includes a CRISPR / Cas domain (also referred to herein as a CRISPR-associated protein). In some embodiments, the CRISPR / Cas domain includes a protein involved in a clustered, regularly arranged short palindromic sequence repeat (CRISPR) system, such as a Cas protein, which optionally binds to a guide RNA, such as a single guide RNA (sgRNA).
[0384] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. The CRISPR system uses RNA-guided nucleases, referred to as CRISPR-associated or "Cas" endonucleases (e.g., Cas9 or Cpf1), to cleave foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is guided to a target nucleotide sequence (e.g., a site in the genome to be sequence-edited) by a sequence-specific, non-coding "guide RNA" that targets a single-stranded or double-stranded DNA sequence. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One class II CRISPR system includes type II Cas endonucleases such as Cas9, CRISPR RNA ("crRNA"), and transactivating crRNA ("tracrRNA"). crRNA typically contains a "spacer" sequence (protospacer), which is an RNA sequence of about 20 nucleotides that corresponds to the target DNA sequence. In the wild-type system and in some modified systems, crRNA forms a partially double-stranded structure that binds to tracrRNA and is cleaved by ribonuclease III. cThis also includes a region that gives rise to the rRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then guides the Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is generally specific to a given Cas endonuclease and is adjacent to a "protospacer-adjacent motif" ("PAM") required for cleavage activity at the target site corresponding to the crRNA spacer. CRISPR endonucleases identified from various prokaryotic species have specific PAM sequence requirements, as listed for example in Table 7 for the Cas enzymes. Examples of PAM sequences include 5'-NGG (Streptococcus pyogenes, SEQ ID NO: 11,019), 5'-NNAGAA (Streptococcus thermophilus CRISPR1, SEQ ID NO: 11,020), 5'-NGGNG (Streptococcus thermophilus CRISPR3, SEQ ID NO: 11,021), and 5'-NNNGATT (Neisseria meningiditis, SEQ ID NO: 11,022). Some endonucleases, such as Cas9 endonuclease, associate with G-rich PAM sites, e.g., 5'-NGG, SEQ ID NO: 11,023, and perform blunt-end cleavage of target DNA at a position of three nucleotides upstream (from 5') of the PAM site. Another class II CRISPR system includes a smaller V-type endonuclease Cpf1 than Cas9, exemplified by AsCpf1 (from Acidaminococcus sp.) and LbCpf1 (from Lachnospiraceae sp.). Cpf1-associated CRISPR arrays do not require tracrRNA and are processed with mature crRNA; in other words, the Cpf1 system, in some embodiments, consists solely of the Cpf1 nuclease and crRNA to cleave the target DNA sequence. Cpf1 endonucleases typically associate with T-rich PAM sites, e.g., 5'-TTN. Cpf1 can also recognize the 5'-CTA PAM motif.Cpf1 typically cleaves target DNA by introducing offset or twisted double-strand breaks in the 5' overhangs of 4- or 5-nucleotides, for example, 18 nucleotides downstream (3' side) from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand. The resulting 5-nucleotide overhangs enable more precise genome editing by DNA insertion via homologous recombination rather than insertion via blunt-end DNA. See, for example, Zetsche et al. (2015) Cell, 163:759-771.
[0385] Various CRISPR-related (Cas) genes or proteins may be used in the techniques provided herein, and the selection of Cas proteins will depend on the specific conditions of the method. Specific examples of Cas proteins include the Class II system comprising Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein, e.g., the Cas9 protein, may originate from any of various prokaryotic species. In some embodiments, a specific Cas protein, e.g., a specific Cas9 protein, is selected to recognize a specific protospacer-adjacent motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain includes a sequence-targeting polypeptide such as a Cas protein, e.g., Cas9. In certain embodiments, the Cas protein, e.g., the Cas9 protein, may be obtained from bacteria or archaea, or synthesized using known methods. In certain embodiments, the Cas protein may originate from Gram-positive or Gram-negative bacteria. In certain embodiments, the Cas protein is found in streptococci (e.g., Streptococcus pyogenes or S. thermophilus), Francisella (e.g., F. nobicida), Staphylococcus (e.g., Staphylococcus aureus), and Acidaminococcus (e.g., Acidaminococcus species BV3L6). It may originate from sp.BV3L6), Neisseria (e.g., Neisseria meningitidis), Cryptococcus, Corynebacterium, Haemophilus, Eubacterium, Pasteurella, Prevotella, Veillonella, or Marinobacter.
[0386] In some embodiments, the gene-modified polypeptide may include the amino acid sequence of SEQ ID NO: 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the amino acid sequence of SEQ ID NO: 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto is located at the N-terminus of the gene-modified polypeptide. In several embodiments, the amino acid sequence of Sequence ID No. 4000 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids from the N-terminus of the gene-modified polypeptide.
[0387] Exemplary N-terminal NLS-Cas9 domain [ka]
[0388] In some embodiments, the gene-modified polypeptide may include the amino acid sequence of SEQ ID NO: 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the amino acid sequence of SEQ ID NO: 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto is located at the C-terminus of the gene-modified polypeptide. In several embodiments, the amino acid sequence of Sequence ID No. 4001 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids from the C-terminus of the gene-modified polypeptide.
[0389] Exemplary C-terminal sequence including NLS AGKRTADGSEFEKRTADGSEFESPKKKAKVE (Sequence ID 4001)
[0390] Exemplary benchmark array [ka] [ka]
[0391] In some embodiments, the gene-modified polypeptide may include a Cas domain or a functional fragment thereof listed in Table 7 or 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0392] [Table 7]
[0393] [Table 8-1]
[0394] [Table 8-2]
[0395] [Table 8-3]
[0396] In some embodiments, the Cas protein is required to have a protospacer-adjacent motif (PAM) located within or adjacent to the target DNA sequence to which the Cas protein binds and / or functions. In some embodiments, the PAM is or includes NGG (SEQ ID NO: 11,024), YG (SEQ ID NO: 11,025), NNGRRT (SEQ ID NO: 11,026), NNNRRT (SEQ ID NO: 11,027), NGA (SEQ ID NO: 11,029), TYCV (SEQ ID NO: 11,030), TATV (SEQ ID NO: 11,031), NTTN (SEQ ID NO: 11,032), or NNNGATT (SEQ ID NO: 11,033), where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A, C, or G. In some embodiments, the Cas protein is a protein listed in Table 7 or 8. In some embodiments, the Cas protein includes one or more mutations that modify its PAM. In some embodiments, the Cas protein includes mutations E1369R, E1449H and R1556A or analogous substitutions to the amino acids corresponding to the said positions. In some embodiments, the Cas protein includes mutations E782K, N968K and R1015H or analogous substitutions to the amino acids corresponding to the said positions. In some embodiments, the Cas protein includes mutations D1135V, R1335Q and T1337R or analogous substitutions to the amino acids corresponding to the said positions. In some embodiments, the Cas protein includes mutations S542R and K607R or analogous substitutions to the amino acids corresponding to the said positions. In some embodiments, the Cas protein includes mutations S542R, K548V and N552R or analogous substitutions to the amino acids corresponding to the said positions. Exemplary advances in the modification of Cas enzymes for recognizing altered PAM sequences are outlined in Collias et al Nature Communications 12:555 (2021), which are incorporated herein by reference in their entirety.
[0397] In some embodiments, the Cas protein is catalytically active and cleaves one or both strands of a target DNA site. In some embodiments, following the cleavage of the target DNA site, modifications, such as insertions or deletions, are formed, for example, by cellular repair mechanisms.
[0398] In some embodiments, Cas proteins are modified to inactivate or partially inactivate nucleases, such as nuclease-deficient Cas9. While wild-type Cas9 produces double-strand breaks (DSBs) at specific DNA sequences targeted by gRNA, several CRISPR endonucleases with modified functionality are available; for example, a “nickase” version of partially inactivated Cas9 produces only single-strand breaks, and catalytically inactive Cas9 ("dCas9") does not cleave target DNA. In some embodiments, binding of dCas9 to a DNA sequence may interfere with transcription at that site due to steric hindrance. In some embodiments, binding of dCas9 to an anchor sequence may interfere with (e.g., reduce or block) the formation and / or maintenance of genomic complexes (e.g., ASMCs). In some embodiments, the DNA-binding domain includes catalytically inactive Cas9, e.g., dCas9. Numerous catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 contains mutations, e.g., D10A and H840A or N863A mutations, within each endonuclease domain of the Cas protein. In some embodiments, catalytically inactive or partially inactive CRISPR / Cas domains contain one or more mutations, e.g., one or more of the mutations listed in Table 7, within the Cas protein. In some embodiments, a Cas protein listed in a given row of Table 7 contains one, two, three, or all of the mutations listed in the same row of Table 7. In some embodiments, for example, a Cas protein not listed in Table 7 contains one, two, three, or all of the mutations listed in a certain row of Table 7, or the corresponding mutation at the corresponding site in that Cas protein.
[0399] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, include a D11 mutation (e.g., a D11A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, include an H969 mutation (e.g., an H969A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, include an N995 mutation (e.g., an N995A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, include mutations at one, two, or three of the D11, H969, and N995 positions (e.g., D11A, H969A, and N995A mutations) or similar substitutions to the amino acids corresponding to the said positions.
[0400] In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein includes a D10 mutation (e.g., a D10A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, or a partially inactivated Cas9 protein includes an H557 mutation (e.g., an H557A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, a catalytically inactive Cas9 protein, e.g., dCas9, includes a D10 mutation (e.g., a D10A mutation) and an H557 mutation (e.g., an H557A mutation) or a similar substitution to the amino acid corresponding to the said position.
[0401] In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include a D839 mutation (e.g., a D839A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include an H840 mutation (e.g., a H840A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include an N863 mutation (e.g., a N863A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, include a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), a H840 mutation (e.g., H840A), and a N863 mutation (e.g., N863A) or a similar substitution to the amino acid corresponding to the said position.
[0402] In some embodiments, a catalytically inactive Cas9 protein, such as dCas9, or a partially inactivated Cas9 protein, includes an E993 mutation (e.g., an E993A mutation) or an analogous substitution to the amino acid corresponding to the said position.
[0403] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins include a D917 mutation (e.g., a D917A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins include an E1006 mutation (e.g., an E1006A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins include a D1255 mutation (e.g., a D1255A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, include a D917 mutation (e.g., D917A), an E1006 mutation (e.g., E1006A), and a D1255 mutation (e.g., D1255A) or a similar substitution to the amino acid corresponding to the said position.
[0404] In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include a D16 mutation (e.g., a D16A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include a D587 mutation (e.g., a D587A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, a partially inactivated Cas domain has nickase activity. In some embodiments, a partially inactivated Cas9 domain is a Cas9 nickase domain. In some embodiments, a catalytically inactive Cas domain or an inactive Cas domain does not form a detectable double-strand break. In some embodiments, catalytically inactive Cas9 proteins, e.g., dCas9, or partially inactivated Cas9 proteins include an H588 mutation (e.g., a H588A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins include an N611 mutation (e.g., an N611A mutation) or a similar substitution to the amino acid corresponding to the said position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, include a D16 mutation (e.g., D16A), a D587 mutation (e.g., D587A), an H588 mutation (e.g., H588A), and an N611 mutation (e.g., N611A) or a similar substitution to the amino acid corresponding to the said position.
[0405] In some embodiments, the DNA-binding domain or endonuclease domain may include a gRNA (e.g., a template nucleic acid containing gRNA, e.g., template RNA) or a Cas molecule linked to it (e.g., covalently).
[0406] In some embodiments, the endonuclease domain or DNA-binding domain includes Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain includes modified SpCas9. In several embodiments, the modified SpCas9 includes modifications that alter the protospacer-adjacent motif (PAM) specificity. In several embodiments, the PAM has specificity for the nucleic acid sequence 5'-NGT-3'. In several embodiments, the modified SpCas9 includes one or more amino acid substitutions at one or more positions of, for example, L1111R, D1135, G1218, E1219, A1322, or R1335, which are selected from, for example, L1111R, D1135V, G1218R, E1219F, A1322R, and R1335V. In several embodiments, the modified SpCas9 includes the amino acid substitution T1337R and one or more other amino acid substitutions selected from L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M or the corresponding amino acid substitutions. In several embodiments, the modified SpCas9 includes (i) one or more amino acid substitutions selected from D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q and T1337, and (ii) one or more other amino acid substitutions selected from L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q and T1337M or the corresponding amino acid substitutions.
[0407] In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas domain, such as a Cas9 domain. In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-inactive Cas (dCas). In several embodiments, the endonuclease domain or DNA-binding domain includes a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-inactive Cas9 (dCas). In some embodiments, the endonuclease domain or DNA-binding domain includes the Cas9 domain of Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain includes Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain includes Streptococcus pyogenes or S. thermophilus Cas9 or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas9 sequence, for example, as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737 (incorporated herein by reference). In some embodiments, the endonuclease domain or DNA-binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of Cas, for example, Cas9 or a variant thereof, as described herein.In some embodiments, the endonuclease domain or DNA-binding domain includes Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain includes a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof. In several embodiments, the Cas polypeptide (e.g., an enzyme) includes Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3 , Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, C sm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Cs a5, selected from type II Cas effector protein, type V Cas effector protein, type VI Cas effector protein, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9(K855A), eSpCas9(1.1), SpCas9-HF1, hyper-accurate Cas9 variant (HypaCas9), their homologs, modified or manipulated versions thereof, and / or their functional fragments.In some embodiments, Cas9 includes one or more substitutions selected from, for example, H840A, D10A, P475A, W476A, N477A, D1125A, W1126A, and D1127A. In some embodiments, Cas9 includes one or more mutations at positions selected from D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, including one or more substitutions selected from, for example, D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. In some embodiments, the endonuclease domain or DNA-binding domain is associated with Corynebacterium ulcerans, Corynebacterium diphtheria, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanense, Streptococcus iniae, Bellilla baltica, Psychroflexus torquis, Streptococcus thermophilus, Listeria innocua, Campylobacter jejuni, and Neisseria meningitidis. It contains a Cas (e.g., Cas9) sequence derived from *Streptococcus meningitidis*, *Streptococcus pyogenes*, or *Staphylococcus aureus*, or a functional fragment or mutant thereof.
[0408] In some embodiments, the endonuclease domain or DNA-binding domain includes a Cpf1 domain, for example, at position: D917, E1006A, D1255 or any combination thereof, comprising one or more substitutions selected from, for example, D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A and D917A / E1006A / D1255A.
[0409] In some embodiments, the endonuclease domain or DNA-binding domain includes spCas9, spCas9-VRQR (SEQ ID NO: 5019), spCas9-VRER (SEQ ID NO: 5020), xCas9(sp), saCas9, saCas9-KKH, spCas9-MQKSER (SEQ ID NO: 5021), spCas9-LRKIQK (SEQ ID NO: 5022), or spCas9-LRVSQL (SEQ ID NO: 5023).
[0410] In some embodiments, the gene-modified polypeptide has an endonuclease domain containing Cas9 niccas, for example, Cas9 H840A. In several embodiments, Cas9 H840A has the following amino acid sequence. Cas9 nickas (H840A): [ka]
[0411] In some embodiments, the gene-modified polypeptide includes a dCas9 sequence containing the D10A and / or H840A mutation, for example, the following sequence. [ka]
[0412] TAL effect pedal and zinc finger nuclease In some embodiments, the endonuclease domain or DNA-binding domain comprises a TAL effector molecule. A TAL effector molecule, such as one that specifically binds to a DNA sequence, typically comprises multiple TAL effector domains or fragments thereof and, optionally, one or more further parts of a naturally occurring TAL effector (e.g., the N-terminus and / or C-terminus of multiple TAL effector domains). Many TAL effectors are known to those skilled in the art and are commercially available, for example, from Thermo Fisher Scientific.
[0413] Naturally occurring TAL effector proteins are naturally occurring effector proteins secreted by a great many bacterial pathogen species, including the plant pathogen Xanthomonas, that regulate gene expression in host plants and promote bacterial colonization and survival. Specific binding of TAL effectors is typically based on a central repeat domain (repeating variable duodenum, RVD domain) of nearly identical repeats arranged in tandem of 33 or 34 amino acids.
[0414] Members of the TAL effector family differ primarily in the number and order of their repeats. The number of repeats typically ranges from 1.5 to 33.5, with the C-terminal repeat usually being shorter (e.g., about 20 amino acids) and commonly referred to as a "half-repeat." Each repeat in a TAL effector generally features a one-repeat-to-one base-pair correlation with different repeat types exhibiting different base-pair specificity (one repeat recognizes one base pair on the target gene sequence). Generally, a decrease in the number of repeats weakens the protein-DNA interaction. Several 6.5-repeat sequences have been shown to be sufficient to activate the transcription of reporter genes (Scholze et al., 2010).
[0415] The variation between repeats primarily occurs at amino acid positions 12 and 13, and is therefore referred to as "hypervariable," as shown in Table 9, which lists exemplary repeat variable two residues (RVDs) and their correspondences to nucleic acid base targets, and is involved in the specificity of interaction with target DNA promoter sequences.
[0416] [Table 9]
[0417] Therefore, it is possible to modify TAL effector repeats and target specific DNA sequences. Furthermore, studies have shown that RVD NK can target G. Additionally, TAL effector target sites tend to contain T adjacent to the 5' base targeted by the first repeat, although the exact mechanism of this recognition is unknown. More than 113 TAL effector sequences are known to date. Non-exclusive examples of TAL effectors from Xanthomonas include Hax2, Hax3, Hax4, AvrXa7, AvrXa10, and AvrBs3.
[0418] Accordingly, the TAL effector domain of the TAL effector molecules described herein may be derived from a TAL effector from any bacterial species (e.g., Xanthomonas species such as the African strain of Xanthomonas oryzae pv. Oryzae (Yu et al. 2011), Xanthomonas campestris pv. raphani strain 756C, and Xanthomonas oryzae pv. oryzicola BLS256 (Bogdanove et al. 2011)). In some embodiments, the TAL effector domain also includes an RVD domain and adjacent sequences (N-terminal and / or C-terminal sequences of the RVD domain) from a naturally occurring TAL effector. It may contain more or fewer repeats than the RVD of a naturally occurring TAL effector. TAL effector molecules can be designed to target a given DNA sequence based on the above-mentioned codes known in the art. The number of TAL effector domains (e.g., repeats (monomers or modules)) and their specific sequences can be selected based on the desired DNA target sequence. For example, TAL effector domains, e.g., repeats, can be removed or added to suit a particular target sequence. In one embodiment, the TAL effector molecule of the present invention comprises 6.5 to 33.5 TAL effector domains, e.g., repeats. In another embodiment, the TAL effector molecule of the present invention comprises 8 to 33.5 TAL effector domains, e.g., repeats, e.g., 10 to 25 TAL effector domains, e.g., repeats, e.g., 10 to 14 TAL effector domains, e.g., repeats.
[0419] In some embodiments, the TAL effector molecule includes a TAL effector domain corresponding to a perfect match with the DNA target sequence. In some embodiments, mismatches between repeats on the DNA target sequence and target base pairs are acceptable as long as they enable the function of the polypeptide containing the TAL effector molecule. Generally, TAL binding is inversely correlated with the number of mismatches. In some embodiments, the TAL effector molecule of the polypeptide of the present invention contains, and optionally does not contain, up to 7, 6, 5, 4, 3, 2, or 1 mismatch with the target DNA sequence. While not intended to be bound by any particular theory, generally, a decrease in the number of mismatches as the number of TAL effector domains of the TAL effector molecule decreases is not only acceptable but also enables the function of the polypeptide containing the TAL effector molecule. Binding affinity is thought to depend on the sum of matching repeat-DNA combinations. For example, a TAL effector molecule with 25 or more TAL effector domains may be able to tolerate up to 7 mismatches.
[0420] In addition to the TAL effector domain, the TAL effector molecule of the present invention may include further sequences derived from naturally occurring TAL effectors. The lengths of the C-terminal and / or N-terminal sequences contained on both sides of the TAL effector domain portion of the TAL effector molecule vary and can be selected by those skilled in the art based, for example, on the tests of Zhang et al. (2011). Zhang et al. characterized several C-terminal and N-terminal cleavage variants in proteins based on the Hax3-derived TAL effector and identified key elements that contribute to optimal binding to the target sequence and thus to transcriptional activation. In general, transcriptional activity was found to be inversely correlated with the length of the N-terminus. With respect to the C-terminus, key elements in the DNA-binding residues of the first 68 amino acids of the Hax3 sequence were identified. Therefore, in some embodiments, the first 68 amino acids on the C-terminal side of the TAL effector domain of a naturally occurring TAL effector are included in the TAL effector molecule. Accordingly, in one embodiment, the TAL effector molecule contains 1) one or more TAL effector domains derived from naturally occurring TAL effectors, 2) at least 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260, 270, 280 or more amino acids from naturally occurring TAL effectors on the N-terminal side of the TAL effector domain, and / or 3) at least 68, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260 or more amino acids from naturally occurring TAL effectors on the C-terminal side of the TAL effector domain.
[0421] In some embodiments, the endonuclease domain or DNA-binding domain is or includes a Zn finger molecule. The Zn finger molecule includes Zn finger proteins, such as naturally occurring Zn finger proteins or modified Zn finger proteins or fragments thereof. Many Zn finger proteins are known to those skilled in the art and are commercially available, for example, from Sigma-Aldrich.
[0422] In some embodiments, the Zn finger molecule includes a non-naturally occurring Zn finger protein that has been modified to bind to a selected target DNA sequence. For example, Beerli, et al. (2002) Nature Biotechnol. 20:135-141; Pabo, et al. (2001) Ann. Rev. Biochem. 70:313-340; Isalan, et al. (2001) Nature Biotechnol. 19:656-660; Segal, et al. (2001) Curr. Opin. Biotechnol. 12:632-637; Choo, et al. al. (2000) Curr. Opin. Struct. Biol. 10:411-416; U.S. Patent No. 6,453,242; U.S. Patent No. 6,534,261; U.S. Patent No. 6,599,692; U.S. Patent No. 6,503,717; U.S. Patent No. 6,689,558; U.S. Patent No. 7,030,215; U.S. Patent No. 6,794,136; U.S. Patent No. 7,067,317 See also U.S. Patent No. 7,262,054; U.S. Patent No. 7,070,934; U.S. Patent No. 7,361,635; U.S. Patent No. 7,253,273 and U.S. Patent Publication No. 2005 / 0064474; U.S. Patent Publication No. 2007 / 0218528; and U.S. Patent Publication No. 2005 / 0267061 (all of which are incorporated herein by reference in their entirety).
[0423] Modified Zn finger proteins may possess novel binding specificity compared to naturally occurring Zn finger proteins. Modification methods include, but are not limited to, rational design and a variety of types of selection. Rational design includes, for example, the use of a database containing triad (or quad) nucleotide sequences and individual Zn finger amino acid sequences, where each triad or quad nucleotide sequence associates with one or more amino acid sequences of Zn fingers that bind to a particular triad or quad sequence. See, for example, U.S. Patent No. 6,453,242 and U.S. Patent No. 6,534,261 (both incorporated herein by reference).
[0424] Exemplary selection methods, including phage displays and two-hybrid systems, are disclosed in U.S. Patent Nos. 5,789,538; 5,925,523; 6,007,988; 6,013,453; 6,410,248; 6,140,466; 6,200,759 and 6,242,568, as well as in International Patent Publications brochures 98 / 37186, 98 / 53057, 00 / 27878 and 01 / 88197 and UK Patent No. 2,338,237. Furthermore, the increased binding specificity in Zn finger proteins is described, for example, in International Patent Publications brochure 02 / 077227.
[0425] Furthermore, as disclosed in these and other references, Zn finger domains and / or polyfingered Zn finger proteins can be linked together using any suitable linker sequence, for example, one containing a linker of five or more amino acid lengths. For example linker sequences of six or more amino acid lengths, see also U.S. Patent No. 6,479,626, U.S. Patent No. 6,903,185, and U.S. Patent No. 7,153,949. The proteins described herein may include any combination of suitable linkers between the individual Zn fingers of the protein. Furthermore, for increasing the binding specificity in the Zn finger binding domain, see, for example, the jointly owned international patent publication, International Publication No. 02 / 077227.
[0426] Proteins and methods for Zn fingers for the design and construction of fusion proteins (and polynucleotides encoding them) are known to those skilled in the art, as described in U.S. Patent Nos. 6,140,0815; 789,538; 6,453,242; 6,534,261; 5,925,523; 6,007,988; 6,013,453 and 6,200,759; International Patent Publications, International Publication No. 95 / 19431; International Publication No. 96 / Further details are provided in Pamphlet No. 06166; International Publication Pamphlet No. 98 / 53057; International Publication Pamphlet No. 98 / 54311; International Publication Pamphlet No. 00 / 27878; International Publication Pamphlet No. 01 / 60970; International Publication Pamphlet No. 01 / 88197; International Publication Pamphlet No. 02 / 099084; International Publication Pamphlet No. 98 / 53058; International Publication Pamphlet No. 98 / 53059; International Publication Pamphlet No. 98 / 53060; International Publication Pamphlet No. 02 / 016536 and International Publication Pamphlet No. 03 / 016496.
[0427] Furthermore, as disclosed in these and other references, Zn finger proteins and / or polyfingered Zn finger proteins can be linked together, for example, as a fusion protein, using any suitable linker sequence including, for example, a linker of five or more amino acid lengths. For example linker sequences of six or more amino acid lengths, see also U.S. Patent No. 6,479,626, U.S. Patent No. 6,903,185, and U.S. Patent No. 7,153,949. The Zn finger molecules described herein may include any suitable combination of linkers between individual Zn finger proteins and / or polyfingered Zn finger proteins of the Zn finger molecule.
[0428] In certain embodiments, the DNA-binding domain or endonuclease domain comprises a Zn finger molecule containing a modified Zn finger protein that binds to a target DNA sequence (in a sequence-specific manner). In some embodiments, the Zn finger molecule comprises one Zn finger protein or a fragment thereof. In other embodiments, the Zn finger molecule comprises multiple Zn finger proteins (or fragments thereof), e.g., 2, 3, 4, 5, 6 or more Zn finger proteins (and optionally up to 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 Zn finger proteins). In some embodiments, the Zn finger molecule comprises at least 3 Zn finger proteins. In some embodiments, the Zn finger molecule comprises 4, 5 or 6 fingers. In some embodiments, the Zn finger molecule comprises 8, 9, 10, 11 or 12 fingers. In some embodiments, a Zn finger molecule comprising 3 Zn finger proteins recognizes a target DNA sequence containing 9 or 10 nucleotides. In some embodiments, a Zn finger molecule containing four Zn finger proteins recognizes a target DNA sequence containing 12–14 nucleotides. In some embodiments, a Zn finger molecule containing six Zn finger proteins recognizes a target DNA sequence containing 18–21 nucleotides.
[0429] In some embodiments, the Zn finger molecule includes a two-handed Zn finger protein. A two-handed Zn finger protein is a protein in which two clusters of Zn finger protein are separated by an intervening amino acid so that two Zn finger domains bind to two discontinuous target DNA sequences. An example of a two-handed Zn finger binding protein is SIP1, in which a cluster of four Zn finger proteins is located at the amino terminus of the protein and a cluster of three Zn finger proteins is located at the carboxyl terminus (see Remade, et al. (1999) EMBO Journal 18(18):5073-5084). Each cluster of Zn fingers in these proteins is capable of binding to a specific target sequence, and the space between the two target sequences may contain a large number of nucleotides.
[0430] Linker In some embodiments, the gene-modified polypeptide may include a linker, such as a peptide linker, such as those listed in Table 10. In some embodiments, the gene-modified polypeptide includes, in the direction from the N-terminus to the C-terminus, a Cas domain (e.g., the Cas domain in Table 8), a linker in Table 10 (or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity thereto), and an RT domain (e.g., the RT domain in Table 6). In some embodiments, the gene-modified polypeptide includes a flexible linker between the endonuclease and the RT domain, such as a linker containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 11,002). In some embodiments, the RT domain of the gene-modified polypeptide may be located C-terminus relative to the endonuclease domain. In some embodiments, the RT domain of the gene-modified polypeptide may be located N-terminus relative to the endonuclease domain.
[0431] [Table 10-1]
[0432] [Table 10-2]
[0433] [Table 10-3]
[0434] [Table 10-4]
[0435] [Table 10-5]
[0436] In some embodiments, the linker of the gene-modified polypeptide is (SGGS) n (Sequence ID 5025), (GGGS) n (Sequence ID 5026), (GGGGS) n (Sequence ID 5027), (G) n , (EAAAK) n (Sequence ID 5028), (GGS) n or (XP) n Includes motifs selected from.
[0437] Selection of gene-modified polypeptides by pooled screening Candidate gene-modified polypeptides can be screened to evaluate their gene-editing capabilities. For example, RNA gene-modification systems designed for targeted editing of coding sequences in the human genome may be used. In certain embodiments, such gene-modification systems may be used in conjunction with pooled screening methods.
[0438] For example, a library of candidate gene-modified polypeptides and template guide RNAs (tgRNAs) may be introduced into mammalian cells by a pooled screening method to test the gene-editing capabilities of the candidates. In certain embodiments, the library of candidate gene-modified polypeptides is introduced into mammalian cells, followed by the introduction of tgRNAs into the cells.
[0439] Representative, non-limiting examples of mammalian cells that may be used for screening include HEK293T cells, U2OS cells, HeLa cells, HepG2 cells, Huh7 cells, K562 cells, or iPS cells.
[0440] Candidate gene-modified polypeptides may include: 1) Cas nucleases, e.g., wild-type Cas nuclease, e.g., wild-type Cas9 nuclease, mutant Cas nucleases, e.g., Cas nickase, e.g., Cas9 N863A nickase, or other Cas9 nickases selected from Table 7 or Table 8; 2) peptide linkers, e.g., sequences from Table D or Table 10 that may exhibit varying degrees of length, flexibility, hydrophobicity, and / or secondary structure; and 3) reverse transcriptase (RT), e.g., RT domains from Table D or Table 6. The candidate gene-modified polypeptide library may include multiple distinct candidate gene-modified polypeptides that differ from each other with respect to one, two, or all three components of the Cas nuclease, peptide linker, or RT domain, or multiple nucleic acid expression vectors encoding such candidate gene-modified polypeptides.
[0441] For screening candidate gene-modified polypeptides, a two-component system comprising a gene-modified polypeptide component and a tgRNA component may be used. The gene-modified component may include, for example, an expression vector, such as an expression plasmid or lentiviral vector encoding a candidate gene-modified polypeptide, such as a human codon-optimized nucleic acid encoding a candidate gene-modified polypeptide, such as the Cas-linker-RT fusion described above. In certain embodiments, a lentiviral cassette may be used that includes (i) a promoter for expression in mammalian cells, such as a CMV promoter; (ii) a candidate gene-modified library, such as a Cas-linker-RT fusion comprising a Cas nuclease from Table 7 or Table 8, a peptide linker from Table 10, and an RT from Table 6, such as a Cas-linker-RT fusion as shown in Table D; (iii) a self-cleaving polypeptide, such as a T2A peptide; (iv) a marker enabling selection in mammalian cells, such as a puromycin resistance gene; and (v) a termination signal, such as a poly-A tail.
[0442] The tgRNA component may include tgRNA or an expression vector, such as an expression plasmid that generates tgRNA and drives tgRNA expression using, for example, a U6 promoter. The tgRNA is a non-coding RNA sequence that is recognized by Cas, localized to a target genomic locus, and also templates the reverse transcription of the desired edit into the genome via its RT domain.
[0443] To prepare a pool of cells expressing candidate gene-modified polypeptide libraries, mammalian cells, such as HEK293T or U2OS cells, can be transduced with a pooled candidate gene-modified polypeptide expression vector preparation, such as a lentiviral preparation of the candidate gene-modified polypeptide library. In certain embodiments, a lentiviral plasmid is used, and HEK293 Lenti-X cells are seeded on a 15 cm plate (approximately 12 × 10⁻¹⁰) prior to lentiviral plasmid transfection. 6(individual cells). In these embodiments, lentiviral plasmid transfection may be performed using Lentiviral Packaging Mix (Biosettia), and the transfection of plasmid DNA for the candidate gene modification library may be performed using Lipofectamine 2000 and Opti-MEM medium according to the manufacturer's protocol. In these embodiments, extracellular DNA may be removed by complete medium change the following day, and the virus-containing medium may be collected after 48 hours. The lentiviral medium may be concentrated using Lenti-X Concentrator (TaKaRa Biosciences), and 5 mL lentiviral aliquots may be prepared and stored at -80°C. Lentiviral titer measurement is performed by counting colony-forming units after selection, for example, after puromycin selection.
[0444] To monitor gene editing of target DNA, mammalian cells, such as HEK293T or U2OS cells carrying the target DNA, may be used. In other embodiments of monitoring gene editing of target DNA, mammalian cells, such as HEK293T or U2OS cells carrying a target DNA genomic landing pad, may be used. In certain embodiments, the target DNA genomic landing pad may contain a gene edited for the treatment of a target disease or disorder. In other specific embodiments, the target DNA is a gene sequence expressing a protein exhibiting detectable characteristics that can be monitored to determine whether gene editing has occurred. For example, in certain embodiments, a blue fluorescent protein (BFP) or green fluorescent protein (GFP) expressing genomic landing pad may be used. In certain embodiments, mammalian cells, such as HEK293T or U2OS cells containing target DNA, such as a target DNA genomic landing pad, are seeded on culture plates at 500× to 3000× cells per candidate gene-modified library and transduced at an infection multiplicity (MOI) of 0.2 to 0.3 to minimize multiple infections per cell. Puromycin (2.5 ug / mL) may be added 48 hours after infection to allow selection of infected cells. In such embodiments, cells may be kept under puromycin selection for at least 7 days and then scaled up for tgRNA introduction, e.g., tgRNA electroporation.
[0445] To confirm whether gene editing occurs, mammalian cells containing the target DNA to be edited may be infected with a candidate gene modification polypeptide library and then transfected with tgRNA designed for use in editing the target DNA. The cells may then be analyzed, for example, using cell sorting and sequence analysis, to determine whether editing of the target locus occurred according to the designed outcome, or whether editing did not occur or only incomplete editing occurred.
[0446] In certain embodiments, to confirm whether genome editing occurs, BFP- or GFP-expressing mammalian cells, e.g., HEK293T or U2OS cells, may be infected with a candidate gene modification library and then transfected or electroporated by electroporation of 250,000 cells / well with a tgRNA plasmid or RNA, e.g., 200 ng of a tgRNA plasmid designed to convert BFP to GFP or GFP to BFP in cell numbers ensuring >250 × ~ 1000 × coverage for the candidate library. In such embodiments, the genome editing capability of various constructs in this assay may be assessed by sorting the cells by fluorescence-activated cell sorting (FACS) for the expression of chromogenic fluorescent proteins (FPs) 4 to 10 days after electroporation. The cells are sorted and collected as different populations of unedited cells (showing the original fluorescent protein signal), edited cells (showing the converted fluorescent protein signal), and incompletely edited cells (showing no fluorescent protein signal). Samples of unsorted cells may also be collected as an input population to determine candidate enrichment during analysis.
[0447] To determine whether a candidate gene-modified library exhibits genome editing capability in an assay, genomic DNA (gDNA) is collected from sorted cell populations and analyzed by sequencing the candidate gene-modified library in each population. Briefly, candidate gene-modified libraries are amplified from the genome using primers specific to a gene-modified polypeptide expression vector, such as a lentiviral cassette, amplified again by a second PCR to dilute the genomic DNA, and then sequenced, for example, by a next-generation sequencing platform. After quality control of the sequencing reads, reads of at least approximately 1500 nucleotides and generally no more than approximately 3200 nucleotides are mapped to gene-modified polypeptide library sequences, and those containing a minimum of approximately 80% agreement with the library sequence are considered to be aligned without issue with a given candidate for this pooled screen. To identify candidates capable of gene editing in the assay, e.g., editing BFP to GFP or GFP to BFP, the read count of each candidate library in the edited population is compared to its read count in the initial unsorted population.
[0448] For pooled screening, gene modification candidates with genome editing capabilities are identified based on the enrichment of the edited (converted FP) population compared to unsorted (input) cells. In some embodiments, enrichment of at least 1.0, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or at least 100 times of the input indicates potentially useful gene editing activity, e.g., at least 2 times enrichment. In some embodiments, enrichment is converted to a log value by taking the log base 2 of the enrichment ratio. In some embodiments, log2 enrichment scores of at least 0, 1, 2, 3, 4, 5, 5.5, 6.0, 6.2, 6.3, 6.4, 6.5, or at least 6.6 indicate potentially useful gene editing activity, e.g., at least 1.0 log2 enrichment score. In certain embodiments, the enrichment values observed for a candidate gene modification can be compared to enrichment values observed under similar conditions using a reference, e.g., element ID number: 17380.
[0449] In some embodiments, multiple tgRNAs may be used to screen a library of candidate gene modification genes. In certain embodiments, multiple tgRNAs may be used to optimize template / Cas-linker-RT fusion pairs for, for example, gene editing of a specific target gene, such as a gene target for disease treatment. In certain embodiments, a pooling method for screening candidate gene modification genes may be performed using many different tgRNAs in sequenced form.
[0450] In some embodiments, multiple types of editing, such as insertions, substitutions, and / or deletions of different lengths, may be used to screen candidate libraries of gene modifications.
[0451] In some embodiments, multiple target sequences, such as different fluorescent proteins, may be used to screen a candidate library of gene modifications. In some embodiments, multiple cell types, such as HEK293T or U2OS, may be used to screen a candidate library of gene modifications. Those skilled in the art will understand that a given candidate may exhibit modified editing ability or an increase or decrease in observable or useful activity across different conditions, including tgRNA sequence (e.g., nucleotide modification, PBS length, RT template length), target sequence, target site, editing type, mutation site relative to the first chain nick of the gene-modified polypeptide, or cell type. Accordingly, in some embodiments, candidate gene-modified libraries are screened across multiple parameters, for example, using at least two different tgRNAs in at least two cell types, and gene editing activity is identified by enrichment under any single condition. In other embodiments, candidates with more robust activity across different tgRNAs and cell types are identified by enrichment under at least two conditions, for example, all conditions being screened. To clarify, candidates found to show little to no enrichment under any given conditions are not presumed to be inactive across all conditions, but can be reconstituted at the polypeptide level by screening with different parameters or by swapping, replacing, or altering, for example, a domain (e.g., RT domain), a linker, or other signal (e.g., NLS).
[0452] Exemplary Cas9-linker-RT fusion sequence In some embodiments, the gene-modified polypeptide includes a linker sequence and an RT sequence. In some embodiments, the gene-modified polypeptide includes a linker sequence listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes an amino acid sequence of the RT domain listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes a linker sequence listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and an amino acid sequence of the RT domain listed in Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes (i) a linker sequence listed in the row of Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, and (ii) an amino acid sequence of an RT domain listed in the same row of Table D or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.
[0453] Exemplary gene-modified polypeptides In some embodiments, the gene-modified polypeptide (for example, a gene-modified polypeptide that is part of the system described herein) includes one amino acid sequence from Sequence ID No. 1 to 7743 of the sequence listing or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from Sequence ID No. 1 to 7743 or an amino acid sequence having at least 80% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from Sequence ID No. 1 to 7743 or an amino acid sequence having at least 90% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from Sequence ID No. 1 to 7743 or an amino acid sequence having at least 95% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from Sequence ID No. 1 to 7743 or an amino acid sequence having at least 99% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from Sequence ID No. 1 to 7743. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from SEQ ID NOs. 6001 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes one amino acid sequence from SEQ ID NOs. 4501 to 4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide described herein includes an RT and linker sequence from any of SEQ ID NOs. 1 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and the St1Cas9 domain described herein. In some embodiments, the gene-modified polypeptide described herein includes an RT and linker sequence from any of SEQ ID NOs. 1 to 7743 and the St1Cas9 domain described herein.
[0454] In some embodiments, the gene-modified polypeptide includes an amino acid sequence as shown in Table A1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0455] In some embodiments, the gene-modified polypeptide includes an amino acid sequence as shown in Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes a linker including a linker sequence as shown in Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes an RT domain including an RT domain sequence as shown in Table T1 or an RT domain including an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes (i) a linker comprising a linker sequence as shown in a certain row of Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (ii) an RT domain comprising an RT domain sequence as shown in the same row of Table T1 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0456] [Table T1]
[0457] In some embodiments, the gene-modified polypeptide includes an amino acid sequence as shown in Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes a linker including a linker sequence as shown in Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes an RT domain including an RT domain sequence as shown in Table T2 or an RT domain including an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide includes (i) a linker comprising a linker sequence as shown in a certain row of Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (ii) an RT domain comprising an RT domain sequence as shown in the same row of Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0458] [Table T2-1]
[0459] [Table T2-2]
[0460] [Table T2-3]
[0461] Exemplary gene-modified polypeptide subsequences In some embodiments, the genetically modified polypeptide comprises, in order from the N-terminus to the C-terminus, one or more N-terminal methionine residues (e.g., 1, 2, 3, 4, 5, or all 6), a first nuclear localization signal (NLS), a DNA-binding domain, a linker, an RT domain, and / or a second NLS. In some embodiments, the genetically modified polypeptide comprises, in order from the N-terminus to the C-terminus, an NLS (e.g., a first NLS), a DNA-binding domain, a linker, and an RT domain, wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide comprises a DNA-binding domain, a linker, an RT domain, and an NLS (e.g., a second NLS) in the order from the N-terminus to the C-terminus, wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide comprises a first NLS, a DNA-binding domain, a linker, an RT domain, and a second NLS in the order from the N-terminus to the C-terminus, wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.
[0462] In some embodiments, the gene-modified polypeptide has, in order from the N-terminus to the C-terminus, one or more N-terminal methionine residues (e.g., 1, 2, 3, 4, 5, or all 6), a first nuclear localization signal (NLS) (e.g., one of any of SEQ ID NOs. 1 to 7743 and / or any of the gene-modified polypeptides listed in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). , DNA-binding domains (e.g., Cas domains, e.g., SpyCas9 domains, e.g., as listed in Table 8, or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, or one of SEQ ID NOs: 1-7743 and / or DNA-binding domains of gene-modified polypeptides as listed in Table A1, Table T1, or Table T2, or at least 70%, 75%, 80%, 85%, 90% identity thereto) (amino acid sequences having 1%, 95%, or 99% identity), linker (e.g., one of the gene-modified polypeptides listed in any one of SEQ ID NOs: 1-7743 and / or in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto), RT domain (e.g., one of the SEQ ID NOs: 1-7743 and / or in Table A1, Table T1, or Table T2) The first is a gene-modified polypeptide as described above, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto) and the second NLS (for example, any one of SEQ ID NOs. 1 to 7743 and / or a gene-modified polypeptide as described in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).In some embodiments, the genetically modified polypeptide further comprises a T2A sequence and / or a puromycin sequence (e.g., one of any of SEQ ID NOs: 1 to 7743 and / or any of the genetically modified polypeptides listed in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto) (e.g., at the C-terminal side of a second NLS). In some embodiments, the nucleic acid encoding the genetically modified polypeptide (e.g., as described herein) encodes a T2A sequence, for example, the T2A sequence is located between a region encoding the genetically modified polypeptide and a second region, which optionally encodes a selectable marker, such as puromycin.
[0463] In certain embodiments, the first NLS includes a first NLS sequence of a gene-modified polypeptide having one of the amino acid sequences of SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the first NLS includes a first NLS sequence of a gene-modified polypeptide as listed in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the first NLS sequence includes a C-myc NLS. In certain embodiments, the first NLS includes the amino acid sequence PAAKRVKLD (SEQ ID NO: 11,095) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0464] In certain embodiments, the gene-modified polypeptide further includes a spacer sequence between the first NLS and the DNA-binding domain. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain includes the amino acid sequence GG.
[0465] In certain embodiments, the DNA-binding domain includes the DNA-binding domain of any one of the gene-modified polypeptides SEQ ID NOs: 1 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the DNA-binding domain includes the DNA-binding domain of a gene-modified polypeptide as listed in Table A1, Table T1, or Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the DNA-binding domain includes a Cas domain (e.g., as listed in Table 8). In certain embodiments, the DNA-binding domain includes the amino acid sequence of a SpyCas9 polypeptide (e.g., as listed in Table 8, e.g., Cas9 N863A polypeptide) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In a particular embodiment, the DNA-binding domain has the amino acid sequence: [ka] Or it includes an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with respect to it.
[0466] In certain embodiments, the gene-modified polypeptide further includes a spacer sequence between the DNA-binding domain and the linker. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker includes the amino acid sequence GG.
[0467] In certain embodiments, the linker includes the linker sequence of any one of the gene-modified polypeptides SEQ ID NOs: 1 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the linker includes the linker sequence of a gene-modified polypeptide as listed in Table A1, Table T1, or Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the linker includes the amino acid sequence as listed in Table D or Table 10 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0468] In certain embodiments, the gene-modified polypeptide further includes a spacer sequence between the linker and the RT domain. In certain embodiments, the spacer sequence between the linker and the RT domain includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the linker and the RT domain includes the amino acid sequence GG.
[0469] In certain embodiments, the RT domain includes the RT domain sequence of any one of the gene-modified polypeptides SEQ ID NOs: 1 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes the RT domain sequence of a gene-modified polypeptide as listed in Table A1, Table T1, or Table T2 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes the amino acid sequence as listed in Table D or Table 6 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain has a length of approximately 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 amino acids.
[0470] In certain embodiments, the gene-modified polypeptide further includes a spacer sequence between the RT domain and the second NLS. In certain embodiments, the spacer sequence between the RT domain and the second NLS includes 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the RT domain and the second NLS includes the amino acid sequence AG.
[0471] In certain embodiments, the second NLS includes a second NLS sequence of any one gene-modified polypeptide from sequence numbers 1 to 7743. In certain embodiments, the second NLS includes a second NLS sequence of a gene-modified polypeptide as listed in Table A1, Table T1, or Table T2. In certain embodiments, the second NLS sequence includes a plurality of partial NLS sequences. In embodiments, the NLS sequence, for example, the second NLS sequence, includes, for example, a first partial NLS sequence containing the amino acid sequence KRTADGSEFE (sequence number 11,097) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the NLS sequence, for example, the second NLS sequence, includes a second partial NLS sequence. In the embodiment, the NLS sequence, for example, the second NLS sequence, includes, for example, an SV40A5 NLS containing the amino acid sequence KRTADGSEFESPKKKAKVE (SEQ ID NO: 11,098), for example, a bipartite SV40A5 NLS, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In a particular embodiment, the NLS sequence, for example, the second NLS sequence, includes the amino acid sequence KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 11,099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0472] In certain embodiments, the gene-modified polypeptide further comprises a spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises the amino acid sequence GSG.
[0473] Linker and RT domain In some embodiments, the gene-modified polypeptide includes a linker (e.g., as described herein) and an RT domain (e.g., as described herein). In certain embodiments, the gene-modified polypeptide includes a linker (e.g., as described herein) and an RT domain (e.g., as described herein) in the order from the N-terminus to the C-terminus.
[0474] In certain embodiments, the linker includes an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to the linker sequence as listed in Table 10. In certain embodiments, the linker includes one linker sequence from SEQ ID NOs. 1 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to it. In certain embodiments, the linker includes one linker sequence from SEQ ID NOs. 6001 to 7743 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to it. In certain embodiments, the linker includes one linker sequence from SEQ ID NOs. 4501 to 4541 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity to it. In certain embodiments, the linker includes a linker sequence of an exemplary genetically modified polypeptide listed in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes an RT domain sequence as listed in Table 6, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes an RT domain sequence of an exemplary genetically modified polypeptide listed in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0475] In some embodiments, the gene-modified polypeptide is a portion of any one of the gene-modified polypeptides SEQ ID NOs: 1 to 7743, comprising a portion including a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the said portion.
[0476] In some embodiments, the gene-modified polypeptide includes a linker of any one gene-modified polypeptide from SEQ ID NOs. 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the gene-modified polypeptide includes a linker of any one gene-modified polypeptide from SEQ ID NOs. 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the gene-modified polypeptide includes a linker of any one gene-modified polypeptide from SEQ ID NOs. 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the gene-modified polypeptide includes a linker of the gene-modified polypeptide as shown in Table A1, Table T1, or Table T2, or a linker containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0477] In some embodiments, the gene-modified polypeptide includes the RT domain of any one gene-modified polypeptide from SEQ ID NOs. 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the gene-modified polypeptide includes the RT domain of any one gene-modified polypeptide from SEQ ID NOs. 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the gene-modified polypeptide includes the RT domain of any one gene-modified polypeptide from SEQ ID NOs. 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the gene-modified polypeptide includes an RT domain of the gene-modified polypeptide as shown in Table A1, Table T1, or Table T2, or an RT domain having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0478] In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains of the genetically modified polypeptide having one amino acid sequence of any one of SEQ ID NOs: 1 to 7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains having at least 80% identity with one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains having at least 90% identity with one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains having at least 95% identity with one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains having at least 99% identity with any one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains of genetically modified polypeptides having one of the amino acid sequences of SEQ ID NOs: 6001 to 7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with respect to thereafter). In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains of genetically modified polypeptides having one of the amino acid sequences of SEQ ID NOs: 4501 to 4541 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with respect to thereafter.In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains from a single row in any of Table A1, Table T1, or Table T2 (for example, from a single exemplary genetically modified polypeptide as listed in any of Table A1, Table T1, or Table T2) (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).
[0479] In certain embodiments, the linker and RT domain of the gene-modified polypeptide include amino acid sequences of the linker and RT domain from two different amino acid sequences selected from SEQ ID NOs: 1 to 7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In certain embodiments, the linker and RT domain of the gene-modified polypeptide include amino acid sequences of the linker and RT domain from a different row in Table A1, Table T1, or Table T2 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).
[0480] In certain embodiments, the gene-modified polypeptide further comprises a first NLS (e.g., a 5' NLS) as described herein. In certain embodiments, the gene-modified polypeptide further comprises a second NLS (e.g., a 3' NLS) as described herein. In certain embodiments, the gene-modified polypeptide further comprises an N-terminal methionine residue.
[0481] RT family and mutants In certain embodiments, the gene-modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, XMRV6, BLVAU, BLVJ, HTL1A, HTL1C, HTL1L, HTL32, HTL3P, HTLV2, JSRV, MLVF5, MLVRD, MMTVB, MPMV, SSFVCP, SMRVH, SRV1, SRV2, and WDSV. In certain embodiments, the gene-modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.
[0482] In certain embodiments, the gene-modified polypeptide includes the amino acid sequence of the RT domain sequence from the MLVMS RT domain. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations or corresponding point mutations as shown in the first column of Table M1. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations or corresponding point mutations as shown in the third column (Gen1 MLVMS) of Table M1. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations at the amino acid positions of the RT domain or corresponding amino acid positions as shown in the first and second columns of Table M2.
[0483] In certain embodiments, the gene-modified polypeptide includes the amino acid sequence of the RT domain sequence from the AVIRE RT domain. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations or corresponding point mutations as shown in column 2 of Table M1. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations or corresponding point mutations as shown in column 4 (Gen2 AVIRE) of Table M1. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations at the amino acid positions of the RT domain or corresponding amino acid positions as shown in columns 3 and 4 of Table M2. In certain embodiments, the RT domain includes IENSSP (e.g., at the C-terminus).
[0484] [Table M1]
[0485] [Table M2]
[0486] In certain embodiments, the genetically modified polypeptide includes an RT domain derived from a gamma retrovirus. In certain embodiments, the gamma retrovirus-derived RT domain of the genetically modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In some embodiments, the gamma retrovirus-derived RT domain of the genetically modified polypeptide does not originate from PERV. In some embodiments, the RT comprises one, two, three, four, five, six or more mutations corresponding to the mutations in the RT domain of mouse leukemia virus reverse transcriptase D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N, as shown in Table 2A. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% identity to any one linker domain of SEQ ID NOs. 1 to 7743. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity to SEQ ID NOs. 5217 or SEQ ID NOs. 11,041.
[0487] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of AVIRE RT (e.g., AVIRE_P03360 sequence, e.g., SEQ ID NO: 8001) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of AVIRE RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, G330P, L605W, T306K, and W313F. In some embodiments, the RT domain includes the amino acid sequence of AVIRE RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D200N, G330P, and L605W.
[0488] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of BAEVM RT (e.g., BAEVM_P10272 sequence, e.g., SEQ ID NO: 8004) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of BAEVM RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D198N, E328P, L602W, T304K, and W311F. In some embodiments, the RT domain includes the amino acid sequence of BAEVM RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D198N, E328P, and L602W.
[0489] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FFV RT (e.g., the FFV_O93209 sequence, e.g., SEQ ID NO: 8012) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further including the corresponding positions of one, two, three, or four mutant or homologous RT domains selected from the group consisting of D21N, T293N, T419P, and L393K. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D21N, T293N, and T419P. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further including the mutant D21N. In some embodiments, the RT domain comprises an amino acid sequence of FFV RT further comprising one, two, or three mutant or homologous RT domains selected from the group consisting of T207N, T333P, and L307K at corresponding positions.
[0490] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FLV RT (e.g., FLV_P10273 sequence, e.g., SEQ ID NO: 8019) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FLV RT further including the corresponding positions of one, two, three, or four mutant or homologous RT domains selected from the group consisting of D199N, L602W, T305K, and W312F. In some embodiments, the RT domain includes the amino acid sequence of FLV RT further including the corresponding positions of one or two mutant or homologous RT domains selected from the group consisting of D199N and L602W.
[0491] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FOAMV RT (e.g., the FOAMV_P14350 sequence, e.g., SEQ ID NO: 8021) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further including the corresponding positions of one, two, three, or four mutant or homologous RT domains selected from the group consisting of D24N, T296N, S420P, and L396K. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D24N, T296N, and S420P. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further including the corresponding position of mutant D24N or homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of FOAMV RT further comprising one, two, or three mutants or homologous RT domains selected from the group consisting of T207N, S331P, and L307K.
[0492] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of GALV RT (e.g., the GALV_P21414 sequence, e.g., SEQ ID NO: 8027) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D198N, E328P, and L600W.
[0493] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of KORV RT (e.g., KORV_Q9TTC1 sequence, e.g., SEQ ID NO: 8047) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further including the corresponding positions of one, two, three, four, five, or six mutant or homologous RT domains selected from the group consisting of D32N, D322N, E452P, L274W, T428K, and W435F. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further including the corresponding positions of one, two, three, or four mutant or homologous RT domains selected from the group consisting of D32N, D322N, E452P, and L274W. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further including the mutant D32N. In some embodiments, the RT domain comprises an amino acid sequence of KORV RT further comprising the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D231N, E361P, L633W, T337K, and W344F.
[0494] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVAV RT (e.g., the MLVAV_P03356 sequence, e.g., SEQ ID NO: 8053) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVAV RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain includes the amino acid sequence of MLVAV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D200N, T330P, and L603W.
[0495] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVBM RT (e.g., the MLVBM_Q7SVK7 sequence, e.g., SEQ ID NO: 8056) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVBM RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D199N, T329P, L602W, T305K, and W312F. In some embodiments, the RT domain includes the amino acid sequence of MLVBM RT further including the corresponding positions of one, two, and three mutant or homologous RT domains selected from the group consisting of D200N, T330P, and L603W.
[0496] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVCB RT (e.g., MLVCB_P08361 sequence, e.g., SEQ ID NO: 8062) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVCB RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain includes the amino acid sequence of MLVCB RT further including the corresponding positions of one, two, and three mutant or homologous RT domains selected from the group consisting of D200N, T330P, and L603W.
[0497] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVFF RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVFF RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain includes the amino acid sequence of MLVFF RT further including the corresponding positions of one, two, and three mutant or homologous RT domains selected from the group consisting of D200N, T330P, and L603W.
[0498] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVMS RT (e.g., MLVMS_reference sequence, e.g., SEQ ID NO: 8137 or MLVMS_P03355 sequence, e.g., SEQ ID NO: 8070) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVMS RT further including the corresponding positions of one, two, three, four, five, or six mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, W313F, and H8Y. In some embodiments, the RT domain includes the amino acid sequence of MLVMS RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain further comprises an amino acid sequence of MLVMS RT that includes one, two, or three mutants selected from the group consisting of D200N, T330P, and L603W, or corresponding positions of homologous RT domains.
[0499] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of PERV RT (e.g., the PERV_Q4VFZ2 sequence, e.g., SEQ ID NO: 8099) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of PERV RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D196N, E326P, L599W, T302K, and W309F. In some embodiments, the RT domain includes the amino acid sequence of PERV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D196N, E326P, and L599W.
[0500] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of SFV1 RT (e.g., SFV1_P23074 sequence, e.g., SEQ ID NO: 8105) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further including the corresponding positions of one, two, three, or four mutants or homologous RT domains selected from the group consisting of D24N, T296N, N420P, and L396K. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further including the corresponding positions of one, two, or three mutants or homologous RT domains selected from the group consisting of D24N, T296N, and N420P. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further including the corresponding positions of D24N or homologous RT domains.
[0501] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of SFV3L RT (e.g., the SFV3L_P27401 sequence, e.g., SEQ ID NO: 8111) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further including the corresponding positions of one, two, three, or four mutant or homologous RT domains selected from the group consisting of D24N, T296N, N422P, and L396K. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D24N, T296N, and N422P. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further including the corresponding position of mutant D24N or homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of SFV3L RT further comprising one, two, or three mutant or homologous RT domains selected from the group consisting of T307N, N333P, and L307K at corresponding positions.
[0502] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of WMSV RT (e.g., WMSV_P03359 sequence, e.g., SEQ ID NO: 8131) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of WMSV RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain includes the amino acid sequence of WMSV RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D198N, E328P, and L600W.
[0503] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of XMRV6 RT (e.g., XMRV6_A1Z651 sequence, e.g., SEQ ID NO: 8134) or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of XMRV6 RT further including the corresponding positions of one, two, three, four, or five mutant or homologous RT domains selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain includes the amino acid sequence of XMRV6 RT further including the corresponding positions of one, two, or three mutant or homologous RT domains selected from the group consisting of D200N, T330P, and L603W.
[0504] In certain embodiments, the RT domain of the gene-modified polypeptide includes an amino acid sequence of the RT domain of AVIRE RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the RT domain includes an amino acid sequence of the RT domain contained in the sequence listed in column 1 of Table A5 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide further includes a linker having at least 99% or 100% identity to SEQ ID NO: 5217 or SEQ ID NO: 11,041.
[0505] In certain embodiments, the RT domain of the gene-modified polypeptide includes the amino acid sequence of the RT domain of MLVMS RT or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the RT domain includes the amino acid sequence of the RT domain contained in any of the sequences listed in columns 2 to 6 of Table A5 or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the gene-modified polypeptide further includes a linker having at least 99% or 100% identity to SEQ ID NO: 5217 or SEQ ID NO: 11,041.
[0506] [Table A5-1]
[0507] [Table A5-2]
[0508] [Table A5-3]
[0509] [Table A5-4]
[0510] [Table A5-5]
[0511] [Table A5-6]
[0512] [Table A5-7]
[0513] [Table A5-8]
[0514] system In some embodiments, the present disclosure relates to a system comprising a nucleic acid molecule encoding a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., template RNA, as described herein). In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide contains one or more silent mutations in its coding region (e.g., in the sequence encoding the RT domain) compared to the nucleic acid molecule as described herein. In certain embodiments, the system further comprises a gRNA (e.g., a gRNA that binds to a polypeptide that induces a nick, for example, on the reverse strand of target DNA to which the genetically modified polypeptide binds).
[0515] In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide encodes a polypeptide as shown in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0516] In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, comprising a portion including a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the said portion. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, comprising a portion including a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the said portion. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, comprising a portion including a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the said portion. In certain embodiments, the nucleic acid molecule encoding a gene-modified polypeptide includes a portion of a polypeptide listed in Table A1, Table T1, or Table T2, comprising a portion including a linker and an RT domain, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the portion.
[0517] In certain embodiments, the nucleic acid molecule encoding a gene-modified polypeptide includes a linker of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding a gene-modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding a gene-modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes a polypeptide linker as shown in Table A1, Table T1, or Table T2, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0518] In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes an RT domain of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the nucleic acid molecule encoding the gene-modified polypeptide includes the RT domain of the polypeptide as shown in Table A1, Table T1, or Table T2, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0519] In some embodiments, the present disclosure relates to a system comprising a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., template RNA, e.g., as described herein).
[0520] In certain embodiments, the gene-modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide comprises a polypeptide as shown in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0521] In certain embodiments, the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, comprising a portion including a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the said portion. In certain embodiments, the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, comprising a portion including a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the said portion. In certain embodiments, the gene-modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, comprising a portion including a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the said portion. In certain embodiments, the gene-modified polypeptide is a portion of a polypeptide listed in Table A1, Table T1, or Table T2, comprising a portion including a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the said portion.
[0522] In certain embodiments, the gene-modified polypeptide includes a linker of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes a polypeptide linker as shown in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0523] In certain embodiments, the gene-modified polypeptide includes an RT domain of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the gene-modified polypeptide includes the RT domain of the polypeptide as shown in Table A1, Table T1, or Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.
[0524] [Table A1-1]
[0525] [Table A1-2]
[0526] [Table A1-3]
[0527] [Table A1-4]
[0528] [Table A1-5]
[0529] Localization sequences for gene modification systems In certain embodiments, the gene editor system RNA further comprises an intracellular localization sequence, such as a nuclear localization sequence (NLS). In some embodiments, the gene-modified polypeptide comprises an NLS contained in SEQ ID NO: 4000 and / or SEQ ID NO: 4001 or an NLS having an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[0530] A nuclear localization sequence can be an RNA sequence that facilitates the entry of RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the genetically modified polypeptide is encoded on a first RNA, the template RNA is a second RNA, and the nuclear localization signal is located on the template RNA and not on the RNA encoding the genetically modified polypeptide. While not intended to be bound by any particular theory, in some embodiments, the RNA encoding the genetically modified polypeptide is targeted primarily to the cytoplasm to facilitate its translation, while the template RNA is targeted primarily to the nucleus to facilitate insertion into the genome. In some embodiments, the nuclear localization signal is located within the 3' end, 5' end, or an internal region of the template RNA. In some embodiments, the nuclear localization signal is at the 3' end of the heterologous sequence (e.g., directly at the 3' end of the heterologous sequence) or at the 5' end of the heterologous sequence (e.g., directly at the 5' end of the heterologous sequence). In some embodiments, the nuclear localization signal is located outside the 5'UTR or outside the 3'UTR of the template RNA. In some embodiments, the nuclear localization signal is located between the 5'UTR and 3'UTR, and optionally, the nuclear localization signal is not transcribed by the transgene (e.g., the nuclear localization signal is antisense directional or downstream of a transcription termination signal or polyadenylation signal). In some embodiments, the nuclear localization sequence is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present in the RNA, for example, in the template RNA. In some embodiments, the nuclear localization signal is less than 5 bp, 10 bp, 25 bp, 50 bp, 75 bp, 100 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, or 1000 bp in length. Various RNA nuclear localization sequences are available. For example, Lubelsky and Ulitsky, Nature 555(107-111), 2018 describe RNA sequences that drive the localization of RNA into the nucleus. In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal.In some embodiments, the nuclear localization signal binds to nuclear-enriched proteins. In some embodiments, the nuclear localization signal binds to HNRNPK proteins. In some embodiments, the nuclear localization signal is abundant in pyrimidines, e.g., C / T-rich, C / U-rich, C-rich, T-rich, or U-rich regions. In some embodiments, the nuclear localization signal originates from long uncoding RNA. In some embodiments, the nuclear localization signal originates from MALAT1 long uncoding RNA or is the 600-nucleotide M region of MALAT1 (as described in Miyagawa et al., RNA 18, (738-751), 2012). In some embodiments, the nuclear localization signal is derived from BORG long uncoding RNA or is an AGCCC motif (as described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.
[0531] In some embodiments, the polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In some embodiments, the NLS is a bipartite NLS. In some embodiments, the NLS facilitates the input of the protein containing the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of the genetically modified polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the genetically modified polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of the Cas domain. In some embodiments, a linker sequence is positioned between the NLS and the adjacent domain of the genetically modified polypeptide.
[0532] In some embodiments, NLS is the amino acid sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 5009), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5010), RKSGKIAAIWKRPRKPKKKRKV (SEQ ID NO: 5011), KRTADGSEFESPKKKRKV (SEQ ID NO: 5012), KKTELQTTNAENKTKKL (SEQ ID NO: 5013), or KRGINDRNFWRGENGRKT This includes R (SEQ ID NO: 5014), KRPAATKKAGQAKKKK (SEQ ID NO: 5015), PAAKRVKLD (SEQ ID NO: 4644), KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4649), KRTADGSEFE (SEQ ID NO: 4001), KRTADGSEFESPKKKAKVE (SEQ ID NO: 4651), AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4651), or functional fragments or variants thereof. Exemplary NLS sequences are also described in PCT / European Patent No. 2000 / 011690, the contents of which are incorporated herein by reference with respect to the disclosure of its exemplary nuclear localization sequences. In some embodiments, the NLS includes amino acid sequences such as those disclosed in Table 11. The NLSs in this table may be used as one, two, three, or more copies of the NLS at one or more positions within the polypeptide, e.g., within the N-terminal domain, between peptide domains, within the C-terminal domain, or in combinations of positions, to improve intracellular localization to the nucleus. Multiple unique sequences may be used within a single polypeptide. Sequences may be naturally monopartite or bipartite, or used as chimeric bipartite sequences, for example, having one or two basic amino acid stretches. Sequence references correspond to UniProt acceptance numbers, unless indicated as SeqNLS for sequences extracted using intracellular localization prediction algorithms (Lin et al BMC Bioinformat 13:157 (2012), the entire work of which is incorporated herein by reference).
[0533] [Table 11-1]
[0534] [Table 11-2]
[0535] [Table 11-3]
[0536] [Table 11-4]
[0537] [Table 11-5]
[0538] [Table 11-6]
[0539] In some embodiments, the NLS is a binocular NLS. A binocular NLS typically contains two basic amino acid clusters (e.g., possibly about 10 amino acids long) separated by a spacer sequence. A mononocular NLS typically lacks a spacer. An example of a binocular NLS is the nucleoplasmin NLS having the sequence KR[PAATKKAGQA]KKKK (SEQ ID NO: 5015) (spacer in parentheses). Another exemplary binocular NLS has the sequence PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5016). Exemplary NLSs are described in International Publication No. 2020051561 (the entire brochure, including the disclosure of the nuclear localization sequence, is incorporated herein by reference).
[0540] In certain embodiments, the gene editor system polypeptide (e.g., the gene-modified polypeptide described herein) further comprises an intracellular localization sequence, e.g., a nuclear localization sequence and / or a nucleolar localization sequence. The nuclear localization sequence and / or nucleolar localization sequence may be an amino acid sequence that facilitates the input of the protein into the nucleus and / or nucleolus, which may facilitate the integration of heterologous sequences into the genome. In certain embodiments, the gene editor system polypeptide (e.g., the gene-modified polypeptide described herein as an example) further comprises a nucleolar localization sequence. In certain embodiments, the gene-modified polypeptide is encoded on a first RNA, the template RNA is a second separated RNA, and the nucleolar localization signal is encoded on the RNA encoding the gene-modified polypeptide and not on the template RNA. In some embodiments, the nucleolar localization signal is located within the N-terminus, C-terminus, or internal region of the polypeptide. In some embodiments, multiple identical or different nucleolar localization signals are used. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, or 100 amino acids long. Nucleolar localization signals of various polypeptides are available. For example, Yang et al., Journal of Biomedical Science 22,33 (2015) describes a nuclear localization signal that also functions as a nucleolar localization signal. In some embodiments, the nucleolar localization signal may also be a nuclear localization signal. In some embodiments, the nucleolar localization signal may overlap with the nuclear localization signal. In some embodiments, the nucleolar localization signal may include stretches of basic residues. In some embodiments, the nucleolar localization signal may be abundant in arginine and lysine residues. In some embodiments, the nucleolar localization signal may originate from proteins that are abundant in the nucleolus. In some embodiments, the nucleolar localization signal may originate from proteins that are abundant in ribosomal RNA loci. In some embodiments, the nucleolar localization signal may originate from a protein that binds to rRNA. In some embodiments, the nucleolar localization signal may originate from MSP58. In some embodiments, the nucleolar localization signal may be a monosegmental motif.In some embodiments, the nucleolus localization signal may be a binocular motif. In some embodiments, the nucleolus localization signal may consist of multiple mononocular or binocular motifs. In some embodiments, the nucleolus localization signal may consist of a mixture of mononocular and binocular motifs. In some embodiments, the nucleolus localization signal may be a double binocular motif. In some embodiments, the nucleolus localization motif may be KRASSQALGTIPKRRSSSRFIKRKK (SEQ ID NO: 5017). In some embodiments, the nucleolus localization signal may originate from an intranuclear factor κB-induced kinase. In some embodiments, the nucleolus localization signal may be the RKKRKKK motif (SEQ ID NO: 5018) (as described in Birbach et al., Journal of Cell Science, 117(3615-3624), 2004).
[0541] Evolutionary variants of gene-modified polypeptides and systems In some embodiments, the present invention provides evolved variants of the gene-modified polypeptides described herein. In some embodiments, evolved variants may be produced by mutagenes of the reference gene-modified polypeptide or one of the fragments or domains contained therein. In some embodiments, one or more domains (e.g., reverse transcriptase domains) evolve. In some embodiments, one or more such evolved variant domains may evolve alone or in combination with other domains. In some embodiments, one or more evolved variant domains may be combined with a non-evolving cognitive component or an evolved variant of a cognitive component (e.g., one that may have evolved in parallel or sequentially).
[0542] In some embodiments, the process of mutagenerating a reference gene-modified polypeptide or a fragment or domain thereof includes mutagenerating a reference gene-modified polypeptide or a fragment or domain thereof. In some embodiments, mutagenesis includes, for example, a progressive evolutionary method (e.g., PACE) or a non-progressive evolutionary method (e.g., PANCE), as described herein. In some embodiments, an evolutionary gene-modified polypeptide or a fragment or domain thereof includes one or more amino acid mutations introduced into its amino acid sequence compared to the amino acid sequence of the reference gene-modified polypeptide or a fragment or domain thereof. In some embodiments, the amino acid sequence mutation may include one or more mutant residues (e.g., conserved substitutions, non-conserved substitutions, or combinations thereof) in the amino acid sequence of the reference gene-modified polypeptide as a result of, for example, a change in the nucleotide sequence encoding the gene-modified polypeptide resulting in a change in any particular codon position in the coding sequence, a deletion of one or more amino acids (e.g., a truncated protein), an insertion of one or more amino acids, or any combination thereof. An evolutionary mutant gene-modified polypeptide may include a mutant in one or more components or domains of the gene-modified polypeptide (e.g., a mutant introduced into the reverse transcriptase domain).
[0543] In some embodiments, the Disclosure provides genetically modified polypeptides, systems, kits and methods that use or include evolutionary variants of genetically modified polypeptides, for example, evolutionary variants of genetically modified polypeptides or genetically modified polypeptides that can be produced by PACE or PANCE. In some embodiments, the non-evolutionary reference genetically modified polypeptide is a genetically modified polypeptide disclosed herein.
[0544] The term "phage-assisted continuous evolution (PACE)," as used herein, generally refers to gradual evolution using a phage as a viral vector. Examples of PACE techniques include, for example, International PCT application number PCT / US2009 / 056194, filed on September 8, 2009, and published on March 11, 2010, as International Publication No. 2010 / 028347; International PCT application, PCT / US2011 / 066747, filed on December 22, 2011, and published on June 28, 2012, as International Publication No. 2012 / 088381; U.S. Patent No. 9,023,594, issued on May 5, 2015; U.S. Patent No. 9,771,574, issued on September 26, 2017; and issued on July 19, 2016. These are described in U.S. Patent No. 9,394,537; International PCT Application, PCT / US2015 / 012022, filed on 20 January 2015 and published on 11 September 2015 as International Publication No. 2015 / 134121; U.S. Patent No. 10,179,911, published on 15 January 2019; and International PCT Application, PCT / US2016 / 027795, filed on 15 April 2016 and published on 20 October 2016 as International Publication No. 2016 / 168631, which are each incorporated herein by reference in their entirety.
[0545] The term "phage-assisted non-gradual evolution (PANCE)," as used herein, generally refers to non-gradual evolution using a phage as a viral vector. An example of PANCE technology is described, for example, in Suzuki T. et al, Crystal structures reveal an elusive functional domain of pyrrolysyl-tRNA synthetase, Nat Chem Biol. 13(12):1261-1266 (2017), which is incorporated herein by reference in its entirety. Briefly, PANCE is a technique for rapid in vivo-directed evolution using sequential flask transfer of evolutionary selection phages (SPs) containing the genes of interest to be evolved into whole fresh host cells (e.g., E. coli cells). The genes inside the host cells can be kept constant while the genes contained in the SPs evolve progressively. After phage proliferation, aliquots of infected cells can be used to transfect the next flask containing host E. coli cells. This process can be iterated and / or continued until the desired phenotype evolves, for example, with respect to a desired number of imports.
[0546] Methods for applying PACE and PANCE to gene-modified polypeptides will be readily apparent to those skilled in the art, particularly by referring to the aforementioned references. For example, evolutionary variants of gene-modified polypeptides or fragments thereof or subdomains can be produced by applying further exemplary methods for directing the gradual evolution of genome-modified proteins or systems, for example, in a population of host cells, using phage particles. Non-limiting examples of such methods are found in International PCT Application, Specification PCT / US2009 / 056194, filed on 8 September 2009 and published on 11 March 2010 as International Publication No. 2010 / 028347; and International PCT Application, filed on 22 December 2011 and published on 28 June 2012 as International Publication No. 2012 / 088381. PCT / US2011 / 066747; US Patent No. 9,023,594 issued on May 5, 2015; US Patent No. 9,771,574 issued on September 26, 2017; US Patent No. 9,394,537 issued on July 19, 2016; filed on January 20, 2015 and published internationally on September 11, 2015, No. 2015 / 134121 These are described in the International PCT Application, Specification PCT / US2015 / 012022, published as a brochure; U.S. Patent No. 10,179,911, issued on January 15, 2019; International Patent Application No. PCT / US2019 / 37216, filed on June 14, 2019, and published as International Publication Brochure No. 2019 / 023680 on January 31, 2019; International PCT Application, Specification PCT / US2016 / 027795, filed on April 15, 2016, and published as International Publication Brochure No. 2016 / 168631 on October 20, 2016; and International Application, Specification PCT / US2019 / 47996, filed on August 23, 2019, each incorporated herein by reference in its entirety.
[0547] In some non-limiting exemplary embodiments, a method for evolving an evolutionary mutant gene-modified pol...
Claims
1. template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) Mutant St1Cas9 scaffold having a partial or complete deletion of stem loop 2, (3) heterologous sequences and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing this RNA.
2. The template RNA according to claim 1, wherein the deletion has a length of 1 to 32 nucleotides (for example, 2 to 29, 2 to 20, 2 to 10, or 10 to 20).
3. The template RNA according to claim 1, wherein the deletion is located at positions 55 to 84.
4. The template RNA according to any one of claims 1 to 3, wherein the mutant St1Cas9 scaffold has either an elongated RAR upper stem or a substitution that generates a G-C base pair in the RAR upper stem, or both.
5. The mutant St1Cas9 scaffold is the template RNA according to any one of claims 1 to 4, having a mutation in the tetraloop.
6. template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having an elongated RAR upper stem or one or both of the substitutions that produce a G-C base pair in the RAR upper stem, (3) heterologous sequences and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing this RNA.
7. The template RNA according to any one of claims 1 to 6, wherein the upper stem of the RAR is extended by 1 to 8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7, or 8 base pairs) compared to the wild-type sequence of Sequence ID No. 25999.
8. template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) Mutant St1Cas9 scaffold with mutation in the tetraloop, (3) heterologous sequences and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing this RNA.
9. The template RNA according to any one of claims 1 to 8, wherein the tetraloop extends, for example, up to 5 nucleotides.
10. The template RNA according to any one of claims 1 to 9, which is the mutant gRNA scaffold containing the sequence shown in Table 23.
11. The mutant gRNA scaffold is a template RNA according to any one of claims 1 to 10, comprising the sequence shown in Table 23.
12. The mutant gRNA scaffold is the template RNA according to any one of claims 1 to 11, comprising the sequence relating to GUCUUUGUACUCUGGUACCAGAAGAAGCUACAAAGAAUAGCUUCAUUGCCGAAAAAUCA (SEQ ID NO: 26000).
13. The template RNA according to any one of claims 1 to 12, wherein the gRNA spacer includes a sequence relating to Table 22 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
14. The gRNA spacer is a template RNA according to any one of claims 1 to 13, comprising the sequence shown in Table 22.
15. The gRNA spacer is a template RNA according to any one of claims 1 to 14, comprising a sequence relating to AAGGCUGUGCUGAACCAUCGA (Sequence ID 26001).
16. The template RNA according to any one of claims 1 to 15, wherein the heterogeneous target sequence includes a sequence relating to Table 24 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
17. The aforementioned heterogeneous target sequence includes the sequence shown in Table 24, wherein the template RNA is according to any one of claims 1 to 16.
18. The template RNA according to any one of claims 1 to 17, wherein the PBS sequence includes a sequence relating to Table 25 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
19. The template RNA according to any one of claims 1 to 18, wherein the PBS sequence includes the sequence shown in Table 25.
20. A template RNA according to any one of claims 1 to 19, comprising a sequence relating to Table 20 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
21. A template RNA according to any one of claims 1 to 20, comprising a sequence relating to Table 21 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
22. A template RNA according to any one of claims 1 to 21, comprising a sequence relating to any one of Table 27, Table E3, Table E3A, Table E7, Table E8, Table E9, Table El1A, Table E11B, Table E12, Table E12A, Table E14, Table E14A, Table E15, or Table E16, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
23. A template RNA according to any one of claims 1 to 22, comprising one or more chemical modifications.
24. The template RNA according to claim 23, comprising one or more phosphorothioate bonds.
25. The template RNA according to claim 23 or 24, comprising one or more 2'-O-methylnucleotides.
26. The template RNA according to any one of claims 23 to 25, comprising a sequence relating to the third column of Table 20 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
27. The template RNA according to any one of claims 23 to 25, comprising a sequence relating to the third column of Table 21 or a sequence having one, two, or three or fewer sequence modifications (e.g., substitutions) compared thereto.
28. It is a gene modification system, A template RNA according to any one of claims 1 to 27, A gene-modified polypeptide or a nucleic acid encoding the gene-modified polypeptide, wherein the gene-modified polypeptide is (1) St1Cas9 domain, (2) Linker, and (3) Reverse transcriptase (RT) domain Genetically modified polypeptides or nucleic acids, including A gene modification system that includes this.
29. The system according to claim 28, wherein the St1Cas9 domain is niccase.
30. The system according to claim 28 or 29, wherein the St1Cas9 domain includes the sequence relating to sequence number 23818 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
31. The system according to any one of claims 28 to 30, wherein the linker includes the sequence relating to sequence number 5006 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
32. The system according to any one of claims 28 to 30, wherein the linker includes the sequence relating to sequence number 5217 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
33. The system according to any one of claims 28 to 30, wherein the linker includes the sequence of Table 10 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
34. The system according to any one of claims 28 to 31, wherein the RT domain includes the sequence relating to sequence number 26006 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
35. The system according to any one of claims 28 to 31, wherein the RT domain includes a sequence relating to any of sequence numbers 8,001 to 8,003 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
36. The system according to any one of claims 28 to 31, wherein the RT domain includes the sequence in Table 6 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
37. The system according to any one of claims 28 to 34, wherein the gene-modified polypeptide comprises the sequence relating to Sequence ID No. 26002 or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity thereto.
38. The system according to any one of claims 28 to 37, further comprising a second nick gRNA (ngRNA), wherein the second nick gRNA optionally directs the second nick to the second strand of the human SERPINA1 gene.
39. The system according to claim 38, wherein the second nick gRNA includes a sequence relating to Table 26 or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity thereto.
40. The system according to any one of claims 28 to 39, wherein the nucleic acid encoding the gene-modified polypeptide includes RNA, for example, mRNA.
41. The template RNA or system according to any one of claims 1 to 40, wherein the nucleic acid molecule is formulated into lipid nanoparticles (LNPs).
42. The system according to any one of claims 28 to 41, wherein the template RNA, the nucleic acid molecule encoding the gene-modified polypeptide and / or the second nick gRNA are formulated into LNPs.
43. A pharmaceutical composition comprising a system according to any one of claims 28 to 42 or one or more nucleic acids encoding the same, and a pharmaceutically acceptable excipient or carrier.
44. The pharmaceutical composition according to claim 43, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).
45. The pharmaceutical composition according to claim 44, wherein the viral vector is an adeno-associated virus.
46. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising a gene modification system or template RNA according to any one of claims 1 to 42.
47. A method for preparing a template RNA according to any one of claims 1 to 27, comprising synthesizing the template RNA in vitro (for example, by in vitro transcription or solid-phase synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that enable the production of the template RNA.
48. A method for modifying a target site in a cell (for example, a target site in the human SERPINA1 gene), comprising contacting the cell with a gene modification system or DNA encoding it as described in any one of claims 28 to 42, or a pharmaceutical composition as described in any one of claims 43 to 45, thereby modifying the target site.
49. A method for treating a subject having a disease or condition related to a mutation of a gene (for example, the human SERPINA1 gene), comprising administering to the subject a gene modification system or DNA encoding the same as described in any one of claims 28 to 42, or a pharmaceutical composition as described in any one of claims 43 to 45, thereby treating the subject having the disease or condition.
50. The method according to claim 49, wherein the disease or condition is alpha-1 antitrypsin deficiency (AATD).
51. The method according to claim 49 or 50, wherein the subject has the E342K mutation.
52. A method for treating a subject having AATD, comprising administering to the subject a gene modification system according to any one of claims 28 to 42 or DNA encoding the same, or a pharmaceutical composition according to any one of claims 43 to 45, thereby treating the subject having AATD.
53. The gene modification system or method according to any one of claims 28 to 42 or 47 to 52, wherein the introduction of the system into target cells results in the modification of a pathogenic mutation in the gene, for example, the SERPINA1 gene.
54. The gene modification system or method according to any one of claims 28 to 42 or 47 to 53, wherein the pathogenic mutation is an E342K mutation, and the modification comprises an amino acid substitution of K342E.
55. The gene modification system or method according to any one of claims 28 to 42 or 47 to 54, wherein the introduction of the system into target cells results in a mutation that causes the restoration of the function of the gene, for example, the SERPINA1 gene.
56. A gene modification system or method according to any one of claims 28 to 42 or 47 to 55, wherein the modification of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more) of the target nucleic acid.
57. The gene modification system or method according to any one of claims 28 to 42 or 47 to 56, wherein the modification of the mutation occurs in at least 10% of the target cells (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, or more).
58. The gene modification system or method according to any one of claims 28 to 42 or 47 to 57, wherein the gene modification system comprises a second nick gRNA, and the modification of the mutation in a population of target cells is increased compared to a population of target cells treated with a gene modification system comprising a template RNA without the second nick gRNA.
59. The gene modification system or method according to any one of claims 28 to 42 or 47 to 58, wherein the template RNA comprises one or more silent substitutions, and the modification of the mutation in a population of target cells is increased compared to a population of target cells treated with a gene modification system comprising template RNA that does not contain one or more silent substitutions.
60. The method according to any one of claims 47 to 59, wherein the cells are mammalian cells such as human cells.
61. The method according to any one of claims 47 to 60, wherein the subject is a human.
62. The method according to any one of claims 47 to 61, wherein the contact is performed ex vivo, and for example, the DNA of the cell or subject is modified ex vivo.
63. The method according to any one of claims 47 to 62, wherein the contact is performed in vivo, and for example, the DNA of the cell or subject is modified in vivo.
64. The method according to any one of claims 47 to 63, wherein bringing the cells or object into contact with the system includes bringing the cells or cells within the object into contact with a nucleic acid (e.g., DNA or RNA) encoding the gene-modified polypeptide under conditions that enable the production of the gene-modified polypeptide.
65. gRNA, (for example, from 5' to 3') (1) gRNA spacer and, (2) A mutant St1Cas9 scaffold, (a) Deletion of part or all of stem loop 2, (b) Extended RAR upper stem, (c) Substitution that generates a G-C base pair in the upper stem of the RAR, or (d) Mutations in tetraloops A mutant St1Cas9 scaffold having one or more of the following: gRNA containing this substance.
66. template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) A mutant St1Cas9 scaffold having a substitution that generates a G-C base pair in the lower RAR stem, (3) heterologous sequences and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing this RNA.
67. template RNA (tgRNA), (for example, from 5' to 3') (1) gRNA spacer, (2) Mutant St1Cas9 scaffold having substitution in the second single-stranded region, (3) heterologous sequences and (4) Primer binding site (PBS) sequence Template RNA (tgRNA) containing this RNA.