Serpin a modulating compositions and methods

The gene modification system, which combines reverse transcriptase and the St1Cas9 domain, solves the problem of low genome insertion efficiency in existing technologies, corrects SERPINA1 gene mutations, improves α-1 antitrypsin activity, and reduces symptoms of AATTD-related liver and lung diseases.

CN122249563APending Publication Date: 2026-06-19FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 221 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FLAGSHIP PIONEERING INNOVATIONS VI LLC
Filing Date
2024-03-14
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies have low frequency and site specificity for inserting target nucleic acids into the genome, and existing methods such as CRISPR/Cas9 are inefficient when integrating longer sequences, making them ineffective in treating liver and lung diseases caused by α-1 antitrypsin deficiency (AATD).

Method used

Using gene-modified peptides containing reverse transcriptase (RT) and St1Cas9 domains, combined with template RNA of a variant gRNA scaffold, the genome can be altered in vivo or in vitro via a gene-modification system to insert, modify, or delete target sequences in order to correct the E342K mutation in the SERPINA1 gene and correct α-1 antitrypsin deficiency.

Benefits of technology

It enables efficient, site-specific insertion or deletion of DNA in the genome, corrects SERPINA1 gene mutations, increases α-1 antitrypsin activity, reduces liver toxicity and lung disease symptoms, and provides a more effective AATD treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

This disclosure provides, for example, compositions, systems, and methods for targeting, editing, modifying, or manipulating the host cell genome at one or more locations within a DNA sequence in a cell, tissue, or subject. An improved gRNA scaffold compatible with St1Cas9 is described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 490,462, filed March 15, 2023. The contents of the above application are incorporated herein by reference in their entirety.

[0003] sequence list

[0004] This application contains a sequence list which has been filed electronically in XML format conforming to WIPO standard ST.26, and is hereby incorporated by reference in its entirety. The XML copy was created on December 12, 2023, named V2065-704901_SL.XML, and is 15,726,626 bytes in size. Background Technology

[0005] Without specific proteins to facilitate insertion events, the integration of target nucleic acids into the genome is infrequent and site-specific. Some existing methods, such as CRISPR / Cas9, are better suited for small edits that rely on host repair pathways and are less efficient at integrating longer sequences. Other existing methods, such as Cre / loxP, require a first step of inserting the loxP site into the genome, followed by a second step of inserting the target sequence into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, altering, or deleting target sequences into the genome.

[0006] AATD is characterized by low circulating levels of AAT. AAT is primarily produced and secreted into the bloodstream by liver cells, but it is also produced by other cell types, including lung epithelial cells and some leukocytes. AAT inhibits several serine proteases secreted by inflammatory cells (most notably neutrophil elastase [NE], protease 3, and cathepsin G), thereby protecting organs such as the lungs from protease-induced damage, especially during inflammation.

[0007] The two most common clinical variants of AAT are the E264V (PiS) and E342K (PiZ) alleles. The clinical single nucleotide variant E342K (PiZ) results in unstable and / or inactive AAT protein structure, leading to liver toxicity and pulmonary inactivation. The inheritance pattern is autosomal codominant. More than half of patients with AATD carry at least one copy of the mutated E342K.

[0008] The most common AAT-related mutation involves a glutamic acid substitution for a lysine residue (E342K) in the SERPINA1 gene, which encodes the AAT protein. The E342K mutation is located at the hinge between the β-sheet and the reaction center loop (RCL) of the AAT protein and leads to the formation of a ring-sheet dimer, which can then extend to form long-chain ring-sheet polymers. These polymers aggregate AAT-Z protein within the rough endoplasmic reticulum (rER) of hepatocytes during biosynthesis. This mutation (called a Z mutation or Z allele) results in misfolded translated protein, thus preventing secretion into the bloodstream. Consequently, circulating AAT levels are significantly reduced in Z-allele homozygous (PiZZ) individuals; only about 15% of the mutant Z-AAT protein is correctly folded and secreted by cells. Another consequence of the Z mutation is reduced secreted Z-AAT activity compared to wild-type protein, which is 40% to 80% of normal antiprotease activity (American Thoracic Society / European Respiratory Society, AmJ Respir Crit Care Med. 2003; 168(7):818-900; and Ogushi et al. J Clin Invest. 1987; 80(5):1366-74).

[0009] Two disease phenotypes are associated with the PiZZ genotype. Accumulation of polymerized Z-AAT protein in hepatocytes leads to gain-of-function cytotoxicity, which can result in cellular stress, inflammation, fibrosis, cirrhosis, hepatocellular carcinoma (HCC), and neonatal liver disease (12% of patients). This accumulation may resolve spontaneously, but can be fatal in a small percentage of children. The loss-of-function phenotype is due to decreased systemic AAT levels leading to increased proteolytic digestion of connective tissue in the lower respiratory tract. Excessive proteolytic digestion of connective tissue and alveolar lining reduces lung elasticity and function, leading to emphysema, a hallmark of chronic obstructive pulmonary disease (COPD). This effect is very pronounced in PiZZ individuals and typically manifests in middle age, resulting in decreased quality of life and shortened lifespan (mean 68 years) (Tanash et al., Int J Chron Obstruct Pulm Dis. [International Journal of Chronic Obstructive Lung Disease] 2016; 11:1663-9). This effect is more pronounced in smoking PiZZ individuals, leading to a further shortened lifespan (58 years). Piitulainen and Tanash, COPD 2015; 12(1):36-41. PiZZ individuals account for the majority of clinically relevant AADTD lung disease patients.

[0010] A milder form of AATD is associated with the SZ genotype, where the Z allele is combined with the S allele. The S allele is associated with a slight decrease in circulating AAT levels but does not cause cytotoxicity in hepatocytes. The result is clinically significant lung disease, but no liver disease. Fregonese and Stolk, Orphanet J Rare Dis. [Rare Disease Journal] 2008;33:16. Similar to the ZZ genotype, a deficiency of circulating AAT in subjects with the SZ genotype leads to unregulated protease activity, which degrades lung tissue over time and may result in emphysema (especially in smokers).

[0011] While limited treatment options exist for AAT deficiency, there is currently no cure. A small percentage of neonatal patients and those with advanced liver disease undergo liver transplantation. For AAT-deficient individuals with severe lung disease or signs of developing severe lung disease, the current standard of care is enhancement therapy or protein replacement therapy. Enhancement therapy involves administering a human AAT protein concentrate purified from pooled donor plasma to enhance the missing AAT. This treatment involves weekly infusions of AAT protein purified from a healthy blood donor. Although plasma protein infusions have been shown to improve survival or slow the progression of emphysema, enhancement therapy is often insufficient under challenging conditions such as active lung infection. Enhancement therapy also fails to restore normal physiological regulation of AAT in patients, and its efficacy is difficult to demonstrate. Furthermore, enhancement therapy cannot address liver disease caused by acquired toxic function of the Z allele. Therefore, new and more effective treatments for AAT deficiency are needed. Summary of the Invention

[0012] This disclosure relates to novel compositions, systems, and methods for altering the genome at one or more sites in a host cell, tissue, or subject, either in vivo or in vitro. This disclosure provides, for example, gene-modifying systems comprising a gene-modifying polypeptide containing a reverse transcriptase (RT) domain and a St1Cas9 domain, and a template RNA containing a variant gRNA scaffold that has been engineered to improve performance, for example, when used in conjunction with the St1Cas9 domain. This disclosure also provides gene-modifying systems capable of modulating (e.g., inserting, altering, or deleting a target sequence) α-1 antitrypsin (AAT) activity, and methods for treating α-1 antitrypsin deficiency (AATD) by administering one or more such systems to alter the genomic sequence at a single nucleotide site.

[0013] In one aspect, this disclosure relates to a system for modifying DNA to correct mutations in the human SERPINA1 gene causing AATD, the system comprising (a) a nucleic acid encoding a gene-modifying polypeptide capable of targeted reverse transcription, the polypeptide comprising (i) a reverse transcriptase domain and (ii) a St1Cas9 nickase that binds to DNA and has endonuclease activity; and (b) a template RNA comprising (i) a gRNA spacer complementary to a first portion of the human SERPINA1 gene, (ii) a gRNA scaffold for binding the polypeptide, (iii) a heterologous target sequence containing a mutation region to correct the mutation, and (iv) a primer binding site (PBS) sequence comprising at least 3, 4, 5, 6, 7, or 8 bases having 100% homology with the target DNA strand at the 3' end of the template RNA. The SERPINA1 gene may contain an E342K mutation (also known as a PiZ mutation). The template RNA sequence may comprise sequences described herein, such as those in Tables 1, 3, 4, 5, 6a, 6B, X3, or X3a.

[0014] The gRNA spacer may contain at least 15 bases that are 100% homologous to the target DNA at the 5' end of the template RNA. The template RNA may further contain a PBS sequence containing at least 5 bases that are at least 80% homologous to the target DNA strand. The template RNA may contain one or more chemical modifications.

[0015] The domains of a genetically modified polypeptide can be linked by peptide linkers. The polypeptide may contain one or more peptide linkers. The genetically modified polypeptide may further contain nuclear localization signals. The polypeptide may contain more than one nuclear localization signal, for example, multiple adjacent nuclear localization signals or one or more nuclear localization signals in different regions of the polypeptide, such as one or more nuclear localization signals at the N-terminus and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding the genetically modified polypeptide may encode one or more inteptide domains.

[0016] Introducing this system into target cells can result in the insertion of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs of exogenous DNA. Introducing this system into target cells can also result in deletions, where the deletion is less than 2, 3, 4, 5, 10, 50, or 100 base pairs of genomic DNA upstream or downstream of the insertion. Introducing this system into target cells can also result in substitutions, such as substitutions of 1, 2, or 3 nucleotides (e.g., consecutive nucleotides).

[0017] The heterologous object sequence can be at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600 or 700 base pairs.

[0018] In one aspect, this disclosure relates to a pharmaceutical composition comprising the above-described system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles. In another aspect, this disclosure relates to a pharmaceutical composition comprising the above-described system and a plurality of pharmaceutically acceptable excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles, for example, wherein the above-described system is delivered by two different excipients or carriers, such as two lipid nanoparticles, two viral vectors, or one lipid nanoparticle and one viral vector. The viral vector may be adeno-associated virus (AAV).

[0019] In one aspect, this disclosure relates to a host cell (e.g., a mammalian cell, such as a human cell) that contains the aforementioned system.

[0020] In one aspect, this disclosure relates to a method for correcting mutations in the human SERPINA1 gene in a cell, tissue, or subject, the method comprising administering the aforementioned system to the cell, tissue, or subject, wherein optionally, correction of the mutated SERPINA1 gene comprises an amino acid substitution of K342E (reversing the pathogenic substitution E342K). The system can be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid of (a) can be integrated into the genome of the host cell. In some embodiments, the nucleic acid of (a) is not integrated into the genome of the host cell. In some embodiments, the heterologous object sequence is inserted into only one target site in the host cell genome. The heterologous object sequence can be inserted into two or more target sites in the host cell genome, for example, into the same corresponding site on two homologous chromosomes, or into two different sites on the same or different chromosomes. The heterologous object sequence can encode a mammalian polypeptide or a fragment or variant thereof. Components of the system can be delivered on 1, 2, 3, 4, or more different nucleic acid molecules. The system can be introduced into host cells by electroporation or by using at least one medium selected from plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

[0021] These compositions or methods may include one or more of the features listed in the examples below.

[0022] Listed Examples

[0023] 1. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0024] (1) gRNA spacers;

[0025] (2) A variant St1Cas9 scaffold with partial or complete absence of stem ring 2;

[0026] (3) Heterogeneous object sequences; and

[0027] (4) Primer binding site (PBS) sequence.

[0028] 2. The template RNA as described in Example 1, wherein the length of the deletion is between 1 and 32 (e.g., 2-29, 2-20, 2-10, or 10-20) nucleotides.

[0029] 3. The template RNA as described in Example 1, wherein the deletion is a deletion of the entire stem-loop 2.

[0030] 4. The template RNA as described in Example 1, wherein the deletion is a deletion at positions 55 to 84.

[0031] 5. The template RNA as described in Example 1, wherein the St1Cas9 scaffold includes a portion of the second single-stranded region (e.g., the deletion of 1, 2, 3, or 4 nucleotides at the 3' end of the single-stranded region).

[0032] 6. The template RNA as described in any of the preceding embodiments, wherein the variant St1Cas9 scaffold has one or both of the following: an elongated RAR upper stem or substitutions that produce GC base pairs in the RAR upper stem.

[0033] 7. The template RNA as described in any of the foregoing embodiments, wherein the variant St1Cas9 scaffold has a mutation in the four-loop.

[0034] 8. The template RNA as described in any of the foregoing embodiments, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to GACUUCGGUC.

[0035] 9. The template RNA as described in any of the foregoing embodiments, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to CUAGAAAUAG.

[0036] 10. The template RNA as described in any of the preceding embodiments, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 12-nucleotide insertion) between positions 14 and 19, and a deletion at positions 15-18, wherein optionally the insertion has a sequence according to CGCGGUAACGCG.

[0037] 11. The template RNA as described in any of the foregoing embodiments, wherein the variant St1Cas9 scaffold has substitutions that produce GC base pairs in the lower stem of the RAR, wherein optionally the substitutions include a G substitution at position 4, and the template further includes a C substitution at position 31.

[0038] 12. The template RNA as described in any of the foregoing embodiments, wherein the template RNA contains a substitution in the second single-stranded region, wherein optionally the substitution is a U substitution at position 51 or a C substitution at position 54.

[0039] 13. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0040] (1) gRNA spacers;

[0041] (2) A variant St1Cas9 scaffold having one or both of the following: an elongated RAR upper stem or substitution of GC base pairs in the RAR upper stem;

[0042] (3) Heterogeneous object sequences; and

[0043] (4) Primer binding site (PBS) sequence.

[0044] 14. The template RNA as described in Example 13, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to GACUUCGGUC.

[0045] 15. The template RNA as described in Example 13, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to CUAGAAAUAG.

[0046] 16. The template RNA as described in Example 13, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 12-nucleotide insertion) between positions 14 and 19, and a deletion at positions 15-18, wherein optionally the insertion has a sequence according to CGCGGUAACGCG.

[0047] 17. The template RNA as described in any of the preceding embodiments, wherein the upper stem of the RAR is extended by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7 or 8 base pairs) relative to the wild-type sequence of SEQ ID NO:25999.

[0048] 18. The template RNA as described in Example 17, wherein at least 50%, 60%, 70%, 80% or 90% of the new base pairs relative to SEQ ID NO: 25999 are GC base pairs.

[0049] 19. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0050] (1) gRNA spacers;

[0051] (2) A variant St1Cas9 scaffold having substitutions that produce GC base pairs in the lower stem of the RAR;

[0052] (3) Heterogeneous object sequences; and

[0053] (4) Primer binding site (PBS) sequence.

[0054] 20. The template RNA as described in Example 19, wherein the template RNA comprises position 4 replaced by G and position 31 replaced by C.

[0055] 21. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0056] (1) gRNA spacers;

[0057] (2) A variant St1Cas9 scaffold with a mutation in the four-ring;

[0058] (3) Heterogeneous object sequences; and

[0059] (4) Primer binding site (PBS) sequence.

[0060] 22. The template RNA as described in any of the foregoing embodiments, wherein one or more nucleotides in the tetracycle are substituted.

[0061] 23. The template RNA as described in any of the foregoing embodiments, wherein the four loops comprise sequences selected from the following: AACA, AAUA, ACCA, ACUA, AGUA, AGCA, AUCA, AUUA, CAAC, CUCG, CUUG, GAAA, GAGA, GCAA, GCGA, GGAA, GGAG, GGGA, GUAA, GUGA, UAAC, UACG, UCAC, UCCG, UGAA, UGAC, UGCG, UUAC, or UUCG.

[0062] 24. The template RNA as described in any of the preceding embodiments, wherein the tetraloop is extended to, for example, 5 nucleotides.

[0063] 25. The template RNA as described in Example 24, wherein the elongated four-loop contains a sequence selected from GAAGA or GACAA.

[0064] 26. The template RNA as described in any one of Examples 21-25, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to GACUUCGGUC.

[0065] 27. The template RNA as described in any one of Examples 21-25, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 10-nucleotide insertion) between positions 15 and 18, and deletions at positions 16 and 17, wherein optionally the insertion has a sequence according to CUAGAAAUAG.

[0066] 28. The template RNA as described in any one of Examples 21-25, wherein the St1Cas9 scaffold comprises an insertion (e.g., a 12-nucleotide insertion) between positions 14 and 19, and a deletion at positions 15-18, wherein optionally the insertion has a sequence according to CGCGGUAACGCG.

[0067] 29. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0068] (1) gRNA spacers;

[0069] (2) A variant St1Cas9 scaffold having a substitution in the second single-chain region;

[0070] (3) Heterogeneous object sequences; and

[0071] (4) Primer binding site (PBS) sequence.

[0072] 30. The template RNA as described in Example 29, wherein the template RNA contains position 51 replaced by U.

[0073] 31. The template RNA as described in Example 29, wherein position 54 is replaced by C.

[0074] 32. The template RNA as described in any of the preceding embodiments, wherein the variant gRNA scaffold comprises a sequence according to Table 23, or a sequence having no more than one, two, or three sequence changes (e.g., substitutions) relative to it.

[0075] 33. The template RNA as described in any of the foregoing embodiments, wherein the template RNA comprises a sequence according to any one of Tables 20, 21, 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15 or E16, or a sequence having at least 80%, 85%, 90%, 95%, 98% or 99% identity with it.

[0076] 34. The template RNA as described in any of the foregoing embodiments, wherein the template RNA comprises the sequence according to SEQ ID NO:27131, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0077] 35. The template RNA as described in any of the foregoing embodiments, wherein the template RNA comprises the sequence according to SEQ ID NO:27132, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0078] 36. The template RNA as described in any of the foregoing embodiments, wherein the template RNA comprises the sequence according to SEQ ID NO:27133, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0079] 37. The template RNA as described in any of the foregoing embodiments, wherein the template RNA comprises the sequence according to SEQ ID NO:27134, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0080] 38. The template RNA as described in any of the preceding embodiments, wherein the length of the variant St1Cas9 scaffold is 50-60, 60-70, 70-80, or 80-84 nucleotides.

[0081] 39. A template RNA (tgRNA) comprising (e.g., from 5' to 3'):

[0082] (1) gRNA spacers;

[0083] (2) A variant gRNA scaffold containing a sequence according to Table 23, or a sequence having no more than one, two or three sequence changes (e.g., substitutions) relative to it;

[0084] (3) Heterogeneous object sequences; and

[0085] (4) Primer binding site (PBS) sequence.

[0086] 40. A template RNA, comprising, for example, from 5' to 3':

[0087] (i) A gRNA spacer complementary to the first part of the human SERPINA1 gene, wherein the gRNA spacer has a sequence comprising the core nucleotide of the gRNA spacer sequence in Table 1, or a sequence having 1, 2 or 3 substitutions thereto, and optionally comprises one or more consecutive nucleotides (e.g., comprising one or more flanking nucleotides adjacent to these core nucleotides) starting from the 3' end of the flanking nucleotides of the gRNA spacer, or wherein the gRNA spacer has the sequence of the gRNA spacer in Tables 6A, 6B, X3 or X3a, or a sequence having 1, 2 or 3 substitutions thereto;

[0088] (ii) gRNA scaffolds that bind to genetically modified peptides (e.g., the Cas domain of the genetically modified peptide).

[0089] (iii) A heterologous object sequence containing a mutation region to introduce a mutation into a second part of the human SERPINA1 gene (e.g., to correct a mutation in that second part) (wherein optionally, the heterologous object sequence contains a post-edited homologous region, a mutation region, and a pre-edited homologous region from 5' to 3'), and

[0090] (iv) Primer binding site (PBS) sequence containing at least 3, 4, 5, 6, 7 or 8 bases that are 100% identical to the third part of the human SERPINA1 gene.

[0091] 41. The template RNA as described in any of the preceding embodiments, wherein the heterologous object sequence comprises the core nucleotide of the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the RT template sequence, or wherein the heterologous object sequence comprises the sequence of the RT template sequence in Table 6A or 6B.

[0092] 42. The template RNA as described in any of the preceding embodiments, wherein the heterologous object sequence comprises a core nucleotide corresponding to the RT template sequence in Table 3 of the gRNA spacer sequence, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides (e.g., comprising one or more flanking nucleotides adjacent to these core nucleotides) starting from the 3' end of the flanking nucleotides of the RT template sequence, or wherein the heterologous object sequence comprises a sequence of the RT template sequence in Table 6A or 6B.

[0093] 43. The template RNA as described in any of the preceding embodiments, wherein the heterologous object sequence has a sequence of a heterologous object sequence from the template RNA shown in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it, or a sequence having 1, 2, or 3 substitutions relative to it.

[0094] 44. The template RNA as described in any of the preceding embodiments, wherein the length of the heterologous object sequence is 6-16 nucleotides (e.g., 6, 8, 10, 12, 14, 15, or 16 nucleotides).

[0095] 45. The template RNA as described in any of the preceding embodiments, wherein the PBS sequence has a sequence comprising a core nucleotide from the PBS sequence in the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides (e.g., comprising one or more flanking nucleotides adjacent to these core nucleotides) starting from the 5' end of the flanking nucleotides of the PBS sequence.

[0096] 46. ​​The template RNA as described in any one of Examples 1-44, wherein the PBS sequence has a sequence comprising a core nucleotide of the PBS sequence in Table 3 corresponding to or having one, two or three substitutions thereto, the gRNA spacer sequence or both, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotides of the PBS sequence, or wherein the PBS sequence has a sequence comprising the PBS sequence in Table 6A or 6B corresponding to or having one, two or three substitutions thereto, the RT template sequence, the gRNA spacer sequence or both.

[0097] 47. The template RNA as described in any of the preceding embodiments, wherein the PBS sequence has a PBS sequence from the template RNA shown in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it, or a sequence having 1, 2, or 3 substitutions relative to it.

[0098] 48. The template RNA as described in any of the preceding embodiments, wherein the PBS sequence is 8-12 nucleotides in length (e.g., 8, 9, 10, 11, or 12 nucleotides).

[0099] 49. The template RNA as described in any of the preceding embodiments, wherein the gRNA scaffold comprises a sequence of the gRNA scaffold in Table 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0100] 50. The template RNA as described in any one of Examples 1-48, wherein the gRNA scaffold comprises a sequence corresponding to the RT template sequence, the gRNA spacer sequence, or both of the gRNA scaffolds in Table 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0101] 51. The template RNA as described in any of the preceding embodiments, wherein the gRNA scaffold has a sequence of the gRNA scaffold derived from the template RNA shown in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0102] 52. The template RNA as described in any of the preceding embodiments, comprising the sequence of the template RNA shown in Table X3 or X3a, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0103] 53. A template RNA, comprising, for example, from 5' to 3':

[0104] (i) A gRNA spacer complementary to the first part of the human SERPINA1 gene.

[0105] (ii) gRNA scaffolds that bind to genetically modified peptides (e.g., the Cas domain of the genetically modified peptide).

[0106] (iii) A heterologous object sequence containing a mutation region to introduce a mutation into a second part of the human SERPINA1 gene (e.g., to correct a mutation in that second part), wherein the heterologous object sequence contains the core nucleotide of the RT template sequence in Table 3, or a sequence having one, two, or three substitutions relative to it, and optionally contains one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the RT template sequence, or wherein the heterologous object sequence contains the RT template sequence of Table 6A or 6B; and

[0107] (iv) A PBS sequence containing at least 3, 4, 5, 6, 7 or 8 bases that are 100% identical to the third part of the human SERPINA1 gene.

[0108] 54. The template RNA as described in any of the preceding embodiments, wherein the gRNA spacer comprises the core nucleotide of the gRNA spacer sequence in Table 1, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the gRNA spacer sequence, or wherein the gRNA spacer comprises the gRNA spacer sequence in Table 6A or 6B.

[0109] 55. The template RNA as described in any of the preceding embodiments, wherein the heterologous object sequence comprises a core nucleotide of the gRNA spacer sequence in Table 1 corresponding to the RT template sequence, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the gRNA spacer sequence, or wherein the heterologous object sequence comprises nucleotides of the gRNA spacer sequence in Table 6A or 6B.

[0110] 56. The template RNA as described in any of the preceding embodiments, wherein the PBS sequence has a sequence comprising a core nucleotide from the PBS sequence in the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotides of the PBS sequence.

[0111] 57. The template RNA as described in any one of Examples 1-55, wherein the PBS sequence has a sequence comprising a core nucleotide corresponding to the RT template sequence or a sequence having one, two, or three substitutions thereto, the gRNA spacer sequence, or both of the PBS sequences in Table 3, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotides of the PBS sequence, or wherein the PBS sequence has a sequence comprising a PBS sequence corresponding to the RT template sequence, the gRNA spacer sequence, or both of the PBS sequences in Table 6A or 6B.

[0112] 58. The template RNA as described in any one of Examples 1-57, wherein the gRNA scaffold comprises a sequence of the gRNA scaffold in Table 6A or 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0113] 59. The template RNA as described in any one of Examples 1-57, wherein the gRNA scaffold comprises a sequence corresponding to the RT template sequence, the gRNA spacer sequence, or both of the gRNA scaffolds in Table 6A or 12, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0114] 60. A template RNA comprising: (iii) a heterologous object sequence containing a mutation region to introduce a mutation into a second part of the human SERPINA1 gene, wherein the heterologous object sequence contains a core nucleotide of the RT template sequence in Table 3, or a sequence having 1, 2, or 3 substitutions relative to it, and optionally contains one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the RT template sequence, and (iv) a PBS sequence containing at least 5, 6, 7, or 8 bases having 100% homology with a third part of the human SERPINA1 gene.

[0115] 61. The template RNA as described in any of the preceding embodiments, wherein the PBS sequence has a sequence comprising a core nucleotide from the PBS sequence in the same row as the RT template sequence in Table 3, or a sequence having one, two, or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotides of the PBS sequence.

[0116] 62. The template RNA as described in any one of Examples 1-60, wherein the PBS sequence has a sequence comprising a core nucleotide of the PBS sequence in Table 3 corresponding to the RT template sequence, or a sequence having one, two or three substitutions thereto, and optionally comprises one or more consecutive nucleotides starting from the 5' end of the flanking nucleotides of the PBS sequence.

[0117] 63. The template RNA as described in any of the preceding embodiments, wherein the mutation introduced by the system is a K342E mutation in the SERPINA1 gene (e.g., to correct a pathogenic E342K mutation).

[0118] 64. The template RNA as described in any of the preceding embodiments, wherein the pre-edit sequence length comprises about 1 nucleotide to about 35 nucleotides (e.g., about 1-5, 5-10, 10-15, 15-20, 20-25, 25-30 or 30-35 nucleotides).

[0119] 65. The template RNA as described in any of the foregoing embodiments, wherein the mutant region comprises a single nucleotide.

[0120] 66. The template RNA as described in any one of Examples 1-64, wherein the length of the mutant region is at least two nucleotides.

[0121] 67. The template RNA as described in any of the preceding embodiments, wherein the length of the mutant region is up to 32 nucleotides (e.g., up to 5, 10, 15, 20, 25, 30, or 32), and the mutant region contains one, two, or three sequence differences relative to the second part of the human SERPINA1 gene.

[0122] 68. The template RNA as described in any of the preceding embodiments, wherein the mutated region contains two sequence differences relative to the second part of the human SERPINA1 gene.

[0123] 69. The template RNA as described in any of the preceding embodiments, wherein the mutation region comprises a first region (e.g., a first nucleotide) designed to correct pathogenic mutations in the SERPINA1 gene and a second region (e.g., a second nucleotide) designed to inactivate the PAM sequence (e.g., the “PAM-kill” mutation as described in Table 5).

[0124] 70. The template RNA as described in any of the preceding embodiments, wherein the mutated region has less than 80%, 70%, 60%, 50%, 40%, or 30% identity with the corresponding portion of the human SERPINA1 gene.

[0125] 71. The template RNA as described in any of the foregoing embodiments, wherein the template RNA contains one or more silent mutations (e.g., silent substitutions), for example, as illustrated in Table 7B.

[0126] 72. The template RNA as described in any of the preceding embodiments, wherein the mutation region comprises a first region designed to correct pathogenic mutations in the SERPINA1 gene and a second region designed to introduce silencing substitutions.

[0127] 73. The template RNA as described in any of the foregoing embodiments, comprising one or more chemically modified nucleotides.

[0128] 74. A gene modification system comprising:

[0129] The template RNA as described in any of the foregoing embodiments, and

[0130] Genetically modified polypeptides or nucleic acids (e.g., RNA) encoding such genetically modified polypeptides.

[0131] 75. The gene modification system as described in Example 74, wherein the gene-modifying polypeptide comprises:

[0132] Reverse transcriptase (RT) domains (e.g., RT domains derived from retroviruses, or polypeptide domains having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with them); and

[0133] A Cas domain (e.g., a Cas9 domain) that binds to the target DNA molecule and is heterologous to the RT domain; and

[0134] Optionally, a connector is arranged between the RT structural domain and the Cas structural domain.

[0135] 76. The gene modification system as described in Example 75, wherein the RT domain comprises:

[0136] (a) The RT structure domain in Table 6; or

[0137] (b) From the following RT domains: murine leukemia virus (MMLV), porcine endogenous retrovirus (PERV), avian reticuloendotheliosis virus (AVIRE), feline leukemia virus (FLV), simmon foam virus (SFV) (e.g., SFV3L), bovine leukemia virus (BLV), Mason-Fischer monkey virus (MPMV), human foam virus (HFV), or bovine foam / syncytial virus (BFV / BSV).

[0138] 77. The gene modification system as described in Example 75 or 76, wherein the Cas domain comprises the Cas domain of Table X1, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with it.

[0139] 78. The gene modification system as described in any one of Examples 75-77, wherein the Cas9 nickase domain comprises the St1Cas9 nickase domain.

[0140] 79. The gene modification system as described in any one of Examples 75-78, wherein the gene-modified polypeptide comprises a first NLS, the St1Cas9 nickase domain, a linker, an RT domain, and a second NLS in the direction from the N-terminus to the C-terminus.

[0141] 80. The gene modification system as described in Example 79, wherein one or both of the following: the first NLS contains the sequence of SEQ ID NO: 11,095, and the second NLS contains the sequence of SEQ ID NO: 11,099.

[0142] 81. The gene modification system as described in any one of Examples 75-80, wherein the adapter comprises the sequence according to SEQ ID NO: 5006.

[0143] 82. The gene modification system as described in any one of Examples 75-81, wherein the spacer comprises a spacer of Table 6A, or a sequence having 1, 2 or 3 substitutions thereto, and the Cas domain comprises a Cas domain of the same row in Table 6A or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with it.

[0144] 83. The gene modification system as described in any one of Examples 75-82, wherein the spacer comprises the spacers of Table 6A, and the Cas domain comprises the Cas domain of the same row in Table 6A.

[0145] 84. The gene modification system as described in any one of Examples 75-83, wherein the spacer comprises a spacer of Table 6B, or a sequence having 1, 2 or 3 substitutions thereto, and the Cas domain comprises a Cas domain of the same row in Table 6B, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with it.

[0146] 85. The gene modification system as described in any one of Examples 75-84, wherein the spacer comprises the spacer of Table 6B, and the Cas domain comprises the Cas domain of the same row in Table 6B.

[0147] 86. The gene modification system as described in any one of Examples 75-85, wherein the Cas domain comprises the Cas domains of Table 7 or Table 8.

[0148] 87. The gene modification system as described in any one of Examples 75-86, wherein the Cas domain:

[0149] (a) is a Cas9 structural domain;

[0150] (b) is a SpCas9, BlatCas9, Nme2Cas9, PnpCas9, SauCas9, SauCas9-KKH, SauriCas9, SauriCas9-KKH, ScaCas9-Sc++, SpyCas9, SpyCas9-NG, SpyCas9-SpRY, or St1Cas9 domain; and / or

[0151] (c) is a Cas9 domain containing the N670A mutation, N611A mutation, N605A mutation, N580A mutation, N588A mutation, N872A mutation, N863 mutation, N622A mutation or H840A mutation.

[0152] 88. The gene modification system as described in Example 87, wherein the Cas9 domain binds to the PAM sequence listed in Table 7 or Table 12.

[0153] 89. The gene modification system as described in Example 88, wherein the second part of the human SERPINA1 gene overlaps with the PAM recognized by the Cas domain, for example, wherein the second part of the human SERPINA1 gene is within the PAM or wherein the PAM is within the second part of the human SERPINA1 gene.

[0154] 90. The gene modification system as described in any one of Examples 75-89, wherein the gRNA spacer is a gRNA spacer according to Table 1, and the Cas domain comprises a Cas domain listed in the same row of Table 1, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0155] 91. The gene modification system as described in any one of Examples 75-90, wherein the template RNA comprises a sequence of the template RNA sequence in Table 6A or 6B or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0156] 92. The gene modification system as described in any one of Examples 75-91, wherein:

[0157] (a) The template RNA contains the sequence of the template RNA in Table 3;

[0158] (b) The Cas structure field contains the Cas structure fields of Table 7 or Table 8;

[0159] (c) The connector contains the connector sequences in Table 10 (e.g., connector sequences of any one of SEQ ID NO: 5217, 5106, 5190, and 5218); and

[0160] (d) The genetically modified polypeptide contains one or two NLS sequences from Table 11 (e.g., the NLS sequence of any one of SEQ ID NO: 5245, 5290, 5323, 5330, 5349, 5350, 5351 and 4001).

[0161] 93. The gene modification system as described in any one of Examples 75-92, wherein the gene modification system produces a first nick in the first strand of the human SERPINA1 gene.

[0162] 94. The gene modification system as described in Example 93, further comprising a second-strand targeted gRNA spacer that directs a second nick to the second strand of the human SERPINA1 gene.

[0163] 95. The gene modification system as described in Example 94, wherein the second-strand targeting gRNA comprises a sequence containing a core nucleotide from the left or right gRNA spacer sequence from Table 2, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the left or right gRNA spacer sequence.

[0164] 96. The gene modification system as described in Example 94, wherein the second-strand targeted gRNA comprises a sequence containing the core nucleotide of the left or right gRNA spacer sequence in Table 2 corresponding to (i), and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the left or right gRNA spacer sequence.

[0165] 97. The gene modification system as described in Example 94, wherein the second-strand targeted gRNA comprises a sequence containing the core nucleotides of the second nick gRNA sequence in Table 4, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the second nick gRNA sequence.

[0166] 98. The gene modification system as described in Example 94, wherein the second-strand targeting gRNA comprises a sequence of core nucleotides of the second nick gRNA sequence in Table 4 corresponding to the gRNA spacer sequence of (i), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, and optionally comprises one or more consecutive nucleotides starting from the 3' end of the flanking nucleotides of the second nick gRNA sequence.

[0167] 99. The gene modification system as described in any of the preceding embodiments, wherein the second-strand targeting gRNA and the template RNA of the gene modification system have a “PAM-in orientation”, for example, as illustrated in Table 4.

[0168] 100. The gene modification system as described in any of the foregoing embodiments, wherein the second-strand targeted gRNA targets a sequence that overlaps with the target mutation of the template RNA.

[0169] 101. The gene modification system as described in Example 100, wherein the second-strand targeting gRNA comprises:

[0170] (i) Sequences complementary to SERPINA1 mutations (e.g., spacer subsequences);

[0171] (ii) A sequence complementary to the wild-type sequence at the target locus (e.g., a spacer sequence).

[0172] (iii) A sequence (e.g., a spacer sequence) that is complementary to a SNP that is close to the target locus, such as a SNP contained in the genomic DNA of a subject (e.g., a patient).

[0173] (iv) A sequence that is complementary to one or more silent substitutions near the target locus or contains one or more silent substitutions (e.g., a spacer subsequence).

[0174] 102. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the gRNA spacer comprises about 1, 2, 3 or more flanking nucleotides of the gRNA spacer.

[0175] 103. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the heterologous object sequence comprises about 2, 3, 4, 5, 10, 20, 30, 40 or more flanking nucleotides of the RT template sequence.

[0176] 104. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the heterologous object sequence comprises about 8-30, 9-25, 10-20, 11-16, or 12-15 (e.g., about 11-16) nucleotides.

[0177] 105. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the mutation region contains a sequence difference of 1, 2 or 3 nucleotide positions relative to the corresponding portion of the human SERPINA1 gene.

[0178] 106. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the mutation region contains a sequence difference of at least two nucleotide positions relative to the corresponding portion of the human SERPINA1 gene.

[0179] 107. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the edited homologous region and / or the unedited homologous region contains 100% identity with the SERPINA1 gene.

[0180] 108. The template RNA or gene modification system as described in any of the foregoing embodiments, wherein the PBS sequence further comprises about 1, 2, 3, 4, 5, 6, 7 or more flanking nucleotides.

[0181] 109. The template RNA or gene modification system as described in any of the preceding embodiments, wherein the PBS sequence comprises about 5-20, 8-16, 8-14, 8-13, 9-13, 9-12 or 10-12 (e.g., about 9-12) nucleotides.

[0182] 110. The template RNA or gene modification system as described in any of the preceding embodiments, wherein the PBS sequence binds to 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 nucleotides at the SERPINA1 gene cleavage site.

[0183] 111. The gene modification system as described in any of the foregoing embodiments, wherein the domains of the gene-modified polypeptide are linked by peptide linkers.

[0184] 112. The gene modification system as described in Example 111, wherein the adapter comprises the adapter sequences of Table 10 (e.g., adapter sequences of any one of SEQ ID NO: 5217, 5106, 5190 and 5218).

[0185] 113. The gene modification system as described in any of the foregoing embodiments, wherein the gene-modified polypeptide further comprises one or more nuclear localization sequences (NLS).

[0186] 114. The gene modification system as described in Example 113, wherein the gene modification polypeptide comprises a first NLS and a second NLS.

[0187] 115. The gene modification system as described in Example 113 or 114, wherein the NLS comprises the NLS sequence of Table 11 (e.g., the sequence of any one of SEQ ID NO: 5245, 5290, 5323, 5330, 5349, 5350, 5351 and 4001).

[0188] 116. A DNA that encodes a template RNA as described in any of the foregoing embodiments.

[0189] 117. A pharmaceutical composition comprising a gene-modifying system as described in any one of Examples 74-115, or one or more nucleic acids encoding the gene-modifying system, and a pharmaceutically acceptable excipient or carrier.

[0190] 118. The pharmaceutical composition as described in Example 117, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

[0191] 119. The pharmaceutical composition as described in Example 118, wherein the viral vector is adeno-associated virus.

[0192] 120. A host cell (e.g., a mammalian cell, such as a human cell) comprising a template RNA or a gene modification system as described in any of the foregoing embodiments.

[0193] 121. A method for preparing template RNA as described in any one of Examples 1-110, the method comprising synthesizing the template RNA by: in vitro transcription (e.g., solid-state synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow the generation of the template RNA.

[0194] 122. A method for modifying a target site in the human SERPINA1 gene in a cell, the method comprising contacting the cell with a gene modification system as described in any one of Examples 74-115 or DNA encoding the gene modification system, thereby modifying the target site in the human SERPINA1 gene in the cell.

[0195] 123. A method for modifying a target site in the human SERPINA1 gene in a cell, the method comprising contacting the cell with: (i) a template RNA as described in any one of Examples 1-73, or DNA encoding thereas; and (ii) a gene-modifying polypeptide or a nucleic acid encoding the gene-modifying polypeptide, thereby modifying the target site in the human SERPINA1 gene in the cell.

[0196] 124. A method for treating a subject suffering from a disease or condition associated with a mutation in the human SERPINA1 gene, the method comprising administering to the subject a gene-modifying system as described in any one of Examples 74-115 or DNA encoding the gene-modifying system, thereby treating the subject suffering from a disease or condition associated with a mutation in the human SERPINA1 gene.

[0197] 125. A method for treating a subject suffering from a disease or condition associated with a mutation in the human SERPINA1 gene, the method comprising administering to the subject a template RNA or DNA encoding such as any one of Examples 1-73; and (ii) a genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, thereby treating the subject suffering from a disease or condition associated with a mutation in the human SERPINA1 gene.

[0198] 126. The method as described in Examples 124 or 125, wherein the disease or condition is α-1 antitrypsin deficiency (AATD).

[0199] 127. The method as described in any one of Examples 124-126, wherein the subject has an E342K mutation (i.e., a PiZ mutation).

[0200] 128. A method for treating a subject suffering from AATD, the method comprising administering to the subject a gene-modifying system as described in any one of Examples 74-115 or DNA encoding the gene-modifying system, thereby treating the subject suffering from AATD.

[0201] 129. A method for treating a subject with AATD, the method comprising administering to the subject (i) a template RNA or DNA encoding the template RNA as described in any one of Examples 74-115, and (ii) a genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, thereby treating the subject with AATD.

[0202] 130. The gene modification system or method as described in any of the foregoing embodiments, wherein introducing the system into target cells results in the correction of pathogenic mutations in the SERPINA1 gene.

[0203] 131. The gene modification system or method as described in any of the foregoing embodiments, wherein the pathogenic mutation is an E342K mutation, and wherein the correction comprises an amino acid substitution of K342E.

[0204] 132. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.

[0205] 133. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target cells.

[0206] 134. The gene modification system or method as described in any of the foregoing embodiments, wherein the gene modification system comprises a second-strand targeting gRNA, and wherein the correction of mutations in the target cell population is increased relative to a target cell population treated with a gene modification system comprising template RNA but without a second-strand targeting gRNA.

[0207] 135. The gene modification system or method as described in any of the foregoing embodiments, wherein the template RNA contains one or more silent substitutions (e.g., as illustrated in Table 7B), and wherein the correction of the mutation in the target cell population is increased relative to a target cell population treated with a gene modification system containing template RNA that does not contain one or more silent substitutions.

[0208] 136. The method as described in any of the foregoing embodiments, wherein the cell is a mammalian cell, such as a human cell.

[0209] 137. The method as described in any of the foregoing embodiments, wherein the subject is a human.

[0210] 138. The method as described in any of the foregoing embodiments, wherein the contact occurs in vitro, for example, wherein the DNA of the cell or subject is modified in vitro.

[0211] 139. The method as described in any of the foregoing embodiments, wherein the contact occurs in vivo, for example, wherein the DNA of the cell or the subject is modified in vivo.

[0212] 140. The method as described in any of the foregoing embodiments, wherein contacting the cell or the subject with the system comprises contacting the cell or a cell in the subject’s body with a nucleic acid (e.g., DNA or RNA) encoding the genetically modified polypeptide under conditions that allow the generation of the genetically modified polypeptide.

[0213] 141. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the gRNA scaffold is a variant gRNA scaffold comprising a sequence according to Table 23, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0214] 142. The template RNA, system, pharmaceutical composition, cell, or method as described in Example 141, wherein the variant gRNA scaffold comprises the sequence according to Table 23.

[0215] 143. The template RNA, system, pharmaceutical composition, cell, or method as described in Examples 141 or 142, wherein the variant gRNA scaffold comprises a sequence according to GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCA (SEQ ID NO: 26000).

[0216] 144. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the gRNA spacer comprises a sequence according to Table 22, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0217] 145. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the gRNA spacer comprises a sequence according to Table 22.

[0218] 146. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the gRNA spacer comprises a sequence according to AAGGCUGUGCUGACCAUCGA (SEQ ID NO: 26001).

[0219] 147. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the preceding embodiments, wherein the heterologous object sequence comprises a sequence according to Table 24, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0220] 148. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the heterologous object sequence comprises the sequence according to Table 24.

[0221] 149. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the preceding embodiments, wherein the PBS sequence comprises a sequence according to Table 25, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0222] 150. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the PBS sequence comprises the sequence according to Table 25.

[0223] 151. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to Table 20, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity with it.

[0224] 152. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises the sequence according to Table 20.

[0225] 153. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to Table 21, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity with it.

[0226] 154. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises the sequence according to Table 21.

[0227] 155. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to any one of Table 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15, or E16, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity with it.

[0228] 156. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to any one of Table 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15, or E16.

[0229] 157. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises one or more chemical modifications.

[0230] 158. The template RNA, system, pharmaceutical composition, cell, or method as described in Example 153, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises one or more phosphate thioester bonds.

[0231] 159. The template RNA, system, pharmaceutical composition, cell, or method as described in Examples 157 or 158, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises one or more 2'-O-methylnucleotides.

[0232] 160. The template RNA, system, pharmaceutical composition, cell, or method as described in any one of Examples 153-157, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to column 3 of Table 20, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0233] 161. The template RNA, system, pharmaceutical composition, cell, or method as described in any one of Examples 153-157, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to column 3 of Table 21, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

[0234] 162. A gene modification system comprising:

[0235] The template RNA as described in any one of Examples 141-161; and

[0236] A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising:

[0237] (1) St1Cas9 structural domain;

[0238] (2) Connector; and

[0239] (3) Reverse transcriptase (RT) domain.

[0240] 163. The system as described in Example 162, wherein the St1Cas9 domain is a nicking enzyme.

[0241] 164. The system as described in Example 162 or 163, wherein the St1Cas9 domain comprises a sequence according to SEQ ID NO:23818, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0242] 165. The system as described in any one of Examples 162-164, wherein the connector comprises a sequence according to SEQ ID NO:5006, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0243] 166. The system as described in any one of Examples 162-165, wherein the connector comprises a sequence according to SEQ ID NO:5217, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0244] 167. The system as described in any one of Examples 162-166, wherein the connector comprises a sequence of Table 10, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0245] 168. The system as described in any one of Examples 162-167, wherein the RT structural domain comprises a sequence according to SEQ ID NO: 26006, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0246] 169. The system as described in any one of Examples 162-167, wherein the RT domain comprises a sequence according to any one of SEQ ID NO: 8,001-8,003, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0247] 170. The system as described in any one of Examples 162-167, wherein the RT structural domain comprises the sequence of Table 6, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0248] 171. The system as described in any one of Examples 162-167, wherein the genetically modified polypeptide comprises the sequence according to SEQ ID NO: 26002, or a sequence having at least 80%, 85%, 90%, 95%, 98% or 99% identity with it.

[0249] 172. The system as described in any one of Examples 162-171, further comprising a second nick gRNA (ngRNA), wherein optionally the second nick gRNA directs a second nick to the second strand of the human SERPINA1 gene.

[0250] 173. The system as described in Example 172, wherein the second nick gRNA comprises a sequence according to Table 26, or a sequence having at least 80%, 85%, 90%, 95% or 98% identity with it.

[0251] 174. The system as described in any one of Examples 162-173, wherein the nucleic acid encoding the gene-modified polypeptide comprises RNA, for example, mRNA.

[0252] 175. The template RNA or system as described in any of the foregoing embodiments, wherein the nucleic acid molecule is formulated in lipid nanoparticles (LNPs).

[0253] 176. The system as described in any of the preceding embodiments, wherein the template RNA, the nucleic acid molecule encoding the gene-modified polypeptide, and / or the gRNA are formulated in an LNP.

[0254] 177. A pharmaceutical composition comprising a system as described in any one of Examples 160-174, or one or more nucleic acids encoding the system, and a pharmaceutically acceptable excipient or carrier.

[0255] 178. The pharmaceutical composition as described in Example 177, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

[0256] 179. The pharmaceutical composition as described in Example 178, wherein the viral vector is adeno-associated virus.

[0257] 180. A host cell (e.g., a mammalian cell, such as a human cell) comprising a gene modification system or template RNA as described in any of the foregoing embodiments.

[0258] 181. A method for preparing template RNA as described in any one of Examples 141-161, the method comprising synthesizing the template RNA by means of: in vitro (e.g., by in vitro transcription or solid-state synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow the generation of the template RNA.

[0259] 182. A method for modifying a target site in a cell (e.g., a target site in the human SERPINA1 gene), the method comprising contacting the cell with the gene modification system, DNA encoding the gene modification system, or a pharmaceutical composition as described in any of the foregoing embodiments, thereby modifying the target site.

[0260] 183. A method for treating a subject suffering from a disease or condition associated with a gene (e.g., human SERPINA1 gene) mutation, the method comprising administering to the subject the gene-modifying system, DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any of the preceding embodiments, thereby treating the subject suffering from the disease or condition.

[0261] 184. The method as described in Example 183, wherein the disease or condition is α-1 antitrypsin deficiency (AATD).

[0262] 185. The method as described in Example 183 or 184, wherein the subject has an E342K mutation.

[0263] 186. A method for treating a subject suffering from AATD, the method comprising administering to the subject the gene-modifying system, or DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any of the preceding embodiments, thereby treating the subject suffering from AATD.

[0264] 187. The gene modification system or method as described in any of the foregoing embodiments, wherein introducing the system into target cells results in the correction of pathogenic mutations in the gene, such as the SERPINA1 gene.

[0265] 188. The gene modification system or method as described in any of the foregoing embodiments, wherein the pathogenic mutation is an E342K mutation, and wherein the correction comprises an amino acid substitution of K342E.

[0266] 189. The gene modification system or method as described in any of the foregoing embodiments, wherein introducing the system into a target cell results in a mutation that leads to the restoration of function of the gene, such as the SERPINA1 gene.

[0267] 190. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.

[0268] 191. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target cells.

[0269] 192. The gene modification system or method as described in any of the foregoing embodiments, wherein the gene modification system comprises a second nick gRNA, and wherein the correction of the mutation is increased in the target cell population relative to a target cell population treated with a gene modification system comprising template RNA but without a second nick gRNA.

[0270] 193. The gene modification system or method as described in any of the foregoing embodiments, wherein the template RNA contains one or more silent substitutions, and wherein the correction of the mutation in the target cell population is increased relative to a target cell population treated with a gene modification system containing template RNA that does not contain one or more silent substitutions.

[0271] 194. The method as described in any of the foregoing embodiments, wherein the cell is a mammalian cell, such as a human cell.

[0272] 195. The method as described in any of the foregoing embodiments, wherein the subject is a human.

[0273] 196. The method as described in any of the foregoing embodiments, wherein the contact occurs in vitro, for example, wherein the DNA of the cell or subject is modified in vitro.

[0274] 197. The method as described in any of the foregoing embodiments, wherein the contact occurs in vivo, for example, wherein the DNA of the cell or the subject is modified in vivo.

[0275] 198. The method as described in any of the foregoing embodiments, wherein contacting the cell or the subject with the system comprises contacting the cell or cells in the subject’s body with nucleic acids (e.g., DNA or RNA) encoding the genetically modified polypeptide under conditions that allow the generation of the genetically modified polypeptide.

[0276] 199. The method as described in any of the foregoing embodiments, wherein the method comprises applying the gene-modification system or the DNA encoding the gene-modification system twice.

[0277] 200. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to any one of Tables 20, 21, 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15, or E16, or a sequence having at least 80%, 85%, 90%, 95%, or 98% identity with it.

[0278] 201. The template RNA, system, pharmaceutical composition, cell, or method as described in any of the foregoing embodiments, wherein the template RNA, system, pharmaceutical composition, cell, or method comprises a sequence according to any one of Tables 20, 21, 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15, or E16.

[0279] 202. A gRNA comprising (e.g., from 5' to 3'):

[0280] (1) gRNA spacers; and

[0281] (2) Variant St1Cas9 scaffold, which has partial or complete absence of stem ring 2.

[0282] 203. The gRNA as described in Example 202, wherein the length of the deletion is between 1 and 32 (e.g., 2-29, 2-20, 2-10, or 10-20) nucleotides.

[0283] 204. The gRNA as described in Example 202, wherein the deletion is a deletion of the entire stem-loop 2.

[0284] 205. The gRNA as described in Example 202, wherein the deletion is a deletion at positions 55 to 84.

[0285] 206. The gRNA as described in Example 202, wherein the St1Cas9 scaffold includes a portion of the second single-stranded region (e.g., 1, 2, 3, or 4 nucleotides at the 3' end of the single-stranded region).

[0286] 207. The gRNA as described in any one of Examples 202-206, wherein the variant St1Cas9 scaffold has one or both of the following: an elongated RAR upper stem or substitutions that produce GC base pairs in the RAR upper stem.

[0287] 208. The gRNA as described in any one of Examples 202-207, wherein the variant St1Cas9 scaffold has a mutation in the four-ring.

[0288] 209. A gRNA comprising (e.g., from 5' to 3'):

[0289] (1) gRNA spacers; and

[0290] (2) A variant St1Cas9 scaffold having one or both of the following: an elongated RAR upper stem or substitution of GC base pairs in the RAR upper stem.

[0291] 210. The gRNA as described in any one of Examples 202-209, wherein the upper stem of the RAR is extended by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7 or 8 base pairs) relative to the wild-type sequence of SEQ ID NO:25999.

[0292] 211. The gRNA as described in Example 210, wherein at least 50%, 60%, 70%, 80% or 90% of the new base pairs relative to SEQ ID NO: 25999 are GC base pairs.

[0293] 212. A gRNA comprising (e.g., from 5' to 3'):

[0294] (1) gRNA spacers; and

[0295] (2) Variant St1Cas9 scaffold with mutations in the four rings.

[0296] 213. The gRNA as described in any one of Examples 202-212, wherein one or more nucleotides in the tetracycle are substituted.

[0297] 214. The gRNA as described in any one of Examples 202-213, wherein the four-loop comprises a sequence selected from the following: AACA, AAUA, ACCA, ACUA, AGUA, AGCA, AUCA, AUUA, CAAC, CUCG, CUUG, GAAA, GAGA, GCAA, GCGA, GGAA, GGAG, GGGA, GUAA, GUGA, UAAC, UACG, UCAC, UCCG, UGAA, UGAC, UGCG, UUAC, or UUCG.

[0298] 215. The gRNA as described in any one of Examples 202-214, wherein the tetracycle is extended to, for example, 5 nucleotides.

[0299] 216. The gRNA as described in Example 215, wherein the elongated tetracycle comprises a sequence selected from GAAGA or GACAA.

[0300] 217. The gRNA as described in any one of Examples 202-216, wherein the variant gRNA scaffold comprises a sequence according to Table 23, or a sequence having no more than one, two or three sequence changes (e.g., substitutions) relative to it.

[0301] 218. The gRNA as described in any one of Examples 202-217, wherein the length of the variant St1Cas9 scaffold is 50-60, 60-70, 70-80, or 80-84 nucleotides.

[0302] 219. A gRNA comprising (e.g., from 5' to 3'):

[0303] (1) gRNA spacers; and

[0304] (2) A variant gRNA scaffold containing a sequence according to Table 23, or a sequence having no more than one, two or three sequence changes (e.g., substitutions) relative to it.

[0305] 220. The gRNA as described in any one of Examples 202-217, wherein the variant gRNA scaffold comprises the sequence according to Table 23.

[0306] 221. The gRNA as described in Examples 202-220, wherein the variant gRNA scaffold comprises the sequence according to GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAU GCCGAAAUCA (SEQ ID NO: 26000). Attached Figure Description

[0307] This patent or application document contains at least one color drawing. Upon request and after payment of the necessary fees, a copy of this patent or application publication with color drawings will be provided by the Patent Office.

[0308] Figure 1 The gene modification system described herein is depicted. The left-hand diagram shows the gene-modifying polypeptide, which contains a Cas nicking enzyme domain (e.g., spCas9 N863A) and a reverse transcriptase domain (RT domain) linked by a linker. The right-hand diagram shows the template RNA, which contains a gRNA spacer, a gRNA scaffold, a heterologous target sequence, and a primer binding site sequence (PBS sequence) from 5' to 3'. The heterologous target sequence may contain a mutant region that contains one or more sequence differences relative to the target site. The heterologous target sequence may also contain pre-edited homologous regions and post-edited homologous regions flanking the mutant region. Without being bound by theory, it is assumed that the gRNA spacer of the template RNA binds to the second strand of the target site in the genome, and that the gRNA scaffold of the template RNA binds to the gene-modifying polypeptide, for example, to locate the gene-modifying polypeptide to the target site in the genome. It is assumed that the Cas domain of the gene-modifying polypeptide creates a nick at the target site (e.g., the first strand of the target site), for example, allowing the PBS sequence to bind to a sequence adjacent to the site to be modified on the first strand of the target site. This theory posits that the RT domain of a gene-modified polypeptide uses the first strand of a target site that binds to a complementary sequence (a PBS sequence containing the template RNA) as a primer, and a heterologous target sequence of the template RNA as a template, such as a sequence complementary to the heterologous target sequence. Not wanting to be bound by theory, it is proposed that reverse transcription can then proceed through the pre-edited homologous region, then through the mutated region, and then through the post-edited homologous region, thereby producing a DNA strand containing the mutation specified by the heterologous target sequence.

[0309] Figure 2 The hypothetical secondary structure of the wild-type St1Cas9 gRNA scaffold is shown, and the description of the variants described in this paper is superimposed.

[0310] Figure 3A A graph showing the rewriting performance of St1Cas9-based gene modification systems containing exemplary template RNAs contained within various scaffolds truncated in the SL2 region is presented.

[0311] Figure 3B The diagram shows a rewritten St1Cas9-based gene modification system containing exemplary template RNA with scaffolds further engineered in the TL and RAR regions using various modified tetracycles.

[0312] Figure 3CA diagram of St1Cas9-based gene modification system rewriting is shown, which contains exemplary template RNAs containing spacers of varying lengths.

[0313] Figure 4A A graph showing the rewriting efficiency of gene modification systems containing different St1Cas9-compatible template RNAs with modified scaffold sequences is presented.

[0314] Figure 4B It shows Figure 4A A graph showing the insertion / deletion percentage levels of the same gene modification system evaluated in the study.

[0315] Figure 4C A graph showing the rewriting efficiency of gene modification systems containing different St1Cas9-compatible template RNAs with modified scaffold sequences is presented.

[0316] Figure 5 A graph showing the rewriting efficiency of a gene modification system containing a St1Cas9-based gene-modifying peptide is presented.

[0317] Figure 6 The graph shows the rewriting efficiency (left) and insertion / deletion % level (right) of gene modification systems containing St1Cas9-based gene-modifying peptides (with and without ngRNA).

[0318] Figure 7 This is a diagram showing the rewriting activity of St1Cas9-based gene modification systems in primary hepatocytes. These gene modification systems contain exemplary template RNAs containing a dSL2 variant gRNA scaffold, PBS sequences of varying lengths, and heterologous object sequences.

[0319] Figure 8A This is a series of graphs showing the percentage of rewrites achieved in primary hepatocytes (left inset), HEK293T cells treated with a high dose (middle inset), or HEK293T cells treated with a low dose (right inset) using gene modification systems containing different St1Cas9 compatible template RNAs, which contain variant scaffolds with various exemplary variant tetracyclic structures.

[0320] Figure 8B The hypothetical secondary structure of the dSL2 truncated St1Cas9 gRNA scaffold is shown, and the description of the variants described herein is superimposed.

[0321] Figure 9AThis is a diagram showing the rewriting activity of exemplary St1Cas9-based gene modification systems, which contain variant template RNAs (with various 2-O'-methyl chemical modifications in the gRNA scaffold region) having the nucleotide sequence of the exemplary template RNA RNACS9201 (containing a dSL2 variant gRNA scaffold).

[0322] Figure 9B This is a graph showing the effect of rewriting activity on three nucleotides of a scaffold modified once with 2'-O-methyl chemical modification.

[0323] Figures 9C-9G Showing Figure 9A The pattern of 2'-O-methyl chemically modified nucleotides in the dSL2 St1Cas9 scaffold sequence. Gray bases represent unmodified nucleotide positions. Bold black bases represent chemically modified nucleotide positions. In this example, the chemically modified nucleotide is a 2'-O-methyl nucleotide.

[0324] Figure 10 This is a graph showing the rewriting efficiency of a gene modification system in primary hepatocytes, which contains different St1Cas9-compatible template RNAs with different patterns of 2'-O-methyl chemical modifications in the dSL2 St1Cas9 scaffold.

[0325] Figure 11 A- Figure 11 C is a series of diagrams showing the rewriting activity (11A), insertion / deletion activity (11B), and hA1AT (11C) of gene modification systems in the liver, which contain different St1Cas9 compatible template RNAs formulated in LNPs and containing dSL2 variant gRNA scaffolds.

[0326] Figures 12A-12C This is a series of graphs showing the percentage of perfect rewrites (12A) and percentage of insertions / deletions (12B) in the liver and serum A1AT levels (12C) using gene-modification systems containing different St1Cas9-compatible template RNAs formulated in LNPs with different patterns of 2'-O-methyl chemical modifications in the dSL2 St1Cas9 scaffold.

[0327] Figure 13A This is a diagram showing the position of the reference dSL2 St1Cas9 stent sequence.

[0328] Figure 13B This is a diagram showing the location of the reference wild-type St1Cas9 scaffold sequence.

[0329] Figure 13CThis is a diagram showing the putative structure of RNACS13597 with a RAR+4_UUCG mutation relative to dSL2.

[0330] Figure 13D This is a diagram showing the putative structure of RNACS17210 with a RAR+4_AGCA mutation relative to dSL2.

[0331] Figures 14A-14B The figure shows the percentage of editing in the liver after one or two doses of gene-modifying peptides and template RNA, as evaluated by Amp-Seq. Figure 14A ) or insertion missing % ( Figure 14B (The image is shown.)

[0332] Figures 15A-15B The percentage of liver editing following administration of the gene-modified peptide and template RNA, as assessed by Amp-Seq, is shown. Figure 15A ) or insertion missing % ( Figure 15B (The image is shown.)

[0333] Figure 16 The rewriting achieved in primary hepatocytes after administration of a genetically modified peptide and template RNA is shown.

[0334] Figure 17 The hypothetical secondary structure of the dSL2 truncated St1Cas9 gRNA scaffold is shown, and the description of the variants described herein is superimposed.

[0335] Figure 18 The rewriting achieved in primary hepatocytes after administration of a genetically modified peptide and template RNA is shown. Detailed Implementation

[0336] definition

[0337] As used herein, the term "expression cassette" refers to a nucleic acid construct containing nucleic acid elements sufficient to express the nucleic acid molecules of the present invention.

[0338] As used in this article, "gRNA spacer" refers to a nucleic acid portion that is complementary to the target nucleic acid and can work with the gRNA scaffold to target the Cas protein to the target nucleic acid.

[0339] As used herein, a "gRNA scaffold" refers to a nucleic acid motif that can bind to the Cas protein and, together with a gRNA spacer, target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold comprises a crRNA sequence, a tetracyclic RNA sequence, and a tracrRNA sequence.

[0340] As used herein, a “variant gRNA scaffold” refers to a gRNA scaffold having a non-naturally occurring sequence. In some embodiments, the variant sequence contains one or more substitutions relative to the nearest naturally occurring sequence. In some embodiments, the variant sequence contains one or more insertions relative to the nearest naturally occurring sequence. In some embodiments, the variant sequence contains one or more deletions relative to the nearest naturally occurring sequence.

[0341] As used herein, the term "St1Cas9 scaffold" refers to a gRNA scaffold capable of binding to the St1Cas9 protein and targeting the St1Cas9 protein to a target nucleic acid together with a gRNA spacer. In some embodiments, the St1Cas9 scaffold comprises a crRNA sequence, a tetracyclic RNA sequence, and a tracerRNA sequence. Exemplary locations of the St1Cas9 scaffold in an exemplary template RNA are shown in... Figure 13A and Figure 13B middle.

[0342] In some embodiments, the St1Cas9 scaffold comprises a full-length wild-type sequence. In some embodiments, the St1Cas9 scaffold comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the sequence GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAU AAGGCUUCAUGCCGAAAUCAACACCCUGUCAUUUUAUGGCAGGGUGUUUU (SEQ ID NO: 25999). In some embodiments, the St1Cas9 scaffold comprises a sequence identical to SEQ ID NO: 25999. In some embodiments, the St1Cas9 scaffold is a truncated mutant. In some embodiments, the St1Cas9 scaffold comprises a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with the sequence GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGGCUUCAUGCCGAAAUCA (SEQ ID NO: 26000). In some embodiments, the St1Cas9 scaffold comprises the same sequence as SEQ ID NO: 26000. In some embodiments, the St1Cas9 scaffold comprises insertions, deletions, or substitutions relative to the reference sequence of SEQ ID NO: 25999 or 26000. In some embodiments, the St1Cas9 scaffold comprises chemically modified nucleotides.

[0343] As used herein, a “genetically modified polypeptide” refers to a polypeptide comprising a retroviral reverse transcriptase, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with a retroviral reverse transcriptase, capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., in mammalian host cells, such as genomic DNA molecules in host cells). In some embodiments, the genetically modified polypeptide is capable of integrating the sequence substantially independently of the host machinery. In some embodiments, the genetically modified polypeptide integrates the sequence into a random location in the genome, and in some embodiments, the genetically modified polypeptide integrates the sequence into a specific target site. In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively facilitate 1) binding to a template nucleic acid, 2) binding to a target DNA molecule, and 3) facilitating the integration of at least a portion of the template nucleic acid into the target DNA. Genetically modified polypeptides include naturally occurring polypeptides and engineered variants of the aforementioned polypeptides, for example, those variants having one or more amino acid substitutions relative to the naturally occurring sequence. Genetically modified peptides also include heterologous constructs, such as those in which one or more of the aforementioned domains are heterologous to each other, whether through heterologous fusion (or other conjugates) of domains that are otherwise wild-type, and through fusion of modified domains, for example, through substitution or fusion of heterologous subdomains or other substituted domains. Exemplary gene-modified peptides, systems comprising them, and methods of using them that can be used in the methods provided herein are described, for example, in PCT / US2021 / 020948, which relates to gene-modified peptides comprising retroviral reverse transcriptase domains and is incorporated herein by reference. In some embodiments, the gene-modified peptide integrates a sequence into a gene. In some embodiments, the gene-modified peptide integrates a sequence into a sequence outside a gene. As used herein, a “gene-modified system” refers to a system comprising a gene-modified peptide and a template nucleic acid.

[0344] As used herein, the term "domain" refers to the structure of a biomolecule that contributes to a specific function of the biomolecule. A domain may contain a continuous region (e.g., a continuous sequence) or distinct non-continuous regions (e.g., non-continuous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcription domains; examples of nucleic acid domains are regulatory domains, such as transcription factor-binding domains. In some embodiments, a domain (e.g., a Cas domain) may contain two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).

[0345] As used herein, the term "exogenous" when used in relation to a biomolecule (such as a nucleic acid sequence or polypeptide) means the artificial introduction of the biomolecule into a host genome, cell, or organism. For example, nucleic acids added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods are exogenous to the existing nucleic acid sequence, cell, tissue, or subject.

[0346] As used herein, the terms "first strand" and "second strand" to describe a single DNA strand of the target DNA distinguish the two DNA strands based on the strand on which the reverse transcriptase domain initiates polymerization, for example, based on the location where target-initiated synthesis begins. The first strand refers to the strand of the target DNA on which the reverse transcriptase domain initiates polymerization, for example, at the location where target-initiated synthesis begins. The second strand refers to the other strand of the target DNA. The names of the first and second strands do not otherwise describe the target site DNA strand; for example, in some embodiments, the first and second strands are cleaved by the polypeptide described herein, but the names "first" and "second" strand are not related to the order in which such cleavage occurs.

[0347] When used herein to describe a first element in reference to a second element, the term "heterogeneous" means that the first and second elements do not exist in nature in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence is (a) a polypeptide, nucleic acid molecule, or part of a polypeptide or nucleic acid molecule sequence that is not native to the cell expressing it, (b) a polypeptide or nucleic acid molecule, or part of a polypeptide or nucleic acid molecule, that has been altered or mutated relative to its native state, or (c) a polypeptide or nucleic acid molecule having altered expression compared to its native expression level under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from how the gene or nucleic acid molecule is normally expressed in nature. In another instance, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., the DNA-binding domain of the polypeptide or a nucleic acid encoding the DNA-binding domain of the polypeptide) may be arranged relative to other domains, or may be different sequences or may originate from different sources relative to other domains or portions of the polypeptide or its encoding nucleic acid. In some embodiments, the heterologous nucleic acid molecule may be present in the natural host cell genome, but may have altered expression levels or different sequences, or both. In other embodiments, the heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but may be introduced into the host cell through transformation (e.g., transfection, electroporation), wherein the added molecule may integrate into the host genome, or may exist transiently as extrachromosomal genetic material (e.g., mRNA) or semi-stable for more than one generation (e.g., cell-free viral vectors, plasmids, or other self-replicating vectors).

[0348] As used herein, “inserting” a sequence into a target site refers to a net addition of DNA sequence at the target site, such as the presence of new nucleotides in a heterologous object sequence that has no homologous position in the unedited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous object sequence with the target nucleic acid sequence will result in alignment gaps in the target nucleic acid sequence.

[0349] As used herein, a “deletion” produced by a heterologous object sequence at a target site refers to a net deletion of the DNA sequence at the target site, such as the presence of nucleotides at an unedited target site where there is no homologous position in the heterologous object sequence. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous object sequence with the target nucleic acid sequence will result in alignment vacancies in the molecule containing the PBS sequence and the heterologous object sequence.

[0350] As used herein, the term "inverted terminal repeat sequence" or "ITR" refers to AAV viral cis-elements, so named because of their symmetry. These elements promote efficient multiplication of the AAV genome. It is assumed that the minimum functional element of an ITR is a Rep binding site (RBS; 5'-GCGCGCTCGCTCGCTC-3', for AAV2; SEQ ID NO: 4601) and a terminal dissociation site (TRS; 5'-AGTTGG-3', for AAV2; SEQ ID NO: 4602) plus a variable palindromic sequence that allows hairpin formation. According to the invention, an ITR comprises at least these three elements (RBS, TRS, and the hairpin-forming sequence). Additionally, in this invention, the term "ITR" refers to the ITR of a known natural AAV serotype (e.g., the ITR of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), a chimeric ITR formed by the fusion of ITR elements derived from different serotypes, and their functional variants. "Functional variant" refers to a sequence that has at least 80%, 85%, 90%, preferably at least 95% sequence identity with a known ITR, allowing the sequence containing the ITR to multiply in the presence of the Rep protein.

[0351] As used herein, the term "mutation region" refers to a region in the template RNA that has one or more sequence differences relative to the corresponding sequence in the target nucleic acid. Sequence differences can include, for example, substitution, insertion, frameshift, or deletion.

[0352] When applied to nucleic acid sequences, the term "mutated" means that nucleotides in the nucleic acid sequence have been inserted, deleted, or altered compared to a reference (e.g., a natural) nucleic acid sequence. A single change (point mutation) can be made at a locus, or multiple nucleotides can be inserted, deleted, or altered at a single locus. Additionally, one or more changes can be made at any number of loci within the nucleic acid sequence. Nucleic acid sequences can be mutated using any method known in the art.

[0353] Nucleic acid molecules refer to both RNA and DNA molecules, including but not limited to complementary DNA (“cDNA”), genomic DNA (“gDNA”), and messenger RNA (“mRNA”), and also include synthetic nucleic acid molecules, such as those produced by chemical synthesis or recombination, such as RNA templates as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule can be sense or antisense. Unless otherwise indicated, and as an example of all sequences described herein in the general format “SEQ ID NO:”, “nucleic acid containing SEQ ID NO:1” means a nucleic acid having (i) the sequence of SEQ ID NO:1 or (ii) a sequence complementary to SEQ ID NO:1, at least a portion thereof. The choice between the two depends on the context of using SEQ ID NO:1. For example, if the nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target. As will be readily understood by those skilled in the art, the nucleic acid sequences disclosed herein may be chemically or biochemically modified or may contain non-natural or derived nucleotide bases. Such modifications include, for example, tagging, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, triphosphates, aminophosphates, carbamates, etc.), charged linkages (e.g., thiophosphates, dithiophosphates, etc.), side chain moieties (e.g., polypeptides), intercalating agents (e.g., acridine, psoralen, etc.), chelating agents, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are chemically modified bases (see, for example, Table 13), main chains (see, for example, Table 14), and modified caps (see, for example, Table 15). Also included are synthetic molecules that mimic the ability of polynucleotides to bind to a specified sequence via hydrogen bonds and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages replace phosphate linkages in the molecular main chain, such as peptide nucleic acids (PNAs). Other modifications may include, for example, analogs in which the ribose ring contains a bridging moiety or other structure (e.g., modifications found in "locked" nucleic acids (LNAs). In various embodiments, nucleic acids are operatively associated with additional genetic elements (e.g., one or more tissue-specific expression control sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences)) and other elements (e.g., inverted repeat sequences (e.g., inverted terminal repeat sequences, such as elements derived from or originating from viruses, such as AAV ITRs) and tandem repeat sequences, inverted repeat sequences / direct repeat sequences, homologous regions (segments with different degrees of homology to the target DNA), untranslated regions (UTRs) (5', 3', or 5' and 3' UTRs)) and various combinations thereof.The nucleic acid elements of the system provided by this invention can be provided in a variety of topologies, including single-stranded, double-stranded, circular, linear, linear with open ends, linear with closed ends, and specific versions of these, such as dog bone DNA (dbDNA) and closed-end DNA (ceDNA).

[0354] As used herein, a “gene expression unit” is a nucleic acid sequence containing at least one regulatory nucleic acid sequence operatively linked to at least one effector sequence. The first nucleic acid sequence is operatively linked to the second nucleic acid sequence when positioned to have a functional relationship with it. For example, if a promoter or enhancer affects the transcription or expression of a coding sequence, then the promoter or enhancer is operatively linked to that coding sequence. The operatively linked DNA sequences can be contiguous or non-contiguous. In cases where it is necessary to link two protein-coding regions, the operatively linked sequences can be within the same reading frame.

[0355] As used herein, the terms “host genome” or “host cell” refer to a cell and / or its genome in which proteins and / or genetic material have been introduced. It should be understood that such terms are intended not only to refer to a specific subject cell and / or genome, but also to the genomes of the offspring of such cells and / or the offspring of such cells. Because certain modifications may occur in offspring due to mutations or environmental influences, such offspring may actually differ from the parent cell, but are still included within the scope of the term “host cell” as used herein. A host genome or host cell can be an isolated cell or cell line grown in a culture, or genomic material isolated from such a cell or cell line, or it can be a host cell or host genome constituting a living tissue or organism. In some cases, the host cell can be an animal cell or plant cell, for example, as described herein. In some cases, the host cell can be a mammalian cell, human cell, avian cell, reptile cell, bovine cell, horse cell, pig cell, goat cell, sheep cell, chicken cell, or turkey cell. In some cases, the host cell can be a corn cell, soybean cell, wheat cell, or rice cell.

[0356] As used herein, “operable association” describes a functional relationship between two nucleic acid sequences (e.g., 1) a promoter and 2) a heterologous object sequence), and in such instances, it means that the orientation of the promoter and the heterologous object sequence (e.g., the target gene) allows the promoter to drive the expression of the heterologous object sequence under appropriate conditions. For example, the template nucleic acid carrying the promoter and the heterologous object sequence can be single-stranded, e.g., with a (+) or (-) orientation. An “operable association” between the promoter and the heterologous object sequence in this template means that, regardless of whether the template nucleic acid is transcribed in a specific state, it is accurately transcribed when it is in the appropriate state (e.g., in a (+) orientation, in the presence of the required catalysts and NTPs, etc.). Operable associations similarly apply to other nucleic acid pairs, including other tissue-specific expression control sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homologous regions and heterologous object sequences or sequences encoding retroviral RT domains.

[0357] As used herein, the term "position" in relation to the St1Cas9 scaffold refers to the nucleotide of the St1Cas9 scaffold that aligns with the corresponding nucleotide of the reference sequence of SEQ ID NO: 25999. The positions of the reference sequence are shown in... Figure 13B In this context, the alignment of nucleic acid or peptide sequences can be performed using sequence analysis tools such as the Basic Local Alignment Search Tool (BLAST) (e.g., NIH megablast with default parameters).

[0358] In some embodiments, the location of the St1Cas9 scaffold can be identified by providing the St1Cas9 scaffold (query sequence) and SEQ ID NO: 25999 (full-length wild-type sequence, see, for example, Figure 13B ) or SEQ ID NO: 26000 (truncated mutant, see, for example, Figure 13A The query sequence is aligned with the reference sequence and the position in the query sequence that corresponds to the position in the reference sequence is identified. For example, in the St1Cas9 scaffold composed of the sequence of SEQ ID NO: 25999, the position substituted is position 1, except that the G at the 5' end is replaced by a single nucleotide other than G.

[0359] As another example, in the St1Cas9 scaffold composed of the sequence SEQ ID NO: 25999, G is still at position 1 except that a new single nucleotide is inserted exactly at the 5' of the most 5'G.

[0360] As yet another example, in the St1Cas9 scaffold composed of the sequence SEQ ID NO: 25999, except for the n-nucleotide sequence inserted between the G at position 1 and the U at position 2, the inserted 3' nucleotides retain their original position numbers. For example, the U at position 2 remains position 2 instead of position n+2. Nucleotides inserted relative to the reference sequence do not need to be assigned position numbers.

[0361] As used herein, the term "primer binding site sequence" or "PBS sequence" refers to a portion of template RNA capable of binding to a region contained in a target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence containing at least 3, 4, 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. Not wishing to be bound by theory, in some embodiments, when the template RNA contains both a PBS sequence and a heterologous object sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence, thereby allowing the reverse transcriptase domain to use that region as a primer for reverse transcription and the heterologous object sequence as a template for reverse transcription.

[0362] As used herein, a "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) that has sufficient self-complementarity to form a stem-loop, for example, having a stem containing at least two (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop having at least three (e.g., four) base pairs. The stem may contain mismatches or protrusions.

[0363] As used herein, a “tissue-specific expression control sequence” refers to a nucleic acid element that, in a tissue-specific manner, preferentially increases or decreases the level of transcripts containing a heterologous object sequence in one or more target tissues relative to one or more off-target tissues. In some embodiments, the tissue-specific expression control sequence preferentially drives or inhibits transcription, activity, or half-life of transcripts containing a heterologous object sequence in a tissue-specific manner, for example, preferentially in one or more target tissues relative to one or more off-target tissues. Exemplary tissue-specific expression control sequences include tissue-specific promoters, repressors, enhancers, or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to target (one or more tissues where expression or activity of the template nucleic acid is desired or tolerated) and off-target (one or more tissues where expression or activity of the template nucleic acid is not desired or tolerated). For example, a tissue-specific promoter preferentially drives expression in target tissues relative to off-target tissues. Conversely, microRNAs binding to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues relative to target tissues, thereby reducing the expression of the template nucleic acid in off-target tissues. Therefore, regarding the transcription, activity, or half-life of associated sequences in tissues, promoters and microRNA recognition sequences specific to the same tissue (e.g., target tissue) have different functions (promoting and inhibiting, respectively, with consistent expression levels, i.e., high levels of microRNA in off-target tissues and low levels in target tissues, while promoters drive high expression in target tissues and low expression in off-target tissues).

[0364] Table of contents

[0365] 1) Introduction

[0366] 2) Gene modification system

[0367] a) Polypeptide components of gene modification systems

[0368] i) Writing structural domain

[0369] ii) Endonuclease domain and DNA-binding domain

[0370] (1) Gene-modified polypeptides containing Cas domains

[0371] (2) TAL effectors and zinc finger nucleases

[0372] iii) Connector

[0373] iv) Localization sequences of gene modification systems

[0374] v) Genetically modified peptides and evolutionary variants of systems

[0375] vi) Contains peptides

[0376] vii) Other structural domains

[0377] b) Template nucleic acid

[0378] i) gRNA spacers and gRNA scaffolds

[0379] ii) Heterogeneous object sequence

[0380] iii) PBS sequence

[0381] iv) Exemplary template sequence

[0382] c) gRNAs with inducible activity

[0383] d) Circular RNA and ribozymes in gene modification systems

[0384] e) Target nucleic acid sites

[0385] f) Second chain cut

[0386] 3) Generation of compositions and systems

[0387] 4) Therapeutic applications

[0388] 5) Application and delivery

[0389] a) Tissue-specific activity / application

[0390] i) Promoter

[0391] ii) microRNA

[0392] b) Viral vectors and their components

[0393] c) AAV administration

[0394] d) Lipid nanoparticles

[0395] 6) Reagent kits, products and pharmaceutical compositions

[0396] 7) Chemistry, Manufacturing and Control (CMC)

[0397] introduction

[0398] This disclosure relates to methods for treating α-1 antitrypsin deficiency (AATD) and compositions for targeting, editing, modifying, or manipulating DNA sequences at one or more locations in a cell, tissue, or subject, for example, in vivo or in vitro (e.g., inserting a heterologous object sequence into a target site in a mammalian genome). The heterologous object DNA sequence may include, for example, substitution.

[0399] More specifically, this disclosure provides methods for treating AATD that use a reverse transcriptase-based system to alter a target genomic DNA sequence, for example, by inserting one or more nucleotides into the target sequence, deleting one or more nucleotides from the target sequence, or replacing one or more nucleotides in the target sequence.

[0400] This disclosure partially provides methods for treating AATD using a gene modification system comprising a genetically modified polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the gene modification system can be used to introduce alterations into a target site in the genome. In some embodiments, the genetically modified polypeptide component comprises a writing domain (e.g., a reverse transcriptase domain), a DNA-binding domain, and a nuclease domain (e.g., a nicking enzyme domain). In some embodiments, the template nucleic acid (e.g., template RNA) comprises a sequence (e.g., a gRNA spacer) that binds to a target site in the genome (e.g., a second strand that binds to the target site), a sequence that binds to the genetically modified polypeptide component (e.g., a gRNA scaffold), a heterologous object sequence, and a PBS sequence. It is not intended to be theoretically constrained to assume that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site in the genome and binds to the genetically modified polypeptide component (e.g., to localize the polypeptide component to the target site in the genome). It is assumed that a nuclease (e.g., a nicking enzyme) of the genetically modified polypeptide component cleaves the target site (e.g., the first strand of the target site), for example, allowing the PBS sequence to bind to a sequence adjacent to the site to be altered on the first strand of the target site. It is assumed that the writing domain of a polypeptide component (e.g., a reverse transcriptase domain) uses a PBS sequence containing a template nucleic acid as a primer and a heterologous object sequence of the template nucleic acid as a template to bind to the first strand of a target site, for example, by polymerizing a sequence complementary to the heterologous object sequence. Not wanting to be bound by theory, it is assumed that choosing a suitable heterologous object sequence can lead to the substitution, deletion, and / or insertion of one or more nucleotides at the target site.

[0401] Gene modification system

[0402] In some embodiments, the gene-modifying system described herein comprises: (A) a gene-modifying polypeptide or a nucleic acid encoding the gene-modifying polypeptide, wherein the gene-modifying polypeptide comprises (i) a reverse transcriptase domain and (x) a DNA-binding endonuclease domain or (y) an endonuclease domain and a separate DNA-binding domain; and (B) a template RNA. In some embodiments, the gene-modifying polypeptide acts as a substantially autonomous protein machine capable of integrating a template nucleic acid sequence into a target DNA molecule (e.g., in mammalian host cells, such as genomic DNA molecules in host cells), substantially independent of the host machine. For example, the gene-modifying protein may comprise a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, the DNA-binding function may involve an RNA component that guides the protein to a DNA sequence (e.g., a gRNA spacer). In other embodiments, the gene-modifying polypeptide may comprise a reverse transcriptase domain and an endonuclease domain. The RNA template element of the gene-modifying system is typically heterologous to the gene-modifying polypeptide element and provides the target sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the gene-modifying polypeptide is capable of targeted reverse transcription. In some embodiments, the genetically modified polypeptide can undergo second-chain synthesis.

[0403] In some embodiments, the gene modification system is combined with a second polypeptide. In some embodiments, the second polypeptide may contain an endonuclease domain. In some embodiments, the second polypeptide may contain a polymerase domain, such as a reverse transcriptase domain. In some embodiments, the second polypeptide may contain a DNA-dependent DNA polymerase domain. In some embodiments, the second polypeptide facilitates genome editing, for example, by facilitating second-strand synthesis or DNA repair dissociation.

[0404] Functional gene-modifying peptides can be composed of unrelated DNA-binding domains, reverse transcription domains, and endonuclease domains. This modular structure allows for the combination of functional domains, such as dCas9 (DNA binding), MMLV reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains can originate from a single protein, such as Cas9 or Cas9 nickase (DNA binding, endonuclease).

[0405] In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively facilitate 1) binding to a template nucleic acid, 2) binding to a target DNA molecule, and 3) facilitating the integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the genetically modified polypeptide is an engineered polypeptide, for example, having one or more amino acid substitutions relative to a naturally occurring sequence. In some embodiments, the genetically modified polypeptide comprises two or more domains that are heterologous to each other, for example, through heterologous fusion (or other conjugates) of domains that are otherwise wild-type, or through the fusion of modified domains, for example, through substitution or fusion of heterologous subdomains or other substituted domains. For example, in some embodiments, one or more of the following are heterologous: the RT domain is heterologous to the DBD; the DBD is heterologous to the endonuclease domain; or the RT domain is heterologous to the endonuclease domain.

[0406] In some embodiments, the template RNA molecule used in this system comprises, from 5′ to 3′, (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous object sequence; and (4) a primer binding site (PBS) sequence. In some embodiments:

[0407] (1) is a gRNA spacer of approximately 18-22 nt (e.g., 20 nt).

[0408] (2) is a gRNA scaffold containing one or more hairpin loops (e.g., 1, 2, or 3 loops) for associating a template with a Cas domain, such as the nickase Cas9 domain. In some embodiments, the gRNA scaffold contains the sequence GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGGACCGAGTCGGTCC (SEQ ID NO: 5008) from 5′ to 3′.

[0409] (3) In some embodiments, the heterologous object sequence length is, for example, 7-74, such as 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, or 70-80 nt or 80-90 nt. In some embodiments, the first (5') base of the sequence is not C.

[0410] (4) In some embodiments, the PBS sequence that binds to the target-initiating sequence after the nick occurs is, for example, 3-20 nt, 7-15 nt, or 12-14 nt. In some embodiments, the PBS sequence has a GC content of 40%-60%.

[0411] In some embodiments, a second gRNA associated with the system may help drive complete integration. In some embodiments, the second gRNA may target a position 0-200 nt from the first strand cut, such as 0-50, 50-100, or 100-200 nt from the first strand cut. In some embodiments, the second gRNA may bind to its target sequence only after editing, for example, the gRNA binds to a sequence present in the heterologous object sequence but not in the initial target sequence.

[0412] In some embodiments, the gene-editing system described herein is used for editing in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the gene-editing system is used for editing in primary cells (e.g., primary cortical neurons from E18.5 mice).

[0413] In some embodiments, the genetically modified polypeptide as described herein comprises a reverse transcriptase or RT domain containing a MoMLV RT sequence or a variant thereof (e.g., as described herein). In embodiments, the MoMLV RT sequence comprises one or more mutations selected from the group consisting of: D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In embodiments, the MoMLV RT sequence comprises a combination of mutations (e.g., D200N, L603W, and T330P), optionally further comprising T306K and / or W313F.

[0414] In some embodiments, the endonuclease domain (e.g., as described herein) nCas9, for example, contains an N863A mutation (e.g., in spCas9) or an H840A mutation.

[0415] In some embodiments, the length of the heterologous object sequence (e.g., the system described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000 or more nucleotides.

[0416] In some embodiments, the RT and endonuclease domains are connected by a flexible linker, for example, containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 5006).

[0417] In some embodiments, the endonuclease domain is located at the N-terminus relative to the RT domain. In some embodiments, the endonuclease domain is located at the C-terminus relative to the RT domain.

[0418] In some embodiments, the system incorporates heterologous object sequences into target sites via TPRT, for example, as described herein.

[0419] In some embodiments, the genetically modified polypeptide includes a DNA-binding domain. In some embodiments, the genetically modified polypeptide includes an RNA-binding domain. In some embodiments, the RNA-binding domain includes an RNA-binding domain of a B-box protein, MS2 coat protein, dCas, or an element of the sequence listed herein. In some embodiments, the RNA-binding domain is capable of binding template RNA with a greater affinity than a reference RNA-binding domain.

[0420] In some embodiments, the gene modification system is capable of generating at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides) of insertion at the target site. In some embodiments, the gene modification system is capable of generating at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides) of insertion at the target site. In some embodiments, the gene modification system is capable of generating insertions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally not exceeding 1, 5, 10, or 20 kilobases) at the target site. In some embodiments, the gene modification system is capable of generating deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene-editing system is capable of producing deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene-editing system is capable of producing deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene modification system is capable of producing deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases (and optionally not exceeding 1, 5, 10, or 20 kilobases).In some embodiments, the gene modification system is capable of generating at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotide substitutions at the target site. In some embodiments, the gene modification system is capable of generating 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotide substitutions at the target site.

[0421] In some embodiments, substitution is a translocation mutation. In some embodiments, substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.

[0422] In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by altering, adding, or deleting sequences in promoters or enhancers (e.g., sequences that bind transcription factors). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene translation (e.g., altering amino acid sequences), inserting or deleting start or stop codons, altering or fixing the translational framework of a gene. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene splicing, for example by inserting, deleting, or altering splice acceptor or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter transcript or protein half-life. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in the cell (e.g., from the cytoplasm to the mitochondria, from the cytoplasm to the extracellular space (e.g., adding secretion tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter (e.g., improve) protein folding (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease gene activity, such as the activity of proteins encoded by the gene.

[0423] Exemplary genetically modified peptides, systems comprising them, and methods of using them are described, for example, in PCT / US2021 / 020948, which relates to the retroviral RT domain (including the amino acid and nucleic acid sequences therein) and is incorporated herein by reference.

[0424] Exemplary gene-modified peptides and retroviral RT domain sequences are also described, for example, in International Application No. PCT / US21 / 20948, filed March 4, 2021, such as Tables 30, 31, and 44 therein; the entire application relating to retroviral RT is incorporated herein by reference, for example, in the sequences and tables described herein. Therefore, the gene-modified peptides described herein may comprise amino acid sequences or their domains (e.g., retroviral RT domains) according to any of the tables mentioned in this paragraph, or functional fragments or variants of any of the foregoing, or amino acid sequences having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with them.

[0425] In some embodiments, the peptide used in any of the systems described herein may be a molecular or genetic reconstruction based on the alignment of peptide sequences with multiple homologous proteins. In some embodiments, the reverse transcriptase domain used in any of the systems described herein may be a molecular or genetic reconstruction, or may be modified at specific residues based on alignments of reverse transcriptase domains from the same or different sources. Based on the accession numbers provided herein, those skilled in the art can align peptide or nucleic acid sequences, for example, using conventional sequence analysis tools such as the Basic Local Alignment Search (BLAST) or CD-search (for conserved domain analysis). Molecular reconstructions may be created based on shared sequences, for example using methods described in Ivics et al., Cell 1997, 501–510; Wagstaff et al., Molecular Biology and Evolution 2013, 88–99.

[0426] polypeptide components of gene modification systems

[0427] In some embodiments, the genetically modified polypeptide has the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing (e.g., reverse transcription). In some embodiments, each function is contained within a different domain. In some embodiments, the function may be attributed to two or more domains (e.g., two or more domains together exhibit the function). In some embodiments, two or more domains may have the same or similar functions (e.g., two or more domains each independently have DNA binding functionality, e.g., for two different DNA sequences). In other embodiments, one or more domains may be capable of performing one or more functions; for example, the Cas9 domain is capable of simultaneously performing DNA binding and target site cleavage. In some embodiments, these domains are all located within a single polypeptide. In some embodiments, a first domain is in one polypeptide and a second domain is in a second polypeptide. For example, in some embodiments, the sequence may be broken between the first polypeptide and the second polypeptide, e.g., where the first polypeptide contains a reverse transcriptase (RT) domain and where the second polypeptide contains a DNA-binding domain and a nuclease domain, such as a nicking enzyme domain. As a further example, in some embodiments, the first polypeptide and the second polypeptide each comprise a DNA-binding domain (e.g., a first DNA-binding domain and a second DNA-binding domain). In some embodiments, the first and second polypeptides can be linked together post-translationally by breaking down introns to form a single genetically modified polypeptide.

[0428] In some embodiments, the genetically modified polypeptide described herein comprises a St1Cas9 domain. The St1Cas9 domain may comprise a naturally occurring St1Cas9 amino acid sequence or a variant thereof. In some embodiments, the St1Cas9 domain is a nicking enzyme. In some embodiments, the St1Cas9 domain comprises the sequence according to SEQ ID NO: 23818, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprising the St1Cas9 domain is used in conjunction with a compatible template RNA comprising a variant gRNA scaffold described herein.

[0429] In some respects, the gene-modifying polypeptides described herein comprise (e.g., the systems described herein comprise gene-modifying polypeptides comprising): 1) a Cas domain (e.g., a Cas nickase domain, such as a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain of Table D, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity therewith, wherein the RT domain is located at the C-terminus of the Cas domain; and a linker located between the RT domain and the Cas domain, wherein the linker has a sequence from the same row of Table D as the RT domain, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity therewith.

[0430] In some embodiments, the RT domain has a sequence that is 100% identical to the RT domain in Table D, and the adapter has a sequence that is 100% identical to the adapter sequence in the same row as the RT domain in Table D. In some embodiments, the Cas domain comprises the sequence in Table 8 or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it. In some embodiments, the gene-modifying polypeptide comprises an amino acid sequence according to any one of SEQ ID NO: 1-3332 in the sequence listing, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it.

[0431] In some embodiments, the gene-modified polypeptide comprises a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain. In some embodiments, the gene-modified polypeptide comprises the sequence of SEQ ID NO: 4000 containing the first NLS and the Cas domain, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises the sequence of SEQ ID NO: 4001 containing the second NLS, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

[0432] Exemplary N-terminal NLS-Cas9 domain

[0433]

[0434] Exemplary C-terminal sequence containing NLS

[0435] AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001)

[0436] Writing a structure field (RT structure field)

[0437] In some aspects of the invention, the writing domain of the gene modification system has reverse transcriptase activity and is also referred to as the reverse transcriptase domain (RT domain). In some embodiments, the RT domain comprises an RT catalytic portion and an RNA-binding region (e.g., a region that binds template RNA).

[0438] In some embodiments, the nucleic acid encoding reverse transcriptase is altered from its native sequence to have modified codon usage, for example, for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase derived from a retrovirus. In some embodiments, the RT domain comprising the genetically modified polypeptide has been mutated from its original amino acid sequence, for example, having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions. In some embodiments, the RT domain is derived from a retroviral RT, such as HIV-1 RT, Moloney murine leukemia virus (MMLV) RT, avian myeloblastoma virus (AMV) RT, or Rous sarcoma virus (RSV) RT.

[0439] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits enhanced strictness in the initiation of target-induced reverse transcription (TPRT), for example, relative to the endogenous RT domain. In some embodiments, the RT domain initiates TPRT when a 3 nt immediately upstream of the first-strand cleavage at the target site, such as the genomic DNA initiating the RNA template, has at least 66% or 100% complementarity to a homologous 3 nt in the RNA template. In some embodiments, the RT domain initiates TPRT when there is less than 5 nt mismatch (e.g., less than 1, 2, 3, 4, or 5 nt mismatch) between the template RNA homology and the target DNA-induced reverse transcription. In some embodiments, the RT domain is modified to increase the strictness of mismatches in the initiation of the TPRT response, for example, wherein the RT domain does not tolerate any mismatches or allows fewer mismatches in the initiation region relative to the wild-type (e.g., unmodified) RT domain. In some embodiments, the RT domain comprises the HIV-1 RT domain. In embodiments, the HIV-1 RT domain initiates low-level synthesis even with a three-nucleotide mismatch relative to an alternative RT domain (e.g., as described by Jamburuthugoda and Eickbush J Mol Biol [Journal of Molecular Biology] 407(5):661-672 (2011); which is incorporated herein by reference in its entirety). In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain is monomeric. In some embodiments, the RT domain naturally functions as a monomer or a dimer (e.g., a heterodimer or a homodimer). In some embodiments, the RT domain naturally functions as a monomer, for example, derived from a virus, where it functions as a monomer.In the embodiments, the RT domain is selected from the following RT domains: murine leukemia virus (MLV; sometimes called MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), avian reticuloendothelial hyperplasia virus (AVIRE) (UniProtKB accession number: P03360), feline leukemia virus (FLV or FeLV) (e.g., UniProtKB accession number: P10273), Mason-Fischer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), and human foamy virus (HFV) (e.g., UniProt...). P14350), simian foamy virus (SFV) (e.g., SFV3L) (e.g., UniProt P23074 or P27401), or bovine foamy / syncytial virus (BFV / BSV) (e.g., UniProt O41894), or functional fragments or variants thereof (e.g., amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with it). In some embodiments, the RT domain is a dimer in its native function. In some embodiments, the RT domain is derived from a virus, wherein it functions as a dimer. In the embodiments, the RT domain is selected from the following RT domains: avian sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Rous sarcoma virus (RSV) (e.g., UniProt P03354), avian myeloblastoma virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci [Cell and Molecular Life Sciences]). 67(16):2717-2747(2010)), or its functional fragments or variants (e.g., amino acid sequences that are at least 70%, 80%, 90%, 95% or 99% identical to it).In some embodiments, the natural heterodimeric RT domain may also function as a homodimer. In some embodiments, the dimer RT domain is expressed as a fusion protein, such as a homodimeric fusion protein or a heterodimeric fusion protein. In some embodiments, the RT function of the system is achieved by multiple RT domains (e.g., as described herein). In other embodiments, the multiple RT domains are fused or separate, for example, they may be on the same polypeptide or on different polypeptides.

[0440] In some embodiments, the gene modification system described herein includes an integrase domain, for example, wherein the integrase domain may be part of an RT domain. In some embodiments, the RT domain (e.g., as described herein) includes an integrase domain. In some embodiments, the RT domain (e.g., as described herein) lacks an integrase domain, or includes an integrase domain that has been inactivated by mutation or deletion. In some embodiments, the gene modification system described herein includes an RNase H domain, for example, wherein the RNase H domain may be part of an RT domain. In some embodiments, the RNase H domain is not part of the RT domain and is covalently linked by a flexible linker. In some embodiments, the RT domain (e.g., as described herein) includes an RNase H domain, for example, an endogenous RNase H domain or a heterologous RNase H domain. In some embodiments, the RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, the RT domain (e.g., as described herein) includes an RNase H domain with the addition, deletion, mutation, or exchange of a heterologous RNase H domain. In some embodiments, the peptide includes an inactivated endogenous RNase H domain. In some embodiments, the endogenous RNase H domain is genetically removed from one of the other domains of the polypeptide, such that it is not included in the polypeptide, for example, the endogenous RNase H domain is partially or completely truncated from the containing domain. In some embodiments, the mutation of the RNase H domain produces a polypeptide exhibiting lower RNase activity, for example, as determined by the method described by Kotewicz et al. Nucleic Acids Res [Nucleic Acids Research] 16(1):265-277 (1988) (which is incorporated herein by reference in its entirety), for example, a reduction of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% compared to a otherwise similar domain without the mutation. In some embodiments, RNase H activity is eliminated.

[0441] In some embodiments, the RT domain is mutated to increase fidelity compared to other similar domains that are not mutated. For example, in some embodiments, the YADD or YMDD motif in the RT domain (e.g., in reverse transcriptase) is replaced by YVDD. In embodiments, the substitution of YADD or YMDD or YVDD results in higher fidelity of retroviral reverse transcriptase activity (e.g., as described by Jamburuthugoda and Eickbush J Mol Biol [Journal of Molecular Biology] 2011; which is incorporated herein by reference in its entirety).

[0442] In some embodiments, the genetically modified polypeptides described herein comprise an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it. In some embodiments, the nucleic acids described herein encode an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it.

[0443] Table 6: Exemplary reverse transcriptase domains from retroviruses

[0444]

[0445]

[0446]

[0447]

[0448]

[0449]

[0450]

[0451]

[0452]

[0453]

[0454]

[0455]

[0456]

[0457]

[0458]

[0459]

[0460]

[0461]

[0462]

[0463]

[0464]

[0465]

[0466]

[0467]

[0468]

[0469]

[0470]

[0471]

[0472]

[0473]

[0474]

[0475]

[0476]

[0477]

[0478]

[0479]

[0480]

[0481]

[0482]

[0483]

[0484] In some embodiments, the reverse transcriptase domain is modified, for example, through site-specific mutations. In some embodiments, the reverse transcriptase domain is engineered to have improved properties, such as SuperScript IV (SSIV) reverse transcriptase derived from MMLV RT. In some embodiments, the reverse transcriptase domain may be engineered to have a lower error rate, for example, as described in WO 2001068895 (incorporated herein by reference). In some embodiments, the reverse transcriptase domain may be engineered to be more heat-resistant. In some embodiments, the reverse transcriptase domain may be engineered to have greater sustained synthetic capacity. In some embodiments, the reverse transcriptase domain may be engineered to be resistant to inhibitors. In some embodiments, the reverse transcriptase domain may be engineered to be faster. In some embodiments, the reverse transcriptase domain may be engineered to be more resistant to modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain may be engineered to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is engineered to bind template RNA. In some embodiments, one or more mutations are selected from D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, H8Y, T306K, or D653N in the RT domain of murine leukemia virus reverse transcriptase, or corresponding mutations at the corresponding positions in another RT domain.

[0485] In some embodiments, the genetically modified polypeptide comprises an RT domain derived from a retroviral reverse transcriptase, such as wild-type M-MLV RT, for example, comprising the following sequence:

[0486] M-MLV (WT): (SEQ ID NO: 5002)

[0487] In some embodiments, the genetically modified polypeptide comprises an RT domain derived from a retroviral reverse transcriptase, such as M-MLV RT, for example, comprising the following sequence:

[0488] (SEQ ID NO: 5003)

[0489] In some embodiments, the genetically modified polypeptide comprises an RT domain derived from a retroviral reverse transcriptase, which contains the sequence of amino acids 659-1329 of NP_057933. In an embodiment, the genetically modified polypeptide further comprises an additional amino acid at the N-terminus of the sequence of amino acids 659-1329 of NP_057933, for example as shown below:

[0490] TLNIEDEHRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDYRPVQDLREVNKRVEDIHPTV PNPYNLLSGLPPSHQWYTVLDLKDAFFCLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFDEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRASAKKA QICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLREFLGTAGFCRLWIPGFAEMAAPLYPLTKTGTLFNWGPDQQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWR RPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPVVALNPATLLPLPEEGLQHNCLDILAEAHGTRPDLTDQPLPDADH TWYTD GSSLLQEGQRKAGAAVTTETEVIWAKALPAGTSAQRAELIALTQALKMAEGKKLNVYTDSRYAFATAHIHGEIYRR RGLLTSEGKEIKNKDEILALLKALFLPKRLSIIHCPGHQKGHSAEARGNRMADQAARKAA (SEQ ID NO: 5004)

[0491] Core RT (bold), as noted above

[0492] RNase H (Underlined), as noted above

[0493] In one embodiment, the gene-modified polypeptide further includes an additional amino acid at the C-terminus of amino acid sequence 659-1329 of NP_057933. In another embodiment, the gene-modified polypeptide includes an RNase H1 domain (e.g., amino acid sequence 1178-1318 of NP_057933).

[0494] In some embodiments, the retroviral reverse transcriptase domain, such as M-MLV RT, may contain one or more mutations of the wild-type sequence, which may improve the characteristics of RT, such as thermostability, sustained synthetic capacity, and / or template binding. In some embodiments, the M-MLV RT domain contains one or more mutations relative to the above-described M-MLV (WT) sequence, such as those selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, K103L, or combinations of mutations, such as D200N, L603W, and T330P, optionally further including T306K and W313F. In some embodiments, the M-MLV RT used herein comprises mutants D200N, L603W, T330P, T306K, and W313F. In the embodiments, the mutant M-MLV RT comprises the following amino acid sequence:

[0495] M-MLV (PE2): (SEQ ID NO: 5005)

[0496] In some embodiments, the writing domain (e.g., the RT domain) includes an RNA-binding domain, which specifically binds to an RNA sequence. In some embodiments, the template RNA includes an RNA sequence specifically bound by the RNA-binding domain of the writing domain.

[0497] In some embodiments, the reverse transcription domain recognizes and reverse transcribes only a specific template, such as the system's template RNA. In some embodiments, the template contains a sequence or structure that can be recognized and reverse transcribed by the reverse transcription domain. In some embodiments, the template contains a sequence or structure that can be associated with an RNA-binding domain of a polypeptide component of the genome engineering system described herein. In some embodiments, the genome engineering system preferably reverse transcribes a template containing the associated sequence, rather than a template lacking the associated sequence.

[0498] The writing domain may also include DNA-dependent DNA polymerase activity, for example, enzymatic activity capable of writing DNA from a template DNA sequence book into the genome. In some embodiments, DNA-dependent DNA polymerization is employed to complete the second-strand synthesis for target site editing. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in a polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, such as second-strand synthesis. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a second polypeptide in the system. In some embodiments, the DNA-dependent DNA polymerase activity is provided by an endogenous host cell polymerase, which is optionally recruited to the target site by a component of the genome engineering system.

[0499] In some embodiments, the reverse transcriptase domain has a lower probability of premature termination (P0) in vitro compared to the reference reverse transcriptase domain. off In some embodiments, the reference reverse transcriptase domain is a viral reverse transcriptase domain, such as the RT domain from M-MLV.

[0500] In some embodiments, the reverse transcriptase domain has a density of less than about 5 x 10⁻⁶. -3 / nt, 5 x 10 -4 / nt or 5 x 10 -6 / nt premature termination rate in vitro (P off The lower probability of premature termination, for example, as measured on 1094 nt RNA. In the examples, the in vitro premature termination rate was determined as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (which is incorporated herein by reference in its entirety).

[0501] In some embodiments, the reverse transcriptase domain is capable of completing at least about 30% or 50% integration in the cell. The percentage of complete integration can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites containing at least 98% of the expected integration sequence) by the total number of integration events in the cell population (including substantially full-length and partial integration events). In embodiments, long-read amplicon sequencing is used to determine integration in the cell (e.g., across integration sites), for example, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0502] In the embodiments, quantifying integration in cells includes counting the integrated portion of a DNA sequence containing at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the template RNA (e.g., template RNA of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 3, 4, or 5 kb in length, such as 0.5–0.6, 0.6–0.7, 0.7–0.8, 0.8–0.9, 1.0–1.2, 1.2–1.4, 1.4–1.6, 1.6–1.8, 1.8–2.0, 2–3, 3–4, or 4–5 kb in length).

[0503] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at a rate of 0.1–50 nt / sec (e.g., 0.1–1, 1–10, or 10–50 nt / sec). In embodiments, the polymerization of dNTPs by the reverse transcriptase domain is measured by a single-molecule assay, for example, as described in Schwartz and Quake (2009) PNAS [Proceedings of the National Academy of Sciences] 106(48):20294–20299 (which is incorporated herein by reference in its entirety).

[0504] In some embodiments, the in vitro error rate of the reverse transcriptase domain (e.g., nucleotide mis-incorporation) is 1 x 10⁻⁶. -3 - 1 x 10 -4 Or 1 x 10 -4 - 1 x 10 -5One substitution / nt, for example, as described in Yasukawa et al. (2017) BiochemBiophys Res Commun [Biochemistry and Biophysics Research Communications] 492(2):147-153 (which is incorporated herein by reference in its entirety). In some embodiments, the error rate (e.g., nucleotide mis-incorporation) of the reverse transcriptase domain in cells (e.g., HEK293T cells) is 1 x 10 -3 - 1 x 10 -4 Or 1 x 10 -4 - 1 x 10 -5 One substitution / nt, for example, by long-read amplicon sequencing, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0505] In some embodiments, the reverse transcriptase domain enables reverse transcription of the target RNA in vitro. In some embodiments, the reverse transcriptase requires a primer of at least 3 nucleotides to initiate reverse transcription of the template. In some embodiments, reverse transcription of the target RNA is determined by detecting cDNA from the target RNA (e.g., when an ssDNA primer is provided, for example, which is annealed to the target at the 3' end by at least 3, 4, 5, 6, 7, 8, 9, or 10 nt), for example as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (which is incorporated herein by reference in its entirety).

[0506] In some embodiments, the reverse transcriptase domain performs reverse transcription at least 5 or 10 times more efficiently than an RNA template lacking a protein-binding motif (e.g., 3'UTR), for example, when its RNA template is converted to cDNA (e.g., via cDNA production). In the embodiments, the reverse transcription efficiency is measured as described in Yasukawa et al. (2017) Biochem Biophys ResCommun [Biochemistry and Biophysics Research Communications] 492(2):147-153 (which is incorporated herein by reference in its entirety).

[0507] In some embodiments, the reverse transcriptase domain binds specifically to a particular RNA template at a higher frequency (e.g., about 5 or 10 times higher) than any endogenous cellular RNA (e.g., when expressed in cells (e.g., HEK293T cells)). In embodiments, the specific binding frequency between the reverse transcriptase domain and the template RNA is measured by CLIP-seq, as described, for example, in Lin and Miles (2019) Nucleic Acids Res [Nucleic Acids Research] 47(11):5490-5501 (which is incorporated herein by reference in its entirety).

[0508] Template nucleic acid binding domain

[0509] Genetically modified peptides typically contain regions capable of being associated with a template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain that can be associated with an RNA molecule containing specific features (e.g., structural motifs). In other embodiments, the template nucleic acid binding domain (e.g., an RNA binding domain) is contained within a reverse transcription domain, for example, a reverse transcriptase-derived component having known RNA preference characteristics.

[0510] In other embodiments, a template nucleic acid binding domain (e.g., an RNA binding domain) is contained within a target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-associated protein that recognizes a structure containing a template nucleic acid (e.g., template RNA) containing gRNA. In some embodiments, the genetically modified polypeptide contains a DNA binding domain that includes a CRISPR-associated protein associated with a gRNA scaffold that allows the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the gRNA scaffold and gRNA spacer are contained within a template nucleic acid (e.g., template RNA), thus the DNA binding domain is also a template nucleic acid binding domain. In some embodiments, the polypeptide has RNA-binding functionality in multiple domains, for example, it can bind a gRNA structure in a CRISPR-associated DNA binding domain and additional sequences or structures in a reverse transcriptase domain.

[0511] In some embodiments, the RNA-binding domain is capable of binding template RNA with a greater affinity than a reference RNA-binding domain. In some embodiments, the reference RNA-binding domain is the RNA-binding domain of Cas9 from *Streptococcus pyogenes*. In some embodiments, the RNA-binding domain is capable of binding template RNA with an affinity of 100 pM–10 nM (e.g., 100 pM–1 nM or 1 nM–10 nM). In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in vitro, e.g., by thermophoresis, e.g., as described in Asmari et al. Methods 146:107–119 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in cells (e.g., by FRET or CLIP-Seq).

[0512] In some embodiments, the RNA-binding domain is associated with the template RNA in vitro at a frequency at least about 5 to 10 times higher than that of the out-of-order RNA. In some embodiments, the binding frequency between the RNA-binding domain and the template RNA or the out-of-order RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) Nucleic Acids Res [Nucleic Acids Research] 47(11):5490-5501 (which is incorporated herein by reference in its entirety). In some embodiments, the RNA-binding domain is associated with the template RNA in cells (e.g., HEK293T cells) at a frequency at least about 5 to 10 times higher than that of the out-of-order RNA. In some embodiments, the association frequency between the RNA-binding domain and the template RNA or the out-of-order RNA is measured by CLIP-seq, for example, as described above in Lin and Miles (2019).

[0513] In some embodiments, the RT domain (e.g., as listed in Table 6) contains one or more mutations listed in Table 2A below. In some embodiments, the RT domain as listed in Table 6 contains one, two, three, four, five, or six mutations listed in the corresponding row of Table 2A below.

[0514] Table 2A. Exemplary RT domain mutations (relative to the corresponding wild-type sequences listed in the corresponding rows of Table 6)

[0515]

[0516]

[0517]

[0518]

[0519]

[0520] Nucleotide endonuclease domain and DNA-binding domain

[0521] In some embodiments, the genetically modified polypeptide has the function of cleaving a DNA target site via a nuclease domain. In some embodiments, the genetically modified polypeptide includes a DNA-binding domain, for example, for binding to a target nucleic acid. In some embodiments, the domains of the genetically modified polypeptide (e.g., a Cas domain) include two or more smaller domains (e.g., a DNA-binding domain and a nuclease domain). It should be understood that when the DNA-binding domain (e.g., a Cas domain) is described as binding to a target nucleic acid sequence, in some embodiments, this binding is mediated by gRNA.

[0522] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-associated endonuclease domain that binds to a template RNA containing gRNA, binds to a target DNA sequence (e.g., complementary to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, the heterologous endonuclease domain or endonuclease / DNA-binding domain may be used or may be modified in the gene modification system described herein (e.g., by inserting, deleting, or substituting one or more residues).

[0523] In some embodiments, the nucleic acid encoding a nuclease domain or a nuclease / DNA binding domain is altered from its native sequence to have a modified codon, for example, for modification targeting human cells. In some embodiments, the nuclease element is a heterologous nuclease element, such as a Cas nuclease (e.g., Cas9), a type II restriction nuclease (e.g., Fok1), a broad-spectrum nuclease (e.g., I-SceI), or other nuclease domains.

[0524] In some aspects, the DNA-binding domain of the genetically modified peptide described herein is selected, designed, or constructed to bind to a desired host DNA target sequence. In some embodiments, the DNA-binding domain of the peptide is a heterologous DNA-binding element. In some embodiments, the heterologous DNA-binding element is a zinc finger element or a TAL effector element, such as a zinc finger or TAL peptide or a functional fragment thereof. In some embodiments, the heterologous DNA-binding element is a sequence-guided DNA-binding element, such as Cas9, Cpf1, or other CRISPR-related proteins that have been modified to lack endonuclease activity. In some embodiments, the heterologous DNA-binding element retains endonuclease activity. In some embodiments, the heterologous DNA-binding element retains partial endonuclease activity to cleave ssDNA, for example, having cleavage enzyme activity. In specific embodiments, the heterologous DNA-binding domain can be any one or more of Cas9, a TAL domain, a ZF domain, a Myb domain, combinations thereof, or multiples thereof.

[0525] In some embodiments, the DNA-binding domain is modified, for example, by site-specific mutations, increasing or decreasing DNA-binding elements (e.g., the number and / or specificity of zinc fingers), to alter DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is altered from its natural sequence to have modified codon usage, for example, for human cells. In embodiments, the DNA-binding domain comprises one or more modifications relative to the wild-type DNA-binding domain, for example, modifications via directed evolution (e.g., phage-assisted sequential evolution (PACE)).

[0526] In some embodiments, the DNA-binding domain comprises a large-scale nuclease domain (e.g., as described herein, for example, within the endonuclease domain portion), or a functional fragment thereof. In some embodiments, the large-scale nuclease domain has endonuclease activity, such as double-strand cleavage and / or nickase activity. In other embodiments, the large-scale nuclease domain has reduced activity, such as lacking endonuclease activity, for example, the large-scale nuclease has no catalytic activity. In some embodiments, a large-scale nuclease without catalytic activity is used as the DNA-binding domain, for example, as described in Fonfara et al. Nucleic Acids Res [Nucleic Acids Research] 40(2):847-860 (2012), which is incorporated herein by reference in its entirety.

[0527] In some embodiments, the genetically modified polypeptide comprises a modification of the DNA-binding domain, for example, relative to the wild-type polypeptide. In some embodiments, the DNA-binding domain comprises an addition, deletion, substitution, or modification of the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence. In some embodiments, the functional domain replaces at least a portion (e.g., all) of the polypeptide's previous DNA-binding domain. In some embodiments, the functional domain comprises a zinc finger (e.g., a zinc finger that specifically binds to a target nucleic acid (e.g., DNA) sequence). In some embodiments, the functional domain comprises a Cas domain (e.g., a Cas domain that specifically binds to a target nucleic acid (e.g., DNA) sequence). In some embodiments, the Cas domain comprises Cas9 or a mutant or variant thereof (e.g., as described herein). In embodiments, the Cas domain is associated with a guide RNA (gRNA), for example, as described herein. In embodiments, the Cas domain is directed by the gRNA to a target nucleic acid (e.g., DNA) sequence. In embodiments, the Cas domain and the gRNA are encoded in the same nucleic acid (e.g., RNA) molecule. In the embodiments, the Cas domain and gRNA are encoded in different nucleic acid (e.g., RNA) molecules.

[0528] In some embodiments, the DNA-binding domain is capable of binding the target sequence (e.g., the dsDNA target sequence) with a greater affinity than a reference DNA-binding domain. In some embodiments, the reference DNA-binding domain is the DNA-binding domain of Cas9 from Streptococcus pyogenes. In some embodiments, the DNA-binding domain is capable of binding the target sequence (e.g., the dsDNA target sequence) with an affinity between 100 pM and 10 nM (e.g., between 100 pM and 1 nM or between 1 nM and 10 nM).

[0529] In some embodiments, the affinity of the DNA-binding domain for its target sequence (e.g., dsDNA target sequence) is measured in vitro, for example by thermophoresis, as described in, for example, Asmari et al. Methods 146:107-119 (2018) (incorporated herein by reference in its entirety).

[0530] In an embodiment, in the presence of, for example, about 100 times molar excess of out-of-order sequence competitor dsDNA, the DNA binding domain is able to bind its target sequence (e.g., dsDNA target sequence) with an affinity, for example, between 100 pM and 10 nM (e.g., between 100 pM and 1 nM or between 1 nM and 10 nM).

[0531] In some embodiments, the DNA-binding domain is found to be associated with its target sequence (e.g., a dsDNA target sequence) more frequently than any other sequence in the genome of the target cell (e.g., a human target cell), as measured by ChIP-seq (e.g., in HEK293T cells), as described in He and Pu (2010) Curr. Protoc Mol Biol [Latest Protocols in Molecular Biology] Chapter 21 (which is incorporated herein by reference in its entirety). In some embodiments, the DNA-binding domain is found to be associated with its target sequence (e.g., a dsDNA target sequence) at a frequency at least about 5 to 10 times more frequently than any other sequence in the genome of the target cell, as measured by ChIP-seq (e.g., in HEK293T cells), as described in He and Pu (2010), as above.

[0532] In some embodiments, the endonuclease domain has nicking enzyme activity and cleaves one strand of the target DNA. In some embodiments, the nicking enzyme activity reduces the formation of double-strand breaks at the target site. In some embodiments, the endonuclease domain creates staggered nicking structures in the first and second strands of the target DNA. In some embodiments, the staggered nicking structures create free 3' overhangs at the target site. In some embodiments, free 3' overhangs at the target site improve editing efficiency, for example, by enhancing access to and annealing of the 3' homologous regions of the template nucleic acid. In some embodiments, the staggered nicking structures reduce the formation of double-strand breaks at the target site.

[0533] In some embodiments, the endonuclease domain cleaves both strands of the target DNA, for example, resulting in blunt-end cleavage of the target, and there are no ssDNA overhangs on either side of the cleavage site. The amino acid sequence of the endonuclease domain of the gene modification system described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domains described herein (e.g., the endonuclease domains in Table 8).

[0534] In some embodiments, the heterologous endonuclease is Fok1 or a functional fragment thereof. In some embodiments, the heterologous endonuclease is a Holliday ligase or a homolog thereof, such as the Holliday ligase from *Sulfolobus solfataricus*-Ssol Hje (Govindaraju et al., *NucleicAcids Research* 44:7, 2016). In some embodiments, the heterologous endonuclease is a large fragment of a spliceosomal protein such as Prp8 (Mahbub et al., *Mobile DNA* 8:16, 2017). In some embodiments, the heterologous endonuclease is derived from a CRISPR-related protein, such as Cas9. In some embodiments, the heterologous endonuclease is engineered to have only ssDNA cleavage activity, for example, only cleavage enzyme activity, such as a Cas9 cleavage enzyme, such as SpCas9 with D10A, H840A, or N863A mutations. Table 8 provides exemplary Cas proteins and mutations associated with nicking enzyme activity. In yet another embodiment, the homologous endonuclease domain is modified, for example, through site-specific mutations, to alter the DNA endonuclease activity. In yet another embodiment, the endonuclease domain is modified to reduce DNA sequence specificity, for example, by truncation to remove the domain that confers DNA sequence specificity or by mutation to inactivate the region that confers DNA sequence specificity.

[0535] In some embodiments, the endonuclease domain has nicking enzyme activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks at a higher frequency than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease substantially does not form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.

[0536] In some embodiments, the endonuclease domain has cleaving enzyme activity for cleaving target site DNA on a first strand; for example, in some embodiments, the endonuclease cleaves target sites of genomic DNA near alteration sites on the strand to which the writing domain is extended. In some embodiments, the endonuclease domain has cleaving enzyme activity for cleaving target site DNA on a first strand but not for cleaving target site DNA on a second strand. For example, when the polypeptide contains a CRISPR-associated endonuclease domain with cleaving enzyme activity, in some embodiments, the CRISPR-associated endonuclease domain cleaves target site DNA strands containing PAM sites (e.g., and does not cleave target site DNA strands without PAM sites). As a further example, when the polypeptide contains a CRISPR-associated endonuclease domain with cleaving enzyme activity, in some embodiments, the CRISPR-associated endonuclease domain cleaves target site DNA strands without PAM sites (e.g., and does not cleave target site DNA strands containing PAM sites).

[0537] In some other embodiments, the endonuclease domain has nicking enzyme activity that nicks the target DNA on both the first and second strands. Without being bound by theory, after the writing domain (e.g., the RT domain) of the polypeptide described herein is polymerized (e.g., reverse transcribed) from a heterologous object sequence of a template nucleic acid (e.g., template RNA), the cellular DNA repair mechanism must repair the nick on the first DNA strand. The target DNA now contains two distinct first DNA strand sequences: one corresponding to the original genomic DNA (e.g., with a free 5′ end), and the second corresponding to the one polymerized from the heterologous object sequence (e.g., with a free 3′ end). It is assumed that these two distinct sequences are mutually balanced, with the first hybridizing to the second strand, and then the other, and that the cellular DNA repair mechanism incorporating the sequence into the target site it repairs can be a random process. Without being bound by theory, it is assumed that introducing an additional nick into the second strand might cause the cellular DNA repair mechanism to favor the heterologous object sequence-based sequence more frequently than the original genomic sequence (Anzalone et al., Nature [Nature] 576:149-157(2019)). In some embodiments, additional nicks are located at at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides at 5' or 3' of a nick on the first strand.

[0538] Alternatively or otherwise, without being bound by theory, it is considered that additional nicks in the second strand can facilitate second-strand synthesis. In some embodiments, when a gene modification system has inserted into or replaced a portion of the first strand, it is necessary to synthesize a new sequence corresponding to the insertion / replacement in the second strand.

[0539] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain) and said domain cleaves a first strand and a second strand. For example, in such an embodiment, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs cleavage of the first strand and another gRNA spacer that directs cleavage of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, and a first endonuclease domain cleaves the first strand and a second endonuclease domain cleaves the second strand (optionally, the first endonuclease domain does not (e.g., cannot) cleave the second strand, and the second endonuclease domain does not (e.g., cannot) cleave the first strand).

[0540] In some embodiments, the endonuclease domain is capable of cleaving the first and second strands. In some embodiments, the first and second strand cleavages occur at the same location in the target site but on opposite strands. In some embodiments, the second strand cleavage occurs at an alternating location with the first cleavage, such as upstream or downstream. In some embodiments, if the second strand cleavage is upstream of the first strand cleavage, the endonuclease domain produces a target site deletion. In some embodiments, if the second strand cleavage is downstream of the first strand cleavage, the endonuclease domain produces a target site duplication. In some embodiments, if the first and second strand cleavages occur at the same location in the target site, the endonuclease domain does not produce duplications and / or deletions. In some embodiments, the endonuclease domain has altered activity depending on protein conformation or RNA binding state, which, for example, promotes cleavage of the first or second strand (e.g., as described by Christensen et al. in PNAS [Proceedings of the National Academy of Sciences] 2006; which is incorporated herein by reference in its entirety).

[0541] In some embodiments, the endonuclease domain comprises a wide range of nucleases or functional fragments thereof. In some embodiments, the endonuclease domain comprises homing endonucleases or functional fragments thereof. In some embodiments, the endonuclease domain comprises a wide range of nucleases from the LAGLIDADG, GIY-YIG, HNH, His-Cys box, or PD-(D / E)XK families, or functional fragments or variants thereof, for example, those functional fragments or variants having conserved amino acid motifs such as those indicated by the family name. In some embodiments, the endonuclease domain comprises a macronuclease or a fragment thereof, selected from, for example, I-SmaMI (Uniprot F7WD42), I-SceI (Uniprot P03882), I-AniI (Uniprot P03880), I-DmoI (Uniprot P21505), I-CreI (Uniprot P05725), I-TevI ​​(Uniprot P13299), I-OnuI (Uniprot Q4VWW5), or I-BmoI (Uniprot Q9ANR6). In some embodiments, the macronuclease, in its functional form, is a natural monomer, such as I-SceI or I-TevI, or a dimer, such as I-CreI. For example, LAGLIDADG macronucleases having a single copy of the LAGLIDADG motif typically form homodimers, while members having two copies of the LAGLIDADG motif are typically found as monomers. In some embodiments, a large-scale nuclease, typically formed in dimer form, is expressed as a fusion, for example, with two subunits expressed as a single ORF and optionally linked by a linker, such as the I-CreI dimer fusion (Rodriguez-Fornes et al. Gene Therapy 2020; incorporated herein by reference in its entirety). In some embodiments, the large-scale nuclease or a functional fragment thereof is modified to favor the cleavage enzyme activity of one strand of a double-stranded DNA molecule, for example, I-SceI (K122I and / or K223I) (Niu et al. J Mol Biol 2008), I-AniI (K227M) (McConnell Smith et al. PNAS 2009), and I-DmoI (Q42A and / or K120M) (Molina et al. J Biol Chem 2015). In some embodiments, a wide range of nucleases or functional fragments thereof having this preference for single-strand cleavage are used as endonuclease domains, for example, having cleavage enzyme activity.In some embodiments, the endonuclease domain comprises a wide range of nucleases or functional fragments thereof that are naturally targeted or engineered to target safe harbor sites, such as I-CreI targeting the SH6 site (Rodriguez-Fornes et al., ibid.). In some embodiments, the endonuclease domain comprises a wide range of nucleases or functional fragments thereof that have sequence-tolerant catalytic domains, such as I-TevI ​​recognizing the minimal motif CNNNG (Kleinstiver et al., PNAS [Proceedings of the National Academy of Sciences] 2012). In some embodiments, the target sequence tolerance catalytic domain is fused to the DNA binding domain, for example to guide activity, such as by fusing I-TevI ​​to: (i) a zinc finger to produce Tev-ZFE (Kleinstiver et al., PNAS [Proceedings of the National Academy of Sciences] 2012), (ii) other large-scale nucleases to produce MegaTevs (Wolfs et al., Nucleic Acids Res [Nucleic Acids Research] 2014), and / or (iii) Cas9 to produce TevCas9 (Wolfs et al., PNAS [Proceedings of the National Academy of Sciences] 2016).

[0542] In some embodiments, the endonuclease domain contains a restriction enzyme, such as an IIS-type or IIP-type restriction enzyme. In some embodiments, the endonuclease domain contains an IIS-type restriction enzyme, such as FokI, or a fragment or variant thereof. In some embodiments, the endonuclease domain contains an IIP-type restriction enzyme, such as PvuII, or a fragment or variant thereof. In some embodiments, the dimer restriction enzyme is expressed as a fusion, thereby functioning as a single strand, for example, the FokI dimer fusion (Minczuk et al. Nucleic Acids Res [Nucleic Acid Research] 36(12):3926-3938 (2008)).

[0543] For example, the use of additional endonuclease domains is described in Guha and Edgell Int J Mol Sci [International Journal of Molecular Sciences] 18(22):2565(2017), which is incorporated herein by reference in its entirety.

[0544] In some embodiments, the genetically modified polypeptide comprises a modification of the endonuclease domain, for example, relative to the wild-type Cas protein. In some embodiments, the endonuclease domain comprises the addition, deletion, substitution, or modification of the amino acid sequence of the wild-type Cas protein. In some embodiments, the endonuclease domain is modified to include a heterologous functional domain that specifically binds to and / or induces endonuclease cleavage of a target nucleic acid (e.g., DNA) sequence. In some embodiments, the endonuclease domain comprises a zinc finger. In embodiments, the endonuclease domain comprising the Cas domain is associated with, for example, a guide RNA (gRNA) as described herein. In some embodiments, the endonuclease domain is modified to include a functional domain that does not target a specific target nucleic acid (e.g., DNA) sequence. In embodiments, the endonuclease domain comprises a Fok1 domain.

[0545] In some embodiments, the endonuclease domain is associated with the target dsDNA in vitro at a frequency at least about 5 to 10 times higher than that of the scrambled dsDNA. In some embodiments, the endonuclease domain is associated with the target dsDNA in vitro at a frequency at least about 5 to 10 times higher than that of the scrambled dsDNA, for example in cells (e.g., HEK293T cells). In some embodiments, the association frequency between the endonuclease domain and the target DNA or the scrambled DNA is measured by ChIP-seq, for example as described in Chapter 21 of He and Pu (2010) Curr. Protoc Mol Biol [Latest Protocols in Molecular Biology] (which is incorporated herein by reference in its entirety).

[0546] In some embodiments, the endonuclease domain can catalyze the formation of a nick at the target sequence, for example, by at least about 5-fold or 10-fold relative to a non-target sequence (e.g., relative to any other genomic sequence in the target cell genome). In some embodiments, the level of nick formation is determined using NickSeq, for example, as described in Elacqua et al. (2019) bioRxivdoi.org / 10.1101 / 867937 (which is incorporated herein by reference in its entirety).

[0547] In some embodiments, the endonuclease domain is capable of cleaving DNA in vitro. In embodiments, cleavage results in exposed bases. In embodiments, the exposed bases can be detected using a nuclease sensitivity assay, for example, as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res [Nucleic Acids Research] 23(19):3805-3809 (which is incorporated herein by reference in its entirety). In embodiments, the level of exposed bases (e.g., detected by a nuclease sensitivity assay) is increased by at least 10%, 50%, or more relative to a reference endonuclease domain. In some embodiments, the reference endonuclease domain is the endonuclease domain of Cas9 from Streptococcus pyogenes.

[0548] In some embodiments, the endonuclease domain is capable of cleaving DNA in cells. In an embodiment, the endonuclease domain is capable of cleaving DNA in HEK293T cells. In an embodiment, unrepaired cleavage undergoing replication in the absence of Rad51 results in an increased NHEJ rate at the cleavage site, which can be detected, for example, by using a Rad51 inhibition assay, as described, for example, as Bothmer et al. (2017) Nat Commun [Nature Communications] 8:13905 (which is incorporated herein by reference in its entirety). In an embodiment, the NHEJ rate increases to more than 0-5%. In an embodiment, for example, after Rad51 inhibition, the NHEJ rate increases to 20%-70% (e.g., in 30%-60% or 40%-50%).

[0549] In some embodiments, the endonuclease domain releases the target upon cleavage. In some embodiments, target release is indirectly indicated by assessing the number of enzyme turnovers, for example, as described in Yourik et al. RNA 25(1):35-44 (2019) (which is incorporated herein by reference in its entirety) and as... Figure 2 As shown. In some embodiments, the k of the nuclease domain, as measured by such a method, exp 1 x 10 -3 – 1 x 10 - 5 min-1.

[0550] In some embodiments, the endonuclease domain has a size greater than about 1 x 10⁻⁶ in vitro. 8 s -1 M -1 catalytic efficiency (k cat / K m In the embodiments, the endonuclease domain has a size greater than approximately 1 x 10⁻⁶ in vitro. 5 1 x 10 61 x 10 7 Or 1 x 10 8 s -1 M -1 The catalytic efficiency. In the examples, the catalytic efficiency was determined as described by Chen et al. (2018) Science [Science] 360(6387):436-439 (which is incorporated herein by reference in its entirety). In some embodiments, the endonuclease domain has a size greater than about 1 x 10^6 ppm in the cell. 8 s -1 M -1 catalytic efficiency (k cat / K m In the examples, the catalytic efficiency of the endonuclease domain in cells is greater than about 1 x 10⁻⁶. 5 1 x 10 6 1 x 10 7 Or 1 x 10 8 s -1 M -1 .

[0551] Gene-modified peptides containing Cas domains

[0552] In some embodiments, the gene-modifying peptides described herein comprise a Cas domain. In some embodiments, the Cas domain can guide the gene-modifying peptide to a target site specified by a gRNA spacer, thereby “cis” modifying the target nucleic acid sequence. In some embodiments, the gene-modifying peptide is fused to a Cas domain. In some embodiments, the gene-modifying peptide comprises a CRISPR / Cas domain (also referred to herein as a CRISPR-associated protein). In some embodiments, the CRISPR / Cas domain comprises a protein (e.g., a Cas protein) involved in clustering of a regulatory short palindromic repeat (CRISPR) system and optionally binds a guide RNA, such as a single guide RNA (sgRNA).

[0553] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases (e.g., Cas9 or Cpf1) called CRISPR-associated or “Cas” endonucleases to cut foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is directed to the target nucleotide sequence (e.g., the site in the genome to be edited) via a sequence-specific non-coding “guide RNA” that targets a single-stranded or double-stranded DNA sequence. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (instead of multiple Cas proteins). One type of Class II CRISPR system includes a type II Cas endonuclease, such as Cas9, CRISPR RNA (“crRNA”), and trans-activating crRNA (“tracrRNA”). The crRNA contains a “spacer” sequence, which is an RNA sequence of about 20 nucleotides that typically corresponds to the target DNA sequence (“protospacer”). In wild-type systems and some engineered systems, crRNA also contains a region that binds to tracrRNA to form a partially double-stranded structure that is cleaved by RNase III, producing a crRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then directs a Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is typically adjacent to a "protospacer adjacent motif" ("PAM"), which is specific to a given Cas endonuclease and is required for cleavage activity at the target site that matches the crRNA spacer. CRISPR endonucleases identified from different prokaryotic species have unique PAM sequence requirements, for example, as listed in Table 7 for exemplary Cas enzymes; examples of PAM sequences include 5´-NGG (Streptococcus pyogenes; SEQ ID NO: 11,019), 5´-NNAGAA (Streptococcus thermophilus CRISPR1; SEQ ID NO: 11,020), 5´-NGGNG (Streptococcus thermophilus CRISPR3; SEQ ID NO: 11,021), and 5´-NNNGATT (Neisseriameningiditis; SEQ ID NO: 11,022). Some endonucleases, such as Cas9 endonuclease, are associated with G-rich PAM sites (e.g., 5'-NGG (SEQ ID NO: 11,023)) and perform blunt-end cleavage of the target DNA at a position 3 nucleotides upstream of the PAM site (the 5' of the site).Another class II CRISPR system includes the type V endonuclease Cpf1, which is smaller than Cas9; examples include AsCpf1 (from the genus *Acidaminococcus* sp.) and LbCpf1 (from the genus *Lachnospiraceae* sp.). Cpf1-associated CRISPR arrays are processed into mature crRNA without the need for tracrRNA; in other words, in some embodiments, the Cpf1 system contains only the Cpf1 endonuclease and crRNA to cleave the target DNA sequence. The Cpf1 endonuclease is typically associated with T-rich PAM sites such as 5'-TTN. Cpf1 can also recognize the 5'-CTA PAM motif. Cpf1 typically cleaves target DNA by introducing misaligned or staggered double-strand breaks with 4 or 5 nucleotides of 5' protrusions, for example, cleaving target DNA in which the 5-nucleotide misaligned or staggered cut is located 18 nucleotides downstream (3') of the PAM site on the coding strand and 23 nucleotides downstream of the PAM site on the complementary strand; the 5-nucleotide protrusions resulting from such misaligned cuts allow for more precise genome editing via DNA insertion through homologous recombination than with DNA inserted at blunt ends. See, for example, Zetsche et al. (2015) Cell, 163:759-771.

[0554] Various CRISPR-related (Cas) genes or proteins can be used in the techniques provided in this disclosure, and the selection of the Cas protein will depend on the specific conditions of the method. Specific examples of Cas proteins include class II systems, including Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein (e.g., Cas9 protein) can be derived from any of a variety of prokaryotic species. In some embodiments, a specific Cas protein (e.g., a specific Cas9 protein) is selected to recognize a specific protospacer neighbor motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain includes the sequence of a targeting polypeptide (e.g., a Cas protein, such as Cas9). In some embodiments, the Cas protein (e.g., Cas9 protein) can be obtained from bacteria or archaea or synthesized using known methods. In some embodiments, the Cas protein can be derived from Gram-positive or Gram-negative bacteria. In some embodiments, the Cas protein may be derived from Streptococcus (e.g., Streptococcus pyogenes or Streptococcus thermophilus), Francisella (e.g., Francisella neoculi), Staphylococcus (e.g., Staphylococcus aureus), Aminococcus (e.g., Aminococcus species BV3L6), Neisseria (e.g., Neisseria meningitidis), Cryptococcus, Corynebacterium, Haemophilus, Eubacterium, Pasteurella, Prevotella, Veillonella, or Marinebacterium.

[0555] In some embodiments, the genetically modified polypeptide may comprise the amino acid sequence of SEQ ID NO: 4000, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In an embodiment, the amino acid sequence of SEQ ID NO: 4000, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, is located at the N-terminus of the genetically modified polypeptide. In an embodiment, the amino acid sequence of SEQ ID NO: 4000, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids at the N-terminus of the genetically modified polypeptide.

[0556] Exemplary N-terminal NLS-Cas9 domain

[0557]

[0558] In some embodiments, the genetically modified polypeptide may comprise the amino acid sequence of SEQ ID NO: 4001, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In an embodiment, the amino acid sequence of SEQ ID NO: 4001, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, is located at the C-terminus of the genetically modified polypeptide. In an embodiment, the amino acid sequence of SEQ ID NO: 4001, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, is located within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, or 30 amino acids at the C-terminus of the genetically modified polypeptide.

[0559] Exemplary C-terminal sequence containing NLS

[0560] AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001)

[0561] Exemplary benchmark sequence

[0562]

[0563] In some embodiments, the genetically modified polypeptide may comprise a Cas domain or a functional fragment thereof as listed in Table 7 or 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0564] Table 7. CRISPR / Cas proteins, species, and mutations

[0565]

[0566] Table 8. Amino acid sequence, species, and mutations of the St1Cas9 protein.

[0567]

[0568]

[0569]

[0570]

[0571] In some embodiments, the Cas protein requires a protospacer neighbor motif (PAM) to be present in or adjacent to the target DNA sequence for the Cas protein to bind and / or function. In some embodiments, the PAM from 5′ to 3′ is or comprises NGG (SEQ ID NO: 11,024), YG (SEQ ID NO: 11,025), NNGRRT (SEQ ID NO: 11,026), NNNRRT (SEQ ID NO: 11,027), NGA (SEQ ID NO: 11,029), TYCV (SEQ ID NO: 11,030), TATV (SEQ ID NO: 11,031), NTTN (SEQ ID NO: 11,032), or NNNGATT (SEQ ID NO: 11,033), where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A, C, or G. In some embodiments, the Cas protein is a protein listed in Table 7 or 8. In some embodiments, the Cas protein comprises one or more mutations that alter its PAM. In some embodiments, the Cas protein contains E1369R, E1449H, and R1556A mutations or similar substitutions of amino acids corresponding to the positions described. In some embodiments, the Cas protein contains E782K, N968K, and R1015H mutations or similar substitutions of amino acids corresponding to the positions described. In some embodiments, the Cas protein contains D1135V, R1335Q, and T1337R mutations or similar substitutions of amino acids corresponding to the positions described. In some embodiments, the Cas protein contains S542R and K607R mutations or similar substitutions of amino acids corresponding to the positions described. In some embodiments, the Cas protein contains S542R, K548V, and N552R mutations or similar substitutions of amino acids corresponding to the positions described. Exemplary advances in engineering Cas enzymes to recognize altered PAM sequences are reviewed in Collias et al., Nature Communications, 12:555 (2021), which is incorporated herein by reference in its entirety.

[0572] In some embodiments, the Cas protein has catalytic activity and cleaves one or both strands of the target DNA site. In some embodiments, cleavage of the target DNA site leads to alterations, such as insertions or deletions, for example, through cellular repair mechanisms.

[0573] In some embodiments, the Cas protein is modified to inactivate or partially inactivate a nuclease, such as a nuclease-deficient Cas9. Wild-type Cas9 produces double-strand breaks (DSBs) at specific DNA sequences targeted by gRNA, and many functionally modified CRISPR endonucleases are available, such as partially inactivated Cas9 “nickase” versions that produce only single-strand breaks; and catalytically inactive Cas9 (“dCas9”) that does not cleave the target DNA. In some embodiments, the binding of dCas9 to the DNA sequence can interfere with transcription at that site through steric hindrance. In some embodiments, the binding of dCas9 to the anchoring sequence can interfere with (e.g., reduce or prevent) the formation and / or maintenance of genomic complexes (e.g., ASMCs). In some embodiments, the DNA-binding domain comprises a catalytically inactive Cas9, such as dCas9. Many catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 comprises mutations in each endonuclease domain of the Cas protein, such as D10A and H840A or N863A mutations. In some embodiments, a catalytically inactive or partially catalytically inactive CRISPR / Cas domain comprises a Cas protein containing one or more mutations, such as those listed in Table 7. In some embodiments, a Cas protein described in a given row of Table 7 contains one, two, three, or all of the mutations listed in the same row of Table 7. In some embodiments, a Cas protein not described in Table 7 contains one, two, three, or all of the mutations listed in the rows of Table 7 or a corresponding mutation at the corresponding site in the Cas protein.

[0574] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain a D11 mutation (e.g., a D11A mutation) or a similar substitution of the amino acid corresponding to that position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an H969 mutation (e.g., an H969A mutation) or a similar substitution of the amino acid corresponding to that position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, or partially inactivated Cas9 proteins, contain an N995 mutation (e.g., an N995A mutation) or a similar substitution of the amino acid corresponding to that position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, contain mutations at one, two, or three of positions D11, H969, and N995 (e.g., D11A, H969A, and N995A mutations) or similar substitutions of the amino acid corresponding to those positions.

[0575] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain a D10 mutation (e.g., a D10A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an H557 mutation (e.g., an H557A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, contain a D10 mutation (e.g., a D10A mutation) and an H557 mutation (e.g., an H557A mutation) or a similar substitution of the amino acid corresponding to the stated position.

[0576] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain a D839 mutation (e.g., a D839A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an H840 mutation (e.g., an H840A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an N863 mutation (e.g., an N863A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, contain a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), an H840 mutation (e.g., H840A), and an N863 mutation (e.g., N863A) or a similar substitution of the amino acid corresponding to the stated position.

[0577] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an E993 mutation (e.g., an E993A mutation) or a similar substitution of an amino acid corresponding to the position.

[0578] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain a D917 mutation (e.g., a D917A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an E1006 mutation (e.g., an E1006A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain a D1255 mutation (e.g., a D1255A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, contain a D917 mutation (e.g., D917A), an E1006 mutation (e.g., E1006A), and a D1255 mutation (e.g., D1255A) or a similar substitution of the amino acid corresponding to the stated position.

[0579] In some embodiments, the catalytically inactive Cas9 protein, such as dCas9, or the partially inactivated Cas9 protein contains a D16 mutation (e.g., a D16A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, the catalytically inactive Cas9 protein, such as dCas9, or the partially inactivated Cas9 protein contains a D587 mutation (e.g., a D587A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, the partially inactivated Cas domain has nicking enzyme activity. In some embodiments, the partially inactivated Cas9 domain is a Cas9 nicking enzyme domain. In some embodiments, the catalytically inactive Cas domain or the inactivated Cas domain does not produce detectable double-strand breaks. In some embodiments, the catalytically inactive Cas9 protein, such as dCas9, or the partially inactivated Cas9 protein contains an H588 mutation (e.g., an H588A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, the catalytically inactive Cas9 protein, such as dCas9, or the partially inactivated Cas9 protein contains an N611 mutation (e.g., an N611A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, non-catalytically active Cas9 proteins, such as dCas9, contain D16 mutations (e.g., D16A), D587 mutations (e.g., D587A), H588 mutations (e.g., H588A), and N611 mutations (e.g., N611A) or similar substitutions of amino acids corresponding to the positions described.

[0580] In some embodiments, the DNA-binding domain or endonuclease domain may contain a Cas molecule that contains or is (e.g., covalently) linked to a gRNA (e.g., a template nucleic acid, such as a template RNA containing gRNA).

[0581] In some embodiments, the endonuclease domain or DNA-binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises modified SpCas9. In an embodiment, the modified SpCas9 comprises a modification that alters the specificity of the protospacer neighbor motif (PAM). In an embodiment, the PAM is specific for the nucleic acid sequence 5′-NGT-3′. In an embodiment, the modified SpCas9 comprises, for example, one or more amino acid substitutions at one or more of positions L1111, D1135, G1218, E1219, A1322, or R1335, for example, the one or more amino acid substitutions being selected from L1111R, D1135V, G1218R, E1219F, A1322R, and R1335V. In the embodiments, the modified SpCas9 comprises an amino acid substitution T1337R and one or more additional amino acid substitutions, for example, the one or more additional amino acid substitutions are selected from L1111, D1135L, S1136R, G1218S, E1219V, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, R1335Q, T1337, T1337L, T1337Q, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337H, T1337Q, and T1337M, or their corresponding amino acid substitutions. In the embodiments, the modified SpCas9 comprises: (i) one or more amino acid substitutions selected from D1135L, S1136R, G1218S, E1219V, A1322R, R1335Q, and T1337; and (ii) one or more amino acid substitutions selected from L1111R, G1218R, E1219F, D1332A, D1332S, D1332T, D1332V, D1332L, D1332K, D1332R, T1337L, T1337I, T1337V, T1337F, T1337S, T1337N, T1337K, T1337R, T1337H, T1337Q, and T1337M, or their corresponding amino acid substitutions.

[0582] In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas domain, such as a Cas9 domain. In embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas domain, a Cas nicking enzyme (nCas) domain, or a nuclease-free Cas (dCas) domain. In embodiments, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nicking enzyme (nCas9) domain, or a nuclease-free Cas9 (dCas9) domain. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 domain (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises Cas9 (e.g., dCas9 and nCas9), Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises Streptococcus pyogenes or Streptococcus thermophilus Cas9, or a functional fragment thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 sequence, for example, as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737; this literature is incorporated herein by reference. In some embodiments, the endonuclease domain or DNA-binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of Cas, for example, Cas9 as described herein, or a variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises Cas12a / Cpfl, Cas12b / C2cl, Cas12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, or Cas12i. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof.In the examples, the Cas polypeptide (e.g., enzyme) is selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5d, Cas5t, Cas5h, Cas5a, Cas6, Cas7, Cas8, Cas8a, Cas8b, Cas8c, Cas9 (e.g., Csn1 or Csx12), Cas10, Cas10d, Cas12a / Cpfl, Cas12b / C2cl, Cas 12c / C2c3, Cas12d / CasY, Cas12e / CasX, Cas12g, Cas12h, Cas12i, Csy1, Csy2, Csy3, Csy4, Cse1, Cse2, Cse3, Cse4, Cse5e, Csc1, Csc2, Csa5, Csn1, Csn2, Csm1, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr 4. Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csx11, Csf1, Csf2, CsO, Csf4, Csd1, Csd2, Cst1, Cst2, Csh1, Csh2, Csa1, Csa2, Csa3, Csa4, Csa5, type II Cas effector proteins, type V C as effector proteins, type VI Cas effector proteins, CARF, DinG, Cpf1, Cas12b / C2c1, Cas12c / C2c3, Cas12b / C2c1, Cas12c / C2c3, SpCas9 (K855A), eSpCas9 (1.1), SpCas9-HF1, the ultra-precise Cas9 variant (HypaCas9), its homologs, its modified or engineered versions, and / or fragments thereof. In embodiments, Cas9 comprises one or more substitutions selected, for example, from H840A, D10A, P475A, W476A, N477A, ​​D1125A, W1126A, and D1127A. In an embodiment, Cas9 includes one or more mutations selected from the following positions: D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987, for example, one or more substitutions selected from D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A.In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas (e.g., Cas9) sequence or a fragment or variant thereof from: Corynebacterium ulcerans, Corynebacterium diphtheriae, Spiroplasma syrphidicola, Prevotella intermedia, Spiroplasma taiwanensis, Streptococcus iniae, Belliella baltica, Psychroflexus torquis, Streptococcus thermophilus, Listeria innocua, Campylobacter jejuni, Neisseria meningitidis, Streptococcus pyogenes, or Staphylococcus aureus.

[0583] In some embodiments, the endonuclease domain or DNA-binding domain comprises, for example, a Cpf1 domain containing one or more substitutions (e.g., at positions D917, E1006A, D1255) or any combination thereof, wherein the one or more substitutions are selected, for example, from D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, and D917A / E1006A / D1255A.

[0584] In some embodiments, the endonuclease domain or DNA-binding domain comprises spCas9, spCas9-VRQR (SEQ ID NO: 5019), spCas9-VRER (SEQ ID NO: 5020), xCas9 (sp), saCas9, saCas9-KKH, spCas9-MQKSER (SEQ ID NO: 5021), spCas9-LRKIQK (SEQ ID NO: 5022), or spCas9-LRVSQL (SEQ ID NO: 5023).

[0585] In some embodiments, the genetically modified polypeptide has an endonuclease domain comprising a Cas9 nickase, such as Cas9 H840A. In an embodiment, Cas9 H840A has the following amino acid sequence:

[0586] Cas9 cleavage enzyme (H840A):

[0587]

[0588] In some embodiments, the genetically modified polypeptide comprises a dCas9 sequence containing a D10A and / or H840A mutation, for example, the following sequence:

[0589]

[0590] TAL effectors and zinc finger nucleases

[0591] In some embodiments, the endonuclease domain or DNA-binding domain comprises a TAL effector molecule. A TAL effector molecule, such as a TAL effector molecule that specifically binds to a DNA sequence, typically comprises multiple TAL effector domains or fragments thereof, and optionally one or more additional portions of a naturally occurring TAL effector (e.g., the N- and / or C-termini of the multiple TAL effector domains). Many TAL effectors are known to those skilled in the art and are commercially available, for example, from Thermo Fisher Scientific.

[0592] Naturally occurring TAL effector proteins are natural effector proteins secreted by various bacterial pathogens, including the plant pathogen Xanthomonas. They regulate gene expression in host plants and promote bacterial colonization and survival. The specific binding of TAL effectors is based on a central repeating domain (repeated variable two residue, RVD domain) of a tandemly arranged, nearly identical 33 or 34 amino acid repeating sequences.

[0593] Members of the TAL effector family differ primarily in the number and order of their repeat sequences. The number of repeat sequences typically ranges from 1.5 to 33.5, and C-terminal repeats are usually shorter (e.g., about 20 amino acids) and are often referred to as “half-repeaters.” Each repeat sequence in a TAL effector typically has a one-to-one base pair association, with different repeat types exhibiting different base pair specificities (one repeat sequence recognizes one base pair on the target gene sequence). Generally, the fewer the number of repeat sequences, the weaker the protein-DNA interaction. A number of 6.5 repeat sequences has been shown to be sufficient to activate reporter gene transcription (Scholze et al., 2010).

[0594] Variations from repetitive sequence to repetitive sequence mainly occur at amino acid positions 12 and 13, hence they are called "highly variable" and are responsible for the specificity of interaction with the target DNA promoter sequence, as shown in Table 9, which lists exemplary repetitive sequence variable double residues (RVD) and their correspondence with nucleic acid base targets.

[0595] Table 9 - RVD and Nucleic Acid Base Specificity

[0596]

[0597] Therefore, it is possible to modify the repeating sequences of TAL effectors to target specific DNA sequences. Further research has shown that RVD NK can target G. TAL effector target sites also tend to include T on the 5' flanking position of the first repeating sequence, but the exact mechanism of this recognition remains unclear. More than 113 TAL effector sequences are known to date. Non-restrictive examples of TAL effectors from the genus Xanthomonas include Hax2, Hax3, Hax4, AvrXa7, AvrXa10, and AvrBs3.

[0598] Accordingly, the TAL effector domain of the TAL effector molecule described herein can be derived from TAL effectors of any bacterial species (e.g., Xanthomonas species, such as African strains of Xanthomonas oryzae pv. Oryzae (Yu et al. 2011), Xanthomonas campestris pv. raphani strain 756C, and Xanthomonas oryzae pv. Oryzicola strain BLS256 (Bogdanove et al. 2011)). In some embodiments, the TAL effector domain comprises an RVD domain and one or more flanking sequences (sequences on the N-terminus and / or C-terminus sides of the RVD domain) also derived from naturally occurring TAL effectors. It can contain more or fewer repetitive sequences than the RVD of naturally occurring TAL effectors. TAL effector molecules can be programmed to target a given DNA sequence based on the above-described coding and other codings known in the art. The number and specific sequence of TAL effector domains (e.g., repetitive sequences (monomers or modules)) can be selected based on the desired DNA target sequence. For example, TAL effector domains, such as repetitive sequences, can be removed or added to accommodate a specific target sequence. In one embodiment, the TAL effector molecule of the present invention comprises 6.5 to 33.5 TAL effector domains, such as repetitive sequences. In another embodiment, the TAL effector molecule of the present invention comprises 8 to 33.5 TAL effector domains, such as repetitive sequences, for example, 10 to 25 TAL effector domains, such as repetitive sequences, for example, 10 to 14 TAL effector domains, such as repetitive sequences.

[0599] In some embodiments, the TAL effector molecule comprises a TAL effector domain corresponding to a perfect match with the DNA target sequence. In some embodiments, mismatches between repetitive sequences and target base pairs on the DNA target sequence are permitted, as long as they allow the function of the polypeptide containing the TAL effector molecule. Typically, TALE binding is negatively correlated with the number of mismatches. In some embodiments, the TAL effector molecule of the polypeptide of the present invention contains no more than 7 mismatches, 6 mismatches, 5 mismatches, 4 mismatches, 3 mismatches, 2 mismatches, or 1 mismatch with the target DNA sequence, and optionally no mismatches. Not wishing to be bound by theory, generally, the fewer the number of TAL effector domains in the TAL effector molecule, the fewer mismatches will be allowed, while still allowing the function of the polypeptide containing the TAL effector molecule. Binding affinity is considered to depend on the sum of the matching repetitive-DNA combinations. For example, a TAL effector molecule having 25 or more TAL effector domains may be able to tolerate up to 7 mismatches.

[0600] In addition to the TAL effector domain, the TAL effector molecule of the present invention may also contain additional sequences derived from naturally occurring TAL effectors. The lengths of one or more C-terminal and / or N-terminal sequences contained on each side of the TAL effector domain portion of the TAL effector molecule can vary and are chosen by those skilled in the art, for example based on the work of Zhang et al. (2011). Zhang et al. have characterized numerous C-terminal and N-terminal truncated mutants in Hax3-derived TAL effector-based proteins and have identified key elements that contribute to optimal binding to target sequences and thus activation of transcription. Typically, transcriptional activity has been found to be negatively correlated with the length of the N-terminus. Regarding the C-terminus, important elements of DNA-binding residues within the first 68 amino acids of the Hax3 sequence have been identified. Therefore, in some embodiments, the first 68 amino acids on the C-terminal side of the TAL effector domain of a naturally occurring TAL effector are included in the TAL effector molecule. Therefore, in the embodiments, the TAL effector molecule comprises 1) one or more TAL effector domains derived from naturally occurring TAL effectors; 2) at least 70, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260, 270, 280 or more amino acids derived from naturally occurring TAL effectors on the N-terminal side of the TAL effector domain; and / or 3) at least 68, 80, 90, 100, 110, 120, 130, 140, 150, 170, 180, 190, 200, 220, 230, 240, 250, 260 or more amino acids derived from naturally occurring TAL effectors on the C-terminal side of the TAL effector domain.

[0601] In some embodiments, the endonuclease domain or DNA-binding domain is or comprises a Zn finger molecule. The Zn finger molecule comprises a Zn finger protein, such as a naturally occurring Zn finger protein or an engineered Zn finger protein, or a fragment thereof. Many Zn finger proteins are known to those skilled in the art and are commercially available, for example, from Sigma-Aldrich.

[0602] In some embodiments, the Zn finger molecule comprises a non-naturally occurring Zn finger protein that is engineered to bind to a selected target DNA sequence. See, for example, Beerli et al. (2002) *Nature Biotechnol* 20:135-141; Pabo et al. (2001) *Ann. Rev. Biochem* 70:313-340; Isalan et al. (2001) *Nature Biotechnol* 19:656-660; Segal et al. (2001) *Curr. Opin. Biotechnol* 12:632-637; Choo et al. (2000) *Curr. Opin. Struct. Biol* 12:632-637; and Choo et al. (2000) *Curr. Opin. Struct. Biol* 13:635-141. U.S. Patent Nos. 10:411-416; U.S. Patent Nos. 6,453,242, 6,534,261, 6,599,692, 6,503,717, 6,689,558, 7,030,215, 6,794,136, 7,067,317, 7,262,054, 7,070,934, 7,361,635, 7,253,273; and U.S. Patent Publications 2005 / 0064474, 2007 / 0218528, and 2005 / 0267061, all of which are incorporated herein by reference in their entirety.

[0603] Engineered Zn finger proteins may exhibit novel binding specificities compared to naturally occurring Zn finger proteins. Engineering methods include, but are not limited to, rational design and various types of selection. Rational design includes, for example, using a database containing triplet (or tetrad) nucleotide sequences and individual Zn finger amino acid sequences, wherein each triplet or tetrad nucleotide sequence is associated with one or more amino acid sequences of a zinc finger that binds to a specific triplet or tetrad sequence. See, for example, U.S. Patent Nos. 6,453,242 and 6,534,261, which are incorporated herein by reference in their entirety.

[0604] Exemplary selection methods (including phage display and two-hybrid systems) are disclosed in the following: U.S. Patent Nos. 5,789,538, 5,925,523, 6,007,988, 6,013,453, 6,410,248, 6,140,466, 6,200,759, and 6,242,568; and International Patent Publications WO 98 / 37186, WO 98 / 53057, WO 00 / 27878, WO 01 / 88197, and GB 2,338,237. Additionally, enhancing the binding specificity of zinc finger proteins has been described, for example, in International Patent Publication WO 02 / 077227.

[0605] Additionally, as disclosed in these and other references, zinc finger domains and / or multifinite zinc finger proteins can be linked together using any suitable linker sequence, including, for example, linkers of 5 or more amino acids in length. See also U.S. Patent Nos. 6,479,626; 6,903,185, and 7,153,949 for exemplary linker sequences of 6 or more amino acids in length. Proteins described herein can include any combination of suitable linkers between individual zinc fingers of the protein. Furthermore, methods to enhance the binding specificity of zinc finger binding domains have been described, for example, in co-owned International Patent Publication No. WO 02 / 077227.

[0606] Zn finger proteins and methods for designing and constructing fusion proteins (and polynucleotides encoding them) are known to those skilled in the art and are described in detail below: U.S. Patent Nos. 6,140,0815, 789,538, 6,453,242, 6,534,261, 5,925,523, 6,007,988, 6,013,453, and 6,200,759; International Patent Publications Nos. WO 95 / 19431, WO96 / 06166, WO 98 / 53057, WO 98 / 54311, WO 00 / 27878, WO 01 / 60970, WO 01 / 88197, WO 02 / 099084, WO 98 / 53058, WO 98 / 53059, WO 98 / 53060, WO 02 / 016536, and WO 03 / 016496.

[0607] Additionally, as disclosed in these and other references, Zn finger proteins and / or polyfinnous Zn finger proteins can be linked together using any suitable linker sequence, including, for example, linkers of 5 or more amino acids in length, as fusion proteins. See also U.S. Patent Nos. 6,479,626; 6,903,185, and 7,153,949 for exemplary linker sequences of 6 or more amino acids in length. The Zn finger molecules described herein may include any combination of suitable linkers between individual zinc finger proteins and / or polyfinnous Zn finger proteins.

[0608] In some embodiments, the DNA-binding domain or endonuclease domain comprises a Zn finger molecule containing an engineered zinc finger protein that binds to a target DNA sequence (in a sequence-specific manner). In some embodiments, the Zn finger molecule comprises one Zn finger protein or a fragment thereof. In other embodiments, the Zn finger molecule comprises multiple Zn finger proteins (or fragments thereof), such as 2, 3, 4, 5, 6 or more Zn finger proteins (and optionally, no more than 12, 11, 10, 9, 8, 7, 6, 5, 4, 3 or 2 Zn finger proteins). In some embodiments, the Zn finger molecule comprises at least three Zn finger proteins. In some embodiments, the Zn finger molecule comprises four, five or six fingers. In some embodiments, the Zn finger molecule comprises 8, 9, 10, 11 or 12 fingers. In some embodiments, a Zn finger molecule comprising three Zn finger proteins recognizes a target DNA sequence comprising 9 or 10 nucleotides. In some embodiments, a Zn finger molecule comprising four Zn finger proteins recognizes a target DNA sequence comprising 12 to 14 nucleotides. In some embodiments, a Zn finger molecule containing six Zn finger proteins recognizes a target DNA sequence containing 18 to 21 nucleotides.

[0609] In some embodiments, the Zn finger molecule comprises biphasic Zn finger proteins. Biphasic zinc finger proteins are proteins in which two clusters of zinc finger proteins are separated by intercalated amino acids, such that two zinc finger domains bind to two discontinuous target DNA sequences. An example of a biphasic zinc finger binding protein is SIP1, in which clusters of four zinc finger proteins are located at the amino terminus of the protein and clusters of three Zn finger proteins are located at the carboxyl terminus (see Remade et al. (1999) EMBO Journal [European Journal of Molecular Biology] 18(18):5073-5084). Each cluster of zinc fingers in these proteins is capable of binding to a unique target sequence, and the space between the two target sequences can contain a number of nucleotides.

[0610] connector

[0611] In some embodiments, the genetically modified polypeptide may include a linker, such as a peptide linker, as described in Table 10. In some embodiments, the genetically modified polypeptide includes a Cas domain (e.g., the Cas domain of Table 8), a linker of Table 10 (or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with it), and an RT domain (e.g., the RT domain of Table 6) in the N-terminal to C-terminal direction. In some embodiments, the genetically modified polypeptide includes a flexible linker between the endonuclease and the RT domain, for example, a linker containing the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 11,002). In some embodiments, the RT domain of the genetically modified polypeptide may be located at the C-terminus of the endonuclease domain. In some embodiments, the RT domain of the genetically modified polypeptide may be located at the N-terminus of the endonuclease domain.

[0612] Table 10 Exemplary Connector Sequences

[0613]

[0614]

[0615]

[0616]

[0617] In some embodiments, the linker of the genetically modified polypeptide comprises a motif selected from the group consisting of (SGGS). n (SEQ ID NO: 5025), (GGGS) n (SEQ ID NO: 5026), (GGGGS) n (SEQ ID NO: 5027), (G) n (EAAAK) n (SEQ ID NO: 5028), (GGS) n Or (XP) n .

[0618] Selecting genetically modified peptides through pooling screening

[0619] Candidate gene-editing peptides can be screened to evaluate their gene-editing capabilities. For example, RNA gene-editing systems designed for targeted editing of coding sequences in the human genome can be used. In some embodiments, such gene-editing systems can be used in conjunction with pooling screening methods.

[0620] For example, a gene-modified peptide candidate library and template guide RNA (tgRNA) can be introduced into mammalian cells to test the gene-editing capabilities of the candidates using a pooling screening method. In a particular embodiment, a gene-modified peptide candidate library is introduced into mammalian cells, followed by the introduction of tgRNA into the cells.

[0621] Representative, non-limiting examples of mammalian cells that can be used for screening include HEK293T cells, U2OS cells, HeLa cells, HepG2 cells, Huh7 cells, K562 cells, or iPS cells.

[0622] Genetically modified peptide candidates may include 1) a Cas nuclease, such as a wild-type Cas nuclease (e.g., wild-type Cas9 nuclease), a mutant Cas nuclease (e.g., a Cas nickase, such as a Cas9 nickase, like the Cas9 N863A nickase), or a Cas nuclease selected from Table 7 or Table 8; 2) a peptide linker, such as a sequence from Table D or Table 10, which may exhibit varying degrees of length, flexibility, hydrophobicity, and / or secondary structure; and 3) a reverse transcriptase (RT), such as an RT domain from Table D or Table 6. A gene-modified peptide candidate library may contain: multiple distinct gene-modified peptide candidates that differ from each other in one, two, or all three of the Cas nuclease, peptide linker, or RT domain components; or multiple nucleic acid expression vectors encoding such gene-modified peptide candidates.

[0623] To screen for genetically modified peptide candidates, a two-component system comprising a genetically modified peptide component and a tgRNA component can be used. The genetic modification component may include, for example, an expression vector, such as an expression plasmid or lentiviral vector, encoding a genetically modified peptide candidate, such as a codon-optimized nucleic acid encoding a genetically modified peptide candidate, such as a Cas-linker-RT fusion as described above. In a particular embodiment, a lentiviral cassette is used, comprising: (i) a promoter for expression in mammalian cells, such as a CMV promoter; (ii) a genetically modified library candidate, such as a Cas-linker-RT fusion containing a Cas nuclease from Table 7 or 8, a peptide linker from Table 10, and an RT from Table 6, such as a Cas-linker-RT fusion as shown in Table D; (iii) a self-cleaving peptide, such as a T2A peptide; (iv) a marker selectable in mammalian cells, such as a puromycin resistance gene; and (v) a termination signal, such as a poly-A tail.

[0624] The tgRNA component may contain tgRNA or expression vectors, such as expression plasmids, which produce tgRNA, for example, by using the U6 promoter to drive the expression of tgRNA. The tgRNA is a non-coding RNA sequence that is recognized by Cas and located to the target genomic locus, and also serves as a template for reverse transcription of the desired edit into the genome via the RT domain.

[0625] To prepare cell pools for expressing gene-modified peptide libraries, mammalian cells (e.g., HEK293T or U2OS cells) can be transduced using pooled gene-modified peptide expression vector formulations (e.g., lentiviral formulations) from gene-modified peptide libraries. In a specific embodiment, lentiviral plasmids were used, and HEK293 Lenti-X cells were seeded in 15 cm plates (approximately 12 x 10 cm²). 6 Lentiviral plasmid transfection is then performed after the cells have been divided into groups. In such an embodiment, lentiviral plasmid transfection can be performed using a lentiviral packaging mixture (Biosettia), and the plasmid DNA of the gene-modified candidate library can be transfected the next day using Lipofectamine 2000 and Opti-MEM medium, according to the manufacturer's protocol. In such an embodiment, extracellular DNA can be removed by replacing the medium with a complete medium the next day, and the virus-containing medium can be harvested after 48 hours. The lentiviral medium can be concentrated using Lenti-X concentrate (TaKaRa Biosciences), which can prepare 5 mL aliquots of lentiviral samples and store them at -80°C. Lentiviral titer determination is performed by counting colony-forming units after selection (e.g., after puromycin selection).

[0626] To monitor gene editing of target DNA, mammalian cells carrying the target DNA, such as HEK293T or U2OS cells, can be used. In other embodiments for monitoring gene editing of target DNA, mammalian cells carrying a target DNA genome landing pad, such as HEK293T or U2OS cells, can be used. In specific embodiments, the target DNA genome landing pad may contain the gene to be edited to treat a target disease or condition. In other specific embodiments, the target DNA is a gene sequence expressing a protein exhibiting detectable characteristics that can be monitored to determine whether gene editing has occurred. For example, in some embodiments, a genome landing pad expressing blue fluorescent protein (BFP) or green fluorescent protein (GFP) is used. In some embodiments, mammalian cells (e.g., HEK293T or U2OS cells) containing the target DNA (e.g., a target DNA genome landing pad) are seeded in culture plates at 500x-3000x cells per gene-modified library candidate and transduced at a multiplicity of infection (MOI) of 0.2-0.3 to minimize multiple infection per cell. Puromycin (2.5 ug / mL) can be added 48 hours post-infection to select infected cells. In such an embodiment, cells can be maintained under puromycin selection for at least 7 days before scaling up to introduce tgRNA, for example, tgRNA electroporation.

[0627] To determine whether gene editing has occurred, mammalian cells containing the target DNA to be edited can be infected with a gene-modified peptide library candidate, followed by transfection with tgRNA designed to edit the target DNA. Subsequently, the cells can be analyzed to determine whether editing of the target locus occurred as designed, or whether no editing occurred or imperfect editing occurred, for example, through cell sorting and sequence analysis.

[0628] In a specific embodiment, to determine whether genome editing has occurred, mammalian cells expressing BFP or GFP (e.g., HEK293T or U2OS cells) can be infected with a gene-modified library candidate, followed by transfection or electroporation with a tgRNA plasmid or RNA, for example, by electroporation at 250,000 cells / well using a 200 ng tgRNA plasmid designed to convert BFP to GFP or GFP to BFP, wherein cell counting ensures coverage of >250x-1000x for each library candidate. In such an embodiment, the genome editing capability of various constructs in the assay can be assessed by sorting cells using fluorescence-activated cell sorting (FACS) for the expression of color-converted fluorescent protein (FP) 4–10 days after electroporation. Cells are sorted and harvested into different populations: unedited cells (showing the original fluorescent protein signal), edited cells (showing the converted fluorescent protein signal), and imperfectly edited cells (showing no fluorescent protein signal). Unsorted cell samples can also be harvested as an input population to determine candidate enrichment during the analysis process.

[0629] To determine which gene-modified library candidates exhibit genome-editing capability in assays, genomic DNA (gDNA) is harvested from sorted cell populations and analyzed by sequencing the gene-modified library candidates in each population. Briefly, gene-modified candidates are amplified from the genome using primers specific to gene-modified peptide expression vectors (e.g., lentiviral boxes), amplified in a second round of PCR to dilute the genomic DNA, and then sequenced, for example, using a next-generation sequencing platform. After quality control of the sequencing reads, reads of at least approximately 1500 nucleotides and typically no more than approximately 3200 nucleotides are mapped to gene-modified peptide library sequences, and those reads that match at least approximately 80% of the library sequences are considered successfully aligned to a given candidate for this pooling screening. To identify candidates capable of gene editing (e.g., BFP to GFP, or GFP to BFP editing) in assays, the read counts of each library candidate in the edited population are compared to the read counts in the initial unsorted population.

[0630] For pooling screening, gene modification candidates with genome editing capabilities are identified based on enrichment of the edited (transformed FP) population relative to unsorted (input) cells. In some embodiments, enrichment of at least 1.0, 1.5, 2.0, 2.5, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10, 15, 20, 25, 30, 40, 50, 60, 70, 80, 90, or at least 100-fold relative to the input indicates potentially useful gene editing activity, such as at least 2-fold enrichment. In some embodiments, enrichment is converted to logarithmic values ​​by taking the logarithm base 2 of the enrichment rate. In some embodiments, log2 enrichment scores of at least 0, 1, 2, 3, 4, 5, 5.5, 6.0, 6.2, 6.3, 6.4, 6.5, or at least 6.6 indicate potentially useful gene editing activity, such as a log2 enrichment score of at least 1.0. In a particular embodiment, the enrichment value of the observed gene modification candidate can be compared with the enrichment value observed under similar conditions using a reference (e.g., element ID number: 17380).

[0631] In some embodiments, multiple tgRNAs can be used to screen gene modification candidate libraries. In a particular embodiment, multiple tgRNAs can be used to optimize template / Cas-connector-RT fusion pairs, for example, for gene editing of specific target genes, such as gene targets for treating diseases. In a particular embodiment, a pooling method for screening gene modification candidates can be used in the form of an array of different tgRNAs.

[0632] In some embodiments, various types of editing, such as insertions, substitutions, and / or deletions of different lengths, can be used to screen gene modification candidate libraries.

[0633] In some embodiments, multiple target sequences (e.g., different fluorescent proteins) may be used to screen gene modification candidate libraries. In some embodiments, multiple cell types, such as HEK293T or U2OS, may be used to screen gene modification candidate libraries. Those skilled in the art will understand that a given candidate may exhibit altered editing capabilities, or even gain or lose any observable or useful activity under different conditions, including tgRNA sequences (e.g., nucleotide modifications, PBS length, RT template length), target sequences, target locations, editing types, mutation locations relative to the first strand cleavage of the gene-modifying polypeptide, or cell types. Therefore, in some embodiments, gene modification library candidates are screened across multiple parameters, for example, using at least two different tgRNAs from at least two cell types, and gene editing activity is identified by enrichment under any single condition. In other embodiments, candidates with stronger activity in different tgRNAs and cell types are identified by enrichment under at least two conditions (e.g., under all screening conditions). For clarity, candidates that show little or no enrichment under any given conditions are not considered inactive under all conditions and can be screened with different parameters or reconfigured at the peptide level, for example by exchanging, shuffling or evolving domains (e.g., RT domains), linkers or other signals (e.g., NLS).

[0634] Exemplary Cas9-connector-RT fusion sequence

[0635] In some embodiments, the gene-modified polypeptide comprises an adapter sequence and an RT sequence. In some embodiments, the gene-modified polypeptide comprises an adapter sequence as listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises an amino acid sequence of an RT domain as listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises an adapter sequence as listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and an amino acid sequence of an RT domain as listed in Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises: (i) an adapter sequence as listed in the row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (ii) an amino acid sequence of an RT domain as listed in the same row of Table D, or an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0636] Exemplary gene-modified peptides

[0637] In some embodiments, the gene-modified polypeptide (e.g., a gene-modified polypeptide as part of the system described herein) comprises the amino acid sequence of any one of SEQ ID NO: 1-7743 in the sequence listing, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 80% identity with it. In some embodiments, the gene-modified polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 90% identity with it. In some embodiments, the gene-modified polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 95% identity with it. In some embodiments, the gene-modified polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 99% identity with it. In some embodiments, the gene-modified polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 1-7743. In some embodiments, the gene-modified polypeptide comprises an amino acid sequence of any one of SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises an amino acid sequence of any one of SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide described herein comprises an RT and adapter sequence from any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it, and the St1Cas9 domain described herein. In some embodiments, the gene-modified polypeptide described herein comprises an RT and adapter sequence from any one of SEQ ID NO: 1-7743, and the St1Cas9 domain described herein.

[0638] In some embodiments, the genetically modified polypeptide comprises an amino acid sequence as listed in Table A1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0639] In some embodiments, the gene-modified polypeptide comprises an amino acid sequence as listed in Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises a linker containing a linker sequence as listed in Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises an RT domain containing an RT domain sequence as listed in Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises: (i) a linker containing a linker sequence as listed in the row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and (ii) an RT domain containing an RT domain sequence as listed in the same row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0640] Table T1. Selection of Exemplary Genetically Modified Peptides

[0641]

[0642] In some embodiments, the gene-modified polypeptide comprises an amino acid sequence as listed in Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises a linker containing a linker sequence as listed in Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the gene-modified polypeptide comprises an RT domain containing an RT domain sequence as listed in Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises: (i) a linker containing a linker sequence as listed in the row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and (ii) an RT domain containing an RT domain sequence as listed in the same row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0643] Table T2. Selection of Exemplary Genetically Modified Peptides

[0644]

[0645]

[0646]

[0647] Subsequence of an exemplary genetically modified polypeptide

[0648] In some embodiments, the gene-modified polypeptide comprises, in N-terminal to C-terminal order, one or more of the following: an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA-binding domain, an adapter, an RT domain, and / or a second NLS (e.g., 1, 2, 3, 4, 5, or all 6). In some embodiments, the gene-modified polypeptide comprises, in N-terminal to C-terminal order, an NLS (e.g., a first NLS), a DNA-binding domain, an adapter, and an RT domain, wherein the adapter and RT domain are the adapter and RT domain of any of the gene-modified polypeptides in SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the adapter and RT domain. In some embodiments, the gene-modified polypeptide comprises, in order from N-terminus to C-terminus, a DNA-binding domain, an adapter, an RT domain, and an NLS (e.g., a second NLS), wherein the adapter and RT domain are the adapter and RT domain of any of the gene-modified polypeptides in SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the adapter and RT domain. In some embodiments, the gene-modified polypeptide comprises, in order from N-terminus to C-terminus, a first NLS, a DNA-binding domain, an adapter, an RT domain, and a second NLS, wherein the adapter and RT domain are the adapter and RT domain of any of the gene-modified polypeptides in SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the adapter and RT domain. In some embodiments, the gene-modified polypeptide further comprises an N-terminal methionine residue.

[0649] In some embodiments, the genetically modified polypeptide comprises, in the order from N-terminus to C-terminus, one or more of the following (e.g., 1, 2, 3, 4, 5, or all 6): an N-terminal methionine residue, a first nuclear localization signal (NLS) (e.g., any one of the following genetically modified polypeptides: SEQ ID NO: 1-7743 and / or an amino acid sequence as listed in any one of Tables A1, T1, or T2, or having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it), and a DNA-binding domain (e.g., a Cas domain, such as a SpyCas9 domain, such as an amino acid sequence as listed in Table 8, or having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; or the DNA-binding domain of the following genetically modified polypeptides: SEQ ID NO: 1-7743). The following amino acid sequences are listed in SEQ ID NO: 1-7743 and / or as listed in any of Tables A1, T1, or T2, or have at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them: SEQ ID NO: 1-7743 and / or as listed in any of Tables A1, T1, or T2, or have at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them: RT domain (e.g., as listed in any of SEQ ID NO: 1-7743 and / or as listed in any of Tables A1, T1, or T2, or have at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them): ... RT domain (e.g., as listed in any of SEQ ID NO: 1-7743 and / or as listed in any of Tables A1, T1, or T2, or have at least 70%, 75%, 80%, 85%, 90%, 95%, or 99%): RT domain (e.g., as listed in any of SEQ ID NO: 1-7743 and / or as listed in any of Tables A1, T The amino acid sequence is any one of SEQ ID NO: 1-7743 and / or any one of Tables A1, T1, or T2, or has at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide further comprises (e.g., the C-terminus of the second NLS) a T2A sequence and / or a puromycin sequence (e.g., an amino acid sequence belonging to any one of SEQ ID NO: 1-7743 and / or any one of Tables A1, T1, or T2, or having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the nucleic acid encoding the genetically modified polypeptide (e.g., as described herein) encodes the T2A sequence, for example, wherein the T2A sequence is located between the region encoding the genetically modified polypeptide and the second region, wherein the second region optionally encodes an optional marker, such as puromycin.

[0650] In some embodiments, the first NLS comprises a first NLS sequence of a genetically modified polypeptide having an amino acid sequence of any one of SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the first NLS comprises a first NLS sequence of a genetically modified polypeptide as listed in any one of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the first NLS sequence comprises a C-myc NLS. In some embodiments, the first NLS comprises the amino acid sequence PAAKRVKLD (SEQ ID NO: 11,095), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0651] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the first NLS and the DNA-binding domain. In some embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises the amino acid sequence GG.

[0652] In some embodiments, the DNA-binding domain comprises the DNA-binding domain of any of the following genetically modified polypeptides: SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them. In some embodiments, the DNA-binding domain comprises the DNA-binding domain of any of the genetically modified polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them. In some embodiments, the DNA-binding domain comprises a Cas domain (e.g., as listed in Table 8). In some embodiments, the DNA-binding domain comprises the amino acid sequence of a SpyCas9 polypeptide (e.g., as listed in Table 8, such as Cas9 N863A polypeptide), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the DNA-binding domain comprises the following amino acid sequences:

[0653]

[0654] Or an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to it.

[0655] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the DNA-binding domain and the adapter. In some embodiments, the spacer sequence between the DNA-binding domain and the adapter comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the DNA-binding domain and the adapter comprises the amino acid sequence GG.

[0656] In some embodiments, the adapter comprises an adapter sequence of any one of the following genetically modified polypeptides: SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the adapter comprises an adapter sequence of any one of the genetically modified polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the adapter comprises an amino acid sequence listed in Table D or 10, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0657] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the adapter and the RT domain. In some embodiments, the spacer sequence between the adapter and the RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the adapter and the RT domain comprises the amino acid sequence GG.

[0658] In some embodiments, the RT domain comprises the RT domain sequence of any of the following genetically modified polypeptides: SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them. In some embodiments, the RT domain comprises the RT domain sequence of any of the genetically modified polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them. In some embodiments, the RT domain comprises the amino acid sequence listed in Table D or 6, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them. In some embodiments, the length of the RT domain is about 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 amino acids.

[0659] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the RT domain and the second NLS. In some embodiments, the spacer sequence between the RT domain and the second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the RT domain and the second NLS comprises the amino acid sequence AG.

[0660] In some embodiments, the second NLS comprises the second NLS sequence of a genetically modified polypeptide of any one of SEQ ID NO: 1-7743. In some embodiments, the second NLS comprises the second NLS sequence of a genetically modified polypeptide as listed in Tables A1, T1, or T2. In some embodiments, the second NLS sequence comprises multiple partial NLS sequences. In embodiments, the NLS sequence (e.g., the second NLS sequence) comprises a first partial NLS sequence, for example, which comprises the amino acid sequence KRTADGSEFE (SEQ ID NO: 11,097), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In embodiments, the NLS sequence (e.g., the second NLS sequence) comprises a second partial NLS sequence. In some embodiments, the NLS sequence (e.g., a second NLS sequence) comprises an SV40A5 NLS, such as a two-component SV40A5 NLS, for example, which comprises the amino acid sequence KRTADGSEFESPKKKAKVE (SEQ ID NO: 11,098), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the NLS sequence (e.g., a second NLS sequence) comprises the amino acid sequence KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 11,099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0661] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence. In some embodiments, the spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence comprises the amino acid sequence GSG.

[0662] Connectors and RT structural domain

[0663] In some embodiments, the gene-modified polypeptide comprises a linker (e.g., as described herein) and an RT domain (e.g., as described herein). In some embodiments, the gene-modified polypeptide comprises a linker (e.g., as described herein) and an RT domain (e.g., as described herein) in an N-terminal to C-terminal order.

[0664] In some embodiments, the adapter comprises an adapter sequence as listed in Table 10, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the adapter comprises any one of the following adapter sequences: SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the adapter comprises any one of the following adapter sequences: SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the adapter comprises any one of the following adapter sequences: SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the linker comprises a linker sequence of any of the exemplary gene-modified polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain comprises an RT domain sequence as listed in Table 6, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain comprises an RT domain sequence of any of the exemplary gene-modified polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0665] In some embodiments, the gene-modified polypeptide comprises a portion of a gene-modified polypeptide of any one of SEQ ID NO: 1-7743 (wherein the portion comprises a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion.

[0666] In some embodiments, the gene-modified polypeptide comprises a linker of any one of the gene-modified polypeptides in SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the gene-modified polypeptide comprises a linker of any one of the gene-modified polypeptides in SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the gene-modified polypeptide comprises a linker of any one of the gene-modified polypeptides in SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the genetically modified polypeptide comprises a linker of a genetically modified polypeptide listed in any of Tables A1, T1 or T2, or comprises a linker having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95% or 99% identical to that of the polypeptide.

[0667] In some embodiments, the gene-modified polypeptide comprises the RT domain of any one of the gene-modified polypeptides in SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said RT domain. In some embodiments, the gene-modified polypeptide comprises the RT domain of any one of the gene-modified polypeptides in SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said RT domain. In some embodiments, the gene-modified polypeptide comprises the RT domain of any one of the gene-modified polypeptides in SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said RT domain. In some embodiments, the genetically modified polypeptide comprises an RT domain of any of the genetically modified polypeptides listed in Tables A1, T1, or T2, or comprises an RT domain having an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to that of the genetically modified polypeptide.

[0668] In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain of the gene-modified polypeptide having the amino acid sequence of any one of SEQ ID NO: 1-7743 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain having at least 80% identity with the linker and RT domain of any one of SEQ ID NO: 1-7743. In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain having at least 90% identity with the linker and RT domain of any one of SEQ ID NO: 1-7743. In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain having at least 95% identity with the linker and RT domain of any one of SEQ ID NO: 1-7743. In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain having at least 99% identity with the linker and RT domain of any one of SEQ ID NO: 1-7743. In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain of the gene-modified polypeptide having an amino acid sequence of any one of SEQ ID NO: 6001-7743 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the linker and RT domain of the gene-modified polypeptide comprises an amino acid sequence of the linker and RT domain of the gene-modified polypeptide having an amino acid sequence of any one of SEQ ID NO: 4501-4541 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the linker and RT domain of the genetically modified polypeptide comprises the amino acid sequence (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with) the linker and RT domain of a single row from any of Tables A1, T1, or T2 (e.g., a single exemplary genetically modified polypeptide listed in any of Tables A1, T1, or T2).

[0669] In some embodiments, the linker and RT domain of the gene-modified polypeptide comprise amino acid sequences from two different amino acid sequences selected from SEQ ID NO: 1-7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them). In some embodiments, the linker and RT domain of the gene-modified polypeptide comprise amino acid sequences from different rows of the linker and RT domain from any one of Tables A1, T1, or T2 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them).

[0670] In some embodiments, the gene-modifying polypeptide further comprises a first NLS (e.g., a 5' NLS), such as that described herein. In some embodiments, the gene-modifying polypeptide further comprises a second NLS (e.g., a 3' NLS), such as that described herein. In some embodiments, the gene-modifying polypeptide further comprises an N-terminal methionine residue.

[0671] RT family and mutants

[0672] In some embodiments, the gene-modified polypeptide comprises an amino acid sequence from an RT domain sequence selected from the following families: AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, XMRV6, BLVAU, BLVJ, HTL1A, HTL1C, HTL1L, HTL32, HTL3P, ​​HTLV2, JSRV, MLVF5, MLVRD, MMTVB, MPMV, SFVCP, SMRVH, SRV1, SRV2, and WDSV. In some embodiments, the gene-modified polypeptide comprises an amino acid sequence from an RT domain sequence selected from the following families: AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.

[0673] In some embodiments, the gene-modified polypeptide comprises an amino acid sequence of the RT domain sequence from the MLVMS RT domain. In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations, or corresponding point mutations, listed in column 1 of Table M1. In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations, or corresponding point mutations, listed in column 3 of Table M1 (Gen1MLVMS). In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations at the amino acid positions of the RT domain listed in columns 1 and 2 of Table M2, or one or more point mutations at the corresponding amino acid positions.

[0674] In some embodiments, the genetically modified polypeptide comprises an amino acid sequence of the RT domain sequence derived from the AVIRE RT domain. In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations, or corresponding point mutations, listed in column 2 of Table M1. In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations, or corresponding point mutations, listed in column 4 (Gen2AVIRE) of Table M1. In some embodiments, the amino acid sequence of the RT domain sequence comprises one or more point mutations at the amino acid positions of the RT domain listed in columns 3 and 4 of Table M2, or one or more point mutations at the corresponding amino acid positions. In some embodiments, the RT domain comprises IENSP (e.g., at the C-terminus).

[0675] Table M1. Exemplary point mutations in the MLVMS and AVIRE RT domains

[0676]

[0677] Table M2. Mutational locations in exemplary MLVMS and AVIRE RT domains

[0678]

[0679] In some embodiments, the genetically modified polypeptide comprises a γ-retrovirus-derived RT domain. In some embodiments, the γ-retrovirus-derived RT domain of the genetically modified polypeptide comprises an amino acid sequence from an RT domain sequence selected from the following families: AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In some embodiments, the γ-retrovirus-derived RT domain of the genetically modified polypeptide is not derived from PERV. In some embodiments, the RT includes one, two, three, four, five, six or more mutations as shown in Table 2A, corresponding to mutations D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K or D653N in the RT domain of murine leukemia virus reverse transcriptase. In some embodiments, the gene-modifying polypeptide further comprises an adapter having at least 99% identity with the adapter domain of any one of SEQ ID NO: 1-7743. In some embodiments, the gene-modifying polypeptide further comprises an adapter having at least 99% or 100% identity with SEQ ID NO: 5217 or SEQ ID NO: 11,041.

[0680] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of AVIRE RT (e.g., the AVIRE_P03360 sequence, such as SEQ ID NO: 8001), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of AVIRE RT containing one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, G330P, L605W, T306K, and W313F. In some embodiments, the RT domain further comprises the amino acid sequence of AVIRE RT containing one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, G330P, and L605W.

[0681] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of BAEVM RT (e.g., the BAEVM_P10272 sequence, e.g., SEQ ID NO: 8004), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of BAEVM RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, L602W, T304K, and W311F. In some embodiments, the RT domain further comprises the amino acid sequence of BAEVM RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, and L602W.

[0682] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of FFV RT (e.g., the FFV_O93209 sequence, e.g., SEQ ID NO: 8012), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of FFV RT with one, two, three, or four mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D21N, T293N, T419P, and L393K. In some embodiments, the RT domain further comprises the amino acid sequence of FFV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D21N, T293N, and T419P. In some embodiments, the RT domain further comprises the amino acid sequence of FFV RT with a mutation of D21N. In some embodiments, the RT domain further comprises an amino acid sequence of FFV RT with one, two, or three mutations selected from the group consisting of T207N, T333P, and L307K, or at corresponding positions in a homologous RT domain. In some embodiments, the RT domain further comprises an amino acid sequence of FFV RT with one or two mutations selected from the group consisting of T207N and T333P, or at corresponding positions in a homologous RT domain.

[0683] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of an FLV RT (e.g., the FLV_P10273 sequence, e.g., SEQ ID NO: 8019), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of an FLV RT with one, two, three, or four mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D199N, L602W, T305K, and W312F. In some embodiments, the RT domain further comprises the amino acid sequence of an FLV RT with one or two mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D199N and L602W.

[0684] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of FOAMV RT (e.g., the FOAMV_P14350 sequence, e.g., SEQ ID NO: 8021), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of FOAMV RT with one, two, three, or four mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D24N, T296N, S420P, and L396K. In some embodiments, the RT domain further comprises the amino acid sequence of FOAMV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D24N, T296N, and S420P. In some embodiments, the RT domain further comprises the amino acid sequence of FOAMV RT with mutation D24N, or at corresponding positions in a homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of FOAMV RT that further comprises one, two, or three mutants selected from the group consisting of T207N, S331P, and L307K, or at corresponding positions in a homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of FOAMV RT that further comprises one or two mutants selected from the group consisting of T207N and S331P, or at corresponding positions in a homologous RT domain.

[0685] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of GALV RT (e.g., the GALV_P21414 sequence, e.g., SEQ ID NO: 8027), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of GALV RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain further comprises the amino acid sequence of GALV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, and L600W.

[0686] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of KORV RT (e.g., the KORV_Q9TTC1 sequence, e.g., SEQ ID NO: 8047), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of GALV RT with one, two, three, four, five, or six mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D32N, D322N, E452P, L274W, T428K, and W435F. In some embodiments, the RT domain further comprises the amino acid sequence of GALV RT with one, two, three, or four mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D32N, D322N, E452P, and L274W. In some embodiments, the RT domain further comprises the amino acid sequence of GALV RT with a mutation of D32N. In some embodiments, the RT domain further comprises an amino acid sequence of KORV RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D231N, E361P, L633W, T337K, and W344F. In some embodiments, the RT domain further comprises an amino acid sequence of KORV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D231N, E361P, and L633W.

[0687] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVAV RT (e.g., the MLVAV_P03356 sequence, e.g., SEQ ID NO: 8053), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of MLVAV RT containing one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain further comprises the amino acid sequence of MLVAV RT containing one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, and L603W.

[0688] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVBM RT (e.g., the MLVBM_Q7SVK7 sequence, such as SEQ ID NO: 8056), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of MLVBM RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D199N, T329P, L602W, T305K, and W312F. In some embodiments, the RT domain further comprises the amino acid sequence of MLVBM RT with one, two, and three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, and L603W.

[0689] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVCB RT (e.g., the MLVCB_P08361 sequence, e.g., SEQ ID NO: 8062), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of MLVCB RT containing one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain further comprises the amino acid sequence of MLVCB RT containing one, two, and three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, and L603W.

[0690] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVFF RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of MLVFF RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain further comprises the amino acid sequence of MLVFF RT with one, two, and three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, and L603W.

[0691] In embodiments, the RT domain comprises the amino acid sequence of the RT domain of MLVMS RT (e.g., the MLVMS_reference sequence, such as SEQ ID NO: 8137; or the MLVMS_P03355 sequence, such as SEQ ID NO: 8070), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of MLVMS RT with one, two, three, four, five, or six mutations at corresponding positions in the homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, W313F, and H8Y. In some embodiments, the RT domain further comprises the amino acid sequence of MLVMS RT with one, two, three, four, or five mutations at corresponding positions in the homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain comprises an amino acid sequence of an MLVMS RT that further comprises one, two, or three mutants selected from the group consisting of D200N, T330P, and L603W, or at corresponding positions in a homologous RT domain.

[0692] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of a PERV RT (e.g., the PERV_Q4VFZ2 sequence, e.g., SEQ ID NO: 8099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of a PERV RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D196N, E326P, L599W, T302K, and W309F. In some embodiments, the RT domain further comprises the amino acid sequence of a PERV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D196N, E326P, and L599W.

[0693] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of SFV1 RT (e.g., the SFV1_P23074 sequence, e.g., SEQ ID NO: 8105), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of SFV1 RT with one, two, three, or four mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D24N, T296N, N420P, and L396K. In some embodiments, the RT domain further comprises the amino acid sequence of SFV1 RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D24N, T296N, and N420P. In some embodiments, the RT domain further comprises the amino acid sequence of SFV1 RT with D24N, or at corresponding positions in a homologous RT domain.

[0694] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of SFV3L RT (e.g., the SFV3L_P27401 sequence, e.g., SEQ ID NO: 8111), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of SFV3L RT with one, two, three, or four mutations selected from the group consisting of D24N, T296N, N422P, and L396K, or at the corresponding position in the homologous RT domain. In some embodiments, the RT domain further comprises the amino acid sequence of SFV3L RT with one, two, or three mutations selected from the group consisting of D24N, T296N, and N422P, or at the corresponding position in the homologous RT domain. In some embodiments, the RT domain further comprises the amino acid sequence of SFV3L RT with mutation D24N, or at the corresponding position in the homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of SFV3L RT that further comprises one, two, or three mutated SFV3L RT sequences selected from the group consisting of T307N, N333P, and L307K, or at corresponding positions in a homologous RT domain. In some embodiments, the RT domain comprises an amino acid sequence of SFV3L RT that further comprises one or two mutated SFV3L RT sequences selected from the group consisting of T307N and N333P, or at corresponding positions in a homologous RT domain.

[0695] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of WMSV RT (e.g., the WMSV_P03359 sequence, e.g., SEQ ID NO: 8131), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of WMSV RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, L600W, T304K, and W311F. In some embodiments, the RT domain further comprises the amino acid sequence of WMSV RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D198N, E328P, and L600W.

[0696] In some embodiments, the RT domain comprises the amino acid sequence of the RT domain of XMRV6 RT (e.g., the XMRV6_A1Z651 sequence, e.g., SEQ ID NO: 8134), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain further comprises the amino acid sequence of XMRV6 RT with one, two, three, four, or five mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, L603W, T306K, and W313F. In some embodiments, the RT domain further comprises the amino acid sequence of XMRV6 RT with one, two, or three mutations at corresponding positions in a homologous RT domain, selected from the group consisting of D200N, T330P, and L603W.

[0697] In some embodiments, the RT domain of the genetically modified polypeptide comprises the amino acid sequence of the RT domain of AVIRE RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain comprises the amino acid sequence of the RT domain included in the sequences listed in column 1 of Table A5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity with SEQ ID NO: 5217 or SEQ ID NO: 11,041.

[0698] In some embodiments, the RT domain of the genetically modified polypeptide comprises the amino acid sequence of the RT domain of MLVMS RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain comprises the amino acid sequence of the RT domain contained in any of the sequences listed in columns 2-6 of Table A5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide further comprises a linker having at least 99% or 100% identity with SEQ ID NO: 5217 or SEQ ID NO: 11,041.

[0699] Table A5. Exemplary gene-modified peptides containing the AVIRE RT domain or the MLVMS RT domain.

[0700]

[0701]

[0702]

[0703]

[0704]

[0705]

[0706]

[0707]

[0708] system

[0709] In one aspect, this disclosure relates to a system comprising a nucleic acid molecule encoding a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., template RNA, e.g., as described herein). In some embodiments, the nucleic acid molecule encoding the genetically modified polypeptide contains one or more silent mutations in the coding region (e.g., in the sequence encoding the RT domain) relative to the nucleic acid molecule described herein. In some embodiments, the system further comprises gRNA (e.g., gRNA that binds to the polypeptide that induces nicking, e.g., in the opposite strand of the target DNA to which the genetically modified polypeptide binds).

[0710] In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide encodes a polypeptide listed in any one of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0711] In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NO: 1-7743 (wherein the portion includes a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NO: 6001-7743 (wherein the portion includes a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide comprises a sequence encoding a portion of an amino acid sequence selected from SEQ ID NO: 4501-4541 (wherein the portion includes a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a sequence encoding a portion of a polypeptide listed in any of Tables A1, T1, or T2 (where the portion includes a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion.

[0712] In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a linker sequence encoding an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a linker sequence encoding a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a linker sequence encoding a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide contains a sequence encoding a linker of a polypeptide listed in any of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0713] In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a sequence encoding an RT domain of an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid molecule encoding the gene-modified polypeptide comprises a sequence encoding the RT domain of any of the polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0714] In one respect, this disclosure relates to a system comprising a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., template RNA, e.g., as described herein).

[0715] In some embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a polypeptide as listed in any one of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0716] In some embodiments, the gene-modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NO: 1-7743 (wherein the portion comprises a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the gene-modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NO: 6001-7743 (wherein the portion comprises a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the gene-modified polypeptide comprises a portion of an amino acid sequence selected from SEQ ID NO: 4501-4541 (wherein the portion comprises a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion. In some embodiments, the genetically modified polypeptide comprises a portion of a polypeptide listed in any one of Tables A1, T1, or T2 (wherein the portion comprises a linker and an RT domain), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with said portion.

[0717] In some embodiments, the genetically modified polypeptide comprises a linker of an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a linker sequence encoding a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a linker sequence encoding a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a linker of a polypeptide listed in any of Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0718] In some embodiments, the genetically modified polypeptide comprises an RT domain of an amino acid sequence selected from SEQ ID NO: 1-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NO: 6001-7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises a sequence encoding an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NO: 4501-4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide comprises the RT domain of any of the polypeptides listed in Tables A1, T1, or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0719] Table A1. Exemplary amino acid and nucleotide sequences of genetically modified peptides

[0720]

[0721]

[0722]

[0723]

[0724]

[0725]

[0726] Location sequences of gene modification systems

[0727] In some embodiments, the gene editor system RNA further comprises an intracellular localization sequence, such as a nuclear localization sequence (NLS). In some embodiments, the gene-modifying polypeptide comprises an NLS as contained in SEQ ID NO: 4000 and / or SEQ ID NO: 4001, or an NLS having an amino acid sequence that is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0728] The nuclear localization sequence can be an RNA sequence that facilitates RNA importation into the cell nucleus. In some embodiments, the nuclear localization signal is located on the template RNA. In some embodiments, the genetically modified polypeptide is encoded on a first RNA, and the template RNA is a second separate RNA, and the nuclear localization signal is located on the template RNA rather than on the RNA encoding the genetically modified polypeptide. While not wishing to be bound by theory, in some embodiments, the RNA encoding the genetically modified polypeptide primarily targets the cytoplasm to facilitate its translation, while the template RNA primarily targets the cell nucleus to facilitate its insertion into the genome. In some embodiments, the nuclear localization signal is located at the 3' end, 5' end, or internal region of the template RNA. In some embodiments, the nuclear localization signal is located at the 3' end of a heterologous sequence (e.g., directly at the 3' end of the heterologous sequence) or at the 5' end of a heterologous sequence (e.g., directly at the 5' end of the heterologous sequence). In some embodiments, the nuclear localization signal is placed outside the 5' UTR or 3' UTR of the template RNA. In some embodiments, the nuclear localization signal is placed between the 5' UTR and the 3' UTR, wherein optionally, the nuclear localization signal is not transcribed with the transgene (e.g., the nuclear localization signal is antisense-oriented or downstream of a transcription termination signal or a polyadenylation signal). In some embodiments, the nuclear localization sequence is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present in RNA, such as in template RNA. In some embodiments, the length of the nuclear localization signal is less than 5, 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 bp. Various RNA nuclear localization sequences can be used. For example, Lubelsky and Ulitsky, Nature [Nature] 555 (107-111), 2018, describe an RNA sequence that drives RNA localization into the cell nucleus. In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, the nuclear localization signal binds to a nuclear enrichment protein. In some embodiments, the nuclear localization signal binds to an HNRNPK protein. In some embodiments, the nuclear localization signal is pyrimidine-rich, such as a C / T-rich, C / U-rich, C-rich, T-rich, or U-rich region. In some embodiments, the nuclear localization signal originates from a long non-coding RNA. In some embodiments, the nuclear localization signal originates from the MALAT1 long non-coding RNA or the M region of MALAT1, which is 600 nucleotides long (described in Miyagawa et al., RNA 18, (738-751), 2012).In some embodiments, the nuclear localization signal is derived from BORG long noncoding RNA or the AGCCC motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.

[0729] In some embodiments, the polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In some embodiments, the NLS is a two-component NLS. In some embodiments, the NLS facilitates the delivery of the protein containing the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of the genetically modified polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the genetically modified polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas domain. In some embodiments, a linker sequence is arranged between the NLS and a neighboring domain of the genetically modified polypeptide.

[0730] In some embodiments, NLS comprises the amino acid sequences MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 5009), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5010), RKSGKIAAIWKRPRKPKKKRKV (SEQ ID NO: 5011), KRTADGSEFESPKKKRKV (SEQ ID NO: 5012), KKTELQTTNAENKTKKL (SEQ ID NO: 5013), or KRGINDRNFWRGENGRKTR (SEQ ID NO: 5014), KRPAATKKAGQAKKKK (SEQ ID NO: 5015), PAAKRVKLD (SEQ ID NO: 4644), KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4649), KRTADGSEFE (SEQ ID NO: 5009), KRTADGSEFE (SEQ ID NO: 5010), KRTADGSEFE (SEQ ID NO: 5011), KRTADGSEFESPKKKRKV (SEQ ID NO: 5012), KKTELQTTNAENKTKKL (SEQ ID NO: 5013), or KRGINDRNFWRGENGRKTR (SEQ ID NO: 5014), KRPAATKKAGQAKKKK (SEQ ID NO: 5015), PAAKRVKLD (SEQ ID NO: 4644), KRTADGSEFE (SEQ ID NO: 5015), KRTADGSEFE (SEQ ID NO: 5016 ... SEQ ID NO: 4650), KRTADGSEFESPKKKAKVE (SEQ ID NO: 4651), AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 4001), or functional fragments or variants thereof. Exemplary NLS sequences are also described in PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences. In some embodiments, the NLS comprises the amino acid sequences disclosed in Table 11. The NLS of this table may be used with one, two, three, or more copies of the peptide at one or more locations within the peptide, such as in the N-terminal domain, between peptide domains, in the C-terminal domain, or in combinations thereof, to improve subcellular localization of the cell nucleus. Multiple unique sequences may be used in a single peptide. The sequences may be naturally monocomponent or bicomponent, for example, having one or two basic amino acid segments, or may be used as chimeric bicomponent sequences. Sequence references correspond to UniProt accession numbers, unless the sequence is indicated as SeqNLS for sequences mined using subcellular localization prediction algorithms (Lin et al., BMC Bioinformat [BMC Bioinformatics] 13:157 (2012), which is incorporated herein by reference in its entirety).

[0731] Table 11. Exemplary nuclear localization signals used in gene modification systems.

[0732]

[0733]

[0734]

[0735]

[0736]

[0737]

[0738] In some embodiments, the NLS is a two-component NLS. A two-component NLS typically comprises two basic amino acid clusters separated by a spacer sequence (which may be, for example, about 10 amino acids in length). A single-component NLS typically lacks a spacer. An example of a two-component NLS is a nucleoplasmin NLS having the sequence KR[PAATKKAGQA]KKKK (SEQ ID NO: 5015), where the spacer is enclosed in parentheses. Another exemplary two-component NLS has the sequence PKKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 5016). Exemplary NLSs are described in International Application WO 2020051561, which is incorporated herein by reference in its entirety, including its disclosure of nuclear localization sequences.

[0739] In some embodiments, the gene editor system peptide (e.g., a gene-modifying peptide as described herein) further comprises intracellular localization sequences, such as nuclear localization sequences and / or nucleolar localization sequences. The nuclear localization sequences and / or nucleolar localization sequences can be amino acid sequences that facilitate protein input into the nucleus and / or nucleolus, where they can facilitate the integration of heterologous sequences into the genome. In some embodiments, the gene editor system peptide (e.g., a gene-modifying peptide as described herein) further comprises a nucleolar localization sequence. In some embodiments, the gene-modifying peptide is encoded on a first RNA, the template RNA is a second separate RNA, and the nucleolar localization signal is encoded on the RNA encoding the gene-modifying peptide, not on the template RNA. In some embodiments, the nucleolar localization signal is located at the N-terminus, C-terminus, or internal region of the peptide. In some embodiments, multiple identical or different nucleolar localization signals are used. In some embodiments, the length of the nuclear localization signal is less than 5, 10, 25, 50, 75, or 100 amino acids. Various peptide nucleolar localization signals can be used. For example, Yang et al., Journal of Biomedical Science 22, 33 (2015) describe a nuclear localization signal that also functions as a nucleolar localization signal. In some embodiments, the nucleolar localization signal can also be a nuclear localization signal. In some embodiments, the nucleolar localization signal can overlap with a nuclear localization signal. In some embodiments, the nucleolar localization signal can contain basic residue segments. In some embodiments, the nucleolar localization signal can be rich in arginine and lysine residues. In some embodiments, the nucleolar localization signal can be derived from a protein enriched in the nucleolus. In some embodiments, the nucleolar localization signal can be derived from a protein enriched at a ribosomal RNA locus. In some embodiments, the nucleolar localization signal can be derived from a protein that binds rRNA. In some embodiments, the nucleolar localization signal can be derived from MSP58. In some embodiments, the nucleolar localization signal can be a single-component motif. In some embodiments, the nucleolar localization signal can be a two-component motif. In some embodiments, the nucleolar localization signal can consist of multiple single-component or two-component motifs. In some embodiments, the nucleolar localization signal can consist of a mixture of single-component and two-component motifs. In some embodiments, the nucleolar localization signal can be a dual two-component motif. In some embodiments, the nucleolar localization motif may be KRASSQALGTIPKRRSSSRFIKRKK (SEQ ID NO: 5017). In some embodiments, the nucleolar localization signal may be derived from nuclear factor-κB-induced kinase. In some embodiments, the nucleolar localization signal may be the RKKRKKK motif (SEQ ID NO: 5018) (described in Birbach et al., Journal of Cell Science, 117 (3615-3624), 2004).

[0740] Evolutionary variants of genetically modified peptides and systems

[0741] In some embodiments, the present invention provides evolutionary variants of gene-modified polypeptides as described herein. In some embodiments, evolutionary variants can be generated by mutagenesis of a reference gene-modified polypeptide or one of the fragments or domains contained therein. In some embodiments, one or more domains (e.g., reverse transcriptase domains) evolve. In some embodiments, one or more such evolutionary variant domains may evolve individually or together with other domains. In some embodiments, one or more evolutionary variant domains may be combined with one or more unevolved homologous components or evolved variants of one or more homologous components, for example, evolved variants of the one or more homologous components that can evolve in a parallel or sequential manner.

[0742] In some embodiments, the process of mutagenesis of a reference genetically modified polypeptide or a fragment or domain thereof includes mutagenesis of the reference genetically modified polypeptide or a fragment or domain thereof. In embodiments, mutagenesis includes sequential evolution methods (e.g., PACE) or discontinuous evolution methods (e.g., PANCE), as described herein. In some embodiments, the evolved genetically modified polypeptide or a fragment or domain thereof comprises one or more amino acid variations introduced into its amino acid sequence relative to the amino acid sequence of the reference genetically modified polypeptide or a fragment or domain thereof. In embodiments, amino acid sequence variations may include one or more mutated residues (e.g., conserved substitutions, non-conserved substitutions, or combinations thereof) within the amino acid sequence of the reference genetically modified polypeptide, for example, the one or more mutated residues being due to a change in the nucleotide sequence encoding the genetically modified polypeptide (e.g., a codon change at any particular position in the coding sequence), which results in the deletion of one or more amino acids (e.g., a truncated protein), the insertion of one or more amino acids, or any combination of the foregoing. Evolutionary variant genetically modified polypeptides may include variants in one or more components or domains of the genetically modified polypeptide (e.g., variants introducing a reverse transcriptase domain).

[0743] In some aspects, this disclosure provides genetically modified peptides, systems, kits, and methods that use or include evolutionary variants of the genetically modified peptide, such as those employing evolutionary variants of the genetically modified peptide or genetically modified peptides produced or that can be produced by PACE or PANCE. In the examples, the unevolved referen...

Claims

1. A template RNA (tgRNA) comprising (e.g., from 5' to 3'): (1) gRNA spacers; (2) A variant St1Cas9 scaffold with partial or complete absence of stem ring 2; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence.

2. The template RNA of claim 1, wherein the length of the deletion is between 1 and 32 (e.g., 2-29, 2-20, 2-10, or 10-20) nucleotides.

3. The template RNA of claim 1, wherein the deletion is at positions 55 to 84.

4. The template RNA as claimed in any of the preceding claims, wherein the variant St1Cas9 scaffold has one or both of the following: an elongated RAR upper stem or substitutions that produce GC base pairs in the RAR upper stem.

5. The template RNA as described in any of the preceding claims, wherein the variant St1Cas9 scaffold has a mutation in the four-loop.

6. A template RNA (tgRNA) comprising (e.g., from 5' to 3'): (1) gRNA spacers; (2) A variant St1Cas9 scaffold having one or both of the following: an elongated RAR upper stem or substitution of GC base pairs in the RAR upper stem; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence.

7. The template RNA as claimed in any of the preceding claims, wherein the upper stem of the RAR is extended by 1-8 base pairs (e.g., 1, 2, 3, 4, 5, 6, 7 or 8 base pairs) relative to the wild-type sequence of SEQ ID NO:25999.

8. A template RNA (tgRNA) comprising (e.g., from 5' to 3'): (1) gRNA spacers; (2) A variant St1Cas9 scaffold with a mutation in the four-ring; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence.

9. The template RNA as described in any of the preceding claims, wherein the tetraloop is extended to, for example, 5 nucleotides.

10. The template RNA as described in any of the preceding claims, wherein the variant gRNA scaffold comprises the sequence according to Table 23.

11. The template RNA as described in any of the preceding claims, wherein the variant gRNA scaffold comprises the sequence according to Table 23.

12. The template RNA as claimed in any of the preceding claims, wherein the variant gRNA scaffold comprises the sequence according to GUCUUUGUACUCUGGUACCAGAAGCUACAAAGAUAAGG CUUCAUGCCGAAAUCA (SEQ ID NO: 26000).

13. The template RNA as claimed in any of the preceding claims, wherein the gRNA spacer comprises a sequence according to Table 22, or a sequence having no more than one, two or three sequence changes (e.g., substitutions) relative to it.

14. The template RNA as described in any of the preceding claims, wherein the gRNA spacer comprises the sequence according to Table 22.

15. The template RNA as claimed in any of the preceding claims, wherein the gRNA spacer comprises a sequence according to AAGGCUGUGCUGACCAUCGA (SEQ ID NO: 26001).

16. The template RNA as claimed in any of the preceding claims, wherein the heterologous object sequence comprises a sequence according to Table 24, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

17. The template RNA as claimed in any of the preceding claims, wherein the heterologous object sequence comprises the sequence according to Table 24.

18. The template RNA as claimed in any of the preceding claims, wherein the PBS sequence comprises a sequence according to Table 25, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it.

19. The template RNA as described in any of the preceding claims, wherein the PBS sequence comprises the sequence according to Table 25.

20. The template RNA as described in any of the preceding claims, wherein the template RNA comprises a sequence according to Table 20, or a sequence having at least 80%, 85%, 90%, 95% or 98% identity with it.

21. The template RNA as described in any of the preceding claims, wherein the template RNA comprises a sequence according to Table 21, or a sequence having at least 80%, 85%, 90%, 95% or 98% identity with it.

22. The template RNA as claimed in any of the preceding claims, wherein the template RNA comprises a sequence according to any one of Table 27, E3, E3A, E7, E8, E9, E11A, E11B, E12, E12A, E14, E14A, E15 or E16, or a sequence having at least 80%, 85%, 90%, 95% or 98% identity with it.

23. The template RNA as described in any of the preceding claims, wherein the template RNA comprises one or more chemical modifications.

24. The template RNA of claim 23, wherein the template RNA comprises one or more phosphate thioester bonds.

25. The template RNA as claimed in claim 23 or 24, wherein the template RNA comprises one or more 2'-O-methyl nucleotides.

26. The template RNA of any one of claims 23-25, wherein the template RNA comprises a sequence according to column 3 of Table 20, or a sequence having no more than one, two or three sequence alterations (e.g., substitutions) relative to it.

27. The template RNA of any one of claims 23-25, wherein the template RNA comprises a sequence according to column 3 of Table 21, or a sequence having no more than one, two or three sequence alterations (e.g., substitutions) relative to it.

28. A gene modification system comprising: Template RNA as claimed in any one of claims 1-27; and A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising: (1) St1Cas9 structural domain; (2) Connector; and (3) Reverse transcriptase (RT) domain.

29. The system of claim 28, wherein the St1Cas9 domain is a nicking enzyme.

30. The system of claim 28 or 29, wherein the St1Cas9 domain comprises a sequence according to SEQ ID NO: 23818, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

31. The system of any one of claims 28-30, wherein the connector comprises a sequence according to SEQ ID NO: 5006, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

32. The system of any one of claims 28-30, wherein the connector comprises a sequence according to SEQ ID NO: 5217, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

33. The system of any one of claims 28-30, wherein the connector comprises a sequence of Table 10, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

34. The system of any one of claims 28-31, wherein the RT domain comprises a sequence according to SEQ ID NO:26006, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

35. The system of any one of claims 28-31, wherein the RT domain comprises a sequence according to any one of SEQ ID NO: 8,001-8,003, or a sequence having at least 80%, 85%, 90%, 95%, 98% or 99% identity with it.

36. The system of any one of claims 28-31, wherein the RT structural domain comprises the sequence of Table 6, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

37. The system of any one of claims 28-34, wherein the genetically modified polypeptide comprises the sequence according to SEQ ID NO:26002, or a sequence having at least 80%, 85%, 90%, 95%, 98%, or 99% identity with it.

38. The system of any one of claims 28-37, further comprising a second nick gRNA (ngRNA), wherein optionally the second nick gRNA directs a second nick to the second strand of the human SERPINA1 gene.

39. The system of claim 38, wherein the second nick gRNA comprises a sequence according to Table 26, or a sequence having at least 80%, 85%, 90%, 95% or 98% identity with it.

40. The system of any one of claims 28-39, wherein the nucleic acid encoding the gene-modified polypeptide comprises RNA, for example, mRNA.

41. The template RNA or system as described in any of the preceding claims, wherein the nucleic acid molecule is formulated in lipid nanoparticles (LNPs).

42. The system according to any one of claims 28-41, wherein the template RNA, the nucleic acid molecule encoding the gene-modified polypeptide, and / or the second nick gRNA are formulated in an LNP.

43. A pharmaceutical composition comprising the system as described in any one of claims 28-42, or one or more nucleic acids encoding the system, and a pharmaceutically acceptable excipient or carrier.

44. The pharmaceutical composition of claim 43, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

45. The pharmaceutical composition of claim 44, wherein the viral vector is adeno-associated virus.

46. ​​A host cell (e.g., a mammalian cell, such as a human cell) comprising a gene modification system or template RNA as described in any one of the preceding claims.

47. A method for preparing template RNA as described in any of the preceding claims, the method comprising synthesizing the template RNA by means of: in vitro (e.g., by in vitro transcription or solid-state synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow the production of the template RNA.

48. A method for modifying a target site in a cell (e.g., a target site in the human SERPINA1 gene), the method comprising contacting the cell with a gene modification system as described in any one of claims 28-42, or DNA encoding the gene modification system, or a pharmaceutical composition as described in any one of claims 43-45, thereby modifying the target site.

49. A method for treating a subject suffering from a disease or condition associated with a gene (e.g., the human SERPINA1 gene) mutation, the method comprising administering to the subject a gene-modifying system as described in any one of claims 28-42, or DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any one of claims 43-45, thereby treating the subject suffering from the disease or condition.

50. The method of claim 49, wherein the disease or condition is α-1 antitrypsin deficiency (AATD).

51. The method of claim 49 or 50, wherein the subject has an E342K mutation.

52. A method for treating a subject with AATD, the method comprising administering to the subject a gene-modifying system as described in any one of claims 28-42, or DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any one of claims 43-45, thereby treating the subject with AATD.

53. The gene modification system or method as claimed in any of the preceding claims, wherein introducing the system into a target cell results in the correction of a pathogenic mutation in the gene, such as the SERPINA1 gene.

54. The gene modification system or method as claimed in any of the preceding claims, wherein the pathogenic mutation is an E342K mutation, and wherein the correction comprises an amino acid substitution of K342E.

55. The gene modification system or method as claimed in any of the preceding claims, wherein introducing the system into a target cell results in a mutation that restores the function of the gene, such as the SERPINA1 gene.

56. The gene modification system or method as claimed in any of the preceding claims, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.

57. The gene modification system or method as claimed in any of the preceding claims, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target cells.

58. The gene modification system or method as claimed in any of the preceding claims, wherein the gene modification system comprises a second nick gRNA, and wherein the correction of the mutation is increased in the target cell population relative to a target cell population treated with a gene modification system comprising template RNA but without a second nick gRNA.

59. The gene modification system or method as claimed in any of the preceding claims, wherein the template RNA comprises one or more silent substitutions, and wherein the correction of the mutation is increased in the target cell population relative to a target cell population treated with a gene modification system comprising template RNA not containing one or more silent substitutions.

60. The method as described in any of the preceding claims, wherein the cell is a mammalian cell, such as a human cell.

61. The method as described in any of the preceding claims, wherein the subject is a human being.

62. The method of any of the preceding claims, wherein the contact occurs in vitro, for example, wherein the DNA of the cell or the subject is modified in vitro.

63. The method of any of the preceding claims, wherein the contact occurs in vivo, for example, wherein the DNA of the cell or the subject is modified in vivo.

64. The method of any of the preceding claims, wherein contacting the cell or the subject with the system comprises contacting the cell or cells in the subject with nucleic acids (e.g., DNA or RNA) encoding the genetically modified polypeptide under conditions that allow for the production of the genetically modified polypeptide.

65. A gRNA comprising (e.g., from 5' to 3'): (1) gRNA spacers; and (2) A variant St1Cas9 stent having one or more of the following: (a) Partial or complete absence of stem ring 2; (b) The elongated upper stem of the RAR; (c) Substitution of GC base pairs in the upper stem of the RAR; or (d) Mutations in the four rings.

66. A template RNA (tgRNA) comprising (e.g., from 5' to 3'): (1) gRNA spacers; (2) A variant St1Cas9 scaffold having substitutions that produce GC base pairs in the lower stem of the RAR; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence.

67. A template RNA (tgRNA) comprising (e.g., from 5' to 3'): (1) gRNA spacers; (2) A variant St1Cas9 scaffold having a substitution in the second single-chain region; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence.

Citation Information

Patent Citations

  • Method for efficient exon (44) skipping in Duchenne Muscular Dystrophy and associated means

    EP2813570A1

  • In vitro peptide or protein expression library

    GB2338237A

  • Amino acid-, peptide- and polypeptide-lipids, isomers, compositions, and uses thereof

    US10086013B2

  • Negative selection and stringency modulation in continuous evolution systems

    US10179911B2

  • Lipids and lipid nanoparticle formulations for delivery of nucleic acids

    US10221127B2