Pah adjustment systems and methods

By precisely modifying the human PAH gene using a gene modification system to correct mutations, the problem of low genome insertion frequency and site specificity in existing technologies has been solved, providing an effective treatment for PKU.

CN121079409APending Publication Date: 2025-12-05FLAGSHIP PIONEERING INNOVATIONS VI LLC
View PDF 163 Cites 0 Cited by

Patent Information

Application Number
CN202480028306.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2024-03-14
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies have low frequency and site specificity for inserting target nucleic acids into the genome, especially when integrating long sequences. Furthermore, existing methods such as CRISPR/Cas9 rely on host repair pathways and are not very efficient, making it difficult to effectively treat human phenylalanine hydroxylase (PAH) gene mutations in phenylketonuria (PKU).

Method used

A gene modification system is employed, comprising a gene-modified polypeptide encoding a Cas9 nickase with reverse transcriptase domain and endonuclease activity, and template RNA. Through the combination of gRNA spacers, gRNA scaffolds, heterologous target sequences, and primer binding site sequences, the human PAH gene is precisely modified to correct mutations and insert or delete specific DNA sequences.

Benefits of technology

This technology enables efficient and precise modification of the human PAH genome, correcting mutations, increasing the frequency and site specificity of genome editing, and providing an effective treatment for PKU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure provides compositions, systems, and methods for targeting, editing, modifying, or manipulating a host cell genome at one or more locations of a DNA sequence in a cell, tissue, or subject, for example. Genetically modified systems for the treatment of phenylketonuria (PKU) are described.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] sequence list

[0002] This application contains a sequence list, which has been filed electronically in XML format conforming to WIPO Standard ST.26 and is hereby incorporated by reference in its entirety. The XML copy was created on March 12, 2024, named V2065-7048WO_SL.xml, and is 8,922,351 bytes in size.

[0003] Cross-references to related applications

[0004] This application claims the benefits of U.S. Provisional Application No. 63 / 490,442, filed March 15, 2023; U.S. Provisional Application No. 63 / 491,490, filed March 21, 2023; and U.S. Provisional Application No. 63 / 600,472, filed November 17, 2023. The contents of the foregoing applications are incorporated herein by reference in their entirety. Background Technology

[0005] Without specific proteins to facilitate insertion events, the integration of target nucleic acids into the genome is infrequent and site-specific. Some existing methods, such as CRISPR / Cas9, are better suited for small edits that rely on host repair pathways and are less efficient at integrating longer sequences. Other existing methods, such as Cre / loxP, require a first step of inserting the loxP site into the genome, followed by a second step of inserting the target sequence into the loxP site. There is a need in the art for improved compositions (e.g., proteins and nucleic acids) and methods for inserting, altering, or deleting target sequences into the genome.

[0006] PKU is a hereditary disorder involving an autosomal recessive congenital metabolic defect caused by a deficiency of the liver enzyme PAH. PAH catalyzes the hydroxylation of phenylalanine to tyrosine, the rate-limiting step in phenylalanine metabolism. This reaction depends on tetrahydrobiopterin (BH4) as a cofactor, molecular oxygen, and iron. Loss-of-function mutations in one or both copies of the PAH gene result in the enzyme losing function or reduced efficiency. This ultimately leads to a phenotypically severe form of PKU, in which phenylalanine can accumulate in the blood to toxic concentrations, and plasma tyrosine levels are impaired. Additionally, this deficiency prevents the normal synthesis of downstream products, including dopamine, norepinephrine, and melanin.

[0007] The PAH genome sequence and its flanking regions span approximately 171 kb and contain 13 exons. Studies of pathogenic allele variants have identified more than 500 different pathogenic mutations in the PAH gene (Mitchell et al. Genet Med. [Medical Genetics] 2011; 13:697-707). Of these mutations, approximately 62% were classified as missense, 13% as deletion, 11% as splicing, 6% as silencing, 5% as nonsense, 2% as insertion, and <1% as exon deletion or duplication. Using enzyme kinetics and crystallography, the effects of several PAH mutations on enzyme activity have been described to characterize these mutations. Mutations affecting the catalytic binding mode (including Y138F, S23A, and Y377F) have been observed to reduce the tendency for tetramer formation (Flydal et al. PNAS. [Proceedings of the National Academy of Sciences] 2019; 116(23):11229-34). Other residues that interact with BH4 in the precatalytic conformation (amino acids 245–255, 286, 322, and 325) also interact with BH4 in the catalytic conformation. Furthermore, these sites are actually associated with severe instability of PAH.

[0008] Naturally occurring N-terminal PAH mutations have been identified as being distributed in a non-random pattern, clustering within residues 46-48 (GAL motif) and 65-69 (IESRP motif (SEQ ID NO: 37634)), both of which are highly conserved in pyruvate dehydrogenase (PDH) (Gjetting et al. Am. / J. Hum. Genet. [American Journal of Human Genetics] 2001; 68:1353-60). Structural-functional studies have shown that mutations in these regions significantly reduce phenylalanine binding. To date, most missense mutations found in PKU have led to phenotypic outcomes associated with PAH enzyme misfolding, increased protein turnover, and loss of enzyme function. Residues in exons 7-9 and the interdomain region within the subunit appear to play important structural roles and constitute unstable hotspots. Furthermore, using recombinant forms of hPAH, mutations within the BH4 responsive domain (including R408W and Y414C) showed residual activity but disrupted allosteric changes, indicating a conformational alteration of the protein (Gersting et al., Hum. Genet. 2008; 83:5-17). Mutational and structural-functional analyses have established a reliable genotype-phenotype profile of the role of PAHs in PKU; however, no successful cure exists other than lifelong symptom management strategies.

[0009] Since its introduction in 1953, phenylalanine (Phe) diet therapy has been a primary treatment for PKU. In the 1970s, combination therapy with tetrahydrobiopterin (BH4) and neurotransmitter precursors (L-DOPA / carbidopa and 5-hydroxytryptophan) showed promise in regulating PKU. Since its establishment as a therapy, compounds such as sapropterin have been formulated as small molecular isomers of BH4. Nevertheless, this form of therapy is generally only effective for patients with a mild PAH-deficient subset of PKU. The response to the therapy is thought to be related to PAH gene mutations, resulting in some residual enzymatic activity. At high blood concentrations, Phe in the blood competes with other large neutral amino acids (LNAAs) for transport across the blood-brain barrier. Although increased plasma Phe levels have been observed, LNAA supplementation has been shown to decrease brain Phe concentrations. Similarly, dietary supplementation with glycomacropeptides (GMP) has been observed to significantly reduce urea production, improve protein retention, and enhance Phe utilization. Nevertheless, these strategies have had little effect on addressing elevated blood levels of Phe or genotype driver factors.

[0010] Modern non-dietary approaches include the development of PAH-based fusion proteins and enzyme replacement therapies. Enzyme replacement therapies may involve administering phenylalanine ammonia-lyase (PAL) to patients. PAL is an enzyme that catalyzes the conversion of Phe into trans-cinnamic acid and a small amount of ammonia. Early studies administering PAL in enteric-coated gelatin capsules to PKU patients showed reduced Phe levels; however, repeated in vivo administration led to an increased immune response. From a clinical perspective, however, these methods are not practical because the limited half-life of the circulating enzyme necessitates multiple intravenous injections. Gene therapy has shown some promise in rescuing PAH function, such as using viral vectors. However, the efficacy of this strategy is hampered by extremely low gene transfer rates and transient transgene expression. Therefore, new and more effective treatments are needed to target PAH in PKU. Summary of the Invention

[0011] This disclosure relates to novel compositions, systems, and methods for altering the genome at one or more locations in a host cell, tissue, or subject, either in vivo or in vitro. This disclosure provides gene-modification systems capable of modulating (e.g., inserting, altering, or deleting a target sequence) phenylalanine hydroxylase (PAH) activity and methods for treating phenylketonuria (PKU), which are performed by administering one or more such systems to alter the genomic sequence within the PAH gene on human chromosome 12q23.2, which is involved as a genetic driver in PKU, for example, by correcting mutations therein.

[0012] In one aspect, this disclosure relates to a system for modifying DNA to correct mutations in the human PAH gene causing PKU, the system comprising (a) a nucleic acid encoding a gene-modifying polypeptide capable of targeted reverse transcription, the polypeptide comprising (i) a reverse transcriptase domain and (ii) a Cas9 nickase that binds to DNA and has endonuclease activity; and (b) a template RNA comprising (i) a gRNA spacer complementary to a first portion of the human PAH gene, (ii) a gRNA scaffold for binding the polypeptide, (iii) a heterologous target sequence comprising a mutation region to correct the mutation, and (iv) a primer-binding site (PBS) sequence comprising at least 3, 4, 5, 6, 7, or 8 bases having 100% homology to the target DNA strand at the 3' end of the template RNA. In some embodiments, the PAH gene may contain an R408W mutation. The template RNA sequence may include the sequences described herein, such as those in Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3.

[0013] The gRNA spacer may contain at least 15 bases that are 100% homologous to the target DNA at the 5' end of the template RNA. The template RNA may further contain a PBS sequence containing at least 5 bases that are at least 80% homologous to the target DNA strand. The template RNA may contain one or more chemical modifications.

[0014] The domains of a genetically modified polypeptide can be linked by peptide linkers. The polypeptide may contain one or more peptide linkers. The genetically modified polypeptide may further contain nuclear localization signals. The polypeptide may contain more than one nuclear localization signal, for example, multiple adjacent nuclear localization signals or one or more nuclear localization signals in different regions of the polypeptide, such as one or more nuclear localization signals at the N-terminus and one or more nuclear localization signals at the C-terminus of the polypeptide. The nucleic acid encoding the genetically modified polypeptide may encode one or more integrin domains.

[0015] Introducing this system into target cells can result in the insertion of at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 500, or 1000 base pairs of exogenous DNA. Introducing this system into target cells can also result in deletions, where the deletion is less than 2, 3, 4, 5, 10, 50, or 100 base pairs of genomic DNA upstream or downstream of the insertion. Introducing this system into target cells can also result in substitutions, such as substitutions of 1, 2, or 3 nucleotides (e.g., consecutive nucleotides).

[0016] The heterologous object sequence can be at least 5, 10, 25, 50, 100, 150, 200, 250, 300, 400, 500, 600 or 700 base pairs.

[0017] In one aspect, this disclosure relates to a pharmaceutical composition comprising the above-described system and a pharmaceutically acceptable excipient or carrier, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles. In another aspect, this disclosure relates to a pharmaceutical composition comprising the above-described system and a plurality of pharmaceutically acceptable excipients or carriers, wherein the pharmaceutically acceptable excipients or carriers are selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles, for example, wherein the above-described system is delivered by two different excipients or carriers, such as two lipid nanoparticles, two viral vectors, or one lipid nanoparticle and one viral vector. The viral vector may be adeno-associated virus (AAV).

[0018] In one aspect, this disclosure relates to a host cell (e.g., a mammalian cell, such as a human cell) that contains the aforementioned system.

[0019] In one aspect, this disclosure relates to a method for correcting human PAH gene mutations in cells, tissues, or subjects, the method comprising administering the aforementioned system to the cell, tissue, or subject, wherein optionally, correction of the mutated PAH gene comprises an amino acid substitution of W408R (reversing the pathogenic substitution R408W). The system can be introduced in vivo, in vitro, ex vivo, or in situ. The nucleic acid of (a) can be integrated into the genome of the host cell. In some embodiments, the nucleic acid of (a) is not integrated into the genome of the host cell. In some embodiments, the heterologous object sequence is inserted into only one target site in the host cell genome. The heterologous object sequence can be inserted into two or more target sites in the host cell genome, for example, into the same corresponding site on two homologous chromosomes, or into two different sites on the same or different chromosomes. The heterologous object sequence can encode a mammalian polypeptide or a fragment or variant thereof. Components of the system can be delivered on 1, 2, 3, 4, or more different nucleic acid molecules. The system can be introduced into host cells by electroporation or by using at least one medium selected from plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

[0020] These compositions or methods may include one or more of the features listed in the examples below.

[0021] Listed Examples

[0022] 1. A gene modification system comprising:

[0023] (a) Template RNA (tgRNA), which contains from 5' to 3':

[0024] (1) gRNA spacers;

[0025] (2) gRNA scaffold;

[0026] (3) Heterogeneous object sequences; and

[0027] (4) Primer binding site (PBS) sequence;

[0028] The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0029] (b) A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising:

[0030] (1) Cas structure domain;

[0031] (2) Connector; and

[0032] (3) Reverse transcriptase (RT) domain;

[0033] The genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

[0034] 2. A gene modification system comprising:

[0035] (a) Template RNA (tgRNA), which contains from 5' to 3':

[0036] (1) gRNA spacer, having the sequence of the gRNA spacer of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or having a sequence relative to it with no more than 1, 2 or 3 sequence changes (e.g., substitutions);

[0037] (2) gRNA scaffold having the gRNA scaffold sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it;

[0038] (3) A heterologous object sequence having a sequence of the template RNA having a heterologous object sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than one, two or three sequence changes relative to it; and

[0039] (4) Primer binding site (PBS) sequence, having a PBS sequence of the template RNA of the form 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having no more than one, two, or three sequence changes relative to it; and

[0040] (b) A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising:

[0041] (1) Cas structure domain;

[0042] (2) Connector; and

[0043] (3) Reverse transcriptase (RT) domain;

[0044] The genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

[0045] 3. The system as described in Example 1 or 2, wherein the template RNA comprises the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 95% identity with it.

[0046] 4. The system as described in Example 1 or 2, wherein the template RNA comprises the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 97%, 98% or 99% identity with it.

[0047] 5. The system as described in Example 1 or 2, wherein the template RNA comprises the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3.

[0048] 6. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 2, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0049] 7. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 2.

[0050] 8. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37638, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0051] 9. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37638.

[0052] 10. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37654, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0053] 11. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37654.

[0054] 12. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37653, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0055] 13. The system as described in any of the foregoing embodiments, wherein the template RNA comprises the nucleotide sequence of SEQ ID NO: 37653.

[0056] 14. The system as described in any one of Examples 1-13, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 95% identity with it.

[0057] 15. The system as described in any one of Examples 1-13, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 99% identity with it.

[0058] 16. The system as described in any one of Examples 1-13, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28.

[0059] 17. The system of any one of Examples 1-16, comprising the nucleic acid encoding the genetically modified polypeptide, wherein the nucleic acid encoding the genetically modified polypeptide has a sequence according to SEQ ID NO: 106, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0060] 18. The system of any one of Examples 1-16, comprising the nucleic acid encoding the genetically modified polypeptide, wherein the nucleic acid encoding the genetically modified polypeptide has the sequence according to SEQ ID NO: 106.

[0061] 19. The system as described in any one of Examples 1-18, further comprising a second nick RNA (ngRNA) that directs the second nick to the second strand of the human PAH gene.

[0062] 20. The system as described in Example 19, wherein the ngRNA comprises the sequence of an ngRNA of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0063] 21. The system as described in Example 19, wherein the ngRNA comprises from 5' to 3':

[0064] (1) A gRNA spacer having the sequence of an ngRNA spacer of any of the following table: E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it; and

[0065] (2) gRNA scaffold having the sequence of the gRNA scaffold of the ngRNA in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it.

[0066] 22. The system as described in Example 19 or 20, wherein the ngRNA and the template RNA of the gene modification system have a “PAM-in-the-inclination”.

[0067] 23. The system as described in any one of Examples 19-22, wherein the ngRNA has a sequence according to SEQ ID NO: 23, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0068] 24. The system as described in any one of Examples 19-22, wherein the ngRNA has the sequence according to SEQ ID NO: 23.

[0069] 25. The system as described in any one of Examples 19-22, wherein the ngRNA has a sequence according to SEQ ID NO: 84, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0070] 26. The system as described in any one of Examples 19-22, wherein the ngRNA has the sequence according to SEQ ID NO: 84.

[0071] 27. The system as described in any one of Examples 1-26, wherein the nucleic acid encoding the gene-modified polypeptide comprises RNA, for example, mRNA.

[0072] 28. The system of any one of Examples 1-27, wherein the nucleic acid encoding the genetically modified polypeptide comprises one or more chemically modified nucleotides.

[0073] 29. A nucleic acid molecule encoding a gene-modified polypeptide, wherein the nucleic acid comprises, for example, from 5' to 3':

[0074] (a) The 5' UTR of SEQ ID NO: 41-44, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0075] (b) The region of the N-terminal NLS encoding SEQ ID NO: 36, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0076] (c) The region encoding the Cas domain of SEQ ID NO: 52, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0077] (d) The region of the connector encoding SEQ ID NO: 54, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0078] (e) The region encoding the reverse transcriptase (RT) domain of SEQ ID NO: 56, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0079] (f) The region of the C-terminal NLS encoding SEQ ID NO: 38, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0080] (g) 3' UTR of SEQ ID NO: 45-48, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0081] (h) Optionally, the expression element of SEQ ID NO: 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0082] (i) The poly(A) tail of SEQ ID NO: 49-51, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

[0083] 30. The nucleic acid molecule as described in Example 29, having a sequence according to SEQ ID NO: 106, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0084] 31. The nucleic acid molecule as described in Example 29, having the sequence according to SEQ ID NO: 106.

[0085] 32. A nucleic acid molecule encoding a gene-modified polypeptide, wherein the nucleic acid comprises, for example, from 5' to 3':

[0086] (a) The 5' UTR of SEQ ID NO: 41-44, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0087] (b) The region of the N-terminal NLS encoding SEQ ID NO: 37, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0088] (c) The region encoding the Cas domain of SEQ ID NO: 53, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0089] (d) The region of the connector encoding SEQ ID NO: 55, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0090] (e) The region encoding the reverse transcriptase (RT) domain of SEQ ID NO: 57, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0091] (f) The region of the C-terminal NLS encoding SEQ ID NO: 39, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0092] (g) 3' UTR of SEQ ID NO: 45-48, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it;

[0093] (h) Optionally, the expression element of SEQ ID NO: 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0094] (i) The poly(A) tail of SEQ ID NO: 49-51, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

[0095] 33. The nucleic acid molecule as described in any one of Examples 29-32, wherein the Cas domain comprises a Cas domain that binds to the target DNA molecule and is heterologous to the RT domain.

[0096] 34. The nucleic acid molecule as described in any one of Examples 29-33, wherein the nucleic acid molecule is mRNA.

[0097] 35. The nucleic acid molecule as described in any one of Examples 29-34, wherein the nucleic acid molecule comprises one or more chemically modified nucleotides.

[0098] 36. A template RNA comprising:

[0099] (1) gRNA spacers;

[0100] (2) gRNA scaffold;

[0101] (3) Heterogeneous object sequences; and

[0102] (4) Primer binding site (PBS) sequence;

[0103] Furthermore, the template RNA contains the nucleotide sequence of the template RNA of SEQ ID NO: 15, 17, 92 or 94, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity with it.

[0104] 37. The nucleic acid molecule as described in Example 36, wherein the template RNA comprises one or more chemically modified nucleotides.

[0105] 38. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37638, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0106] 39. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37638.

[0107] 40. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37653, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

[0108] 41. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37653.

[0109] 42. A gene modification system comprising:

[0110] (a) Template RNA (tgRNA), which contains from 5' to 3':

[0111] (1) gRNA spacers;

[0112] (2) gRNA scaffold;

[0113] (3) Heterogeneous object sequences; and

[0114] (4) Primer binding site (PBS) sequence;

[0115] The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0116] (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising:

[0117] (1) Cas structure domain;

[0118] (2) Connector; and

[0119] (3) Reverse transcriptase (RT) domain;

[0120] The nucleotide encoding the gene-modified polypeptide comprises a nucleic acid as described in any one of Examples 29-35.

[0121] 43. A gene modification system comprising:

[0122] (a) Template RNA (tgRNA), which contains from 5' to 3':

[0123] (1) gRNA spacers;

[0124] (2) gRNA scaffold;

[0125] (3) Heterogeneous object sequences; and

[0126] (4) Primer binding site (PBS) sequence;

[0127] The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and

[0128] (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising:

[0129] (1) Cas structure domain;

[0130] (2) Connector; and

[0131] (3) Reverse transcriptase (RT) domain;

[0132] The nucleotide encoding the genetically modified polypeptide comprises a nucleic acid sequence of a genetically modified polypeptide listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0133] 44. A gene modification system comprising:

[0134] (a) Template RNA (tgRNA), which contains from 5' to 3':

[0135] (1) gRNA spacer, having the sequence of the gRNA spacer of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or having a sequence relative to it with no more than 1, 2 or 3 sequence changes (e.g., substitutions);

[0136] (2) gRNA scaffold having the gRNA scaffold sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it;

[0137] (3) A heterologous object sequence having a sequence of the template RNA having a heterologous object sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than one, two or three sequence changes relative to it; and

[0138] (4) Primer binding site (PBS) sequence, having a PBS sequence of the template RNA of the form 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having no more than one, two, or three sequence changes relative to it; and

[0139] (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising:

[0140] (1) Cas structure domain;

[0141] (2) Connector; and

[0142] (3) Reverse transcriptase (RT) domain;

[0143] The nucleotide encoding the genetically modified polypeptide comprises a nucleic acid sequence of a genetically modified polypeptide listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0144] 45. The system as described in any one of Examples 42-44, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 95% identity with it.

[0145] 46. ​​The system as described in any one of Examples 42-44, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 97%, 98% or 99% identity with it.

[0146] 47. The system as described in any one of Examples 42-44, wherein the template RNA comprises the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3.

[0147] 48. The system of any one of Examples 42-47, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15, or a sequence having at least 95% identity with it.

[0148] 49. The system of any one of Examples 42-47, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15, or a sequence having at least 99% identity with it.

[0149] 50. The system as described in any one of Examples 42-47, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15.

[0150] 51. The system as described in any one of Examples 42-50, further comprising a second nick RNA (ngRNA) that directs the second nick to the second strand of the human PAH gene.

[0151] 52. The system as described in Example 51, wherein the ngRNA comprises the sequence of an ngRNA of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0152] 53. The system as described in Example 51, wherein the ngRNA comprises from 5' to 3':

[0153] (1) A gRNA spacer having the sequence of an ngRNA spacer of any of the following table: E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it; and

[0154] (2) gRNA scaffold having the sequence of the gRNA scaffold of the ngRNA in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it.

[0155] 54. The system as described in Example 51 or 52, wherein the ngRNA and the template RNA of the gene modification system have a “PAM-in-internal orientation”.

[0156] 55. The system as described in any one of Examples 42-54, wherein the nucleic acid encoding the gene-modified polypeptide comprises RNA, for example, mRNA.

[0157] 56. The system of any one of Examples 42-55, wherein the nucleic acid encoding the genetically modified polypeptide comprises one or more chemically modified nucleotides.

[0158] 57. The nucleic acid molecule as described in any one of Examples 29-35 or the template RNA as described in any one of Examples 36-41, wherein the nucleic acid molecule is formulated in lipid nanoparticles (LNPs).

[0159] 58. The system as described in any one of Examples 1-28 or 42-56, wherein the tgRNA, the nucleic acid molecule encoding the gene-modified polypeptide, and / or the ngRNA are formulated in an LNP.

[0160] 59. A pharmaceutical composition comprising a system as described in any one of Examples 1-28 or 42-56, or one or more nucleic acids encoding the system, and a pharmaceutically acceptable excipient or carrier.

[0161] 60. The pharmaceutical composition as described in Example 59, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

[0162] 61. The pharmaceutical composition as described in Example 60, wherein the viral vector is adeno-associated virus.

[0163] 62. A host cell (e.g., a mammalian cell, such as a human cell) comprising a nucleic acid molecule, a gene modification system, or a template RNA as described in any of the foregoing embodiments.

[0164] 63. A method for preparing a nucleic acid molecule or template RNA as described in any of the foregoing embodiments, the method comprising synthesizing the nucleic acid molecule or template RNA by means of: in vitro (e.g., by in vitro transcription or solid-state synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow the generation of the template RNA.

[0165] 64. A method for modifying a target site in a human PAH gene in a cell, the method comprising contacting the cell with a gene modification system as described in any one of Examples 1-28 or 42-56, or DNA encoding the gene modification system, or a pharmaceutical composition as described in any one of Examples 59-61, thereby modifying the target site in the human PAH gene in the cell.

[0166] 65. A method for treating a subject suffering from a disease or condition associated with a human PAH gene mutation, the method comprising administering to the subject a gene-modifying system, or DNA encoding the gene-modifying system, or a pharmaceutical composition, as described in any one of Examples 1-28, 42-56, or 58, or any one of Examples 59-61, thereby treating the subject suffering from a disease or condition associated with a human PAH gene mutation.

[0167] 66. The method as described in Example 65, wherein the disease or condition is phenylketonuria (PKU) or hyperphenylalaninemia (e.g., mild or severe hyperphenylalaninemia).

[0168] 67. The method as described in Example 65 or 66, wherein the subject has the R408W mutation.

[0169] 68. A method for treating a subject with PKU, the method comprising administering to the subject a gene-modifying system, or DNA encoding the gene-modifying system, or a pharmaceutical composition, as described in any one of Examples 1-28, 42-56 or 58, or any one of Examples 59-61, thereby treating the subject with PKU.

[0170] 69. The gene modification system or method as described in any of the foregoing embodiments, wherein introducing the system into target cells can correct pathogenic mutations in the PAH gene.

[0171] 70. The gene modification system or method as described in any of the foregoing embodiments, wherein the pathogenic mutation is an R408W mutation, and wherein the correction comprises an amino acid substitution of W408R.

[0172] 71. The gene modification system or method as described in any of the foregoing embodiments, wherein introducing the system into a target cell results in a mutation that leads to the restoration of function of the PAH gene.

[0173] 72. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.

[0174] 73. The gene modification system or method as described in any of the foregoing embodiments, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target cells.

[0175] 74. The gene modification system or method as described in any of the foregoing embodiments, wherein the gene modification system comprises a second-strand targeting gRNA, and wherein the correction of mutations in the target cell population is increased relative to a target cell population treated with a gene modification system comprising template RNA but without a second-strand targeting gRNA.

[0176] 75. The method as described in any of the foregoing embodiments, wherein the cell is a mammalian cell, such as a human cell.

[0177] 76. The method as described in any of the foregoing embodiments, wherein the subject is a human.

[0178] 77. The method as described in any of the foregoing embodiments, wherein the contact occurs in vitro, for example, wherein the DNA of the cell or the subject is modified in vitro.

[0179] 78. The method as described in any of the foregoing embodiments, wherein the contact occurs in vivo, for example, wherein the DNA of the cell or the subject is modified in vivo.

[0180] 79. The method as described in any of the foregoing embodiments, wherein contacting the cell or the subject with the system comprises contacting the cell or cells in the subject’s body with nucleic acids (e.g., DNA or RNA) encoding the genetically modified polypeptide under conditions that allow the generation of the genetically modified polypeptide.

[0181] 80. The template RNA or system as described in any of the foregoing embodiments, wherein the heterologous object sequence comprises a region of at least 5 consecutive nucleotides, the region comprising 2'-fluorine modifications on alternating nucleotides.

[0182] 81. The template RNA or system as described in Example 80, wherein the heterologous object sequence comprises a 2'-fluorine modification on alternating nucleotides starting at position +4 of the heterologous object sequence.

[0183] 82. The template RNA or system as described in Example 81, wherein the heterologous object sequence comprises a 2'-fluorine modification on alternating nucleotides at positions +4 to +8, +10, +12, +14, +16, or +18 of the heterologous object sequence.

[0184] 83. The template RNA or system as described in Example 80, wherein the heterologous object sequence comprises 2'-fluorine modifications on alternating nucleotides in a region of length 2-5, 5-10, 10-15, 15-20, 20-25, 25-30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, or 80-81 nucleotides.

[0185] 84. The template RNA or system as described in Example 80, wherein the heterologous object sequence comprises 2'-fluorine modifications on alternating nucleotides in a region of length 80-90, 90-100, 100-150, 150-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1500, 1500-2000, 2000-2500, 2500-3000, 3000-3500, 3500-4000, 4000-4500, or 4500-5000 nucleotides.

[0186] 85. The template RNA or system as described in Example 80, wherein the heterologous object sequence comprises a 2'-fluorine modification on alternating nucleotides starting at position +5 of the heterologous object sequence.

[0187] 86. The template RNA or system as described in Example 81, wherein the heterologous object sequence comprises 2'-fluorine modification on alternating nucleotides at positions +5 to +9, +11, +13, +15, or +17 of the heterologous object sequence.

[0188] 87. The template RNA or system as described in any one of Examples 80-86, wherein the template RNA or system comprises a 2'-fluorinated nucleotide at the 5' end of the heterologous object sequence (e.g., at position +10 or +11 of the heterologous object sequence).

[0189] 88. The template RNA or system as described in Example 80, wherein the template RNA or system comprises 2'-fluorine modified nucleotides at positions +4, +6, +8 and / or +10 of the heterologous object sequence.

[0190] 89. The template RNA or system as described in Example 80, wherein the template RNA or system comprises 2'-fluorine modified nucleotides at positions +5, +7, +9 and / or +11 of the heterologous object sequence.

[0191] 90. The template RNA or system as described in Example 80, wherein the second nucleotide from the 5' end of the heterologous object sequence contains a 2'-fluorine modification (e.g., position +9 or +10 of the heterologous object sequence).

[0192] 91. The template RNA or system as described in Example 80, wherein the template RNA or system comprises 2'-fluorine modified nucleotides at positions +5, +7 and / or +9 of the heterologous object sequence.

[0193] 92. The template RNA or system as described in Example 80, wherein the template RNA or system comprises 2'-fluorine modified nucleotides at positions +4, +6, +8 and / or +10 of the heterologous object sequence. Attached Figure Description

[0194] Figure 1 This is a graph showing the Phe levels in the plasma of mice treated with an exemplary gene-modification system.

[0195] Figure 2A This is a graph showing the rewrite level % of the template RNA and its spacer corresponding to the second cut for each test.

[0196] Figure 2B This is a graph showing the insertion / deletion activity percentage of the template RNA and its spacer corresponding to the second nick for each test.

[0197] Figure 3A This is a graph showing the rewriting activity in primary mouse hepatocytes given the indicated LNP (containing the indicated dose of template RNA, second nick guide, and mRNA encoding the RNA IVT338 gene modification polypeptide), as measured by Amp-SEQ.

[0198] Figure 3B This is a graph showing the percentage of insertion / deletion activity in primary mouse hepatocytes treated with the indicated LNP (containing the indicated dose of template RNA, second nick guide, and mRNA encoding the RNA IVT338 gene modification polypeptide).

[0199] Figure 4A This is a graph showing rewriting activity in primary mouse hepatocytes transfected with the indicated dose of template RNA, second nick guide, and RNAIVT338 gene-modified peptide, as measured by Amp-SEQ.

[0200] Figure 4BThis is a graph showing the percentage of insertion / deletion activity in primary mouse hepatocytes transfected with the indicated dose of template RNA, second nick guide, and RNAIVT338 gene-modified peptide.

[0201] Figure 5 This is a graph showing the level of Phe in the plasma of mice treated with a gene modification system containing the indicated template RNA and a second nick.

[0202] Figure 6A This is a graph showing the percentage of rewrite levels achieved by treatment with a gene modification system containing the indicated template RNA and a second nick.

[0203] Figure 6B This is a graph showing the percentage of insertion / deletion activity obtained by treating hPAH mice with a gene modification system containing the indicated template RNA and a second nick.

[0204] Figure 7A This is a graph showing the percentage of rewrite level of each test template RNA in mouse livers without second-incision guidance (left) and with second-incision guidance (right).

[0205] Figure 7B This is a graph showing the percentage of insertion / deletion activity of each test template RNA in mouse livers without second-incision guidance (left) and with second-incision guidance (right).

[0206] Figure 8A This is a graph showing the percentage of rewrite levels obtained using gene modification systems containing various mRNAs containing RNACS4134 and RNACS1810, as well as gene-modifying polypeptides.

[0207] Figure 8B This is a graph showing the percentage of insertion / deletion activity obtained using gene modification systems containing various mRNAs encoding gene-modifying peptides.

[0208] Figure 9A These are a pair of graphs showing the percentage rewrite level of various mRNAs containing RNACS4134 template RNA and RNACS1809 or RNACS1810 second nick guides and gene-modifying polypeptides.

[0209] Figure 9B These are a pair of graphs showing the insertion / deletion activity percentage of gene modification systems for various mRNAs containing RNACS4134 template RNA and RNACS1809 or RNACS1810 second nick guides and gene modification polypeptides.

[0210] Figure 10AThese are a pair of graphs showing the levels of Phe in the plasma of treated mice when treated with the indicated dose of the gene-modifying system (containing RNACS4134 template RNA and RNACS1809 or RNACS1810 second nick guides, as well as various mRNAs encoding gene-modifying peptides).

[0211] Figure 10B These are a pair of graphs showing the levels of Phe in the brains of mice treated with the indicated dose of a gene-modification system containing RNACS4134 template RNA and RNACS1809 or RNACS1810 second nick guides, as well as various mRNAs encoding gene-modification peptides.

[0212] Figure 11 The gene modification system described herein is depicted. The left-hand diagram shows the gene-modifying polypeptide, which contains a Cas nicking enzyme domain (e.g., spCas9 N863A) and a reverse transcriptase domain (RT domain) linked by a linker. The right-hand diagram shows the template RNA, which contains a gRNA spacer, a gRNA scaffold, a heterologous target sequence, and a primer binding site sequence (PBS sequence) from 5' to 3'. The heterologous target sequence may contain a mutant region that contains one or more sequence differences relative to the target site. The heterologous target sequence may also contain pre-edited homologous regions and post-edited homologous regions flanking the mutant region. Without being bound by theory, it is assumed that the gRNA spacer of the template RNA binds to the second strand of the target site in the genome, and that the gRNA scaffold of the template RNA binds to the gene-modifying polypeptide, for example, to locate the gene-modifying polypeptide to the target site in the genome. It is assumed that the Cas domain of the gene-modifying polypeptide creates a nick at the target site (e.g., the first strand of the target site), for example, allowing the PBS sequence to bind to a sequence adjacent to the site to be modified on the first strand of the target site. This theory posits that the RT domain of a gene-modified polypeptide uses the first strand of a target site that binds to a complementary sequence (a PBS sequence containing the template RNA) as a primer, and a heterologous target sequence of the template RNA as a template, such as a sequence complementary to the heterologous target sequence. Not wanting to be bound by theory, it is proposed that reverse transcription can then proceed through the pre-edited homologous region, then through the mutated region, and then through the post-edited homologous region, thereby producing a DNA strand containing the mutation specified by the heterologous target sequence.

[0213] Figure 12This is a series of graphs showing the heterologous target sequence and PBS (priming) sequence of a series of template RNA variants (each containing a 2'-fluorine modification, as shown in the gray box in the top table). Two of the variants include alternating patterns of 2'-fluorine modification, with every other nucleotide in the subsequence of the heterologous target sequence containing a 2'-fluorine modification. These variants further contain, in 5' to 3' order, a 2'-fluorine modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides each containing a 2'-OMe and a phosphate thioester modification at the 3' end of the priming region. The ability of these template RNA variants to introduce alterations to the target nucleic acid sequence was tested, with the resulting rewriting efficiency and the percentage of insertions or deletions (insertions / deletions) for each variant shown in the lower left and lower right graphs, respectively. The graphs disclose SEQ ID NOs 37647-37649 in the order they appear.

[0214] Figure 13 This is a series of graphs showing the heterologous target sequence and PBS (priming) sequence of a series of template RNA variants (each containing a 2'-fluorine modification, as shown in the gray box in the top table). Two of the variants include alternating patterns of 2'-fluorine modification, with every other nucleotide in the subsequence of the heterologous target sequence containing a 2'-fluorine modification. These variants further contain, in 5' to 3' order, a 2'-fluorine modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides each containing 2'-OMe and / or phosphate thioester modifications at the 3' end of the priming region. The ability of these template RNA variants to introduce alterations to the target nucleic acid sequence was tested, with the resulting rewriting efficiency and percentage of insertions / deletions for each variant shown in the lower left and lower right graphs, respectively. The graphs disclose SEQ ID NOs 37647-37649 in the order they appear.

[0215] Figure 14 This is a series of graphs showing the heterologous target sequence and PBS (priming) sequence of template RNA variants (each containing a 2'-fluorine modification, as shown in the gray box in the top table). The RNACS6874 variant includes an alternating pattern of 2'-fluorine modifications, with every other nucleotide in the subsequence of the heterologous target sequence containing a 2'-fluorine modification. This variant further contains, at the 3' end of the priming region, a 2'-fluorine modified nucleotide, three 2'-OMe modified nucleotides, and three nucleotides each containing a 2'-OMe and a phosphate thioester modification in a 5' to 3' sequence. The ability of these template RNA variants to introduce alterations to the target nucleic acid sequence was tested, with the resulting rewriting efficiency and percentage of insertions / deletions for each variant shown in the lower left and lower right graphs, respectively. SEQ ID NOs 37647-37648 are disclosed in the order of appearance.

[0216] Figure 15This is a series of graphs showing the heterologous target sequence and PBS (priming) sequence of a series of template RNA variants (each containing a 2'-fluorine modification, as shown in the gray box in the top table). Two of the variants include alternating patterns of 2'-fluorine modification, with every other nucleotide in the subsequence of the heterologous target sequence containing a 2'-fluorine modification. These variants further contain, in 5' to 3' order, a 2'-fluorine modified nucleotide, a 2'-OMe modified nucleotide, and three nucleotides, each containing a 2'-OMe and / or phosphate thioester modification, at the 3' end of the priming region. The ability of these template RNA variants to introduce alterations to the target nucleic acid sequence was tested, with the resulting rewriting efficiency and percentage of insertions / deletions for each variant shown in the lower left and lower right graphs, respectively. SEQ ID NOs 37650-37652 are disclosed in the order they appear.

[0217] Figure 16 This is a series of graphs showing the heterologous target sequence and PBS (priming) sequence of a series of template RNA variants (each containing a 2'-fluorine modification, as shown in the gray box in the top table). Two of the variants include alternating patterns of 2'-fluorine modification, with every other nucleotide in the subsequence of the heterologous target sequence containing a 2'-fluorine modification. These variants further contain, in 5' to 3' order, a 2'-fluorine modified nucleotide, a 2'-OMe modified nucleotide, and three nucleotides, each containing a 2'-OMe and / or phosphate thioester modification, at the 3' end of the priming region. The ability of these template RNA variants to introduce alterations to the target nucleic acid sequence was tested, with the resulting rewriting efficiency and percentage of insertions / deletions for each variant shown in the lower left and lower right graphs, respectively. SEQ ID NOs 37650-37652 are disclosed in the order they appear.

[0218] Figure 17 A- Figure 17 B is a pair of graphs showing the rewrite level (17A) or insertion / deletion (17B) of gene modification systems containing RNACS6874 template RNA and various mRNAs encoding gene modification polypeptides (where the mRNAs contain different versions of WPRE or lack WPRE). Detailed Implementation

[0219] definition

[0220] As used herein, the term "alternating nucleotide" in relation to chemical modifications refers to a pattern of nucleotides in which all odd-numbered nucleotides in a region have the same chemical modification, while all even-numbered nucleotides do not, or vice versa: all even-numbered nucleotides in a region have the same chemical modification, while all odd-numbered nucleotides do not. For example, in a region of five nucleotides in length with alternating nucleotides, the first, third, and fifth positions of the region may all contain a 2'F chemical modification, while the second and fourth positions may contain unmodified nucleotides or chemical modifications other than 2'F. The second and fourth positions may be the same or different. Furthermore, among all nucleotides containing the same chemical modification, one or more may contain a second chemical modification. As a non-limiting example, in the aforementioned regions with 2'F chemical modifications at the first, third, and fifth positions, if only one nucleotide at those positions further contains a main chain modification, the region still contains alternating nucleotides with respect to the 2'F chemical modification. Alternating nucleotides can be found in regions of larger nucleic acids that contain one or more other non-alternating regions.

[0221] As used herein, the term "expression cassette" refers to a nucleic acid construct containing nucleic acid elements sufficient to express the nucleic acid molecules of the present invention.

[0222] As used in this article, "gRNA spacer" refers to a nucleic acid portion that is complementary to the target nucleic acid and can work with the gRNA scaffold to target the Cas protein to the target nucleic acid.

[0223] As used herein, a "gRNA scaffold" refers to a nucleic acid motif that can bind to the Cas protein and, together with a gRNA spacer, target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold comprises a crRNA sequence, a tetracyclic RNA sequence, and a tracrRNA sequence.

[0224] As used herein, a “genetically modified polypeptide” refers to a polypeptide comprising a retroviral reverse transcriptase, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with a retroviral reverse transcriptase, capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template nucleic acid) into a target DNA molecule (e.g., in mammalian host cells, such as genomic DNA molecules in host cells). In some embodiments, the genetically modified polypeptide is capable of integrating the sequence substantially independently of the host machinery. In some embodiments, the genetically modified polypeptide integrates the sequence into a random location in the genome, and in some embodiments, the genetically modified polypeptide integrates the sequence into a specific target site. In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively facilitate 1) binding to a template nucleic acid, 2) binding to a target DNA molecule, and 3) facilitating the integration of at least a portion of the template nucleic acid into the target DNA. Genetically modified polypeptides include naturally occurring polypeptides and engineered variants of the aforementioned polypeptides, for example, those variants having one or more amino acid substitutions relative to the naturally occurring sequence. Genetically modified peptides also include heterologous constructs, such as those in which one or more of the aforementioned domains are heterologous to each other, whether through heterologous fusion (or other conjugates) of domains that are otherwise wild-type, and through fusion of modified domains, for example, through substitution or fusion of heterologous subdomains or other substituted domains. Exemplary gene-modified peptides, systems comprising them, and methods of using them that can be used in the methods provided herein are described, for example, in PCT / US2021 / 020948, which relates to gene-modified peptides comprising retroviral reverse transcriptase domains and is incorporated herein by reference. In some embodiments, the gene-modified peptide integrates a sequence into a gene. In some embodiments, the gene-modified peptide integrates a sequence into a sequence outside a gene. As used herein, a “gene-modified system” refers to a system comprising a gene-modified peptide and a template nucleic acid.

[0225] As used herein, the term "domain" refers to the structure of a biomolecule that contributes to a specific function of the biomolecule. A domain may contain a continuous region (e.g., a continuous sequence) or distinct non-continuous regions (e.g., non-continuous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcription domains; examples of nucleic acid domains are regulatory domains, such as transcription factor-binding domains. In some embodiments, a domain (e.g., a Cas domain) may contain two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).

[0226] As used herein, the term "exogenous" when used in relation to a biomolecule (such as a nucleic acid sequence or polypeptide) means the artificial introduction of the biomolecule into a host genome, cell, or organism. For example, nucleic acids added to an existing genome, cell, tissue, or subject using recombinant DNA technology or other methods are exogenous to the existing nucleic acid sequence, cell, tissue, or subject.

[0227] As used herein, the terms "first strand" and "second strand" to describe a single DNA strand of the target DNA distinguish the two DNA strands based on the strand on which the reverse transcriptase domain initiates polymerization, for example, based on the location where target-initiated synthesis begins. The first strand refers to the strand of the target DNA on which the reverse transcriptase domain initiates polymerization, for example, at the location where target-initiated synthesis begins. The second strand refers to the other strand of the target DNA. The names of the first and second strands do not otherwise describe the target site DNA strand; for example, in some embodiments, the first and second strands are cleaved by the polypeptide described herein, but the names "first" and "second" strand are not related to the order in which such cleavage occurs.

[0228] When used herein to describe a first element in reference to a second element, the term "heterogeneous" means that the first and second elements do not exist in nature in the arrangement described. For example, a heterologous polypeptide, nucleic acid molecule, construct, or sequence is (a) a polypeptide, nucleic acid molecule, or part of a polypeptide or nucleic acid molecule sequence that is not native to the cell expressing it, (b) a polypeptide or nucleic acid molecule, or part of a polypeptide or nucleic acid molecule, that has been altered or mutated relative to its native state, or (c) a polypeptide or nucleic acid molecule having altered expression compared to its native expression level under similar conditions. For example, heterologous regulatory sequences (e.g., promoters, enhancers) can be used to regulate the expression of a gene or nucleic acid molecule in a manner different from how the gene or nucleic acid molecule is normally expressed in nature. In another instance, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., the DNA-binding domain of the polypeptide or a nucleic acid encoding the DNA-binding domain of the polypeptide) may be arranged relative to other domains, or may be different sequences or may originate from different sources relative to other domains or portions of the polypeptide or its encoding nucleic acid. In some embodiments, the heterologous nucleic acid molecule may be present in the natural host cell genome, but may have altered expression levels or different sequences, or both. In other embodiments, the heterologous nucleic acid molecule may not be endogenous to the host cell or host genome, but may be introduced into the host cell through transformation (e.g., transfection, electroporation), wherein the added molecule may integrate into the host genome, or may exist transiently as extrachromosomal genetic material (e.g., mRNA) or semi-stable for more than one generation (e.g., cell-free viral vectors, plasmids, or other self-replicating vectors).

[0229] As used herein, “inserting” a sequence into a target site refers to a net addition of DNA sequence at the target site, such as the presence of new nucleotides in a heterologous object sequence that has no homologous position in the unedited target site. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous object sequence with the target nucleic acid sequence will result in alignment gaps in the target nucleic acid sequence.

[0230] As used herein, a “deletion” produced by a heterologous object sequence at a target site refers to a net deletion of the DNA sequence at the target site, such as the presence of nucleotides at an unedited target site where there is no homologous position in the heterologous object sequence. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous object sequence with the target nucleic acid sequence will result in alignment vacancies in the molecule containing the PBS sequence and the heterologous object sequence.

[0231] As used herein, the term "inverted terminal repeat sequence" or "ITR" refers to AAV viral cis-elements, so named because of their symmetry. These elements promote efficient multiplication of the AAV genome. It is assumed that the minimum functional element of an ITR is a Rep binding site (RBS; 5'-GCGCGCTCGCTCGCTC-3', for AAV2; SEQ ID NO: 4601) and a terminal dissociation site (TRS; 5'-AGTTGG-3', for AAV2) plus a variable palindromic sequence that allows hairpin formation. According to the invention, an ITR comprises at least these three elements (RBS, TRS, and the hairpin-forming sequence). Additionally, in this invention, the term "ITR" refers to the ITR of a known natural AAV serotype (e.g., the ITR of serotypes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 AAV), a chimeric ITR formed by the fusion of ITR elements derived from different serotypes, and their functional variants. "Functional variant" refers to a sequence that has at least 80%, 85%, 90%, preferably at least 95% sequence identity with a known ITR, allowing the sequence containing the ITR to multiply in the presence of the Rep protein.

[0232] As used herein, the term "mutation region" refers to a region in the template RNA that has one or more sequence differences relative to the corresponding sequence in the target nucleic acid. Sequence differences can include, for example, substitution, insertion, frameshift, or deletion.

[0233] When applied to nucleic acid sequences, the term "mutated" means that nucleotides in the nucleic acid sequence have been inserted, deleted, or altered compared to a reference (e.g., a natural) nucleic acid sequence. A single change (point mutation) can be made at a locus, or multiple nucleotides can be inserted, deleted, or altered at a single locus. Additionally, one or more changes can be made at any number of loci within the nucleic acid sequence. Nucleic acid sequences can be mutated using any method known in the art.

[0234] Nucleic acid molecules refer to both RNA and DNA molecules, including but not limited to complementary DNA (“cDNA”), genomic DNA (“gDNA”), and messenger RNA (“mRNA”), and also include synthetic nucleic acid molecules, such as those produced by chemical synthesis or recombination, such as RNA templates as described herein. Nucleic acid molecules can be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule can be sense or antisense. Unless otherwise indicated, and as an example of all sequences described herein in the general format “SEQ ID NO:”, “nucleic acid containing SEQ ID NO:1” means a nucleic acid having (i) the sequence of SEQ ID NO:1 or (ii) a sequence complementary to SEQ ID NO:1, at least a portion thereof. The choice between the two depends on the context of using SEQ ID NO:1. For example, if the nucleic acid is used as a probe, the choice between the two depends on the requirement that the probe is complementary to the desired target. As will be readily understood by those skilled in the art, the nucleic acid sequences disclosed herein may be chemically or biochemically modified or may contain non-natural or derived nucleotide bases. Such modifications include, for example, tagging, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methylphosphonates, triphosphates, aminophosphates, carbamates, etc.), charged linkages (e.g., thiophosphates, dithiophosphates, etc.), side chain moieties (e.g., polypeptides), intercalating agents (e.g., acridine, psoralen, etc.), chelating agents, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). Also included are chemically modified bases (see, for example, Table 13), main chains (see, for example, Table 14), and modified caps (see, for example, Table 15). Also included are synthetic molecules that mimic the ability of polynucleotides to bind to a specified sequence via hydrogen bonds and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages replace phosphate linkages in the molecular main chain, such as peptide nucleic acids (PNAs). Other modifications may include, for example, analogs in which the ribose ring contains a bridging moiety or other structure (e.g., modifications found in "locked" nucleic acids (LNAs). In various embodiments, nucleic acids are operatively associated with additional genetic elements (e.g., one or more tissue-specific expression control sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences)) and other elements (e.g., inverted repeat sequences (e.g., inverted terminal repeat sequences, such as elements derived from or originating from viruses, such as AAV ITRs) and tandem repeat sequences, inverted repeat sequences / direct repeat sequences, homologous regions (segments with different degrees of homology to the target DNA), untranslated regions (UTRs) (5', 3', or 5' and 3' UTRs)) and various combinations thereof.The nucleic acid elements of the system provided by this invention can be provided in a variety of topologies, including single-stranded, double-stranded, circular, linear, linear with open ends, linear with closed ends, and specific versions of these, such as dog bone DNA (dbDNA) and closed-end DNA (ceDNA).

[0235] As used herein, a “gene expression unit” is a nucleic acid sequence containing at least one regulatory nucleic acid sequence operatively linked to at least one effector sequence. The first nucleic acid sequence is operatively linked to the second nucleic acid sequence when positioned to have a functional relationship with it. For example, if a promoter or enhancer affects the transcription or expression of a coding sequence, then the promoter or enhancer is operatively linked to that coding sequence. The operatively linked DNA sequences can be contiguous or non-contiguous. In cases where it is necessary to link two protein-coding regions, the operatively linked sequences can be within the same reading frame.

[0236] As used herein, the terms “host genome” or “host cell” refer to a cell and / or its genome in which proteins and / or genetic material have been introduced. It should be understood that such terms are intended not only to refer to a specific subject cell and / or genome, but also to the genomes of the offspring of such cells and / or the offspring of such cells. Because certain modifications may occur in offspring due to mutations or environmental influences, such offspring may actually differ from the parent cell, but are still included within the scope of the term “host cell” as used herein. A host genome or host cell can be an isolated cell or cell line grown in a culture, or genomic material isolated from such a cell or cell line, or it can be a host cell or host genome constituting a living tissue or organism. In some cases, the host cell can be an animal cell or plant cell, for example, as described herein. In some cases, the host cell can be a mammalian cell, human cell, avian cell, reptile cell, bovine cell, horse cell, pig cell, goat cell, sheep cell, chicken cell, or turkey cell. In some cases, the host cell can be a corn cell, soybean cell, wheat cell, or rice cell.

[0237] As used herein, “operable association” describes a functional relationship between two nucleic acid sequences (e.g., 1) a promoter and 2) a heterologous object sequence), and in such instances, it means that the orientation of the promoter and the heterologous object sequence (e.g., the target gene) allows the promoter to drive the expression of the heterologous object sequence under appropriate conditions. For example, the template nucleic acid carrying the promoter and the heterologous object sequence can be single-stranded, e.g., with a (+) or (-) orientation. An “operable association” between the promoter and the heterologous object sequence in this template means that, regardless of whether the template nucleic acid is transcribed in a specific state, it is accurately transcribed when it is in the appropriate state (e.g., in a (+) orientation, in the presence of the required catalysts and NTPs, etc.). Operable associations similarly apply to other nucleic acid pairs, including other tissue-specific expression control sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homologous regions and heterologous object sequences or sequences encoding retroviral RT domains.

[0238] As used herein, the term "primer binding site sequence" or "PBS sequence" refers to a portion of template RNA capable of binding to a region contained in a target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence containing at least 3, 4, 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 bases that are 100% identical to a region contained in the target nucleic acid sequence. Not wishing to be bound by theory, in some embodiments, when the template RNA contains both a PBS sequence and a heterologous object sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence, thereby allowing the reverse transcriptase domain to use that region as a primer for reverse transcription and the heterologous object sequence as a template for reverse transcription.

[0239] As used herein, a "stem-loop sequence" refers to a nucleic acid sequence (e.g., an RNA sequence) that has sufficient self-complementarity to form a stem-loop, for example, having a stem containing at least two (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs and a loop having at least three (e.g., four) base pairs. The stem may contain mismatches or protrusions.

[0240] As used herein, a “tissue-specific expression control sequence” refers to a nucleic acid element that, in a tissue-specific manner, preferentially increases or decreases the level of transcripts containing a heterologous object sequence in one or more target tissues relative to one or more off-target tissues. In some embodiments, the tissue-specific expression control sequence preferentially drives or inhibits transcription, activity, or half-life of transcripts containing a heterologous object sequence in a tissue-specific manner, for example, preferentially in one or more target tissues relative to one or more off-target tissues. Exemplary tissue-specific expression control sequences include tissue-specific promoters, repressors, enhancers, or combinations thereof, and tissue-specific microRNA recognition sequences. Tissue specificity refers to target (one or more tissues where expression or activity of the template nucleic acid is desired or tolerated) and off-target (one or more tissues where expression or activity of the template nucleic acid is not desired or tolerated). For example, a tissue-specific promoter preferentially drives expression in target tissues relative to off-target tissues. Conversely, microRNAs binding to tissue-specific microRNA recognition sequences are preferentially expressed in off-target tissues relative to target tissues, thereby reducing the expression of the template nucleic acid in off-target tissues. Therefore, regarding the transcription, activity, or half-life of associated sequences in tissues, promoters and microRNA recognition sequences specific to the same tissue (e.g., target tissue) have different functions (promoting and inhibiting, respectively, with consistent expression levels, i.e., high levels of microRNA in off-target tissues and low levels in target tissues, while promoters drive high expression in target tissues and low expression in off-target tissues).

[0241] Unless otherwise specified, the following numbering system is used to describe the positions of chemically modified nucleotides in the PBS and / or heterologous object sequences of the template RNA. Nucleotides in the heterologous object sequence are numbered +1, +2, +3, etc., starting from the 3' end of the heterologous object sequence. Nucleotides in the PBS sequence are numbered -1, -2, -3, etc., starting from the 5' end of the PBS sequence. Therefore, positions +1 and -1 are directly adjacent to each other.

[0242] introduction

[0243] This disclosure relates to methods for treating phenylketonuria (PKU) and compositions for, for example, targeting, editing, modifying, or manipulating DNA sequences at one or more locations in a cell, tissue, or subject, either in vivo or in vitro (e.g., inserting a heterologous object sequence into a target site in a mammalian genome). The heterologous object DNA sequence may include, for example, substitution.

[0244] More specifically, this disclosure provides methods for treating PKU that use reverse transcriptase-based systems to alter a target genomic DNA sequence, for example, by inserting one or more nucleotides into the target sequence, deleting one or more nucleotides from the target sequence, or replacing one or more nucleotides in the target sequence.

[0245] This disclosure partially provides methods for treating PKU using a gene modification system comprising a genetically modified polypeptide component and a template nucleic acid (e.g., template RNA) component. In some embodiments, the gene modification system can be used to introduce alterations into a target site in the genome. In some embodiments, the genetically modified polypeptide component comprises a writing domain (e.g., a reverse transcriptase domain), a DNA-binding domain, and a nuclease domain (e.g., a nicking enzyme domain). In some embodiments, the template nucleic acid (e.g., template RNA) comprises a sequence (e.g., a gRNA spacer) that binds to a target site in the genome (e.g., a second strand that binds to the target site), a sequence that binds to the genetically modified polypeptide component (e.g., a gRNA scaffold), a heterologous object sequence, and a PBS sequence. It is not intended to be theoretically construed that the template nucleic acid (e.g., template RNA) binds to the second strand of the target site in the genome and binds to the genetically modified polypeptide component (e.g., to localize the polypeptide component to the target site in the genome). It is assumed that a nuclease (e.g., a nicking enzyme) of the genetically modified polypeptide component cleaves the target site (e.g., the first strand of the target site), for example, allowing the PBS sequence to bind to a sequence adjacent to the site to be altered on the first strand of the target site. It is assumed that the writing domain of a polypeptide component (e.g., a reverse transcriptase domain) uses a PBS sequence containing a template nucleic acid as a primer and a heterologous object sequence of the template nucleic acid as a template to bind to the first strand of a target site, for example, by polymerizing a sequence complementary to the heterologous object sequence. Not wanting to be bound by theory, it is assumed that choosing a suitable heterologous object sequence can lead to the substitution, deletion, and / or insertion of one or more nucleotides at the target site.

[0246] Gene modification system

[0247] In some embodiments, the gene-modifying system described herein comprises: (A) a gene-modifying polypeptide or a nucleic acid encoding the gene-modifying polypeptide, wherein the gene-modifying polypeptide comprises (i) a reverse transcriptase domain and a nuclease domain having DNA-binding functionality; and (B) a template RNA. In some embodiments, the gene-modifying polypeptide acts as a substantially autonomous protein machine capable of integrating a template nucleic acid sequence into a target DNA molecule (e.g., in mammalian host cells, such as genomic DNA molecules in host cells), substantially independent of the host machine. For example, the gene-modifying protein may comprise a DNA-binding domain, a reverse transcriptase domain, and a nuclease domain. In some embodiments, the DNA-binding function may involve an RNA component that guides the protein to a DNA sequence (e.g., a gRNA spacer). In other embodiments, the gene-modifying polypeptide may comprise a reverse transcriptase domain and a nuclease domain. The RNA template element of the gene-modifying system is typically heterologous to the gene-modifying polypeptide element and provides the target sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the gene-modifying polypeptide is capable of targeted reverse transcription. In some embodiments, the gene-modifying polypeptide is capable of second-strand synthesis.

[0248] Functional gene-modifying peptides can be composed of unrelated DNA-binding domains, reverse transcription domains, and endonuclease domains. This modular structure allows for the combination of functional domains, such as dCas9 (DNA binding), AVIRE reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains can originate from a single protein, such as Cas9 or Cas9 nickase (DNA binding, endonuclease).

[0249] In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively facilitate 1) binding to a template nucleic acid, 2) binding to a target DNA molecule, and 3) facilitating the integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the genetically modified polypeptide is an engineered polypeptide, for example, having one or more amino acid substitutions relative to a naturally occurring sequence. In some embodiments, the genetically modified polypeptide comprises two or more domains that are heterologous to each other, for example, through heterologous fusion (or other conjugates) of domains that are otherwise wild-type, or through the fusion of modified domains, for example, through substitution or fusion of heterologous subdomains or other substituted domains. For example, in some embodiments, one or more of the following are heterologous: the RT domain is heterologous to the DBD; the DBD is heterologous to the endonuclease domain; or the RT domain is heterologous to the endonuclease domain.

[0250] In some embodiments, the template RNA molecule used in this system comprises, from 5′ to 3′, (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous object sequence; and (4) a primer binding site (PBS) sequence. In some embodiments:

[0251] (1) is a gRNA spacer of approximately 18-22 nt (e.g., 20 nt).

[0252] (2) is a gRNA scaffold containing one or more hairpin loops (e.g., 1, 2, or 3 loops) for associating a template with a Cas domain, such as the nickase Cas9 domain. In some embodiments, the gRNA scaffold contains the sequence GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21) from 5′ to 3′.

[0253] (3) In some embodiments, the length of the heterogeneous object sequence is, for example, 7-74, such as 10-20, 20-30, 30-40, 40-50, 50-60, 60-70, or 70-80 nt or 80-90 nt.

[0254] (4) In some embodiments, the PBS sequence that binds to the target-initiating sequence after the nick occurs is, for example, 3-20 nt, 7-15 nt, or 12-14 nt. In some embodiments, the PBS sequence has a GC content of 40%-60%.

[0255] In some embodiments, a second gRNA associated with the system may help drive complete integration. In some embodiments, the second gRNA may target a position 0-200 nt from the first strand cut, such as 0-50, 50-100, or 100-200 nt from the first strand cut. In some embodiments, the second gRNA may bind to its target sequence only after editing, for example, the gRNA binds to a sequence present in the heterologous object sequence but not in the initial target sequence.

[0256] In some embodiments, the gene-editing system described herein is used for editing in HEK293, K562, U2OS, or HeLa cells. In some embodiments, the gene-editing system is used for editing in primary cells (e.g., human primary cells).

[0257] In some embodiments, the genetically modified polypeptide as described herein comprises a reverse transcriptase or RT domain containing an AVIRE RT sequence or a variant thereof (e.g., as described herein). In some embodiments, the endonuclease domain (e.g., as described herein) comprises a Cas9 domain (e.g., SpCas9, for example, containing an N863A mutation (e.g., in spCas9)).

[0258] In some embodiments, the length of the heterologous object sequence (e.g., the system described herein) is about 1-50, 50-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000 or more nucleotides.

[0259] In some embodiments, the RT and endonuclease domains are connected by a flexible linker, for example, containing the amino acid sequence AEAAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAA AKEAAAKA (SEQ ID NO: 54).

[0260] In some embodiments, the endonuclease domain is located at the N-terminus relative to the RT domain. In some embodiments, the endonuclease domain is located at the C-terminus relative to the RT domain.

[0261] In some embodiments, the system incorporates heterologous object sequences into target sites via TPRT, for example, as described herein.

[0262] In some embodiments, the gene modification system is capable of generating at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides) of insertion at the target site. In some embodiments, the gene modification system is capable of generating at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides) of insertion at the target site. In some embodiments, the gene modification system is capable of generating insertions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, or 10 kilobases (and optionally not exceeding 1, 5, 10, or 20 kilobases) at the target site. In some embodiments, the gene modification system is capable of generating deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene-editing system is capable of producing deletions of at least 81, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene-editing system is capable of producing deletions of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, or 200 nucleotides (and optionally not exceeding 500, 400, 300, or 200 nucleotides). In some embodiments, the gene modification system is capable of producing deletions of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases (and optionally not exceeding 1, 5, 10, or 20 kilobases).In some embodiments, the gene modification system is capable of generating at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotide substitutions at the target site. In some embodiments, the gene modification system is capable of generating 1-2, 2-3, 3-4, 4-5, 5-10, 10-15, 15-20, 20-30, 30-40, 40-50, 50-60, 60-70, 70-80, 80-90, or 90-100 nucleotide substitutions at the target site.

[0263] In some embodiments, substitution is a translocation mutation. In some embodiments, substitution converts adenine to thymine, adenine to guanine, adenine to cytosine, guanine to thymine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.

[0264] In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by altering, adding, or deleting sequences in promoters or enhancers (e.g., sequences that bind transcription factors). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene translation (e.g., altering amino acid sequences), inserting or deleting start or stop codons, altering or fixing the translational framework of a gene. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene splicing, for example by inserting, deleting, or altering splice acceptor or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter transcript or protein half-life. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in the cell (e.g., from the cytoplasm to the mitochondria, from the cytoplasm to the extracellular space (e.g., adding secretion tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter (e.g., improve) protein folding (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease gene activity, such as the activity of proteins encoded by the gene.

[0265] Exemplary genetically modified peptides, systems comprising them, and methods of using them are described, for example, in PCT / US2021 / 020948, which relates to the retroviral RT domain (including the amino acid and nucleic acid sequences therein) and is incorporated herein by reference.

[0266] Exemplary gene-modified peptides and retroviral RT domain sequences are also described, for example, in International Application No. PCT / US21 / 20948, filed March 4, 2021, such as Tables 30, 31, and 44 therein; the entire application relating to retroviral RT is incorporated herein by reference, for example, in the sequences and tables described herein. Therefore, the gene-modified peptides described herein may comprise amino acid sequences or their domains (e.g., retroviral RT domains) according to any of the tables mentioned in this paragraph, or functional fragments or variants of any of the foregoing, or amino acid sequences having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with them.

[0267] In some embodiments, the peptide used in any of the systems described herein may be a molecular or genetic reconstruction based on the alignment of peptide sequences with multiple homologous proteins. In some embodiments, the reverse transcriptase domain used in any of the systems described herein may be a molecular or genetic reconstruction, or may be modified at specific residues based on alignments of reverse transcriptase domains from the same or different sources. Based on the accession numbers provided herein, those skilled in the art can align peptide or nucleic acid sequences, for example, using conventional sequence analysis tools such as the Basic Local Alignment Search (BLAST) or CD-search (for conserved domain analysis). Molecular reconstructions may be created based on shared sequences, for example using methods described in Ivics et al., Cell 1997, 501–510; Wagstaff et al., Molecular Biology and Evolution 2013, 88–99.

[0268] polypeptide components of gene modification systems

[0269] In some embodiments, the genetically modified polypeptide has the functions of DNA target site binding, template nucleic acid (e.g., RNA) binding, DNA target site cleavage, and template nucleic acid (e.g., RNA) writing (e.g., reverse transcription). In some embodiments, each function is contained within a different domain. In some embodiments, the function may be attributed to two or more domains (e.g., two or more domains together exhibit the function). In some embodiments, two or more domains may have the same or similar functions (e.g., two or more domains each independently have DNA binding functionality, e.g., for two different DNA sequences). In other embodiments, one or more domains may be capable of performing one or more functions; for example, the Cas9 domain is capable of simultaneously performing DNA binding and target site cleavage. In some embodiments, these domains are all located within a single polypeptide. In some embodiments, a first domain is in one polypeptide and a second domain is in a second polypeptide. For example, in some embodiments, the sequence may be broken between the first polypeptide and the second polypeptide, e.g., where the first polypeptide contains a reverse transcriptase (RT) domain and where the second polypeptide contains a DNA-binding domain and a nuclease domain, such as a nicking enzyme domain. As a further example, in some embodiments, the first polypeptide and the second polypeptide each comprise a DNA-binding domain (e.g., a first DNA-binding domain and a second DNA-binding domain). In some embodiments, the first and second polypeptides can be linked together post-translationally by breaking down introns to form a single genetically modified polypeptide.

[0270] In some respects, the gene-modified polypeptide described herein comprises (e.g., the system described herein comprises a gene-modified polypeptide comprising): 1) a Cas domain (e.g., a Cas nickase domain, such as a Cas9 nickase domain); 2) a reverse transcriptase (RT) domain, wherein the RT domain is located at the C-terminus of the Cas domain; and a linker disposed between the RT domain and the Cas domain.

[0271] In some embodiments, the gene-modified polypeptide comprises a GG amino acid sequence between the Cas domain and the linker, an AG amino acid sequence between the RT domain and the second NLS, and / or a GG amino acid sequence between the linker and the RT domain.

[0272] Writing a structure field (RT structure field)

[0273] In some aspects of the invention, the writing domain of the gene modification system has reverse transcriptase activity and is also referred to as the reverse transcriptase domain (RT domain). In some embodiments, the RT domain comprises an RT catalytic portion and an RNA-binding region (e.g., a region that binds template RNA).

[0274] In some embodiments, the nucleic acid encoding reverse transcriptase is altered from its native sequence to have altered codon usage, for example, modified for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase derived from a retrovirus. In some embodiments, the RT domain comprising the genetically modified polypeptide has been mutated from its original amino acid sequence, for example, having at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 substitutions. In some embodiments, the RT domain is derived from a retroviral RT, for example, avian reticuloendotheliosis virus (AVIRE) (e.g., UniProtKB accession number: P03360) RT.

[0275] In some embodiments, the retroviral reverse transcriptase (RT) domain exhibits enhanced strictness in target-induced reverse transcription (TPRT) initiation, for example, relative to the endogenous RT domain. In some embodiments, the RT domain initiates TPRT when a 3 nt immediately upstream of the first-strand cleavage at the target site, such as the genomic DNA initiating the RNA template, has at least 66% or 100% complementarity to a homologous 3 nt in the RNA template. In some embodiments, the RT domain initiates TPRT when there is less than 5 nt mismatch (e.g., less than 1, 2, 3, 4, or 5 nt mismatch) between the template RNA homology and the target DNA-initiated reverse transcription. In some embodiments, the RT domain is modified to increase the strictness of mismatches in TPRT response initiation, for example, wherein the RT domain does not tolerate any mismatches or allows fewer mismatches in the initiation region relative to the wild-type (e.g., unmodified) RT domain.

[0276] In some embodiments, the natural heterodimeric RT domain may also function as a homodimer. In some embodiments, the dimer RT domain is expressed as a fusion protein, such as a homodimeric fusion protein or a heterodimeric fusion protein. In some embodiments, the RT function of the system is achieved by multiple RT domains. In other embodiments, the multiple RT domains are fused or separate, for example, they may be on the same polypeptide or on different polypeptides.

[0277] In some embodiments, the gene modification system described herein includes an integrase domain, for example, wherein the integrase domain may be part of an RT domain. In some embodiments, the RT domain (e.g., as described herein) includes an integrase domain. In some embodiments, the RT domain (e.g., as described herein) lacks an integrase domain, or includes an integrase domain that has been inactivated by mutation or deletion. In some embodiments, the gene modification system described herein includes an RNase H domain, for example, wherein the RNase H domain may be part of an RT domain. In some embodiments, the RNase H domain is not part of the RT domain and is covalently linked by a flexible linker. In some embodiments, the RT domain (e.g., as described herein) includes an RNase H domain, for example, an endogenous RNase H domain or a heterologous RNase H domain. In some embodiments, the RT domain (e.g., as described herein) lacks an RNase H domain. In some embodiments, the RT domain (e.g., as described herein) includes an RNase H domain with the addition, deletion, mutation, or exchange of a heterologous RNase H domain. In some embodiments, the peptide includes an inactivated endogenous RNase H domain. In some embodiments, the endogenous RNase H domain is genetically removed from one of the other domains of the polypeptide, such that it is not included in the polypeptide, for example, the endogenous RNase H domain is partially or completely truncated from the containing domain. In some embodiments, the mutation of the RNase H domain produces a polypeptide exhibiting lower RNase activity, for example, as determined by the method described by Kotewicz et al. Nucleic Acids Res [Nucleic Acids Research] 16(1):265-277 (1988) (which is incorporated herein by reference in its entirety), for example, a reduction of at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90% compared to a otherwise similar domain without the mutation. In some embodiments, RNase H activity is eliminated.

[0278] In some embodiments, the RT domain is mutated to increase fidelity compared to other similar domains that are not mutated. For example, in some embodiments, the YADD (SEQ ID NO: 37635) or YMDD motif (SEQ ID NO: 37636) in the RT domain (e.g., in reverse transcriptase) is replaced by YVDD (SEQ ID NO: 37637). In embodiments, replacing YADD (SEQ ID NO: 37635) or YMDD (SEQ ID NO: 37636) or YVDD (SEQ ID NO: 37637) results in higher fidelity of retroviral reverse transcriptase activity (e.g., as described in Jamburuthugoda and Eickbush J Mol Biol [Journal of Molecular Biology] 2011; which is incorporated herein by reference in its entirety).

[0279] In some embodiments, the RT domain used in the gene modification system described herein comprises an RT domain derived from AVIRE. In some embodiments, the RT domain comprises one or more mutations listed in Table 2 below, compared to a wild-type or classic RT domain (e.g., from AVIRE). In some embodiments, the RT domain comprises one, two, three, four, five, or six mutations listed in the corresponding rows of Table 2 below.

[0280] Table 2. Exemplary RT domain mutations (relative to the corresponding wild-type sequence)

[0281]

[0282] In some embodiments, the genetically modified polypeptides described herein comprise an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it. In some embodiments, the genetically modified polypeptides described herein comprise an RT domain encoded by a nucleic acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it. In some embodiments, the nucleic acid described herein encodes an RT domain having an amino acid sequence according to Table 6, or a sequence having at least 70%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with it.

[0283] Table 6: Exemplary reverse transcriptase domains from retroviruses

[0284] RT name SEQ ID NO: RT amino acid sequence AVIRE_P03360_3mutA 56 TAPLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPVRVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLREVNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIF AFEWADAEEGESGQLTWTRLPQGFKNSPTLFNEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQAACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQIPVPKTKRQVREFLGKIGYCRLFIPGFAELAQPLYAATRPGNDPLV WGEKEEEAFQSLKLALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAGWPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSL TSTRPDLTDQPLAQAEATLFTDGSSYIRDGKRYAGAAVVTTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNIYTDSRYAFATLHVHGMIYRERGWLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKGHQKDDAPTSTGNRRADEVAREVAIRPLSTQATIS RT name SEQ ID NO: RT nucleic acid sequence AVIRE_P03360_3mutA 57 ACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGC

[0285] In some embodiments, the reverse transcriptase domain is modified, for example, through site-specific mutations. In some embodiments, the reverse transcriptase domain is engineered to have improved properties. In some embodiments, the reverse transcriptase domain may be engineered to have a lower error rate, for example, as described in WO 2001068895 (incorporated herein by reference). In some embodiments, the reverse transcriptase domain may be engineered to be more heat-resistant. In some embodiments, the reverse transcriptase domain may be engineered to have greater sustained synthetic capacity. In some embodiments, the reverse transcriptase domain may be engineered to be resistant to inhibitors. In some embodiments, the reverse transcriptase domain may be engineered to be faster. In some embodiments, the reverse transcriptase domain may be engineered to better tolerate modified nucleotides in the RNA template. In some embodiments, the reverse transcriptase domain may be engineered to insert modified DNA nucleotides. In some embodiments, the reverse transcriptase domain is engineered to bind template RNA.

[0286] In some embodiments, the retroviral reverse transcriptase domain may contain one or more mutations in the wild-type sequence, which may improve RT characteristics such as thermostability, sustained synthetic capacity, and / or template binding.

[0287] In some embodiments, the writing domain (e.g., the RT domain) includes an RNA-binding domain, which specifically binds to an RNA sequence. In some embodiments, the template RNA includes an RNA sequence specifically bound by the RNA-binding domain of the writing domain.

[0288] In some embodiments, the reverse transcription domain recognizes and reverse transcribes only a specific template, such as the system's template RNA. In some embodiments, the template contains a sequence or structure that can be recognized and reverse transcribed by the reverse transcription domain. In some embodiments, the template contains a sequence or structure that can be associated with an RNA-binding domain of a polypeptide component of the genome engineering system described herein. In some embodiments, the genome engineering system preferably reverse transcribes a template containing the associated sequence, rather than a template lacking the associated sequence.

[0289] The writing domain may also include DNA-dependent DNA polymerase activity, for example, enzymatic activity capable of writing DNA from a template DNA sequence book into the genome. In some embodiments, DNA-dependent DNA polymerization is employed to complete the second-strand synthesis for target site editing. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in a polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, such as second-strand synthesis. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a second polypeptide in the system. In some embodiments, the DNA-dependent DNA polymerase activity is provided by an endogenous host cell polymerase, which is optionally recruited to the target site by a component of the genome engineering system.

[0290] In some embodiments, the reverse transcriptase domain has a lower probability of premature termination (P0) in vitro compared to the reference reverse transcriptase domain. off In some embodiments, the reference reverse transcriptase domain is a viral reverse transcriptase domain, such as the RT domain from AVIRE.

[0291] In some embodiments, the reverse transcriptase domain has a density of less than about 5 x 10⁻⁶. -3 / nt, 5 x 10 -4 / nt or 5 x 10 -6 / nt premature termination rate in vitro (P off The lower probability of premature termination, for example, as measured on 1094 nt RNA. In the examples, the in vitro premature termination rate was determined as described in Bibillo and Eickbush (2002) J Biol Chem [Journal of Biochemistry] 277(38):34836-34845 (which is incorporated herein by reference in its entirety).

[0292] In some embodiments, the reverse transcriptase domain is capable of completing at least about 30% or 50% integration in the cell. The percentage of complete integration can be measured by dividing the number of substantially full-length integration events (e.g., genomic sites containing at least 98% of the expected integration sequence) by the total number of integration events in the cell population (including substantially full-length and partial integration events). In embodiments, long-read amplicon sequencing is used to determine integration in the cell (e.g., across integration sites), for example, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0293] In the embodiments, quantifying integration in cells includes counting the integrated portion of a DNA sequence containing at least about 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the template RNA (e.g., template RNA of at least 0.05, 0.1, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 3, 4, or 5 kb in length, such as 0.5–0.6, 0.6–0.7, 0.7–0.8, 0.8–0.9, 1.0–1.2, 1.2–1.4, 1.4–1.6, 1.6–1.8, 1.8–2.0, 2–3, 3–4, or 4–5 kb in length).

[0294] In some embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro. In embodiments, the reverse transcriptase domain is capable of polymerizing dNTPs in vitro at a rate of 0.1–50 nt / sec (e.g., 0.1–1, 1–10, or 10–50 nt / sec). In embodiments, the polymerization of dNTPs by the reverse transcriptase domain is measured by a single-molecule assay, for example, as described in Schwartz and Quake (2009) PNAS [Proceedings of the National Academy of Sciences] 106(48):20294–20299 (which is incorporated herein by reference in its entirety).

[0295] In some embodiments, the in vitro error rate of the reverse transcriptase domain (e.g., nucleotide mis-incorporation) is 1 x 10⁻⁶. -3 - 1 x 10 -4 Or 1 x 10 -4 - 1 x 10 -5 One substitution / nt, for example, as described in Yasukawa et al. (2017) BiochemBiophys Res Commun [Biochemistry and Biophysics Research Communications] 492(2):147-153 (which is incorporated herein by reference in its entirety). In some embodiments, the error rate (e.g., nucleotide mis-incorporation) of the reverse transcriptase domain in cells (e.g., HEK293T cells) is 1 x 10 -3 - 1 x 10 -4 Or 1 x 10 -4 - 1 x 10 -5 One substitution / nt, for example, by long-read amplicon sequencing, as described in Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (which is incorporated herein by reference in its entirety).

[0296] In some embodiments, the reverse transcriptase domain enables reverse transcription of the target RNA in vitro. In some embodiments, the reverse transcriptase requires a primer of at least 3 nucleotides to initiate reverse transcription of the template. In some embodiments, reverse transcription of the target RNA is determined by detecting cDNA from the target RNA (e.g., when an ssDNA primer is provided, for example, which is annealed to the target at the 3' end by at least 3, 4, 5, 6, 7, 8, 9, or 10 nt), for example as described in Bibillo and Eickbush (2002) J Biol Chem 277(38):34836-34845 (which is incorporated herein by reference in its entirety).

[0297] In some embodiments, the reverse transcriptase domain performs reverse transcription at least 5 or 10 times more efficiently than an RNA template lacking a protein-binding motif (e.g., 3'UTR), for example, when its RNA template is converted to cDNA (e.g., via cDNA production). In the embodiments, the reverse transcription efficiency is measured as described in Yasukawa et al. (2017) Biochem Biophys ResCommun [Biochemistry and Biophysics Research Communications] 492(2):147-153 (which is incorporated herein by reference in its entirety).

[0298] In some embodiments, the reverse transcriptase domain binds specifically to a particular RNA template at a higher frequency (e.g., about 5 or 10 times higher) than any endogenous cellular RNA (e.g., when expressed in cells (e.g., HEK293T cells)). In embodiments, the specific binding frequency between the reverse transcriptase domain and the template RNA is measured by CLIP-seq, as described, for example, in Lin and Miles (2019) Nucleic Acids Res [Nucleic Acids Research] 47(11):5490-5501 (which is incorporated herein by reference in its entirety).

[0299] Template nucleic acid binding domain

[0300] Genetically modified peptides typically contain regions capable of being associated with a template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain that can be associated with an RNA molecule containing specific features (e.g., structural motifs). In other embodiments, the template nucleic acid binding domain (e.g., an RNA binding domain) is contained within a reverse transcription domain, for example, a reverse transcriptase-derived component having known RNA preference characteristics.

[0301] In other embodiments, a template nucleic acid binding domain (e.g., an RNA binding domain) is contained within a target DNA binding domain. For example, in some embodiments, the DNA binding domain is a CRISPR-associated protein that recognizes a structure containing a template nucleic acid (e.g., template RNA) containing gRNA. In some embodiments, the genetically modified polypeptide contains a DNA binding domain that includes a CRISPR-associated protein associated with a gRNA scaffold that allows the DNA binding domain to bind to a target genomic DNA sequence. In some embodiments, the gRNA scaffold and gRNA spacer are contained within a template nucleic acid (e.g., template RNA), thus the DNA binding domain is also a template nucleic acid binding domain. In some embodiments, the polypeptide has RNA-binding functionality in multiple domains, for example, it can bind a gRNA structure in a CRISPR-associated DNA binding domain and additional sequences or structures in a reverse transcriptase domain.

[0302] In some embodiments, the RNA-binding domain is capable of binding template RNA with a greater affinity than a reference RNA-binding domain. In some embodiments, the reference RNA-binding domain is the RNA-binding domain of Cas9 from *Streptococcus pyogenes*. In some embodiments, the RNA-binding domain is capable of binding template RNA with an affinity of 100 pM – 10 nM (e.g., 100 pM – 1 nM or 1 nM – 10 nM). In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in vitro, for example by thermophoresis, as described, for example, in Asmari et al. Methods 146:107-119 (2018) (which is incorporated herein by reference in its entirety). In some embodiments, the affinity of the RNA-binding domain for its template RNA is measured in cells (e.g., by FRET or CLIP-Seq).

[0303] In some embodiments, the RNA-binding domain is associated with the template RNA in vitro at a frequency at least about 5 to 10 times higher than that of the out-of-order RNA. In some embodiments, the binding frequency between the RNA-binding domain and the template RNA or the out-of-order RNA is measured by CLIP-seq, for example, as described in Lin and Miles (2019) Nucleic Acids Res [Nucleic Acids Research] 47(11):5490-5501 (which is incorporated herein by reference in its entirety). In some embodiments, the RNA-binding domain is associated with the template RNA in cells (e.g., HEK293T cells) at a frequency at least about 5 to 10 times higher than that of the out-of-order RNA. In some embodiments, the association frequency between the RNA-binding domain and the template RNA or the out-of-order RNA is measured by CLIP-seq, for example, as described above in Lin and Miles (2019).

[0304] Nucleotide endonuclease domain and DNA-binding domain

[0305] In some embodiments, the genetically modified polypeptide has the function of cleaving a DNA target site via a nuclease domain. In some embodiments, the genetically modified polypeptide includes a DNA-binding domain, for example, for binding to a target nucleic acid. In some embodiments, the domains of the genetically modified polypeptide (e.g., a Cas domain) include two or more smaller domains (e.g., a DNA-binding domain and a nuclease domain). It should be understood that when the DNA-binding domain (e.g., a Cas domain) is described as binding to a target nucleic acid sequence, in some embodiments, this binding is mediated by gRNA.

[0306] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-associated endonuclease domain that binds to a template RNA containing gRNA, binds to a target DNA sequence (e.g., complementary to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, the heterologous endonuclease domain or endonuclease / DNA-binding domain may be used or may be modified in the gene modification system described herein (e.g., by inserting, deleting, or substituting one or more residues).

[0307] In some embodiments, the nucleic acid encoding the endonuclease domain or the endonuclease / DNA binding domain is altered from its natural sequence to have a modified codon, for example, for modification targeting human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element, such as a Cas endonuclease (e.g., Cas9).

[0308] In some aspects, the DNA-binding domain of the genetically modified peptide described herein is selected, designed, or constructed to bind to a desired host DNA target sequence. In some embodiments, the DNA-binding domain of the peptide is a heterologous DNA-binding element. In some embodiments, the heterologous DNA-binding element is a sequence-guided DNA-binding element, such as Cas9, Cpf1, or other CRISPR-related proteins that have been modified to lack endonuclease activity. In some embodiments, the heterologous DNA-binding element retains endonuclease activity. In some embodiments, the heterologous DNA-binding element retains partial endonuclease activity to cleave ssDNA, for example, having nicking enzyme activity.

[0309] In some embodiments, the DNA-binding domain is modified, for example, by site-specific mutations, increasing or decreasing DNA-binding elements (e.g., the number and / or specificity of zinc fingers), to alter DNA-binding specificity and affinity. In some embodiments, the nucleic acid sequence encoding the DNA-binding domain is altered from its natural sequence to have modified codon usage, for example, for human cells. In embodiments, the DNA-binding domain comprises one or more modifications relative to the wild-type DNA-binding domain, for example, modifications via directed evolution (e.g., phage-assisted sequential evolution (PACE)).

[0310] In some embodiments, the genetically modified polypeptide comprises a modification of the DNA-binding domain, for example, relative to the wild-type polypeptide. In some embodiments, the DNA-binding domain comprises an addition, deletion, substitution, or modification of the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to include a heterologous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence. In some embodiments, the functional domain replaces at least a portion (e.g., all) of the previous DNA-binding domain of the polypeptide. In some embodiments, the Cas domain comprises Cas9 or a mutant or variant thereof (e.g., as described herein). In embodiments, the Cas domain is associated with a guide RNA (gRNA), for example, as described herein. In embodiments, the Cas domain is directed by the gRNA to a target nucleic acid (e.g., DNA) sequence. In embodiments, the Cas domain and the gRNA are encoded in the same nucleic acid (e.g., RNA) molecule. In embodiments, the Cas domain and the gRNA are encoded in different nucleic acid (e.g., RNA) molecules.

[0311] In some embodiments, the DNA-binding domain is capable of binding the target sequence (e.g., the dsDNA target sequence) with a greater affinity than a reference DNA-binding domain. In some embodiments, the reference DNA-binding domain is the DNA-binding domain of Cas9 from Streptococcus pyogenes. In some embodiments, the DNA-binding domain is capable of binding the target sequence (e.g., the dsDNA target sequence) with an affinity between 100 pM and 10 nM (e.g., between 100 pM and 1 nM or between 1 nM and 10 nM).

[0312] In some embodiments, the affinity of the DNA-binding domain for its target sequence (e.g., dsDNA target sequence) is measured in vitro, for example by thermophoresis, as described in, for example, Asmari et al. Methods 146:107-119 (2018) (incorporated herein by reference in its entirety).

[0313] In an embodiment, in the presence of, for example, about 100 times molar excess of out-of-order sequence competitor dsDNA, the DNA binding domain is able to bind its target sequence (e.g., dsDNA target sequence) with an affinity, for example, between 100 pM and 10 nM (e.g., between 100 pM and 1 nM or between 1 nM and 10 nM).

[0314] In some embodiments, the DNA-binding domain is found to be associated with its target sequence (e.g., a dsDNA target sequence) more frequently than any other sequence in the genome of the target cell (e.g., a human target cell), as measured by ChIP-seq (e.g., in HEK293T cells), as described in He and Pu (2010) Curr. Protoc Mol Biol [Latest Protocols in Molecular Biology] Chapter 21 (which is incorporated herein by reference in its entirety). In some embodiments, the DNA-binding domain is found to be associated with its target sequence (e.g., a dsDNA target sequence) at a frequency at least about 5 to 10 times more frequently than any other sequence in the genome of the target cell, as measured by ChIP-seq (e.g., in HEK293T cells), as described in He and Pu (2010), as above.

[0315] In some embodiments, the endonuclease domain has nicking enzyme activity and cleaves one strand of the target DNA. In some embodiments, the nicking enzyme activity reduces the formation of double-strand breaks at the target site. In some embodiments, the endonuclease domain creates staggered nicking structures in the first and second strands of the target DNA. In some embodiments, the staggered nicking structures create free 3' overhangs at the target site. In some embodiments, free 3' overhangs at the target site improve editing efficiency, for example, by enhancing access to and annealing of the 3' homologous regions of the template nucleic acid. In some embodiments, the staggered nicking structures reduce the formation of double-strand breaks at the target site.

[0316] In some embodiments, the endonuclease domain cleaves both strands of the target DNA, for example, resulting in blunt-end cleavage of the target, and without ssDNA overhangs on either side of the cleavage site. The amino acid sequence of the endonuclease domain of the gene modification system described herein may be at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to the amino acid sequence of the endonuclease domain described herein.

[0317] In some embodiments, the heterologous endonuclease is derived from a CRISPR-related protein, such as Cas9. In some embodiments, the heterologous endonuclease is engineered to possess only ssDNA cleavage activity, for example, only nicking enzyme activity, such as a Cas9 nicking enzyme, such as SpCas9 with D10A, H840A, or N863A mutations. In still other embodiments, the homologous endonuclease domain is modified, for example, through site-specific mutations, to alter the DNA endonuclease activity. In still other embodiments, the endonuclease domain is modified to reduce DNA sequence specificity, for example, by truncation to remove the DNA sequence-specific domain or by mutation to inactivate the DNA sequence-specific region.

[0318] In some embodiments, the endonuclease domain has nicking enzyme activity and does not form double-strand breaks. In some embodiments, the endonuclease domain forms single-strand breaks at a higher frequency than double-strand breaks, for example, at least 90%, 95%, 96%, 97%, 98%, or 99% of the breaks are single-strand breaks, or less than 10%, 5%, 4%, 3%, 2%, or 1% of the breaks are double-strand breaks. In some embodiments, the endonuclease substantially does not form double-strand breaks. In some embodiments, the endonuclease does not form detectable levels of double-strand breaks.

[0319] In some embodiments, the endonuclease domain has cleaving enzyme activity for cleaving target site DNA on a first strand; for example, in some embodiments, the endonuclease cleaves target sites of genomic DNA near alteration sites on the strand to which the writing domain is extended. In some embodiments, the endonuclease domain has cleaving enzyme activity for cleaving target site DNA on a first strand but not for cleaving target site DNA on a second strand. For example, when the polypeptide contains a CRISPR-associated endonuclease domain with cleaving enzyme activity, in some embodiments, the CRISPR-associated endonuclease domain cleaves target site DNA strands containing PAM sites (e.g., and does not cleave target site DNA strands without PAM sites). As a further example, when the polypeptide contains a CRISPR-associated endonuclease domain with cleaving enzyme activity, in some embodiments, the CRISPR-associated endonuclease domain cleaves target site DNA strands without PAM sites (e.g., and does not cleave target site DNA strands containing PAM sites).

[0320] In some other embodiments, the endonuclease domain has nicking enzyme activity that nicks the target DNA on both the first and second strands. Without being bound by theory, after the writing domain (e.g., the RT domain) of the polypeptide described herein is polymerized (e.g., reverse transcribed) from a heterologous object sequence of a template nucleic acid (e.g., template RNA), the cellular DNA repair mechanism must repair the nick on the first DNA strand. The target DNA now contains two distinct first DNA strand sequences: one corresponding to the original genomic DNA (e.g., with a free 5′ end), and the second corresponding to the one polymerized from the heterologous object sequence (e.g., with a free 3′ end). It is assumed that these two distinct sequences are mutually balanced, with the first hybridizing to the second strand, and then the other, and that the cellular DNA repair mechanism incorporating the sequence into the target site it repairs can be a random process. Without being bound by theory, it is assumed that introducing an additional nick into the second strand might cause the cellular DNA repair mechanism to favor the heterologous object sequence-based sequence more frequently than the original genomic sequence (Anzalone et al., Nature [Nature] 576:149-157(2019)). In some embodiments, additional nicks are located at at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, or 150 nucleotides at 5' or 3' of a nick on the first strand.

[0321] Alternatively or otherwise, without being bound by theory, it is considered that additional nicks in the second strand can facilitate second-strand synthesis. In some embodiments, when a gene modification system has inserted into or replaced a portion of the first strand, it is necessary to synthesize a new sequence corresponding to the insertion / replacement in the second strand.

[0322] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain) and said domain cleaves a first strand and a second strand. For example, in such an embodiment, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs cleavage of the first strand and another gRNA spacer that directs cleavage of the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, and a first endonuclease domain cleaves the first strand and a second endonuclease domain cleaves the second strand (optionally, the first endonuclease domain does not (e.g., cannot) cleave the second strand, and the second endonuclease domain does not (e.g., cannot) cleave the first strand).

[0323] In some embodiments, the endonuclease domain is capable of cleaving the first and second strands. In some embodiments, the first and second strand cleavages occur at the same location in the target site but on opposite strands. In some embodiments, the second strand cleavage occurs at an alternating location with the first cleavage, such as upstream or downstream. In some embodiments, if the second strand cleavage is upstream of the first strand cleavage, the endonuclease domain produces a target site deletion. In some embodiments, if the second strand cleavage is downstream of the first strand cleavage, the endonuclease domain produces a target site duplication. In some embodiments, if the first and second strand cleavages occur at the same location in the target site, the endonuclease domain does not produce duplications and / or deletions. In some embodiments, the endonuclease domain has altered activity depending on protein conformation or RNA binding state, which, for example, promotes cleavage of the first or second strand (e.g., as described in Christensen et al., PNAS [Proceedings of the National Academy of Sciences] 2006; which is incorporated herein by reference in its entirety).

[0324] In some embodiments, the genetically modified polypeptide comprises a modification of the endonuclease domain, for example, relative to the wild-type Cas protein. In some embodiments, the endonuclease domain comprises the addition, deletion, substitution, or modification of the amino acid sequence of the wild-type Cas protein. In some embodiments, the endonuclease domain is modified to include a heterologous functional domain that specifically binds to and / or induces endonuclease cleavage of a target nucleic acid (e.g., DNA) sequence.

[0325] In some embodiments, the endonuclease domain is associated with the target dsDNA in vitro at a frequency at least about 5 to 10 times higher than that of the scrambled dsDNA. In some embodiments, the endonuclease domain is associated with the target dsDNA in vitro at a frequency at least about 5 to 10 times higher than that of the scrambled dsDNA, for example in cells (e.g., HEK293T cells). In some embodiments, the association frequency between the endonuclease domain and the target DNA or the scrambled DNA is measured by ChIP-seq, for example as described in Chapter 21 of He and Pu (2010) Curr. Protoc Mol Biol [Latest Protocols in Molecular Biology] (which is incorporated herein by reference in its entirety).

[0326] In some embodiments, the endonuclease domain can catalyze the formation of a nick at the target sequence, for example, by at least about 5-fold or 10-fold relative to a non-target sequence (e.g., relative to any other genomic sequence in the target cell genome). In some embodiments, the level of nick formation is determined using NickSeq, for example, as described in Elacqua et al. (2019) bioRxivdoi.org / 10.1101 / 867937 (which is incorporated herein by reference in its entirety).

[0327] In some embodiments, the endonuclease domain is capable of cleaving DNA in vitro. In embodiments, cleavage results in exposed bases. In embodiments, the exposed bases can be detected using a nuclease sensitivity assay, for example, as described in Chaudhry and Weinfeld (1995) Nucleic Acids Res [Nucleic Acids Research] 23(19):3805-3809 (which is incorporated herein by reference in its entirety). In embodiments, the level of exposed bases (e.g., detected by a nuclease sensitivity assay) is increased by at least 10%, 50%, or more relative to a reference endonuclease domain. In some embodiments, the reference endonuclease domain is the endonuclease domain of Cas9 from Streptococcus pyogenes.

[0328] In some embodiments, the endonuclease domain is capable of cleaving DNA in cells. In an embodiment, the endonuclease domain is capable of cleaving DNA in HEK293T cells. In an embodiment, unrepaired cleavage undergoing replication in the absence of Rad51 results in an increased NHEJ rate at the cleavage site, which can be detected, for example, by using a Rad51 inhibition assay, as described, for example, as Bothmer et al. (2017) Nat Commun [Nature Communications] 8:13905 (which is incorporated herein by reference in its entirety). In an embodiment, the NHEJ rate increases to more than 0-5%. In an embodiment, for example, after Rad51 inhibition, the NHEJ rate increases to 20%-70% (e.g., in 30%-60% or 40%-50%).

[0329] In some embodiments, the endonuclease domain releases the target upon cleavage. In some embodiments, target release is indirectly indicated by assessing the number of enzyme turnovers, for example, as described in Yourik et al. RNA 25(1):35-44 (2019) (which is incorporated herein by reference in its entirety) and shown in Figure 2. In some embodiments, the k-value of the endonuclease domain is measured as such. exp 1 x 10 -3 – 1 x 10 - 5 min-1.

[0330] In some embodiments, the endonuclease domain has a size greater than about 1 x 10⁻⁶ in vitro. 8 s -1 M -1 catalytic efficiency (k cat / K m In the embodiments, the endonuclease domain has a size greater than approximately 1 x 10⁻⁶ in vitro. 5 1 x 10 6 1 x 10 7 Or 1 x 10 8 s -1 M -1 The catalytic efficiency. In the examples, the catalytic efficiency was determined as described by Chen et al. (2018) Science [Science] 360(6387):436-439 (which is incorporated herein by reference in its entirety). In some embodiments, the endonuclease domain has a size greater than about 1 x 10^6 ppm in the cell. 8 s -1 M -1 catalytic efficiency (k cat / K m In the examples, the catalytic efficiency of the endonuclease domain in cells is greater than about 1 x 10⁻⁶.5 1 x 10 6 1 x 10 7 Or 1 x 10 8 s -1 M -1 .

[0331] Gene-modified peptides containing Cas domains

[0332] In some embodiments, the gene-modifying peptides described herein comprise a Cas domain. In some embodiments, the Cas domain can guide the gene-modifying peptide to a target site specified by a gRNA spacer, thereby “cis” modifying the target nucleic acid sequence. In some embodiments, the gene-modifying peptide is fused to a Cas domain. In some embodiments, the gene-modifying peptide comprises a CRISPR / Cas domain (also referred to herein as a CRISPR-associated protein). In some embodiments, the CRISPR / Cas domain comprises a protein (e.g., a Cas protein) involved in clustering of a regulatory short palindromic repeat (CRISPR) system and optionally binds a guide RNA, such as a single guide RNA (sgRNA).

[0333] The CRISPR system is an adaptive defense system first discovered in bacteria and archaea. CRISPR systems use RNA-guided nucleases (e.g., Cas9 or Cpf1) called CRISPR-associated or “Cas” endonucleases to cut foreign DNA. For example, in a typical CRISPR-Cas system, the endonuclease is directed to the target nucleotide sequence (e.g., the site in the genome to be edited) via a sequence-specific non-coding “guide RNA” that targets a single-stranded or double-stranded DNA sequence. Three classes (I-III) of CRISPR systems have been identified. Class II CRISPR systems use a single Cas endonuclease (instead of multiple Cas proteins). One type of Class II CRISPR system includes a type II Cas endonuclease, such as Cas9, CRISPR RNA (“crRNA”), and trans-activating crRNA (“tracrRNA”). The crRNA contains a “spacer” sequence, which is an RNA sequence (“protospacer”) that typically corresponds to about 20 nucleotides of the target DNA sequence. In wild-type systems and some engineered systems, crRNA also contains a region that binds to tracrRNA to form a partially double-stranded structure that is cleaved by RNase III, producing a crRNA / tracrRNA hybrid molecule. The crRNA / tracrRNA hybrid then directs a Cas endonuclease to recognize and cleave the target DNA sequence. The target DNA sequence is typically adjacent to a "protospacer adjacent motif" ("PAM"), which is specific to a given Cas endonuclease and is required for cleavage activity at the target site that matches the crRNA spacer. CRISPR endonucleases identified from different prokaryotic species have unique PAM sequence requirements, for example, as listed in Table 7 for exemplary Cas enzymes; examples of PAM sequences include 5´-NGG (Streptococcus pyogenes), 5´-NNAGAA (Streptococcus thermophilus CRISPR1), 5´-NGGNG (Streptococcus thermophilus CRISPR3), and 5´-NNNGATT (Neisseria meningiditis). Some endonucleases (e.g., Cas9 endonuclease) are associated with G-rich PAM sites (e.g., 5´-NGG) and cleave the target DNA at the blunt end three nucleotides upstream (5´) of the PAM site. Another class II CRISPR system includes the type V endonuclease Cpf1, which is smaller than Cas9; examples include AsCpf1 (from the genus Acidaminococcus sp.) and LbCpf1 (from the genus Lachnospiraceae sp.).Cpf1-associated CRISPR arrays are processed into mature crRNA without the need for tracrRNA; in other words, in some embodiments, the Cpf1 system contains only the Cpf1 nuclease and crRNA to cleave the target DNA sequence. The Cpf1 endonuclease is typically associated with T-rich PAM sites, such as 5'-TTN. Cpf1 can also recognize 5'-CTA PAM motifs. Cpf1 typically cleaves target DNA by introducing misaligned or staggered double-strand breaks with 4 or 5 nucleotides of 5' protrusions, for example, cleaving target DNA in which the 5-nucleotide misaligned or staggered cut is located 18 nucleotides downstream (3') of the PAM site on the coding strand and 23 nucleotides downstream of the PAM site on the complementary strand; the 5-nucleotide protrusions resulting from such misaligned cuts allow for more precise genome editing via DNA insertion through homologous recombination than with DNA cut at blunt ends. See, for example, Zetsche et al. (2015) Cell, 163:759-771.

[0334] Various CRISPR-related (Cas) genes or proteins can be used in the techniques provided in this disclosure, and the selection of the Cas protein will depend on the specific conditions of the method. Specific examples of Cas proteins include class II systems, including Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, the Cas protein (e.g., Cas9 protein) can be derived from any of a variety of prokaryotic species. In some embodiments, a specific Cas protein (e.g., a specific Cas9 protein) is selected to recognize a specific protospacer neighbor motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain includes the sequence of a targeting polypeptide (e.g., a Cas protein, such as Cas9). In some embodiments, the Cas protein (e.g., Cas9 protein) can be obtained from bacteria or archaea or synthesized using known methods. In some embodiments, the Cas protein can be derived from Gram-positive or Gram-negative bacteria. In some embodiments, the Cas protein can be derived from the genus Streptococcus (e.g., Streptococcus pyogenes).

[0335] In some embodiments, the genetically modified polypeptide may comprise a Cas domain or a functional fragment thereof as listed in Table 7 or 8, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0336] Table 7. CRISPR / Cas proteins, species, and mutations

[0337] name enzymes Species The number of AA PAM Mutations that alter PAM recognition Mutations leading to catalytic inactivation SpCas9 Cas9 Streptococcus pyogenes 1368 5'-NGG-3' Wt D10A / D839A / H840A / N863A

[0338] Table 8. Sequences of exemplary CRISPR / Cas proteins

[0339] name Cas amino acids SEQ ID NO Exemplary nucleic acids SEQ ID NO SpCas9 DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKARGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD 52 GACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGAC 53

[0340] In some embodiments, the Cas protein requires the presence of a protospacer neighbor motif (PAM) in or near the target DNA sequence for the Cas protein to bind and / or function. In some embodiments, the PAM, from 5′ to 3′, is or comprises NGG, YG, NNGRRT, NNNRRT, NGA, TYCV, TATV, NTTN, or NNNGATT, where N represents any nucleotide, Y represents C or T, R represents A or G, and V represents A, C, or G. In some embodiments, the Cas protein is a protein listed in Table 7 or 8. In some embodiments, the Cas protein comprises one or more mutations that alter its PAM. Exemplary advances in engineering Cas enzymes to recognize altered PAM sequences are reviewed in Collias et al., Nature Communications [Nature Communications] 12:555 (2021), which is incorporated herein by reference in its entirety.

[0341] In some embodiments, the Cas protein has catalytic activity and cleaves one or both strands of the target DNA site. In some embodiments, cleavage of the target DNA site leads to alterations, such as insertions or deletions, for example, through cellular repair mechanisms.

[0342] In some embodiments, the Cas protein is modified to inactivate or partially inactivate a nuclease, such as a nuclease-deficient Cas9. Wild-type Cas9 produces double-strand breaks (DSBs) at specific DNA sequences targeted by gRNA, and many functionally modified CRISPR endonucleases are available, such as partially inactivated Cas9 “nickase” versions that produce only single-strand breaks; and catalytically inactive Cas9 (“dCas9”) that does not cleave the target DNA. In some embodiments, the binding of dCas9 to the DNA sequence can interfere with transcription at that site through steric hindrance. In some embodiments, the binding of dCas9 to the anchoring sequence can interfere with (e.g., reduce or prevent) the formation and / or maintenance of genomic complexes (e.g., ASMCs). In some embodiments, the DNA-binding domain comprises a catalytically inactive Cas9, such as dCas9. Many catalytically inactive Cas9 proteins are known in the art. In some embodiments, dCas9 comprises mutations in each endonuclease domain of the Cas protein, such as the N863A mutation. In some embodiments, a catalytically inactive or partially catalytically inactive CRISPR / Cas domain comprises a Cas protein containing one or more mutations, such as those listed in Table 7. In some embodiments, a Cas protein described in a given row of Table 7 contains one, two, three, or all of the mutations listed in the same row of Table 7. In some embodiments, a Cas protein not described in Table 7 contains one, two, three, or all of the mutations listed in the rows of Table 7 or a corresponding mutation at the corresponding site in the Cas protein.

[0343] In some embodiments, catalytically inactive Cas9 proteins, such as dCas9 or partially inactivated Cas9 proteins, contain an N863 mutation (e.g., an N863A mutation) or a similar substitution of the amino acid corresponding to the stated position. In some embodiments, catalytically inactive Cas9 proteins, such as dCas9, contain a D10 mutation (e.g., D10A), a D839 mutation (e.g., D839A), an H840 mutation (e.g., H840A), and an N863 mutation (e.g., N863A) or a similar substitution of the amino acid corresponding to the stated position.

[0344] In some embodiments, the DNA-binding domain or endonuclease domain may contain a Cas molecule that contains or is (e.g., covalently) linked to a gRNA (e.g., a template nucleic acid, such as a template RNA containing gRNA).

[0345] In some embodiments, the endonuclease domain or DNA-binding domain comprises Streptococcus pyogenes Cas9 (SpCas9) or a functional fragment or variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a modified SpCas9. In an embodiment, the modified SpCas9 comprises a modification that alters the specificity of the protospacer neighbor motif (PAM). In an embodiment, the PAM is specific for the nucleic acid sequence 5′-NGT-3′. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas domain, such as a Cas9 domain. In an embodiment, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas domain, a Cas nickase (nCas) domain, or a nuclease-free Cas (dCas) domain. In an embodiment, the endonuclease domain or DNA-binding domain comprises a nuclease-active Cas9 domain, a Cas9 nickase (nCas9) domain, or a nuclease-free Cas9 (dCas9) domain.

[0346] In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas9 sequence, for example, as described in Chylinski, Rhun, and Charpentier (2013) RNA Biology 10:5, 726-737; this literature is incorporated herein by reference. In some embodiments, the endonuclease domain or DNA-binding domain comprises the HNH nuclease subdomain and / or RuvC1 subdomain of Cas, for example, Cas9 as described herein, or a variant thereof. In some embodiments, the endonuclease domain or DNA-binding domain comprises a Cas polypeptide (e.g., an enzyme) or a functional fragment thereof. In embodiments, Cas9 comprises one or more substitutions and / or one or more mutations.

[0347] Additional components of an exemplary genetically modified polypeptide

[0348] In addition to the writing / RT domain and the endonuclease and DNA-binding domain, exemplary gene-modifying peptides for use in gene-modifying systems may also contain additional components as described herein.

[0349] Adapter

[0350] In some embodiments, the genetically modified polypeptide may include a linker, such as a peptide linker, as described in Table 10. In some embodiments, the genetically modified polypeptide includes a Cas domain (e.g., the Cas domain of Table 8), a linker of Table 10 (or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with it), and an RT domain (e.g., the RT domain of Table 6) in the N-terminal to C-terminal direction. In some embodiments, the genetically modified polypeptide includes a flexible linker between the endonuclease and the RT domain, for example, a linker containing the amino acid sequence AEAAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 54) or a sequence having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain of the genetically modified polypeptide may be located at the C-terminus of the endonuclease domain. In some embodiments, the RT domain of the genetically modified polypeptide may be located at the N-terminus of the endonuclease domain.

[0351] Table 10 Exemplary Connector Sequences

[0352] Amino acid sequence SEQ ID NO Exemplary nucleic acid sequence SEQ ID NO AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA 54 GCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCC 55

[0353] In some embodiments, the genetically modified polypeptide comprises: (i) a linker containing a linker sequence as listed in the row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and (ii) an RT domain containing an RT domain sequence as listed in the same row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0354] Table T1. Selection of Exemplary Genetically Modified Peptides

[0355] Adapter sequence SEQ ID NO: For the adapter RT sequence SEQ ID NO: of RT AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA 54 TAPLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPVRVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLREVNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIFAFEWADAEEGESGQLTWTRLPQGFKNSPTLFNEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQAACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQIPVPKTKRQVREFLGKIGYCRLFIPGFAELAQPLYAATRPGNDPLVWGEKEEEAFQSLKLALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAGWPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSLTSTRPDLTDQPLAQAEATLFTDGSSYIRDGKRYAGAAVVTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNIYTDSRYAFATLHVHGMIYRERGWLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKGHQKDDAPTSTGNRRADEVAREVAIRPLSTQATIS 56

[0356] In some embodiments, the genetically modified polypeptide comprises: (i) a Cas domain comprising a Cas sequence as listed in the row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; (ii) a linker comprising a linker sequence as listed in the same row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and (iii) an RT domain comprising an RT domain sequence as listed in the same row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0357] Table T2. Selection of Exemplary Genetically Modified Peptides

[0358] Cas domain sequence SEQ ID NO of the Cas domain: Connector sequence SEQ ID NO: For connectors RT sequence RT's SEQ ID NO: DKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKARGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGD 52 AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA 54 TAPLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPVRVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLREVNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIF AFEWADAEEGESGQLTWTRLPQGFKNSPTLFNEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQAACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQIPVPKTKRQVREFLGKIGYCRLFIPGFAELAQPLYAATRPGNDPLV WGEKEEEAFQSLKLALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAGWPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSL TSTRPDLTDQPLAQAEATLFTDGSSYIRDGKRYAGAAVVTTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNIYTDSRYAFATLHVHGMIYRERGWLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKGHQKDDAPTSTGNRRADEVAREVAIRPLSTQATIS 56

[0359] Location sequence

[0360] In some embodiments, the gene-modified system RNA further comprises an intracellular localization sequence, such as a nuclear localization sequence (NLS). In some embodiments, the gene-modified polypeptide or the nucleic acid (e.g., RNA) encoding the gene-modified polypeptide comprises an NLS. The nuclear localization sequence may be an RNA sequence that facilitates RNA entry into the cell nucleus. In some embodiments, the nuclear localization signal is located on the template RNA. In some embodiments, the gene-modified polypeptide is encoded on a first RNA, and the template RNA is a second separate RNA, and the nuclear localization signal is located on the template RNA rather than on the RNA encoding the gene-modified polypeptide. While not wishing to be bound by theory, in some embodiments, the RNA encoding the gene-modified polypeptide primarily targets the cytoplasm to facilitate its translation, while the template RNA primarily targets the cell nucleus to facilitate its insertion into the genome. In some embodiments, the nuclear localization signal is located at the 3' end, 5' end, or internal region of the template RNA. In some embodiments, the nuclear localization signal is located at the 3' end of a heterologous sequence (e.g., directly at the 3' end of the heterologous sequence) or at the 5' end of a heterologous sequence (e.g., directly at the 5' end of the heterologous sequence). In some embodiments, the nuclear localization signal is located outside the 5' UTR or outside the 3' UTR of the template RNA. In some embodiments, the nuclear localization signal is placed between the 5' UTR and the 3' UTR, wherein optionally, the nuclear localization signal is not transcribed with the transgene (e.g., the nuclear localization signal is antisense-oriented or downstream of a transcription termination signal or a polyadenylation signal). In some embodiments, the nuclear localization sequence is located inside an intron. In some embodiments, multiple identical or different nuclear localization signals are present in RNA, such as in template RNA. In some embodiments, the length of the nuclear localization signal is less than 5, 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 bp. Various RNA nuclear localization sequences can be used. For example, Lubelsky and Ulitsky, Nature 555 (107-111), 2018, describe an RNA sequence that drives RNA localization into the cell nucleus. In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, the nuclear localization signal binds to nuclear enrichment proteins. In some embodiments, the nuclear localization signal binds to the HNRNPK protein. In some embodiments, the nuclear localization signal is pyrimidine-rich, such as a C / T-rich, C / U-rich, C-rich, T-rich, or U-rich region. In some embodiments, the nuclear localization signal originates from a long non-coding RNA. In some embodiments, the nuclear localization signal originates from the MALAT1 long non-coding RNA or the 600-nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738-751), 2012).In some embodiments, the nuclear localization signal is derived from BORG long noncoding RNA or the AGCCC motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014)). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a retrovirus.

[0361] In some embodiments, the polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, such as nuclear localization sequences (NLS). In some embodiments, the NLS is a two-component NLS. In some embodiments, the NLS facilitates the delivery of the protein containing the NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of the genetically modified polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the genetically modified polypeptide. In some embodiments, the NLS is fused to the N-terminus or C-terminus of a Cas domain. In some embodiments, a linker sequence is arranged between the NLS and a neighboring domain of the genetically modified polypeptide.

[0362] In some embodiments, the NLS comprises an amino acid sequence as disclosed in Table 11. The NLS may be used with one or more copies of the peptide at one or more locations within the peptide, such as in the N-terminal domain, between peptide domains, in the C-terminal domain, or in combinations of multiple locations, to improve subcellular localization to the cell nucleus. Multiple unique sequences may be used in a single peptide. The sequence may be naturally monocomponent or bicomponent, for example, having one or two basic amino acid segments, or may be used as a chimeric bicomponent sequence. Sequence references correspond to UniProt accession numbers unless the sequence is indicated as SeqNLS for a sequence mined using a subcellular localization prediction algorithm (Lin et al., BMC Bioinformat [BMC Bioinformatics] 13:157 (2012), which is incorporated herein by reference in its entirety).

[0363] Table 11 Exemplary nuclear localization signals used in gene modification systems

[0364] amino acid sequence SEQ ID NO: Exemplary nucleic acid sequence SEQ ID NO: PAAKRVKLDGG 36 CCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGC 37 KRTADGSEFEKRTADGSEFESPKKKAKVE 38 AAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAG 39

[0365] In some embodiments, the genetically modified polypeptide disclosed herein comprises an N-terminal NLS containing the amino acid sequence of PAAKRVKLDGG (SEQ ID NO: 36) or a functional fragment thereof (e.g., an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the genetically modified polypeptide disclosed herein comprises a C-terminal NLS containing the amino acid sequence of KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 38) or a functional fragment thereof (e.g., an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it).

[0366] Other exemplary NLS sequences are also described in PCT / EP2000 / 011690, the contents of which are incorporated herein by reference for their disclosure of exemplary nuclear localization sequences.

[0367] In some embodiments, the gene-modified polypeptide comprises, in the order from N-terminus to C-terminus, one or more of the following: an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA-binding domain, an adapter, an RT domain, and / or a second NLS (e.g., 1, 2, 3, 4, 5, or all 6). In some embodiments, the gene-modified polypeptide comprises, in the direction from N-terminus to C-terminus, a first NLS, a DNA-binding domain, an adapter, an RT domain, and a second NLS.

[0368] In some embodiments, the genetically modified polypeptide comprises: (i) an N-terminal NLS containing the NLS sequence of PAAKRVKLDGG (SEQ ID NO: 36) or a functional fragment thereof (e.g., an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it); (ii) a Cas domain containing the Cas sequence of SEQ ID NO: 52, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; (iii) a linker containing the linker sequence of AEAAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 54), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and (iv) an RT domain containing the NLS sequence of SEQ ID NO: 52. The RT domain sequence of NO:56, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it; and the C-terminal NLS of (v) KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 38) or a functional fragment thereof (e.g., an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it). In some embodiments, the gene-modifying polypeptide comprises the amino acid sequence of SEQ ID NO 28 (e.g., as shown in Table T3 below), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0369] Table T3. Exemplary Genetically Modified Peptides

[0370] sequence SEQ ID NO PAAKRVKLDGGDKKYSIGLDIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFLKDDSIDNKVLTRSDKARGKSDNVPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYTSTKEVLDATLIHQSITGLYETRIDLSQLGGDGGAEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKAGGTAPLEEEYRLFLEAPIQNVTLLEQWKREIPKVWAEINPPGLASTQAPIHVQLLSTALPVRVRQYPITLEAKRSLRETIRKFRAAGILRPVHSPWNTPLLPVRKSGTSEYRMVQDLREVNKRVETIHPTVPNPYTLLSLLPPDRIWYSVLDLKDAFFCIPLAPESQLIFAFEWADAEEGESGQLTWTRLPQGFKNSPTLFNEALNRDLQGFRLDHPSVSLLQYVDDLLIAADTQAACLSATRDLLMTLAELGYRVSGKKAQLCQEEVTYLGFKIHKGSRSLSNSRTQAILQIPVPKTKRQVREFLGKIGYCRLFIPGFAELAQPLYAATRPGNDPLVWGEKEEEAFQSLKLALTQPPALALPSLDKPFQLFVEETSGAAKGVLTQALGPWKRPVAYLSKRLDPVAAGWPRCLRAIAAAALLTREASKLTFGQDIEITSSHNLESLLRSPPDKWLTNARITQYQVLLLDPPRVRFKQTAALNPATLLPETDDTLPIHHCLDTLDSLTSTRPDLTDQPLAQAEATLFTDGSSYIRDGKRYAGAAVVTLDSVIWAEPLPIGTSAQKAELIALTKALEWSKDKSVNIYTDSRYAFATLHVHGMIYRERGWLTAGGKAIKNAPEILALLTAVWLPKRVAVMHCKGHQKDDAPTSTGNRRADEVAREVAIRPLSTQATISAGKRTADGSEFEKRTADGSEFESPKKKAKVE 28

[0371] In some embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.

[0372] In some embodiments, the genetically modified polypeptide further comprises (e.g., the C-terminus of a second NLS) a T2A sequence and / or a puromycin sequence. In some embodiments, the nucleic acid encoding the genetically modified polypeptide (e.g., as described herein) encodes the T2A sequence, for example, wherein the T2A sequence is located between the region encoding the genetically modified polypeptide and a second region, wherein the second region optionally encodes an optional marker, such as puromycin.

[0373] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the first NLS and the DNA-binding domain. In some embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises the amino acid sequence GG.

[0374] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the DNA-binding domain and the adapter. In some embodiments, the spacer sequence between the DNA-binding domain and the adapter comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the DNA-binding domain and the adapter comprises the amino acid sequence GG.

[0375] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the adapter and the RT domain. In some embodiments, the spacer sequence between the adapter and the RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the adapter and the RT domain comprises the amino acid sequence GG.

[0376] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the RT domain and the second NLS. In some embodiments, the spacer sequence between the RT domain and the second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the RT domain and the second NLS comprises the amino acid sequence AG.

[0377] In some embodiments, the genetically modified polypeptide further comprises a spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence. In some embodiments, the spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In some embodiments, the spacer sequence between the second NLS and T2A sequences and / or the puromycin sequence comprises the amino acid sequence GSG.

[0378] Other structural domains

[0379] Gene-modifying peptides can bind to a target DNA sequence and a template nucleic acid (e.g., template RNA), nick the target site, and write the template into DNA (e.g., reverse transcription), thereby producing modifications at the target site. In some embodiments, additional domains may be added to the peptide to improve the efficiency of the process. In some embodiments, the gene-modifying peptide may include an additional DNA-linking domain to link the reverse-transcribed DNA to the target site DNA. In some embodiments, the peptide may include a heteroRNA-binding domain. In some embodiments, the peptide may include a domain having 5' to 3' exonuclease activity (e.g., where 5' to 3' exonuclease activity increases the repair of alterations at the target site, such as those beneficial to alterations to the original genome sequence). In some embodiments, the peptide may include a domain having 3' to 5' exonuclease activity, such as proofreading activity. In some embodiments, a writing domain, such as an RT domain, has 3' to 5' exonuclease activity, such as proofreading activity.

[0380] Nucleic acids encoding gene-modified polypeptides

[0381] This article provides nucleic acids encoding gene-modified polypeptides for use in the gene-modification systems disclosed herein.

[0382] In some embodiments, the nucleic acid encoding the gene-modifying polypeptide is altered from its classical sequence to have altered codon usage, for example, for modification targeting human cells. In some embodiments, the nucleic acid molecule encoding the gene-modifying polypeptide contains one or more silent mutations in the coding region (e.g., in the sequence encoding the RT domain) relative to the nucleic acid molecule described herein.

[0383] In some embodiments, the nucleic acid (e.g., RNA) encoding the genetically modified polypeptide contains a post-transcriptional regulatory element that enhances nuclear output. In some embodiments, the post-transcriptional regulatory element is a post-transcriptional regulatory element of hepatitis B virus (HPRE) or marmot hepatitis virus (WPRE). The nucleic acid sequences of exemplary WPRE regulatory elements are shown in Table N1 below. In some embodiments, the WPRE regulatory element contained in the nucleic acid encoding the genetically modified polypeptide contains one or more mutations relative to the wild-type WPRE regulatory element, such as those listed in Zanta-Boussif, MA et al. Gene Ther. [Gene Therapy] 16,605–619 (2009).

[0384] In some embodiments, the nucleic acid flanking the gene-modifying polypeptide has an untranslated region (UTR) that modifies protein expression levels. Various 5' and 3' UTRs can influence protein expression. For example, in some embodiments, a 5' UTR modifying RNA stability or protein translation may precede the coding sequence. In some embodiments, a 3' UTR modifying RNA stability or translation may follow the sequence. In some embodiments, a 5' UTR may precede the sequence, followed by a 3' UTR modifying RNA stability or translation. In some embodiments, 5' and / or 3' UTRs may be selected to enhance protein expression. In some embodiments, 5' and / or 3' UTRs may be selected to modify protein expression, thereby minimizing overproduction inhibition. In some embodiments, the UTR is located around the coding sequence, e.g., outside the coding sequence, and in other embodiments, close to the coding sequence.

[0385] In some embodiments, the system described herein comprises DNA encoding a transcript, wherein the DNA contains corresponding 5' UTR and 3' UTR sequences, wherein T replaces U in the sequences listed above. In some embodiments, the DNA vector used to generate the RNA component of the system further comprises a promoter upstream of the 5' UTR for initiating in vitro transcription, such as the T7, T3, or SP6 promoter. The above 5' UTRs begin with GGG, which is a suitable starting point for optimizing transcription using T7 RNA polymerase. For adjusting transcription levels and altering transcription start site nucleotides to accommodate alternative 5' UTRs, the teachings of Davidson et al., Pac Symp Biocomput [Pac Symp Biocomput] 433-443 (2010) describe T7 promoter variants that satisfy both of these characteristics and their discovery methods.

[0386] Exemplary sequences of the 5' UTR and 3' UTR used in nucleic acids encoding gene-modified polypeptides described herein are shown in Tables N1, E12, and E16.

[0387] In some embodiments, the nucleic acid (e.g., RNA) encoding the gene-modified polypeptide further includes a poly(A) tail for enhancing polypeptide expression.

[0388] In some embodiments, the poly(A) tail consists only of adenosine ribonucleotides. In some embodiments, the poly(A) tail contains ribonucleotides other than adenosine, for example, scattered within / between the polyadenosine regions.

[0389] In some embodiments, the poly(A) tail consists of the following nucleic acid sequences, which include patterns of 12-18 adenosine nucleotides followed by at least two non-adenosine nucleotides, 22-28 adenosine nucleotides followed by at least two non-adenosine nucleotides, 32-38 adenosine nucleotides followed by at least two non-adenosine nucleotides, and 42-48 adenosine nucleotides.

[0390] In some embodiments, the poly(A) tail consists of the following nucleic acid sequences, which include patterns of 13-17 adenosine nucleotides followed by 1-4 non-adenosine nucleotides, 23-27 adenosine nucleotides followed by 1-5 non-adenosine nucleotides, 33-37 adenosine nucleotides followed by 2-6 non-adenosine nucleotides, and 43-47 adenosine nucleotides.

[0391] In some embodiments, the poly(A) tail consists of nucleic acid sequences comprising the following patterns: 14-16 adenosine nucleotides followed by at least two (e.g., 2-6) non-adenosine nucleotides; 24-26 adenosine nucleotides followed by at least two (e.g., 2-6) non-adenosine nucleotides; 34-36 adenosine nucleotides followed by at least two (e.g., 2-6) non-adenosine nucleotides; and 44-46 adenosine nucleotides.

[0392] In some embodiments, the poly(A) tail consists of the following nucleic acid sequences, which include patterns of 15 adenosine nucleotides followed by at least two (e.g., 2-4) non-adenosine nucleotides, 25 adenosine nucleotides followed by at least two (e.g., 2-4) non-adenosine nucleotides, 35 adenosine nucleotides followed by at least two (e.g., 2-4) non-adenosine nucleotides, and 45 adenosine nucleotides.

[0393] In some embodiments, the poly(A) tail consists of the following nucleic acid sequences, which include patterns of 15 adenosine nucleotides followed by 1-4 non-adenosine nucleotides, 25 adenosine nucleotides followed by 1-5 non-adenosine nucleotides, 35 adenosine nucleotides followed by 2-6 non-adenosine nucleotides, and 45 adenosine nucleotides.

[0394] In some embodiments, there are two to ten nonadenosine nucleotides. In some embodiments, the number of nonadenosine nucleotides in each group of nonadenosine nucleotides increases from 5' to 3'. In some embodiments, the number of nonadenosine nucleotides in each group of nonadenosine nucleotides is 2, 3, and 4, respectively, from 5' to 3'. In some embodiments, the nonadenosine nucleotides are selected from cytosine nucleotides or uridine nucleotides.

[0395] Exemplary sequences for use in nucleic acids encoding gene-modified polypeptides described herein are shown in Tables N1, E12, and E16.

[0396] Table N1: Exemplary Adjustment Elements, 5' UTR, 3' UTR and Poly(A) Tail

[0397] Component type Sequence SEQ ID NO WPRE element AAUCAACCUCUGGAUUACAAAAUUUGUGAAAGAUUGACUGAUAUUCUUAACUAUGUUGCUCCUUUUACGCUGUGUGGAUAUGCUGCUUUAAUGCCUCUGUAUCAUGCUAUUGCUUCCCGUACGGCUUUCGUUUUCUCCUCCUUGUAUAAAUCCUGGUUGCUGUCUCUUUAUGAGGAGUUGUGGCCCGUUGUCCGUCAACGUGGCGUGGUGUGCUCUGUGUUUGCUGACGCAACCCCCACUGGCUGGGGCAUUGCCACCACCUGUCAACUCCUUUCUGGGACUUUCGCUUUCCCCCUCCCGAUCGCCACGGCAGAACUCAUCGCCGCCUGCCUUGCCCGCUGCUGGACAGGGGCUAGGUUGCUGGGCACUGAUAAUUCCGUGGUGUUGUCGGGGAAGCUGACGUCCUUUCCAUGGCUGCUCGCCUGUGUUGCCAACUGGAUCCUGCGCGGGACGUCCUUCUGCUACGUCCCUUCGGCUCUCAAUCCAGCGGACCUCCCUUCCCGAGGCCUUCUGCCGGUUCUGCGGCCUCUCCCGCGUCUUCGCUUUCGGCCUCCGACGAGUCGGAUCUCCCUUUGGGCCGCCUCCCCGCCUG 40 Mutant WPRE element, variant 1 AATCAACCTCTGGATTACAAAATTTGTGAAAGATTGACTGATATTCTTAACTATGTTGCTCCTTTTACGCTGTGTGGATATGCTGCTTTAATGCCTCTGTATCATGCTATTGCTTCCCGTACGGCTTTCGTTTTCTCCTCCTTGTATAAATCCTGGTTGCTGTCTCTTTATGAGGAGTTGTGGCCCGTTGTCCGTCAACGTGGCGTGGTGTGCTCTGTGTTTGCTGACGCAACCCCCACTGGCTGGGGCATTGCCACCACCTGTCAACTCCTTTCTGGGACTTTCGCTTTCCCCCTCCCGATCGCCACGGCAGAACTCATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGGCTAGGTTGCTGGGCACTGATAATTCCGTGGTGTTGTCGGGGAAGCTGACGTCCTTTCCAGGGCTGCTCGCCTGTGTTGCCAACTGGATCCTGCGCGGGACGTCCTTCTGCTACGTCCCTTCGGCTCTCAATCCAGCGGACCTCCCTTCCCGAGGCCTTCTGCCGGTTCTGCGGCCTCTCCCGCGTCTTCGCTTTCGGCCTCCGACGAGTCGGATCTCCCTTTGGGCCGCCTCCCCGCCTG 101 Mutant WPRE element, variant 2 AATCAACCTCTGGATTACAAAATTTGTGAAAGATTGACTGATATTCTTAACTATGTTGCTCCTTTTACGCTGTGTGGATATGCTGCTTTAATGCCTCTGTATCATGCTATTGCTTCCCGTACGGCTTTCGTTTTCTCCTCCTTGTATAAATCCTGGTTGCTGTCTCTTTATGAGGAGTTGTGGCCCGTTGTCCGTCAACGTGGCGTGGTGTGCTCTGTGTTTGCTGACGCAACCCCCACTGGCTGGGGCATTGCCACCACCTGTCAACTCCTTTCTGGGACTTTCGCTTTCCCCCTCCCGATCGCCACGGCAGAACTCATCGCCGCCTGCCTTGCCCGCTGCTGGACAGGGGCTAGGTTGCTGGGCACTGATAATTCCGTGGTGTTGTCGGGGAAATCATCGTCCTTTCCTTGGCTGCTCGCCTGTGTTGCCAACTGGATCCTGCGCGGGACGTCCTTCTGCTACGTCCCTTCGGCTCTCAATCCAGCGGACCTCCCTTCCCGAGGCCTTCTGCCGGTTCTGCGGCCTCTCCCGCGTCTTCGCTTTCGGCCTCCGACGAGTCGGATCTCCCTTTGGGCCGCCTCCCCGCCTG 102 5’ UTR 1 AGGAUAAUAUACUUACAUACUUACUAAUUAAUACUAAACUCAACGCCACC 41 5’ UTR 2 AGGACAACAAUAACUCUAAUAAACAAACGAAUUCUAUUAUAUCACCACAUUAUAUCUCUAAUCUGCCACC 42 5’ UTR 3 AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC 43 5’ UTR 4 AGGCAAAAAUCAAAAUCAAUCAUCAUCACAACAUCAACAAUCAAUCAUCAACACAUCAUCAAGACACCACC 44 3’ UTR 1 UUGCCAUGUGUAUGUGGGUUUUUUUUUUCCCACAUACUCUGAUGAUCCUUUUUUUUUUGGAUCAUUCAUGGCAA 45 3’ UTR 2 UUGCGUUCGCGCAA 46 3’ UTR 3 GCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA 47 3' UTR 4 GCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGACUAG 48 A-tail 1 AAAAAAAAAAAAAAACCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 49 A-tail 2 AAAAAAAAAAAAAAAUUAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 50 A-tail 3 AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 51

[0398] In some embodiments, the nucleic acid encoding the gene-modified polypeptide comprises a nucleic acid sequence listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the nucleic acid encoding the gene-modified polypeptide comprises the nucleic acid sequence of SEQ ID NO: 40, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the nucleic acid encoding the gene-modified polypeptide comprises the nucleic acid sequence of SEQ ID NO: 101, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the nucleic acid encoding the gene-modified polypeptide comprises the nucleic acid sequence of SEQ ID NO: 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it. In some embodiments, the nucleic acid encoding the gene-modified polypeptide comprises the nucleic acid sequence of SEQ ID NO:106, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0399] Table N2: Exemplary nucleic acid sequences encoding gene-modified polypeptides

[0400] name RNA sequence SEQIDNO RNAIVT1299 AGGAUAAUAUACUUACAUACUUACUAAUUAAUACUAAACUCAACGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAUUGCCAUGUGUAUGUGGGUUUUUUUUUUCCCACAUACUCUGAUGAUCCUUUUUUUUUUGGAUCAUUCAUGGCAAAAUCAACCUCUGGAUUACAAAAUUUGUGAAAGAUUGACUGAUAUUCUUAACUAUGUUGCUCCUUUUACGCUGUGUGGAUAUGCUGCUUUAAUGCCUCUGUAUCAUGCUAUUGCUUCCCGUACGGCUUUCGUUUUCUCCUCCUUGUAUAAAUCCUGGUUGCUGUCUCUUUAUGAGGAGUUGUGGCCCGUUGUCCGUCAACGUGGCGUGGUGUGCUCUGUGUUUGCUGACGCAACCCCCACUGGCUGGGGCAUUGCCACCACCUGUCAACUCCUUUCUGGGACUUUCGCUUUCCCCCUCCCGAUCGCCACGGCAGAACUCAUCGCCGCCUGCCUUGCCCGCUGCUGGACAGGGGCUAGGUUGCUGGGCACUGAUAAUUCCGUGGUGUUGUCGGGGAAGCUGACGUCCUUUCCAUGGCUGCUCGCCUGUGUUGCCAACUGGAUCCUGCGCGGGACGUCCUUCUGCUACGUCCCUUCGGCUCUCAAUCCAGCGGACCUCCCUUCCCGAGGCCUUCUGCCGGUUCUGCGGCCUCUCCCGCGUCUUCGCUUUCGGCCUCCGACGAGUCGGAUCUCCCUUUGGGCCGCCUCCCCGCCUGAAAAAAAAAAAAAAACCAAAAAAAAAAAAAAAAAAAAAAAAACCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACCCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 29 RNAIVT1350 AGGACAACAAUAACUCUAAUAAACAAACGAAUUCUAUUAUAUCACCACAUUAUAUCUCUAAUCUGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAUUGCGUUCGCGCAAAAUCAACCUCUGGAUUACAAAAUUUGUGAAAGAUUGACUGAUAUUCUUAACUAUGUUGCUCCUUUUACGCUGUGUGGAUAUGCUGCUUUAAUGCCUCUGUAUCAUGCUAUUGCUUCCCGUACGGCUUUCGUUUUCUCCUCCUUGUAUAAAUCCUGGUUGCUGUCUCUUUAUGAGGAGUUGUGGCCCGUUGUCCGUCAACGUGGCGUGGUGUGCUCUGUGUUUGCUGACGCAACCCCCACUGGCUGGGGCAUUGCCACCACCUGUCAACUCCUUUCUGGGACUUUCGCUUUCCCCCUCCCGAUCGCCACGGCAGAACUCAUCGCCGCCUGCCUUGCCCGCUGCUGGACAGGGGCUAGGUUGCUGGGCACUGAUAAUUCCGUGGUGUUGUCGGGGAAGCUGACGUCCUUUCCAUGGCUGCUCGCCUGUGUUGCCAACUGGAUCCUGCGCGGGACGUCCUUCUGCUACGUCCCUUCGGCUCUCAAUCCAGCGGACCUCCCUUCCCGAGGCCUUCUGCCGGUUCUGCGGCCUCUCCCGCGUCUUCGCUUUCGGCCUCCGACGAGUCGGAUCUCCCUUUGGGCCGCCUCCCCGCCUGAAAAAAAAAAAAAAAUUAAAAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAUUUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 30 RNAIVT1425 AGGACAACAAUAACUCUAAUAAACAAACGAAUUCUAUUAUAUCACCACAUUAUAUCUCUAAUCUGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAUUGCGUUCGCGCAAAAAAAAAAAAAAAAAUUAAAAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAUUUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 31 RNAIVT1535 AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGAAAAAAAAAAAAAAAACCAAAAAAAAAAAAAAAAAAAAAAAAACCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACCCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 32 RNAIVT1536 AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGAAAAAAAAAAAAAAAAUUAAAAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAUUUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 33 RNAIVT1044 AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 34 RNAIVT1540 AGGCAAAAAUCAAAAUCAAUCAUCAUCACAACAUCAACAAUCAAUCAUCAACACAUCAUCAAGACACCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGAAAAAAAAAAAAAAAACCAAAAAAAAAAAAAAAAAAAAAAAAACCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACCCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 35 RNAIVT6650 AGGAUAAUAUACUUACAUACUUACUAAUUAAUACUAAACUCAACGCCACCAUGCCCGCCGCCAAGCGGGUGAAGCUGGACGGCGGCGACAAGAAGUACAGCAUCGGCCUGGACAUCGGCACCAACAGCGUGGGCUGGGCCGUGAUCACCGACGAGUACAAGGUGCCCAGCAAGAAGUUCAAGGUGCUGGGCAACACCGACCGGCACAGCAUCAAGAAGAACCUGAUCGGCGCCCUGCUGUUCGACAGCGGCGAGACCGCCGAGGCCACCCGGCUGAAGCGGACCGCCCGGCGGCGGUACACCCGGCGGAAGAACCGGAUCUGCUACCUGCAGGAGAUCUUCAGCAACGAGAUGGCCAAGGUGGACGACAGCUUCUUCCACCGGCUGGAGGAGAGCUUCCUGGUGGAGGAGGACAAGAAGCACGAGCGGCACCCCAUCUUCGGCAACAUCGUGGACGAGGUGGCCUACCACGAGAAGUACCCCACCAUCUACCACCUGCGGAAGAAGCUGGUGGACAGCACCGACAAGGCCGACCUGCGGCUGAUCUACCUGGCCCUGGCCCACAUGAUCAAGUUCCGGGGCCACUUCCUGAUCGAGGGCGACCUGAACCCCGACAACAGCGACGUGGACAAGCUGUUCAUCCAGCUGGUGCAGACCUACAACCAGCUGUUCGAGGAGAACCCCAUCAACGCCAGCGGCGUGGACGCCAAGGCCAUCCUGAGCGCCCGGCUGAGCAAGAGCCGGCGGCUGGAGAACCUGAUCGCCCAGCUGCCCGGCGAGAAGAAGAACGGCCUGUUCGGCAACCUGAUCGCCCUGAGCCUGGGCCUGACCCCCAACUUCAAGAGCAACUUCGACCUGGCCGAGGACGCCAAGCUGCAGCUGAGCAAGGACACCUACGACGACGACCUGGACAACCUGCUGGCCCAGAUCGGCGACCAGUACGCCGACCUGUUCCUGGCCGCCAAGAACCUGAGCGACGCCAUCCUGCUGAGCGACAUCCUGCGGGUGAACACCGAGAUCACCAAGGCCCCACUGAGCGCCAGCAUGAUCAAGCGGUACGACGAGCACCACCAGGACCUGACCCUGCUGAAGGCCCUGGUGCGGCAGCAGCUGCCCGAGAAGUACAAGGAGAUCUUCUUCGACCAGAGCAAGAACGGCUACGCCGGCUACAUCGACGGCGGCGCCAGCCAGGAGGAGUUCUACAAGUUCAUCAAGCCCAUCCUGGAGAAGAUGGACGGCACCGAGGAGCUGCUGGUGAAGCUGAACCGGGAGGACCUGCUGCGGAAGCAGCGGACCUUCGACAACGGCAGCAUCCCCCACCAGAUCCACCUGGGCGAGCUGCACGCCAUCCUGCGGCGGCAGGAGGACUUCUACCCCUUCCUGAAGGACAACCGGGAGAAGAUCGAGAAGAUCCUGACCUUCCGGAUCCCCUACUACGUGGGCCCACUGGCCCGGGGCAACAGCCGGUUCGCCUGGAUGACCCGGAAAAGCGAGGAGACCAUCACCCCCUGGAACUUCGAGGAGGUGGUGGACAAGGGCGCCAGCGCCCAGAGCUUCAUCGAGCGGAUGACCAACUUCGACAAGAACCUGCCCAACGAGAAGGUGCUGCCCAAGCACAGCCUGCUGUACGAGUACUUCACCGUGUACAACGAGCUGACCAAGGUGAAGUACGUGACCGAGGGCAUGCGGAAGCCAGCCUUCCUGAGCGGCGAGCAGAAGAAGGCCAUCGUGGACCUGCUGUUCAAGACCAACCGGAAGGUGACCGUGAAGCAGCUGAAGGAGGACUACUUCAAGAAGAUCGAGUGCUUCGACAGCGUGGAGAUCAGCGGCGUGGAGGACCGGUUCAACGCCAGCCUGGGCACCUACCACGACCUGCUGAAGAUCAUCAAGGACAAGGACUUCCUGGACAACGAGGAGAACGAGGACAUCCUGGAGGACAUCGUGCUGACCCUGACCCUGUUCGAGGACCGGGAGAUGAUCGAGGAGCGGCUGAAGACCUACGCCCACCUGUUCGACGACAAGGUGAUGAAGCAGCUGAAGCGGCGGCGGUACACCGGCUGGGGCCGGCUGAGCCGGAAGCUGAUCAACGGCAUCCGGGACAAGCAGAGCGGCAAGACCAUCCUGGACUUCCUGAAAAGCGACGGCUUCGCCAACCGGAACUUCAUGCAGCUGAUCCACGACGACAGCCUGACCUUCAAGGAGGACAUCCAGAAGGCCCAGGUGAGCGGCCAGGGCGACAGCCUGCACGAGCACAUCGCCAACCUGGCCGGCAGCCCCGCCAUCAAGAAGGGCAUCCUGCAGACCGUGAAGGUGGUGGACGAGCUGGUGAAGGUGAUGGGCCGGCACAAGCCCGAGAACAUCGUGAUCGAGAUGGCCCGGGAGAACCAGACCACCCAGAAGGGCCAGAAGAACAGCCGGGAGCGGAUGAAGCGGAUCGAGGAGGGCAUCAAGGAGCUGGGCAGCCAGAUCCUGAAGGAGCACCCCGUGGAGAACACCCAGCUGCAGAACGAGAAGCUGUACCUGUACUACCUGCAGAACGGCCGGGACAUGUACGUGGACCAGGAGCUGGACAUCAACCGGCUGAGCGACUACGACGUGGACCACAUCGUGCCCCAGAGCUUCCUGAAGGACGACAGCAUCGACAACAAGGUGCUGACCCGGAGCGACAAGGCCCGGGGCAAGAGCGACAACGUGCCCAGCGAGGAGGUGGUGAAGAAGAUGAAGAACUACUGGCGGCAGCUGCUGAACGCCAAGCUGAUCACCCAGCGGAAGUUCGACAACCUGACCAAGGCCGAGCGGGGCGGCCUGAGCGAGCUGGACAAGGCCGGCUUCAUCAAGCGGCAGCUGGUGGAGACCCGGCAGAUCACCAAGCACGUGGCCCAGAUCCUGGACAGCCGGAUGAACACCAAGUACGACGAGAACGACAAGCUGAUCCGGGAGGUGAAGGUGAUCACCCUCAAGAGCAAGCUGGUGAGCGACUUCCGGAAGGACUUCCAGUUCUACAAGGUGCGGGAGAUCAACAACUACCACCACGCCCACGACGCCUACCUGAACGCCGUGGUGGGCACCGCCCUGAUCAAGAAGUACCCCAAGCUGGAGAGCGAGUUCGUGUACGGCGACUACAAGGUGUACGACGUGCGGAAGAUGAUCGCCAAGAGCGAGCAGGAGAUCGGCAAGGCCACCGCCAAGUACUUCUUCUACAGCAACAUCAUGAACUUCUUCAAGACCGAGAUCACCCUGGCCAACGGCGAGAUCCGGAAGCGGCCCCUGAUCGAGACCAACGGCGAGACCGGCGAGAUCGUGUGGGACAAGGGCCGGGACUUCGCCACCGUGCGGAAGGUGCUGAGCAUGCCCCAGGUGAACAUCGUGAAGAAGACCGAGGUGCAGACCGGCGGCUUCAGCAAGGAGAGCAUCCUGCCCAAGCGGAACAGCGACAAGCUGAUCGCCCGGAAGAAGGACUGGGACCCAAAGAAGUACGGCGGCUUCGACAGCCCCACCGUGGCCUACAGCGUGCUGGUGGUGGCCAAGGUGGAGAAGGGCAAGAGCAAGAAGCUCAAGAGCGUGAAGGAGCUGCUGGGCAUCACCAUCAUGGAGCGGAGCAGCUUCGAGAAGAACCCCAUCGACUUCCUGGAGGCCAAGGGCUACAAGGAGGUGAAGAAGGACCUGAUCAUCAAGCUGCCCAAGUACAGCCUGUUCGAGCUGGAGAACGGCCGGAAGCGGAUGCUGGCCAGCGCCGGCGAGCUGCAGAAGGGCAACGAGCUGGCCCUGCCCAGCAAGUACGUGAACUUCCUGUACCUGGCCAGCCACUACGAGAAGCUGAAGGGCAGCCCCGAGGACAACGAGCAGAAGCAGCUGUUCGUGGAGCAGCACAAGCACUACCUGGACGAGAUCAUCGAGCAGAUCAGCGAGUUCAGCAAGCGGGUGAUCCUGGCCGACGCCAACCUGGACAAGGUGCUGAGCGCCUACAACAAGCACCGGGACAAGCCUAUCCGGGAGCAGGCCGAGAACAUCAUCCACCUGUUCACCCUGACCAACCUGGGCGCCCCCGCCGCCUUCAAGUACUUCGACACCACCAUCGACCGGAAGCGGUACACCAGCACCAAGGAGGUGCUGGACGCCACCCUGAUCCACCAGAGCAUCACCGGCCUGUACGAGACCCGGAUCGACCUGAGCCAGCUGGGCGGCGACGGCGGCGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCCUGGAGGCCGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGAGGCCGCCGCCAAGGCCGGCGGCACCGCCCCCCUGGAGGAGGAGUACCGGCUGUUCCUGGAGGCCCCCAUCCAGAACGUGACCCUGCUGGAGCAGUGGAAGCGGGAGAUCCCCAAGGUGUGGGCCGAGAUCAACCCUCCCGGCCUGGCCAGCACCCAGGCCCCCAUCCACGUGCAGCUGCUGAGCACCGCCCUGCCCGUGCGGGUGCGGCAGUACCCAAUCACCCUGGAGGCCAAGCGGAGCCUGCGGGAGACCAUCCGGAAGUUCCGGGCCGCCGGCAUCCUGCGGCCCGUGCACAGCCCCUGGAACACCCCUCUGCUGCCCGUGCGGAAAAGCGGCACCAGCGAGUACCGGAUGGUGCAGGACCUGCGGGAGGUGAACAAGCGGGUGGAGACCAUCCACCCUACCGUGCCCAACCCCUACACCCUGCUGAGCCUGCUGCCCCCCGACCGGAUCUGGUACAGCGUGCUGGACCUGAAGGACGCCUUCUUCUGCAUCCCUCUGGCCCCCGAGAGCCAGCUGAUCUUCGCCUUCGAGUGGGCCGACGCCGAGGAGGGCGAGAGCGGCCAGCUGACCUGGACCCGGCUGCCUCAGGGCUUCAAGAACAGCCCCACCCUGUUCAACGAGGCCCUGAACCGGGACCUGCAGGGCUUCCGGCUGGACCACCCCAGCGUGAGCCUGCUGCAGUACGUGGACGACCUGCUGAUCGCCGCCGACACCCAGGCCGCCUGCCUGAGCGCCACCCGGGACCUGCUGAUGACCCUGGCCGAGCUGGGCUACCGGGUGAGCGGCAAGAAGGCCCAGCUGUGCCAGGAGGAGGUGACCUACCUGGGCUUCAAGAUCCACAAGGGCAGCCGGAGCCUGAGCAACAGCCGGACCCAGGCCAUCCUGCAGAUCCCUGUGCCCAAGACCAAGCGGCAGGUGCGGGAGUUCCUGGGCAAGAUCGGCUACUGCCGGCUGUUCAUCCCCGGCUUCGCCGAGCUGGCCCAGCCCCUGUACGCCGCCACCCGGCCCGGCAACGACCCCCUGGUGUGGGGCGAGAAGGAGGAGGAGGCCUUCCAGAGCCUGAAGCUGGCCCUGACCCAGCCCCCAGCCCUGGCCCUGCCCAGCCUGGACAAGCCCUUCCAGCUGUUCGUGGAGGAGACCAGCGGCGCCGCCAAGGGCGUGCUGACCCAGGCCCUGGGCCCCUGGAAGCGGCCCGUGGCCUACCUGAGCAAGCGGCUGGACCCCGUGGCCGCCGGCUGGCCCCGGUGCCUGCGGGCCAUCGCCGCCGCCGCCCUGCUGACCCGGGAGGCCAGCAAGCUGACCUUCGGCCAGGACAUCGAGAUCACCAGCAGCCACAACCUGGAGAGCCUGCUGCGGAGCCCCCCCGACAAGUGGCUGACCAACGCCCGGAUCACCCAGUACCAGGUGCUGCUGCUGGACCCCCCCCGGGUGCGGUUCAAGCAGACCGCCGCCCUGAACCCCGCCACCCUGCUGCCCGAGACCGACGACACCCUGCCCAUCCACCACUGCCUGGACACCCUGGACAGCCUGACCAGCACCCGGCCCGACCUGACCGACCAGCCCCUGGCCCAGGCCGAGGCCACCCUGUUCACCGACGGCAGCAGCUACAUCCGGGACGGCAAGCGGUACGCGGGCGCCGCCGUGGUGACCCUGGACAGCGUGAUCUGGGCCGAGCCCCUGCCCAUCGGCACCAGCGCCCAGAAGGCCGAGCUGAUCGCCCUGACCAAGGCCCUGGAGUGGAGCAAGGACAAGAGCGUGAACAUCUACACCGACAGCCGGUACGCCUUCGCCACCCUGCACGUGCACGGCAUGAUCUACCGGGAGCGGGGCUGGCUGACCGCGGGCGGCAAGGCCAUCAAGAACGCCCCCGAGAUCCUGGCCCUGCUGACCGCCGUGUGGCUGCCCAAGCGGGUGGCCGUGAUGCACUGCAAGGGCCACCAGAAGGACGACGCCCCCACCAGCACCGGCAACCGGCGGGCCGACGAGGUGGCCCGGGAGGUGGCCAUCCGGCCCCUGAGCACCCAGGCCACCAUCAGCGCCGGCAAGCGGACCGCCGACGGCAGCGAGUUCGAGAAGCGGACCGCCGACGGCAGCGAGUUCGAGAGCCCCAAGAAGAAGGCCAAGGUGGAGUAAUAGUGAUUGCCAUGUGUAUGUGGGUUUUUUUUUUCCCACAUACUCUGAUGAUCCUUUUUUUUUUGGAUCAUUCAUGGCAAAAUCAACCUCUGGAUUACAAAAUUUGUGAAAGAUUGACUGAUAUUCUUAACUAUGUUGCUCCUUUUACGCUGUGUGGAUAUGCUGCUUUAAUGCCUCUGUAUCAUGCUAUUGCUUCCCGUACGGCUUUCGUUUUCUCCUCCUUGUAUAAAUCCUGGUUGCUGUCUCUUUAUGAGGAGUUGUGGCCCGUUGUCCGUCAACGUGGCGUGGUGUGCUCUGUGUUUGCUGACGCAACCCCCACUGGCUGGGGCAUUGCCACCACCUGUCAACUCCUUUCUGGGACUUUCGCUUUCCCCCUCCCGAUCGCCACGGCAGAACUCAUCGCCGCCUGCCUUGCCCGCUGCUGGACAGGGGCUAGGUUGCUGGGCACUGAUAAUUCCGUGGUGUUGUCGGGGAAGCUGACGUCCUUUCCAgGGCUGCUCGCCUGUGUUGCCAACUGGAUCCUGCGCGGGACGUCCUUCUGCUACGUCCCUUCGGCUCUCAAUCCAGCGGACCUCCCUUCCCGAGGCCUUCUGCCGGUUCUGCGGCCUCUCCCGCGUCUUCGCUUUCGGCCUCCGACGAGUCGGAUCUCCCUUUGGGCCGCCUCCCCGCCUGAAAAAAAAAAAAAAACCAAAAAAAAAAAAAAAAAAAAAAAAACCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAACCCCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA 106

[0401] Template Nucleic Acid

[0402] The gene modification system described herein can modify host target DNA sites using a template nucleic acid sequence. In some embodiments, the gene modification system described herein transcribes an RNA sequence template into the host target DNA site via target-induced reverse transcription (TPRT). By modifying one or more DNA sequences by directly reverse transcribing the RNA sequence template into the host genome, the gene modification system can insert the target sequence into the target genome without introducing a foreign DNA sequence into the host cell (unlike, for example, CRISPR systems) and eliminating the foreign DNA insertion step. The gene modification system can also remove sequences from the target genome or introduce substitutions using the target sequence. Therefore, the gene modification system provides a platform for using a customized RNA sequence template containing the target sequence, for example, a sequence containing heterologous gene coding and / or functional information.

[0403] In some embodiments, the template nucleic acid comprises one or more sequences (e.g., two sequences) that bind to the genetically modified polypeptide.

[0404] In some embodiments, the systems or methods described herein comprise a single template nucleic acid (e.g., template RNA). In some embodiments, the systems or methods described herein comprise multiple template nucleic acids (e.g., template RNA). For example, the systems described herein comprise a first RNA and a second RNA (e.g., template RNA), the first RNA comprising (e.g., from 5' to 3') a sequence that binds a genetically modified polypeptide (e.g., a DNA-binding domain and / or a nuclease domain, e.g., gRNA) and a sequence that binds a target site (e.g., a second strand of a site in the target genome), the second RNA comprising (e.g., from 5' to 3') a sequence that optionally binds a genetically modified polypeptide (e.g., a specific binding RT domain), a heterologous object sequence, and a PBS sequence. In some embodiments, when the system comprises multiple nucleic acids, each nucleic acid comprises a conjugation domain. In some embodiments, the conjugation domain enables nucleic acid molecules to be associated, e.g., through hybridization of complementary sequences. For example, in some embodiments, the first RNA comprises a first conjugation domain and the second RNA comprises a second conjugation domain, and the first and second conjugation domains are capable of hybridizing with each other, for example, under stringent conditions. In some embodiments, the stringent conditions for hybridization include hybridization at about 65°C in 4x sodium chloride / sodium citrate (SSC), followed by washing at about 65°C in 1x SSC.

[0405] In some embodiments, the template nucleic acid comprises RNA. In some embodiments, the template nucleic acid comprises DNA (e.g., single-stranded or double-stranded DNA).

[0406] In some embodiments, the template nucleic acid includes one or more (e.g., two) homologous domains that are homologous to the target sequence. In some embodiments, the length of the homologous domain is approximately 10-20, 20-50, or 50-100 nucleotides.

[0407] In some embodiments, the template RNA may comprise a gRNA sequence, for example, to guide a genetically modified polypeptide to a target site. In some embodiments, the template RNA comprises (e.g., from 5′ to 3′): (i) a gRNA spacer that optionally binds to a target site (e.g., the second strand of a site in the target genome), (ii) a gRNA scaffold that optionally binds to a polypeptide described herein (e.g., a genetically modified polypeptide or a Cas polypeptide), (iii) a heterologous object sequence comprising a mutant region (optionally, the heterologous object sequence from 5′ to 3′ comprises a first homologous region, a mutant region, and a second homologous region), and (iv) a primer binding site (PBS) sequence comprising a 3′ target homologous domain.

[0408] The template nucleic acid (e.g., template RNA) component of the genome editing system described herein is typically capable of binding the system's genetically modified peptide. In some embodiments, the template nucleic acid (e.g., template RNA) has a 3′ region capable of binding the genetically modified peptide. The binding region (e.g., the 3′ region) can be a structured RNA region, for example, having at least one, two, or three hairpin loops capable of binding the system's genetically modified peptide. The binding region can associate the template nucleic acid (e.g., template RNA) with any peptide module. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) can be associated with an RNA-binding domain in the peptide. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) can be associated with a reverse transcription domain of the genetically modified peptide (e.g., specifically binding RT domain). In some embodiments, the template nucleic acid (e.g., template RNA) can be associated with a DNA-binding domain of the peptide, for example, gRNA associated with a Cas9-derived DNA-binding domain. In some embodiments, the binding region can also provide DNA target recognition, for example, gRNA hybridizing with a target DNA sequence and binding the peptide, such as a Cas9 domain. In some embodiments, the template nucleic acid (e.g., template RNA) may be associated with multiple components of the polypeptide (e.g., DNA-binding domain and reverse transcription domain).

[0409] In some embodiments, the template RNA has a poly-A tail at the 3' end. In some embodiments, the template RNA does not have a poly-A tail at the 3' end.

[0410] In some embodiments, the template nucleic acid is template RNA. In some embodiments, the template RNA comprises one or more modified nucleotides. For example, in some embodiments, the template RNA comprises one or more deoxyribonucleotides. In some embodiments, regions of the template RNA are replaced with DNA nucleotides, for example, to enhance the stability of the molecule. For example, the 3' end of the template may contain DNA nucleotides, while the remainder of the template contains RNA nucleotides that can be reverse transcribed. For example, in some embodiments, the heterologous object sequence is primarily or entirely composed of RNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% RNA nucleotides). In some embodiments, the PBS sequence is primarily or entirely composed of DNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% DNA nucleotides). In other embodiments, the heterologous object sequence for writing into the genome may contain DNA nucleotides. In some embodiments, the DNA nucleotides in the template are copied into the genome via domains capable of having DNA-dependent DNA polymerase activity. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain in the polypeptide. In some embodiments, the DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that is also capable of DNA-dependent DNA polymerization, such as second-strand synthesis. In some embodiments, the template molecule consists only of DNA nucleotides.

[0411] In some embodiments, the system described herein comprises two nucleic acids that together constitute the sequence of the template RNA described herein. In some embodiments, the two nucleic acids are associated with each other in a non-covalent manner, for example, directly (e.g., by base pairing) or indirectly as part of a complex comprising one or more other molecules.

[0412] The template RNA described herein, from 5' to 3', may contain: (1) a gRNA spacer; (2) a gRNA scaffold; (3) a heterologous object sequence; and (4) a primer binding site (PBS) sequence. Each of these components will now be described in more detail.

[0413] gRNA spacers and gRNA scaffolds

[0414] The template RNA described herein may include a gRNA spacer that guides the gene-modifying system to the target nucleic acid, and a gRNA scaffold that facilitates association between the template RNA and the Cas domain of the gene-modifying polypeptide. The system described herein may also include gRNA that is not part of the template nucleic acid. For example, gRNA containing a gRNA spacer and a gRNA scaffold but not containing a heterologous target sequence or PBS sequence may be used, for example, to induce a second-strand nick, as described in the section entitled “Second-Strand Nick” herein.

[0415] In some embodiments, gRNA is a short synthetic RNA consisting of a scaffold sequence involved in the binding of CRISPR-related proteins and a user-defined target sequence of approximately 20 nucleotides targeting a genomic target. The structure of a complete gRNA is described by Nishimasu et al. in Cell [Cell] 156, pp. 935-949 (2014). gRNA (also known as sgRNA, a single guide RNA) consists of sequences derived from crRNA and tracrRNA, linked by an artificial four-loop. The crRNA sequence is divided into a guide region (20 nt) and a repetitive sequence region (12 nt), while the tracrRNA sequence is divided into an anti-repetitive sequence region (14 nt) and three tracrRNA stem-loops (Nishimasu et al., Cell [Cell] 156, pp. 935-949 (2014)). In practice, guide RNA sequences are typically designed to be 17-24 nucleotides long (e.g., 19, 20, or 21 nucleotides) and complementary to the target nucleic acid sequence. Custom gRNA generators and algorithms are commercially available for designing efficient guide RNAs. In some embodiments, the gRNA comprises two RNA components from the natural CRISPR system, such as crRNA and tracrRNA. As is well known in the art, the gRNA may also comprise a chimeric single guide RNA (sgRNA) containing sequences from tracrRNA (for binding nucleases) and at least one crRNA (for guiding the nuclease to the targeted editing / binding sequence). Chemically modified sgRNAs have also been shown to be effective with CRISPR-associated proteins; see, for example, Hendel et al. (2015) Nature Biotechnol., 985-991. In some embodiments, the gRNA spacer comprises a nucleic acid sequence complementary to the DNA sequence associated with the target gene.

[0416] In some embodiments, the region of the template nucleic acid (e.g., template RNA) containing gRNA adopts a band-like structure with the gRNA binding to the target DNA (e.g., as described below: Mulepati et al. Science [Science] 2014 Sep 19: Vol. 345, No. 6203, pp. 1479-1484). Not wishing to be bound by theory, this atypical structure is thought to be facilitated by rotating the RNA-DNA hybrid every six nucleotides. Therefore, in some embodiments, the region of the template nucleic acid (e.g., template RNA) containing gRNA can tolerate increased mismatches with the target site at certain intervals (e.g., every six bases). In some embodiments, the region of the template nucleic acid (e.g., template RNA) containing gRNA homologous to the target site can have a wobbly position at regular intervals (e.g., every six bases) that does not require base pairing with the target site.

[0417] In some embodiments, the template nucleic acid (e.g., template RNA) has at least 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 bases that are at least 80%, 85%, 90%, 95%, 99%, or 100% homologous to the target site, for example at the 5' end, such as a gRNA spacer sequence containing a Cas9 domain (Table 8) of a length suitable for a gene-modifying polypeptide.

[0418] In some embodiments, a Cas9 derivative with enhanced activity may be used in the genetically modified peptide. In some embodiments, the Cas9 derivative may contain a mutation that improves the activity of the HNH endonuclease domain (see, for example, Spencer and Zhang Sci Rep [Scientific Reports] 7:16836 (2017), the Cas9 derivative and the mutations it contains are incorporated herein by reference). In some embodiments, the Cas9 derivative may include one or more types of mutations described herein, such as PAM-modifying mutations, protein-stabilizing mutations, activity-enhancing mutations, and / or mutations that partially or completely inactivate one or both endonuclease domains relative to the parent enzyme (e.g., one or more mutations to eliminate endonuclease activity against one or both strands of the target DNA, such as nicking enzymes or catalytic inactivation enzymes). In some embodiments, the Cas9 enzyme used in the systems described herein may contain mutations that confer nicking enzyme activity in addition to mutations that improve catalytic efficiency. Table 12 provides parameters defining the components used to design the gRNA and / or template RNA for applying the Cas variants listed in Table 8 to genetic modifications. The cleavage site indicates the validated or predicted protospacer neighbor motif (PAM) requirement and the validated or predicted cleavage site location (the upstream base relative to the PAM site). A gRNA for a given enzyme can be assembled by linking crRNA, tetra-loop, and tracrRNA sequences, and further by adding 5′ spacers of length matching the protospacer at the target site within the spacers (min) and (max). Furthermore, the predicted location of the ssDNA cleavage on the target is important for designing PBS sequences for the template RNA (which can immediately anneal to the 5′ cleavage sequence to initiate target-induced reverse transcription). In some embodiments, the gRNA scaffold described herein comprises nucleic acid sequences or sequences having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with the following sequences in the 5′ to 3′ direction: crRNA from Table 12, tetra-loop from the same row of Table 12, and tracrRNA from the same row of Table 12. In some embodiments, the scaffold-containing gRNA or template RNA further comprises gRNA spacers of length within the spacer (min) and spacer (max) indicated in the same row of Table 12. In some embodiments, the system further comprising a genetically modified polypeptide comprises gRNA or template RNA having a sequence according to Table 12, wherein the genetically modified polypeptide comprises the Cas domain described in the same row of Table 12.

[0419] Table 12 defines the parameters used to design components of gRNA and / or template RNA to apply the Cas variants listed in Table 8 to the gene modification system.

[0420] variants PAM Cutting grade Spacer (min) Max spacer crRNA SEQIDNO: Fourth Ring Road tracrRNA SEQIDNO: SpyCas9 NGG -3 1 20 20 GTTTTAGAGCTA 10,058 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,158 SpyCas9_i_v1 NGG -3 1 20 20 GTTTTAGAGCTA 10,058 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGGACTTCGGTCCAAGTGGCACCGAGTCGGTGC 10,193 SpyCas9_i_v2 NGG -3 1 20 20 GTTTTAGAGCTA 10,058 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGGAGCTTGCTCCAAGTGGGCACCGAGTCGGTGC 10,194 SpyCas9_i_v3 NGG -3 1 20 20 GTTTTAGAGCTA 10,058 GAAA GTTTTAGAGCTAGAAATAGCAAGTTAAATAAGGCTAGTCCGTTATCGACTTGAAAAAGTCGCACCGAGTCGGTGC 10,195 SpyCas9-NG NG(NGG=NG=NGT>NGC) -3 1 20 20 GTTTTAGAGCTA 10,059 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,159 SpyCas9-SpRY NRN>NO -3 1 20 20 GTTTTAGAGCTA 10,060 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,160 SpyCas9-3var-NRRH NRRH -3 2 20 20 GTTTAAGCTATGCTG 10,074 GAAA CAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,174 SpyCas9-3var-NRTH NRTH -3 2 20 20 GTTTAAGCTATGCTG 10,075 GAAA CAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,175 SpyCas9-3var-NRCH NRCH -3 2 20 20 GTTTAAGCTATGCTG 10,076 GAAA CAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,176 SpyCas9-HF1 NGG -3 2 20 20 GTTTTAGAGCTA 10,077 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,177 SpyCas9-QQR1 NAAG -3 2 20 20 GTTTTAGAGCTA 10,078 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,178 SpyCas9-SpG NGN -3 2 20 20 GTTTTAGAGCTA 10,079 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,179 SpyCas9-VQR HERE -3 2 20 20 GTTTTAGAGCTA 10,080 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,180 SpyCas9-VRER NGCG -3 2 20 20 GTTTTAGAGCTA 10,081 GAAA TAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,181 SpyCas9-xCas NG;GAA;GAT -3 2 20 20 GTTTAAGAGCTATGCTG 10,082 GAAA CAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,182 SpyCas9-xCas-NG NG -3 2 20 20 GTTTAAGAGCTATGCTG 10,083 GAAA CAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGC 10,183

[0421] In this document, when an RNA sequence (e.g., a template RNA sequence) is considered to contain a specific sequence containing thymine (T) (e.g., the sequences in Table 12 or portions thereof), it should be understood that the RNA sequence may (and indeed often) contain uracil (U) in place of T. For example, an RNA sequence may contain U at each position shown as T in the sequences in Table 12. More specifically, this disclosure provides RNA sequences according to each gRNA scaffold sequence in Table 12, wherein the RNA sequence has U in place of each T in the sequences in Table 12. Furthermore, it should be understood that terminal U and T may optionally be added or removed from the tracrRNA sequence, and may be modified or unmodified when provided as RNA. Alternative versions of the gRNA scaffold sequences illustrated in Table 12 may also function with the different Cas9 enzymes or derivatives thereof illustrated in Table 8, for example, alternative gRNA scaffold sequences with nucleotide additions, substitutions, or deletions, such as sequences with added or removed stem-loop structures. This paper anticipates that gRNA scaffold sequences represent components of gene modification systems and can be similarly optimized for a given system, Cas-RT fusion peptide, indication, target mutation, template RNA, or delivery medium.

[0422] Heterogeneous object sequence

[0423] The template RNA described herein may contain a heterologous object sequence, which a genetically modified polypeptide may use as a template for reverse transcription to write the desired sequence into the target nucleic acid. In some embodiments, the heterologous object sequence comprises a post-edited homologous region, a mutant region, and a pre-edited homologous region from 5' to 3'. Not wishing to be bound by theory, the RT that reverse transcribes the template RNA first reverse transcribes the pre-edited homologous region, then the mutant region, and then the post-edited homologous region, thereby producing a DNA strand containing the desired mutation, flanked by homologous regions.

[0424] In some embodiments, the length of the heterogeneous object sequence is at least 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80. 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500 or 1,000 nucleotides (nt), or a length of at least 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases. In some embodiments, the length of the heterogeneous object sequence is no more than 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 7 6, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 120, 140, 160, 180, 200, 500, 1,000 or 2,000 nucleotides (nt), or a length not exceeding 20, 15, 10, 9, 8, 7, 6, 5, 4 or 3 kilobases.In some embodiments, the length of the heterogeneous object sequence is 30-1000, 40-1000, 50-1000, 60-1000, 70-1000, 74-1000, 75-1000, 76-1000, 77-1000, 78-1000, 79-1000, 80-1000, 85-1000, 90-1000, 100-1000, 120-1000, 140-1000, 160-1000, 180-1000, 200-1000, 500-1000, 30-500, 40-500, 50-500, 60-500. 70-500, 74-500, 75-500, 76-500, 77-500, 78-500, 79-500, 80-500, 85-500, 90-500, 100-500, 120-500, 140-500, 160-500, 180-500, 200-500, 30-200, 40-200, 50-200, 60-200, 70-200, 74-200, 75-200, 76-200, 77-200, 78-200, 79-200, 80-200, 85-200, 90-200, 100-20 0, 120-200, 140-200, 160-200, 180-200, 30-100, 40-100, 50-100, 60-100, 70-100, 74-100, 75-100, 76-100, 77-100, 78-100, 79-100, 80-100, 85-100, or 90-100 nucleotides (nt), or lengths of 1-20, 1-15, 1-10, 1-9, 1-8, 1-7, 1-6, 1-5, 1-4, 1-3, 1-2, 2-20, 2-15, 2-10, 2-9, 2-8, 2-7, 2-6, 2-5, 2-4, 2-3, 3-20, 3-15, 3-10, 3-9, 3-8, 3-7, 3-6, 3-5, 3-4, 4-20, 4-15, 4-10, 4-9, 4-8, 4-7, 4-6, 4-5, 5-20, 5-15, 5-10, 5-9, 5-8, 5-7, 5-6, 6-20, 6-15, 6-10, 6-9, 6-8, 6-7, 7-20, 7-15, 7-10, 7-9, 7-8, 8-20, 8-15, 8-10, 8-9, 9-20, 9-15, 9-10, 10-15, 10-20 or 15-20 kJ bases.In some embodiments, the length of the heteroobject sequence is 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10-40, 10-30, or 10-20 nt, for example, 10-80, 10-50, or 10-20 nt, such as about 10-20 nt. In some embodiments, the length of the heteroobject sequence is 8-30, 9-25, 10-20, 11-16, or 12-15 nucleotides, for example, 11-16 nt. Not wishing to be bound by theory, in some embodiments, a larger insert size, a larger edited region (e.g., the distance between the first and second edits / substitutions in the target region), and / or more desired edits (e.g., mismatches between the heteroobject sequence and the target genome) may produce a longer optimal heteroobject sequence.

[0425] In some embodiments, the template nucleic acid comprises a custom RNA sequence template that can be identified, designed, engineered, and constructed to contain sequences that alter or specify host genome function, such as by introducing a heterologous coding region into the genome; affecting or inducing exon structure / alternative splicing, such as causing exon skipping of one or more exons; inducing endogenous gene disruption, such as causing gene knockout; inducing transcriptional activation of endogenous genes; inducing epigenetic regulation of endogenous DNA; inducing upregulation of one or more operatively linked genes, such as causing gene activation or overexpression; inducing upregulation of one or more operatively linked genes, such as causing gene knockout; and so on. In some embodiments, the custom RNA sequence template can be engineered to contain sequences encoding exons and / or transgenes, providing binding sites for transcription factor activators, repressors, enhancers, and combinations thereof. In some embodiments, the custom template can be engineered to encode nucleic acid or peptide tags for expression in endogenous RNA transcripts or endogenous proteins operatively linked to target sites. In other embodiments, the coding sequence can be further customized with splice donor sites, splice acceptor sites, or poly-A tails.

[0426] The template nucleic acid (e.g., template RNA) of the system typically contains an object sequence (e.g., a heterologous object sequence) for writing a desired sequence into target DNA. The object sequence can be coding or non-coding. The template nucleic acid (e.g., template RNA) can be programmed to produce insertions, mutations, or deletions at target DNA loci. In some embodiments, the template nucleic acid (e.g., template RNA) can be programmed to result in insertion into the target DNA. For example, the template nucleic acid (e.g., template RNA) can contain a heterologous sequence, wherein reverse transcription will result in the insertion of the heterologous sequence into the target DNA. In other embodiments, the RNA template can be programmed to introduce deletions into the target DNA. For example, the template nucleic acid (e.g., template RNA) can match the target DNA upstream and downstream of the desired deletion, wherein reverse transcription will result in replication from the upstream and downstream sequences of the template nucleic acid (e.g., template RNA) without intercalation sequences, such as resulting in the deletion of intercalation sequences. In other embodiments, the template nucleic acid (e.g., template RNA) can be programmed to introduce edits into the target DNA. For example, the template RNA can match the target DNA sequence except for one or more nucleotides, wherein reverse transcription will result in the replication of these edits into the target DNA, such as resulting in mutations, such as translocation or transversion mutations.

[0427] In some embodiments, writing the object sequence to the target site results in nucleotide substitution, for example, where the full length of the object sequence corresponds to the matching length of the target site having one or more mismatched bases. In some embodiments, the heterologous object sequence may be designed such that combinations of sequence changes can occur, such as simultaneous addition and deletion, addition and substitution, or deletion and substitution.

[0428] In some embodiments, the heterologous object sequence may comprise an open reading frame or a fragment of an open reading frame. In some embodiments, the heterologous object sequence has a Kozak sequence. In some embodiments, the heterologous object sequence has an internal ribosome entry site. In some embodiments, the heterologous object sequence has a self-cleaving peptide, such as a T2A or P2A site. In some embodiments, the heterologous object sequence has a start codon. In some embodiments, the template RNA has a splice acceptor site. In some embodiments, the template RNA has a splice donor site. Exemplary splice acceptor and splice donor sites are described in WO 2016044416, which is incorporated herein by reference in its entirety. Exemplary splice acceptor site sequences are known to those skilled in the art. In some embodiments, the template RNA has a microRNA binding site downstream of a stop codon. In some embodiments, the template RNA has a poly-A tail downstream of a stop codon in an open reading frame. In some embodiments, the template RNA comprises one or more exons. In some embodiments, the template RNA comprises one or more introns. In some embodiments, the template RNA comprises a eukaryotic transcription terminator. In some embodiments, the template RNA comprises an enhanced translation element or a translation-enhancing element. In some embodiments, the RNA comprises a human T-cell leukemia virus (HTLV-1) R region. In some embodiments, the RNA contains post-transcriptional regulatory elements that enhance nuclear output, such as post-transcriptional regulatory elements of hepatitis B virus (HPRE) or marmot hepatitis virus (WPRE).

[0429] In some embodiments, the heterologous object sequence may contain non-coding sequences. For example, the template nucleic acid (e.g., template RNA) may contain regulatory elements, such as promoter or enhancer sequences or miRNA binding sites. In some embodiments, integration of the object sequence at the target site results in upregulation of endogenous genes. In some embodiments, integration of the object sequence at the target site results in downregulation of endogenous genes. In some embodiments, the template nucleic acid (e.g., template RNA) contains tissue-specific promoters or enhancers, each of which may be unidirectional or bidirectional. In some embodiments, the promoter is an RNA polymerase I promoter, an RNA polymerase II promoter, or an RNA polymerase III promoter. In some embodiments, the promoter contains a TATA element. In some embodiments, the promoter contains a B recognition element. In some embodiments, the promoter has one or more binding sites for transcription factors.

[0430] In some embodiments, the template nucleic acid (e.g., template RNA) contains sites that coordinate epigenetic modifications. In some embodiments, the template nucleic acid (e.g., template RNA) contains chromatin insulators. For example, the template nucleic acid (e.g., template RNA) contains CTCF sites or sites targeting DNA methylation.

[0431] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a gene expression unit consisting of at least one regulatory region operatively linked to an effector sequence. The effector sequence may be a sequence transcribed into RNA (e.g., a coding sequence or a non-coding sequence, such as a sequence encoding microRNA).

[0432] In some embodiments, a heterologous object sequence of a template nucleic acid (e.g., template RNA) is inserted into an endogenous intron of the target genome. In some embodiments, the heterologous object sequence of a template nucleic acid (e.g., template RNA) is inserted into the target genome to serve as a novel exon. In some embodiments, the insertion of the heterologous object sequence into the target genome results in the substitution or skipping of a natural exon.

[0433] A template nucleic acid (e.g., template RNA) can be programmed to produce insertions, mutations, or deletions at target DNA loci. In some embodiments, the template nucleic acid (e.g., template RNA) can be programmed to result in insertion into the target DNA. For example, the template nucleic acid (e.g., template RNA) may contain a heterologous object sequence, wherein reverse transcription will result in the insertion of the heterologous object sequence into the target DNA. In other embodiments, the RNA template can be programmed to write deletions into the target DNA. For example, the template nucleic acid (e.g., template RNA) can match the target DNA upstream and downstream of the desired deletion, wherein reverse transcription will result in replication from the upstream and downstream sequences of the template nucleic acid (e.g., template RNA) without intercalation sequences, such as resulting in the deletion of intercalation sequences. In other embodiments, the template nucleic acid (e.g., template RNA) can be programmed to write edits into the target DNA. For example, the template RNA can match the target DNA sequence except for one or more nucleotides, wherein reverse transcription will result in the replication of these edits into the target DNA, such as resulting in mutations, such as translocation or inversion mutations.

[0434] In some embodiments, the pre-edit homology domain comprises a nucleic acid sequence that has at least 100% sequence identity with the nucleic acid sequence contained in the target nucleic acid molecule.

[0435] In some embodiments, the edited homology domain comprises a nucleic acid sequence that has at least 100% sequence identity with the nucleic acid sequence contained in the target nucleic acid molecule.

[0436] PBS sequence

[0437] In some embodiments, the template nucleic acid (e.g., template RNA) comprises a PBS sequence. In some embodiments, the PBS sequence is located at the 3′ of the heterologous target sequence and is complementary to a sequence adjacent to a site to be modified by the system described herein, or to a sequence adjacent to a site to be modified by the system / gene-modifying polypeptide, and the sequence contains no more than 1, 2, 3, 4, or 5 mismatches. In some embodiments, the PBS sequence binds within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides at a cleavage site in the target nucleic acid molecule. In some embodiments, the binding of the PBS sequence to the target nucleic acid molecule allows the initiation of target-induced reverse transcription (TPRT), for example, the 3′ homologous domain acts as a primer for TPRT. In some embodiments, the length of the PBS sequence is 3-5, 5-10, 10-30, 10-25, 10-20, 10-19, 10-18, 10-17, 10-16, 10-15, 10-14, 10-13, 10-12, 10-11, 11-30, 11-25, 11-20, 11-19, 11-18, 11- 17, 11-16, 11-15, 11-14, 11-13, 11-12, 12-30, 12-25, 12-20, 12-19, 12-18, 12-17, 12-16, 12-15, 12-14, 12-13, 13-30, 13-25, 13-20, 13-19, 13-18, 13-17, 13 -16, 13-15, 13-14, 14-30, 14-25, 14-20, 14-19, 14-18, 14-17, 14-16, 14-15, 15-30, 15-25, 15-20, 15-19, 15-18, 15-17, 15-16, 16-30, 16-25, 16-20, 16-19, 1 The length of the PBS sequence is 6-18, 16-17, 17-30, 17-25, 17-20, 17-19, 17-18, 18-30, 18-25, 18-20, 18-19, 19-30, 19-25, 19-20, 20-30, 20-25, or 25-30 nucleotides, for example, 10-17, 12-16, or 12-14 nucleotides. In some embodiments, the length of the PBS sequence is 5-20, 8-16, 8-14, 8-13, 9-13, 9-12, or 10-12 nucleotides, for example, 9-12 nucleotides.

[0438] The template nucleic acid (e.g., template RNA) may have some homology with the target DNA. In some embodiments, the PBS sequence domain of the template nucleic acid (e.g., template RNA) can be used as an annealing region of the target DNA, such that the target DNA is localized to initiate reverse transcription of the template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid (e.g., template RNA) has at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases at the 3′ end of the RNA that are completely homologous to the target DNA. In some embodiments, the template nucleic acid (e.g., template RNA) has, for example, at the 5′ end of the template nucleic acid (e.g., template RNA) at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 175, 200 or more bases that are at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 97%, 98%, 99% or 100% homologous to the target DNA.

[0439] Exemplary template sequence

[0440] In some embodiments of the systems and methods described herein, the template RNA comprises a gRNA spacer containing nucleotides of the gRNA spacer sequence from Table 1A. In some embodiments, the heterologous object sequence comprises nucleotides of the RT template sequence from Table 1A corresponding to the gRNA spacer sequence. In the context of the sequence listing, when two components are in the same row of the reference table, the first component “corresponds” to the second component. In some embodiments, the primer binding site (PBS) sequence has a sequence comprising nucleotides from the PBS sequence in the same row as the RT template sequence in Table 1A.

[0441] Table 1A: Exemplary template RNAs (tgRNAs) for correcting the pathogenic R408W mutation

[0442] Table 1A provides an exemplary component design for a gene modification system used to correct the pathogenic R408W mutation in the PAH gene to its wild-type form. The table details the sequence of a complete template RNA (tgRNA) comprising (1) a gRNA spacer (e.g., for targeting a first-strand nick), (2) a gRNA scaffold, (3) a RT (heterologous target sequence) sequence, and (4) a PBS sequence (e.g., for initiating TPRT at the first-strand nick). The gRNA spacer, gRNA scaffold, heterologous target sequence / RT template sequence, and PBS sequence are designed to pair with the gene modification system to generate a nick at the appropriate location, thereby enabling the installation of the desired genome edit. The template in this table uses a scaffold of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACU UGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21).

[0443] RNACS Name Full sequence SEQ ID NO Spacer SEQ ID NO RNACS4135 hPKU5_R10_P9 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCAAUACCUCGGCCCUUCUCA 1 [[ID= 19 ​ ​ ​ 2 ​ 19 ​ ​ ​ 3 ​ 19 ​ ​ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUGGCACCGAGUCGGUGCACAAUACCUCGGGCCCUUCUCAG 4 UAGCGAACUGAGAAGGGCCA 19 RNACS4142 hPKU5_R12_P11 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUGGCACCGAGUCGGUGCACAAUACCUCGGCCCUUCUCAGU 5 UAGCGAACUGAGAAGGGCCA 19 RNACS4173 hPKU5_R18_P8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUGGCACCGAGUCGGUGCCGCUGCCACAAUACCUCGGCCCUUCUC 6 UAGCGAACUGAGAAGGGCCA 19 RNACS3643 hPKU6_R11_P11 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCCGAGUCGGUGCAAGGGCCGAGGUAUUGUGGCAG 7 ACUUUGCUGCCACAAUACCU 20 RNACS1792 hPKU6_R9_P11 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGGGCCGAGGUAUUGUGGCAG 8 ACUUUGCUGCCACAAUACCU 20 RNACS4933 hPKU5_R18_P10_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGCACCGAGUCGGUGCCUGCCACAAUACCUCGCCCCUUCAC 9 UAGCGAACUGAGAAGGGCCA 19 RNACS4939 hPKU5_R16_P10_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUGCCACAAUACCUCGCCCCUUCUCAG 10 UAGCGAACUGAGAAGGGCCA 19 RNACS4945 hPKU5_R14_P10_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCCCACAAUACCUCGCCCCUUCUCAG 11 UAGCGAACUGAGAAGGGCCA 19 RNACS4951 hPKU5_R12_P10_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCACAAUACCUCGCCCCUUCUCAG 12 UAGCGAACUGAGAAGGGCCA 19 RNACS4969 hPKU5_R16_P9_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCUGCCACAAUACCUCGCCCCUUCUCA 13 UAGCGAACUGAGAAGGGCCA 19 RNACS4975 hPKU5_R14_P9_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCCCACAAUACCUCGCCCCUUCUCA 14 UAGCGAACUGAGAAGGGCCA 19 RNACS5534 hPKU5_R10_P10_sub4 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCAAUACCUCGCCCCUUCUCAG 15 UAGCGAACUGAGAAGGGCCA 19 RNACS5096 hPKU6_R14_P9_sub7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGAGAAGGGACGCGGUAUUGUGGC 16 ACUUUGCUGCCACAAUACCU 20 RNACS5533 hPKU6_R9_P11_sub7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGGGACGCGGUAUUGUGGCAG 17 ACUUUGCUGCCACAAUACCU 20 RNACS5094 hPKU6_R14_P9_sub5 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGAGAACGGACGGGGUAUUGUGGC 18 ACUUUGCUGCCACAAUACCU 20

[0444] In this document, when an RNA sequence (e.g., a template RNA sequence) is considered to contain a specific sequence containing uracil (U) (e.g., the sequence in Table 1A or a portion thereof), it should be understood that the RNA sequence may (and indeed often) contain thymine (T) in place of U. For example, an RNA sequence may contain T at each position shown as U in the sequences in Table 1A. More specifically, this disclosure provides RNA sequences according to each gRNA spacer subsequence shown in Table 1A, wherein the RNA sequence has T in place of each U in the sequences in Table 1A.

[0445] In some embodiments of the systems and methods described herein, the system further includes a second-strand targeting gRNA (ngRNA) that directs the nick to the second strand of the human PAH gene. In some embodiments, the second-strand targeting gRNA comprises a spacer sequence from the ngRNAs in Table 2A, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with them.

[0446] Table 2A: Exemplary second-cut gRNAs (ngRNAs)

[0447] Table 2A provides a list of exemplary second nick gRNAs (ngRNAs) for optional use in correcting pathogenic R408W mutations in PAHs. These ngRNAs can be used in combination with the tgRNAs listed in Table 1A. The templates in this table use the scaffold of GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGC (SEQ ID NO: 21).

[0448] gRNA name full sequence SEQ ID NO spacer SEQ ID NO RNACS1809 hPKU_ngRNA_79+ GUGCCCUUCACUCAAGCCUGGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUUUU 22 GUGCCCUUCACUCAAGCCUG 25 RNACS1810 hPKU_ngRNA_85+ UUCACUCAAGCCUGUGGUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCCGGUGCUUUU 23 UUCACUCAAGCCUGUGGUUU 26 RNACS1812 hPKU_ngRNA_170- GUCCAAGACCUCAAUCCUUUGUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAGUGGCACCGAGUCCGGUGCUUUU 24 GUCCAAGACCUCAAUCCUUU 27

[0449] In some embodiments, the systems and methods provided herein may include template sequences listed in Table 1A and optional second nick gRNA sequences listed in Table 2A, which are designed to pair with gene-modifying peptides to correct mutations in the PAH gene. The template RNA sequences shown in Tables 1A and 2A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3 may be customized depending on the targeted cell. For example, in some embodiments, it is desirable to inactivate the PAM sequence during editing (e.g., using a “PAM-kill” modification) to reduce the likelihood of further gene editing after the initial editing (e.g., retargeting via Cas). Therefore, certain template RNAs described herein are designed to write mutations (e.g., substitutions) into the PAM at the target site such that, after editing, the PAM site will mutate into a sequence that the gene-modifying peptide no longer recognizes. Thus, the mutated region within the heterologous target sequence of the template RNA may contain a PAM-kill sequence. To avoid being bound by theory, in some embodiments, the PAM-killing sequence may prevent the re-entry of the genetically modified peptide after gene modification is complete, or reduce re-entry relative to template RNA lacking the PAM-killing sequence. In some embodiments, the PAM-killing sequence does not alter the amino acid sequence encoded by the gene; for example, the PAM-killing sequence produces a silent mutation. In other embodiments, it is desirable to maintain the integrity of the PAM sequence (without PAM-killing).

[0450] Similarly, in some embodiments, to reduce the likelihood of further gene editing after initial editing (e.g., retargeting via Cas), it may be desirable to alter the first three nucleotides of the RT template sequence via a "seed-killing" motif. Therefore, some template RNAs described herein are designed to incorporate mutations (e.g., substitutions) into the target site portion corresponding to the first three nucleotides of the RT template sequence, such that after editing, the target site will mutate into a sequence with low homology to the RT template sequence. Thus, the mutated region within the heterologous target sequence of the template RNA may contain a seed-killing sequence. Not wishing to be bound by theory, in some embodiments, the seed-killing sequence may prevent the re-entry of the gene-modified polypeptide after gene modification is complete, or reduce re-entry relative to template RNAs that are otherwise similar but lack a seed-killing sequence. In some embodiments, the seed-killing sequence does not alter the amino acid sequence encoded by the gene; for example, the seed-killing sequence produces a silent mutation. In other embodiments, it is desirable to maintain the integrity of the seed region and not use a seed-killing sequence.

[0451] In other embodiments, to optimize or improve gene editing efficiency, it may be necessary to evade the mismatch repair or nucleotide repair pathways of the target cells, or to bias the target cells' repair pathways towards preserving the edited strand. In some embodiments, multiple silent mutations (e.g., silent substitutions) may be introduced within the RT template sequence to evade the target cells' mismatch repair or nucleotide repair pathways, or to bias the target cells' repair pathways towards preserving the edited strand.

[0452] Table 7A provides exemplary silencing mutations at different locations within the PAH gene that can be used with a template to correct the R408W mutation.

[0453] Table 7A. Exemplary silent mutation codons of the PAH gene used as templates for correcting the R408W mutation.

[0454]

[0455]

[0456] In some embodiments, the template RNA contains one or more silent mutations.

[0457] It should be understood that the silent mutations shown in Table 7A can be used alone or in any combination in the template RNA sequences described herein.

[0458] gRNA with inducible activity

[0459] In some embodiments, the gRNA described herein (e.g., gRNA as part of a template RNA or gRNA for second-strand cleavage) has inducible activity. Inducible activity can be achieved via a template nucleic acid, such as template RNA, which (in addition to the gRNA) also contains a blocking domain, wherein the sequence of some or all of the blocking domain is at least partially complementary to a portion or all of the gRNA. Thus, the blocking domain is capable of hybridizing or substantially hybridizing with a portion or all of the gRNA. In some embodiments, the blocking domain and the inducible gRNA are arranged on a template nucleic acid, such as template RNA, such that the gRNA can be in a first conformation (where the blocking domain hybridizes or substantially hybridizes with the gRNA) and a second conformation (where the blocking domain does not hybridize or substantially does not hybridize with the gRNA). In some embodiments, in the first conformation, the gRNA cannot bind gene-modified peptides (e.g., template nucleic acid binding domains, DNA binding domains, or endonuclease domains (e.g., CRISPR / Cas proteins)) or binds in a manner with significantly reduced affinity compared to template RNA that lacks other aspects similar to the blocking domain. In some embodiments, in the second conformation, the gRNA can bind to the genetically modified polypeptide (e.g., a template nucleic acid binding domain, a DNA binding domain, or a nuclease domain (e.g., a CRISPR / Cas protein)). In some embodiments, whether the gRNA is in the first or second conformation can affect the DNA binding or nuclease activity of the genetically modified polypeptide (e.g., a CRISPR / Cas protein contained in the genetically modified polypeptide).

[0460] In some embodiments, the coordinating gRNA for the second nick has inducible activity. In some embodiments, the coordinating gRNA for the second nick is induced after the template is reverse transcribed. In some embodiments, hybridization of the gRNA with the blocking domain can be disrupted using an open molecule. In some embodiments, the open molecule comprises an agent that binds to part or all of the gRNA or the blocking domain and inhibits hybridization of the gRNA with the blocking domain. In some embodiments, the open molecule comprises a nucleic acid, for example, a sequence that is partially or completely complementary to the gRNA, the blocking domain, or both. By selecting or designing a suitable open molecule, the provided open molecule can facilitate a conformational change in the gRNA, enabling it to be associated with a CRISPR / Cas protein and provide the relevant functions of the CRISPR / Cas protein (e.g., DNA binding and / or endonuclease activity). Providing the open molecule at a selected time and / or location, without being bound by theory, can allow for spatial and temporal control over the activity of the gRNA, the CRISPR / Cas protein, or the gene modification system containing them. In some embodiments, the open molecule comprises an exogenous substance to the cell containing the gene-modified peptide and / or template nucleic acid. In some embodiments, the open molecule comprises an endogenous agent (e.g., endogenous to the cell containing the genetically modified polypeptide and / or template nucleic acid (which contains gRNA and a blocking domain)). For example, the inducible gRNA, blocking domain, and open molecule may be selected such that the open molecule is an endogenous agent expressed in the target cell or tissue, thereby ensuring the activity of the genetic modification system in the target cell or tissue. As a further example, the inducible gRNA, blocking domain, and open molecule may be selected such that the open molecule is absent or substantially not expressed in one or more non-target cells or tissues, thereby ensuring that the activity of the genetic modification system does not occur or substantially does not occur in one or more non-target cells or tissues, or occurs at a reduced level compared to the target cells or tissues. Exemplary blocking domains, open molecules, and their uses are described in PCT application publication WO 2020044039 A1, which is incorporated herein by reference in its entirety. In some embodiments, the template nucleic acid (e.g., template RNA) may comprise one or more sequences or structures for binding by one or more components of the genetically modified polypeptide (e.g., reverse transcriptase or RNA-binding domain and gRNA). In some embodiments, gRNA promotes interaction with template nucleic acid binding domains (e.g., RNA binding domains) of the genetically modified peptide. In some embodiments, gRNA guides the genetically modified peptide to a matching target sequence, such as in the genome of a target cell.

[0461] Target nucleic acid sites

[0462] In some embodiments, after gene modification, the target site surrounding the edited sequence contains a limited number of insertions or deletions, for example, in less than about 50% or 10% of the edit events, as determined by long-read amplicon sequencing of the target site, as described in, for example, Karst et al. (2020) bioRxiv doi.org / 10.1101 / 645903 (incorporated herein by reference in its entirety). In some embodiments, the target site does not exhibit multiple consecutive edit events, such as head-to-tail or head-to-head repetitions, as determined by long-read amplicon sequencing of the target site, as described in, for example, Karst et al. bioRxiv doi.org / 10.1101 / 645903 (2020) (incorporated herein by reference in its entirety). In some embodiments, the target site contains an integration sequence corresponding to the template RNA. In some embodiments, the target site does not contain an insertion generated by endogenous RNA in more than about 1% or 10% of events, for example, as determined by long-read amplicon sequencing of the target site, as described, for example, in Karst et al. bioRxiv doi.org / 10.1101 / 645903 (2020) (which is incorporated herein by reference in its entirety). In some embodiments, the target site contains an integration sequence corresponding to the template RNA.

[0463] In some aspects of the invention, the host DNA binding site integrated by the gene modification system can be in a gene, in an intron, in an exon, in an ORF, outside the coding region of any gene, within the regulatory region of a gene, or outside the regulatory region of a gene. In other aspects, the polypeptide can bind to one or more host DNA sequences.

[0464] In some embodiments, the gene modification system is used to edit target loci in multiple alleles. In some embodiments, the gene modification system is designed to edit a specific allele. For example, the gene-modifying peptide may target a specific sequence present only in one allele, such as a template RNA containing homology to a target allele (e.g., gRNA or annealing domain), but not a second homologous allele. In some embodiments, the gene modification system may alter haplotype-specific alleles. In some embodiments, a gene modification system targeting a specific allele preferentially targets that allele, for example, having a preference of at least 2, 4, 6, 8, or 10-fold for the target allele.

[0465] Second chain cut

[0466] In some embodiments, the gene-modifying system described herein includes cleavage enzyme activity for cleaving a first strand (e.g., in the gene-modifying peptide) and cleavage enzyme activity for cleaving a second strand of the target DNA (e.g., in a peptide separate from the gene-modifying peptide). As discussed herein, it is not intended to be theoretically constrained, but cleavage of the first strand of the target site DNA is thought to provide a 3' OH that can be used by the RT domain to reverse transcribe template RNA (e.g., a heterologous object sequence). It is not intended to be theoretically constrained, but it is thought that introducing additional cleavage into the second strand may bias cellular DNA repair mechanisms toward using heterologous object sequence-based sequences more frequently than the original genomic sequence. In some embodiments, the additional cleavage of the second strand is generated by the same endonuclease domain (e.g., a cleavage enzyme domain) as the cleavage of the first strand. In some embodiments, the same gene-modifying peptide cleaves both the first and second strands. In some embodiments, the gene-modifying peptide contains a CRISPR / Cas domain and the additional cleavage of the second strand is guided by another nucleic acid (e.g., containing a second gRNA that guides the CRISPR / Cas domain to cleave the second strand). In other embodiments, an additional second-strand nick is generated by a nuclease domain (e.g., a nicking enzyme domain) that is different from the nick of the first strand. In some embodiments, this different nuclease domain is located in a separate polypeptide (e.g., the system of the present invention further comprises a separate polypeptide), separate from the gene-modifying polypeptide. In some embodiments, the separate polypeptide comprises the nuclease domain (e.g., a nicking enzyme domain) described herein. In some embodiments, the separate polypeptide comprises, for example, a DNA-binding domain described herein.

[0467] In this paper, it is anticipated that the location of the second-strand nick relative to the first-strand nick can affect the degree to which one or more of the following occur: desired genetic modification, unwanted double-strand breaks (DSBs), unwanted insertions, or unwanted deletions. Without being bound by theory, the second-strand nick may occur in two general orientations: inward and outward.

[0468] In some embodiments, in an inward nick orientation, the RT domain is polymerized (e.g., using template RNA (e.g., a heterologous object sequence)) away from the second-strand nick. In some embodiments, in an inward nick orientation, the location of the first-strand nick and the location of the second-strand nick are located between the first PAM site and the second PAM site (e.g., where both nicks are generated by a polypeptide containing a CRISPR / Cas domain (e.g., a genetically modified polypeptide)). This inward nick orientation can also be referred to as "PAM-outside" when there are two PAMs on the outside and two nicks on the inside. In some embodiments, in an inward nick orientation, the location of the first-strand nick and the location of the second-strand nick are between the sites where the polypeptide and another polypeptide bind to the target DNA. In some embodiments, in an inward nick orientation, the location of the second-strand nick is between the binding sites of the polypeptide and another polypeptide, and the first-strand nick is also between the binding sites of the polypeptide and another polypeptide. In some embodiments, in an inward nick orientation, the location of the first-strand nick and the location of the second-strand nick are between the PAM site and the binding site of a second polypeptide at a distance from the target site.

[0469] Examples of gene modification systems with inward cleavage orientation include: a gene-modifying polypeptide containing a CRISPR / Cas domain, a template RNA containing gRNA guiding cleavage of the target DNA on the first strand, and additional nucleic acid (containing additional gRNA guiding cleavage at a site a certain distance from the first cleavage site), wherein the location of the first cleavage and the location of the second cleavage are between the PAM sites of the sites guided by the two gRNAs to which the gene-modifying polypeptide is located. As a further example, another gene modification system with inward cleavage orientation includes a gene-modifying polypeptide containing a zinc finger molecule and a first cleavage enzyme domain, wherein the zinc finger molecule binds to the target DNA in a manner that guides the first cleavage enzyme domain to cleave the first strand of the target site; an additional polypeptide containing a CRISPR / Cas domain; and additional nucleic acid containing gRNA guiding the additional polypeptide to cleave the target DNA at a site a certain distance from the target DNA on the second strand, wherein the location of the first cleavage and the location of the second cleavage are between the PAM site and the zinc finger binding site. As a further example, another gene modification system with inward nick orientation is provided, comprising a gene-modifying polypeptide containing a zinc finger molecule and a first nicking enzyme domain, wherein the zinc finger molecule binds to the target DNA in a manner that guides the first nicking enzyme domain to nick the first strand of the target site; and an additional polypeptide containing a TAL effector molecule and a second nicking enzyme domain, wherein the TAL effector molecule binds to a site at a distance from the target site in a manner that guides the additional polypeptide to nick the second strand, wherein the positions of the first nick and the second nick are between the binding site of the TAL effector molecule and the binding site of the zinc finger molecule.

[0470] In some embodiments, in an outward nick orientation, the RT domain is polymerized (e.g., using template RNA (e.g., a heterologous object sequence)) toward the second-strand nick. In some embodiments, in an outward nick orientation, when both the first and second nicks are generated by a polypeptide containing a CRISPR / Cas domain (e.g., a genetically modified polypeptide), the first PAM site and the second PAM site are located between the locations of the first-strand nick and the second-strand nick. This outward nick orientation can also be referred to as "PAM-in" when there are two PAMs on the inside and two nicks on the outside. In some embodiments, in an outward nick orientation, the polypeptide (e.g., a genetically modified polypeptide) and another polypeptide bind to sites on the target DNA located between the locations of the first-strand nick and the second-strand nick. In some embodiments, in an outward nick orientation, the location of the second-strand nick is located opposite the location of the first-strand nick and the binding sites of the other polypeptide. In some embodiments, in an outward orientation, the PAM site and the binding site of the second polypeptide at a distance from the target site are located between the locations of the first-strand nick and the second-strand nick.

[0471] Examples of gene modification systems providing outward nick orientation include: a gene modification polypeptide containing a CRISPR / Cas domain, a template RNA containing gRNA guiding nicks to target DNA on the first strand, and additional nucleic acid (containing additional gRNA guiding nicks at a site a certain distance from the first nick site), wherein the location of the first nick and the location of the second nick are outside the PAM site at the site guided by the two gRNAs to the gene modification polypeptide (i.e., the PAM site is located between the location of the first nick and the location of the second nick). As a further example, another gene modification system with outward nicking orientation is provided, comprising a gene-modifying polypeptide containing a zinc finger molecule and a first nicking enzyme domain, wherein the zinc finger molecule binds to the target DNA in a manner that guides the first nicking enzyme domain to nick the first strand of the target site; an additional polypeptide containing a CRISPR / Cas domain; and an additional nucleic acid containing gRNA that guides the additional polypeptide to nick the target DNA at a site on the second strand at a distance from the target site, wherein the location of the first nick and the second nick are outside the PAM site and the site where the zinc finger molecule binds (i.e., the PAM site and the site where the zinc finger molecule binds are between the locations of the first nick and the second nick). As a further example, another gene modification system with outward nicking orientation is provided, comprising a gene-modifying polypeptide containing a zinc finger molecule and a first nicking enzyme domain, wherein the zinc finger molecule binds to the target DNA in a manner that guides the first nicking enzyme domain to nick the first strand of the target site; and an additional polypeptide containing a TAL effector molecule and a second nicking enzyme domain, wherein the TAL effector molecule binds to a site at a distance from the target site in a manner that guides the additional polypeptide to nick the second strand, wherein the positions of the first and second nicks are outside the binding sites of the TAL effector molecule and the zinc finger molecule (i.e., between the positions of the first and second nicks).

[0472] Without being bound by theory, it is considered that for gene modification systems providing a second-strand nick, an outward nick orientation is preferred in some embodiments. As described herein, an inward nick can generate a greater number of double-strand breaks (DSBs) compared to an outward nick orientation. DSBs can be recognized by DSB repair pathways in the cell nucleus, which can lead to undesirable insertions and deletions. An outward nick orientation provides a reduced risk of DSB formation and a correspondingly smaller number of undesirable insertions and deletions. In some embodiments, undesirable insertions and deletions are insertions and deletions not encoded by the foreign object sequence, such as insertions or deletions generated by double-strand break repair pathways unrelated to modifications encoded by the foreign object sequence. In some embodiments, the desired gene modification comprises alterations (e.g., substitutions, insertions, or deletions) to the target DNA encoded by the foreign object sequence (e.g., and achieved by writing the foreign object sequence into the target site via gene modification). In some embodiments, the first-strand nick and the second-strand nick are in an outward orientation.

[0473] Furthermore, the distance between the first and second strand cuts may affect the degree to which one or more of the following occur: the desired genetic modification system DNA modification is achieved, undesired double-strand breaks (DSBs) occur, undesired insertions occur, or undesired deletions occur. It is not desirable to be theoretically bound by the assumption that the benefit of the second strand cut, i.e., DNA repair biased towards incorporating heterologous object sequences into the target DNA, increases as the distance between the first and second strand cuts decreases. However, the risk of DSB formation is also considered to increase as the distance between the first and second strand cuts decreases. Accordingly, the number of undesired insertions and / or deletions is considered to increase as the ...

Claims

1. A nucleic acid molecule encoding a gene-modified polypeptide, wherein the nucleic acid comprises, for example, from 5' to 3': (a) The 5' UTR of SEQ ID NO: 41-44, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (b) The region of the N-terminal NLS encoding SEQ ID NO: 36, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (c) The region encoding the Cas domain of SEQ ID NO: 52, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (d) The region of the connector encoding SEQ ID NO: 54, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (e) The region encoding the reverse transcriptase (RT) domain of SEQ ID NO: 56, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (f) The region of the C-terminal NLS encoding SEQ ID NO: 38, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (g) 3' UTR of SEQ ID NO: 45-48, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (h) Optionally, the expression element of SEQ ID NO: 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (i) The poly(A) tail of SEQ ID NO: 49-51, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

2. A nucleic acid molecule encoding a gene-modified polypeptide, wherein the nucleic acid comprises, for example, from 5' to 3': (a) The 5' UTR of SEQ ID NO: 41-44, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (b) The region of the N-terminal NLS encoding SEQ ID NO: 37, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (c) The region encoding the Cas domain of SEQ ID NO: 53, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (d) The region of the connector encoding SEQ ID NO: 55, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (e) The region encoding the reverse transcriptase (RT) domain of SEQ ID NO: 57, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (f) The region of the C-terminal NLS encoding SEQ ID NO: 39, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (g) 3' UTR of SEQ ID NO: 45-48, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it; (h) Optionally, the expression element of SEQ ID NO: 40, 101, or 102, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (i) The poly(A) tail of SEQ ID NO: 49-51, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

3. The nucleic acid molecule of claim 1 or 2, wherein the Cas domain comprises a Cas domain that binds to the target DNA molecule and is heterologous to the RT domain.

4. The nucleic acid molecule according to any one of claims 1-3, wherein the nucleic acid molecule is mRNA.

5. A template RNA comprising: (1) gRNA spacers; (2) gRNA scaffold; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence; Furthermore, the template RNA contains the nucleotide sequence of the template RNA of SEQ ID NO: 15, 17, 92 or 94, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identity with it.

6. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37638, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

7. A template RNA comprising the nucleotide sequence of SEQ ID NO: 37653, or a sequence having at least 95%, 97%, 98% or 99% identity with it.

8. A gene modification system comprising: (a) Template RNA (tgRNA), which contains from 5' to 3': (1) gRNA spacers; (2) gRNA scaffold; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence; The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (b) A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising: (1) Cas structure domain; (2) Connector; and (3) Reverse transcriptase (RT) domain; The genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

9. A gene modification system comprising: (a) Template RNA (tgRNA), which contains from 5' to 3': (1) gRNA spacer, having the sequence of the gRNA spacer of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or having a sequence relative to it with no more than 1, 2 or 3 sequence changes (e.g., substitutions); (2) gRNA scaffold having the gRNA scaffold sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it; (3) A heterologous object sequence having a sequence of the template RNA having a heterologous object sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than one, two or three sequence changes relative to it; and (4) Primer binding site (PBS) sequence, having a PBS sequence of the template RNA of the form 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having no more than one, two, or three sequence changes relative to it; and (b) A genetically modified polypeptide or a nucleic acid encoding the genetically modified polypeptide, the genetically modified polypeptide comprising: (1) Cas structure domain; (2) Connector; and (3) Reverse transcriptase (RT) domain; The genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 90%, 95%, 96%, 97%, 98% or 99% identity with it.

10. The system of claim 8 or 9, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 95% identity with it.

11. The system of claim 8 or 9, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 97%, 98% or 99% identity with it.

12. The system of claim 8 or 9, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3.

13. The system of any one of claims 8-12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 95% identity with it.

14. The system of any one of claims 8-12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO: 28, or a sequence having at least 99% identity with it.

15. The system of any one of claims 8-12, wherein the genetically modified polypeptide comprises the amino acid sequence of the genetically modified polypeptide of SEQ ID NO:

28.

16. A gene modification system comprising: (a) Template RNA (tgRNA), which contains from 5' to 3': (1) gRNA spacers; (2) gRNA scaffold; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence; The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising: (1) Cas structure domain; (2) Connector; and (3) Reverse transcriptase (RT) domain; The nucleotide encoding the gene-modified polypeptide comprises a nucleic acid as described in any one of claims 1-4.

17. A gene modification system comprising: (a) Template RNA (tgRNA), which contains from 5' to 3': (1) gRNA spacers; (2) gRNA scaffold; (3) Heterogeneous object sequences; and (4) Primer binding site (PBS) sequence; The tgRNA contains the nucleotide sequence of the template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it; and (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising: (1) Cas structure domain; (2) Connector; and (3) Reverse transcriptase (RT) domain; The nucleotide encoding the genetically modified polypeptide comprises a nucleic acid sequence of a genetically modified polypeptide listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

18. A gene modification system comprising: (a) Template RNA (tgRNA), which contains from 5' to 3': (1) gRNA spacer, having the sequence of the gRNA spacer of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or having a sequence relative to it with no more than 1, 2 or 3 sequence changes (e.g., substitutions); (2) gRNA scaffold having the gRNA scaffold sequence of the template RNA of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it; (3) A heterologous object sequence having a sequence of the template RNA having a heterologous object sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having no more than one, two or three sequence changes relative to it; and (4) Primer binding site (PBS) sequence, having a PBS sequence of the template RNA of the form 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A, or X3, or a sequence having no more than one, two, or three sequence changes relative to it; and (b) Nucleic acid (e.g., mRNA) encoding a gene-modifying polypeptide, comprising: (1) Cas structure domain; (2) Connector; and (3) Reverse transcriptase (RT) domain; The nucleotide encoding the genetically modified polypeptide comprises a nucleic acid sequence of a genetically modified polypeptide listed in Table N2, E11, or E15, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

19. The system of any one of claims 16-18, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 95% identity with it.

20. The system of any one of claims 16-18, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3, or a sequence having at least 97%, 98% or 99% identity with it.

21. The system of any one of claims 16-18, wherein the template RNA comprises the nucleotide sequence of a template RNA sequence of Table 1A, E1, E1A, E3, E3A, E5, E5A, E7, E7A, E9, E9A, E13, E13A or X3.

22. The system of any one of claims 16-21, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15, or a sequence having at least 95% identity with it.

23. The system of any one of claims 16-21, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15, or a sequence having at least 99% identity with it.

24. The system of any one of claims 16-21, wherein the nucleotide encoding the genetically modified polypeptide comprises the nucleic acid sequence of the genetically modified polypeptide of Table N2, E11 or E15.

25. The system of any one of claims 16-24, further comprising a second nick RNA (ngRNA) that directs the second nick to the second strand of the human PAH gene.

26. The system of claim 25, wherein the ngRNA comprises the sequence of an ngRNA of Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

27. The system of claim 25, wherein the ngRNA comprises from 5' to 3': (1) A gRNA spacer having the sequence of an ngRNA spacer of any of the following table: E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14, or E14A, or a sequence having no more than one, two, or three sequence alterations (e.g., substitutions) relative to it; and (2) gRNA scaffold having the sequence of the gRNA scaffold of the ngRNA in Table 2A, E2, E2A, E4, E4A, E6, E6A, E8, E8A, E10, E10A, E14 or E14A, or a sequence having no more than 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 sequence changes relative to it.

28. The system of any one of claims 25-27, wherein the ngRNA and the template RNA of the gene modification system have a "PAM-in-the-inclination".

29. The system of any one of claims 8-28, wherein the nucleic acid encoding the gene-modified polypeptide comprises RNA, for example, mRNA.

30. The nucleic acid molecule of any one of claims 1-4 or the template RNA of claims 5-7, wherein the nucleic acid molecule is formulated in lipid nanoparticles (LNPs).

31. The system according to any one of claims 8-29, wherein the tgRNA, the nucleic acid molecule encoding the gene-modified polypeptide, and / or the ngRNA are formulated in an LNP.

32. A pharmaceutical composition comprising the system as described in any one of claims 8-29, or one or more nucleic acids encoding the system, and a pharmaceutically acceptable excipient or carrier.

33. The pharmaceutical composition of claim 32, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles (LNPs).

34. The pharmaceutical composition of claim 33, wherein the viral vector is adeno-associated virus.

35. A host cell (e.g., a mammalian cell, such as a human cell) comprising a nucleic acid molecule, a gene modification system, or a template RNA as described in any one of the preceding claims.

36. A method for preparing a nucleic acid molecule or template RNA as described in any of the preceding claims, the method comprising synthesizing the nucleic acid molecule or template RNA by means of: in vitro (e.g., by in vitro transcription or solid-state synthesis) or by introducing DNA encoding the template RNA into a host cell under conditions that allow the generation of the template RNA.

37. A method for modifying a target site in a human PAH gene in a cell, the method comprising contacting the cell with a gene modification system as described in any one of claims 8-29 or 31, or DNA encoding the gene modification system, or a pharmaceutical composition as described in any one of claims 32-34, thereby modifying the target site in the human PAH gene in the cell.

38. A method for treating a subject suffering from a disease or condition associated with a human PAH gene mutation, the method comprising administering to the subject a gene-modifying system as described in any one of claims 8-29 or 31, or DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any one of claims 32-34, thereby treating the subject suffering from a disease or condition associated with a human PAH gene mutation.

39. The method of claim 38, wherein the disease or condition is phenylketonuria (PKU) or hyperphenylalaninemia (e.g., mild or severe hyperphenylalaninemia).

40. The method of claim 38 or 39, wherein the subject has the R408W mutation.

41. A method for treating a subject with PKU, the method comprising administering to the subject a gene-modifying system as described in any one of claims 8-29 or 31, or DNA encoding the gene-modifying system, or a pharmaceutical composition as described in any one of claims 32-34, thereby treating the subject with PKU.

42. The gene modification system or method as described in any of the preceding claims, wherein introducing the system into target cells can correct pathogenic mutations in the PAH gene.

43. The gene modification system or method as claimed in any of the preceding claims, wherein the pathogenic mutation is an R408W mutation, and wherein the correction comprises an amino acid substitution of W408R.

44. The gene modification system or method as described in any of the preceding claims, wherein introducing the system into a target cell results in a mutation that prompts the restoration of the function of the PAH gene.

45. The gene modification system or method as claimed in any of the preceding claims, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target nucleic acid.

46. ​​The gene modification system or method as claimed in any of the preceding claims, wherein the correction of the mutation occurs in at least 10% (e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70% or more) of the target cells.

47. The gene modification system or method as claimed in any of the preceding claims, wherein the gene modification system comprises a second-strand targeting gRNA, and wherein the correction of mutations in the target cell population is increased relative to a target cell population treated with a gene modification system comprising template RNA but without a second-strand targeting gRNA.

48. The method as described in any of the preceding claims, wherein the cell is a mammalian cell, such as a human cell.

49. The method as described in any of the preceding claims, wherein the subject is a human being.

50. The method of any of the preceding claims, wherein the contact occurs in vitro, for example, wherein the DNA of the cell or the subject is modified in vitro.

51. The method of any of the preceding claims, wherein the contact occurs in vivo, for example, wherein the DNA of the cell or the subject is modified in vivo.

52. The method of any of the preceding claims, wherein contacting the cell or the subject with the system comprises contacting the cell or cells in the subject with nucleic acid (e.g., DNA or RNA) encoding the genetically modified polypeptide under conditions that allow for its production.

Citation Information

Patent Citations

  • Method for efficient exon (44) skipping in Duchenne Muscular Dystrophy and associated means

    EP2813570A1

  • Amino acid-, peptide- and polypeptide-lipids, isomers, compositions, and uses thereof

    US10086013B2

  • Lipids and lipid nanoparticle formulations for delivery of nucleic acids

    US10221127B2

  • Isolation of novel AAV's and uses thereof

    US10300146B2

  • Lipid-based formulations

    US20030077829A1