Untranslated regions and poly(A) tail sequences for use in methods and compositions for genome modulation

Artificial RNA molecules with specific UTR and poly(A) tail sequences enhance polypeptide expression, addressing low frequency and site-specificity issues in genome integration, thereby improving genome editing efficiency.

JP2026511445APending Publication Date: 2026-04-14TESSERA THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TESSERA THERAPEUTICS INC
Filing Date
2024-03-14
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Integration of a nucleic acid of interest into the genome occurs with low frequency and low site-specificity in the absence of special proteins that promote the insertion event, and is limited by the expression level of a gene-modified polypeptide.

Method used

Artificial RNA molecules comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, with specific 3'-untranslated region (3'UTR), 5'-untranslated region (5'UTR), and poly(A) tail sequences to enhance polypeptide expression, facilitating targeted genome modification.

Benefits of technology

Enhances the expression and site-specific integration of nucleic acids into the host genome, improving the efficiency and accuracy of genome editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026511445000025
    Figure 2026511445000025
  • Figure 2026511445000026
    Figure 2026511445000026
  • Figure 2026511445000027
    Figure 2026511445000027
Patent Text Reader

Abstract

(1) An artificial nucleic acid molecule comprising (a) at least one of a 3'-untranslated region (3'UTR) element and / or (b) a 5'-untranslated region (5'UTR) element, and (2) a poly(A) tail is described. The artificial nucleic acid molecule may further comprise a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain for genome modification. A system comprising the artificial nucleic acid molecule and a method of using the system are also described.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims the priority of U.S. Provisional Patent Application No. 63 / 490,413, filed on March 15, 2023, the disclosure of which is incorporated herein by reference in its entirety.

[0002] The present invention belongs to the field of genome editing. In particular, the present invention relates to untranslated regions (UTRs) and poly(A) tail sequences for use in genome modification.

[0003] Reference to the Sequence Listing Submitted by Electronic Means This application contains a sequence listing submitted by electronic means. The content of the electronic sequence listing (070992.3WO1 Sequence Listing.xml, size: 16,449,880 bytes, creation date February 23, 2024) is incorporated herein by reference in its entirety.

Background Art

[0004] Integration of a nucleic acid of interest into the genome occurs with low frequency and low site - specificity in the absence of special proteins that promote the insertion event. The integration of a nucleic acid of interest into the genome can be limited by the expression level of a gene - modified polypeptide that delivers the nucleic acid of interest. Thus, there is a need in the art for improved methods and constructs for enhancing the expression of gene - modified polypeptides capable of editing the host genome.

Summary of the Invention

Means for Solving the Problems

[0005] Artificial ribonucleic acid (RNA) molecules for enhancing the expression of polypeptides comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain are provided herein. The RNA molecule may comprise, for example, (1) a nucleotide sequence encoding a polypeptide; (2) at least one of a 3'-untranslated region (3'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281; and / or a 5'-untranslated region (5'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263; and / or (3) a poly(A) tail for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282-8314.

[0006] In a particular embodiment, the RNA molecule may include, for example, (1) a nucleotide sequence encoding a polypeptide, (2) (a) a 3'-untranslated region (3'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (b) a 5'-untranslated region (5'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263.

[0007] In a particular embodiment, the RNA molecule may include, for example, (1) a nucleotide sequence encoding a polypeptide, and (2) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314.

[0008] Furthermore, an artificial ribonucleic acid (RNA) molecule is provided, comprising (1) at least one of (a) a 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 to 8211, and / or (b) a 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212 to 8247, and / or (2) a poly(A) tail for improving polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314.

[0009] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises at least one of the following: (a) a 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269 and 8281; and / or (b) a 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212-8247 and 8259-8263.

[0010] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314.

[0011] Furthermore, a system for modifying DNA is provided. The system includes, for example, (a) an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) a nucleotide sequence encoding a polypeptide, (2) (i) a 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) a 5'-untranslated region (5'UTR) element for improving polypeptide expression, comprising SEQ ID NOs. 8212-8247 and An artificial RNA molecule comprising at least one 5'UTR element containing a nucleic acid sequence selected from the group consisting of 8259-8263, and / or (3) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282-8314; and (b) (e.g., 5' to 3') (i) optionally a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) a sequence that binds to a polypeptide, (iii) a heterologous object sequence, and (iv) optionally a template RNA (or DNA encoding the template RNA) containing a 3' target homologous domain. The heterologous object sequence may include, for example, modifications compared to the corresponding original sequence (e.g., the wild-type sequence), where the modifications improve the speed, fidelity, or speed and fidelity of reverse transcription primed by the target by reverse transcriptase.

[0012] Furthermore, a system for modifying DNA is provided. The system is, for example, (a) an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising (1) a nucleotide sequence encoding a polypeptide, (2) (i) a 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) a 5'-untranslated region (5'UTR) element for improving polypeptide expression, comprising SEQ ID NOs. 8212-8247 and An artificial nucleic acid molecule comprising (b) at least one of 5'UTR elements comprising a nucleic acid sequence selected from the group consisting of 8259 to 8263, and / or (3) a poly(A) tail for enhancing polypeptide expression, comprising a poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314; and (b) (e.g., 5' to 3') (i) optionally a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) a sequence that binds to a polypeptide, (iii) a heterogeneous object sequence, and (iv) optionally a template RNA (or DNA encoding the template RNA) comprising a 3' target homologous domain.Preferably, heterogeneous object sequences have one, two, or all of the following characteristics: i) they do not contain self-complementary sequences, e.g., self-complementary sequences that form a hairpin structure under stringent conditions, or if self-complementary sequences are present, they have one, two, or all of the following characteristics: (1) each self-complementary sequence is no longer than 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides; (2) the self-complementary sequences form a hairpin including an arm of no longer than 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides; or (3) the self-complementary sequences include at least 1, 2, 3, 4, or 5 positions that are incomplementary (e.g., mismatched or bulge) with their partner sequence; and (4) they do not contain repetitive sequences (e.g., single, di, or trinucleotide repetitive sequences); or if repetitive sequences are present, they are sequences of no longer than 12, 11, 10, 9, 8, 7, or 6 nucleotides.

[0013] In a particular embodiment, the 3'UTR element for improving polypeptide expression consists of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269, and 8281. In a particular embodiment, the 5'UTR element for improving polypeptide expression consists of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212-8247 and 8259-8263. In a particular embodiment, the poly(A) tail for improving polypeptide expression consists of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282-8314.

[0014] In a particular embodiment, the 3'UTR element of the RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281. In a particular embodiment, the 3'UTR element of the RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

[0015] In a particular embodiment, the 5'UTR element of the RNA molecule comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243.

[0016] In a particular embodiment, the poly(A)tail of the RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314. In a particular embodiment, the poly(A)tail of the RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0017] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8236; (2) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236; (3) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214; (4) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8243; (5) the 3'UTR element comprises SEQ ID NO: 8281 and the 5'UTR element comprises SEQ ID NO: 8259; or (6) the 3'UTR element comprises SEQ ID NO: 8281 and the 5'UTR element comprises SEQ ID NO: 8263. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8236; (2) the 3'UTR element is sequence number 8201 and the 5'UTR element is sequence number 8236; (3) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8214; (4) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8243; (5) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8259; or (6) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8263.

[0018] In a particular embodiment, the artificial RNA molecule includes a poly(A)tail, where the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0019] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8236, (2) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236, (3) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214, or (4) the 3'UTR element comprises (5) The 3'UTR element contains sequence number 8209, and the 5'UTR element contains sequence number 8243, or (6) the 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8259, or (7) the 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263; the poly(A)tail contains the nucleic acid sequence of sequence number 8294, sequence number 8299, sequence number 8306, sequence number 8307, sequence number 8309, sequence number 8311, sequence number 8312, or sequence number 8314. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8236, (2) the 3'UTR element is sequence number 8201 and the 5'UTR element is sequence number 8236, (3) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8214, or (4) the 3'UTR element is sequence number (5) The sequence number is 8209, and the 5'UTR element is sequence number 8243, or (6) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8259, or (7) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8263, and the poly(A)tail consists of the nucleic acid sequences of sequence number 8294, 8299, 8306, 8307, 8309, 8311, 8312, or 8314.

[0020] In certain embodiments, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (3) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail (4) The 3'UTR element includes sequence number 8209, the 5'UTR element includes sequence number 8243, and the poly(A) tail includes sequence number 8306 or sequence number 8307, (5) The 3'UTR element includes sequence number 8281, the 5'UTR element includes sequence number 8259, and the poly(A) tail includes sequence number 8306 or sequence number 8307, or (6) The 3'UTR element includes sequence number 8281, the 5'UTR element includes sequence number 8263, and the poly(A) tail includes sequence number 8306 or sequence number 8307. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (3) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; or (4) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8243, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0021] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail comprises SEQ ID NO: 8306. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element consists of SEQ ID NO: 8201, the 5'UTR element consists of SEQ ID NO: 8236, and the poly(A) tail consists of SEQ ID NO: 8307; (2) the 3'UTR element consists of SEQ ID NO: 8209, the 5'UTR element consists of SEQ ID NO: 8214, and the poly(A) tail consists of SEQ ID NO: 8306.

[0022] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises (1) at least one of (a) a 3'-untranslated region (3'UTR) element comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, and / or (b) a 5'-untranslated region (5'UTR) element comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212-8247, and / or (2) a poly(A) tail for enhancing polypeptide expression, comprising a poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282-8314.

[0023] In certain embodiments, the 3'UTR element of the artificial RNA molecule comprises the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209 or SEQ ID NO: 8211. In certain embodiments, the 3'UTR element of the artificial RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209 or SEQ ID NO: 8211.

[0024] In certain embodiments, the 5'UTR element of the artificial RNA molecule comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236 or SEQ ID NO: 8243. In certain embodiments, the 5'UTR element of the artificial RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236 or SEQ ID NO: 8243.

[0025] In certain embodiments, the poly(A) tail of the RNA molecule comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314. In certain embodiments, the poly(A) tail of the RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0026] In certain embodiments, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236, or (2) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214. In certain embodiments, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element consists of SEQ ID NO: 8201 and the 5'UTR element consists of SEQ ID NO: 8236, or (2) the 3'UTR element consists of SEQ ID NO: 8209 and the 5'UTR element consists of SEQ ID NO: 8214.

[0027] In certain embodiments, the artificial RNA molecule comprises a poly(A) tail, where the poly(A) tail comprises a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314. In certain embodiments, the artificial RNA molecule comprises a poly(A) tail, where the poly(A) tail consists of a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0028] In certain embodiments, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236, or (2) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214; the poly(A) tail comprises a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314. In certain embodiments, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element consists of SEQ ID NO: 8201 and the 5'UTR element consists of SEQ ID NO: 8236, or (2) the 3'UTR element consists of SEQ ID NO: 8209 and the 5'UTR element consists of SEQ ID NO: 8214; the poly(A) tail consists of a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0029] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail comprises SEQ ID NO: 8306. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element consists of SEQ ID NO: 8201, the 5'UTR element consists of SEQ ID NO: 8236, and the poly(A) tail consists of SEQ ID NO: 8307; (2) the 3'UTR element consists of SEQ ID NO: 8209, the 5'UTR element consists of SEQ ID NO: 8214, and the poly(A) tail consists of SEQ ID NO: 8306.

[0030] In certain embodiments, the polypeptide is a heterogeneously modified polypeptide or a retrotransposon genetically modified polypeptide. In certain embodiments, the polypeptide comprises a reverse transcriptase domain and an endonuclease domain. The endonuclease domain may be, for example, a nickasase domain, or a Cas9 domain selected from, for example, the SpCas9 domain, BlatCas9 domain, Nme2 Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain. In a particular embodiment, the Cas9 domain includes the N670A mutation, N611A mutation, N605A mutation, N580A mutation, N588A mutation, N872A mutation, N863A mutation, N622A mutation, or H840A mutation.

[0031] The reverse transcriptase domain may be selected from, for example, retroviral transcriptase domains. In certain embodiments, the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus. The gamma retrovirus-derived reverse transcriptase domain may include, for example, amino acid sequences of reverse transcriptase domain sequences from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In certain embodiments, the gamma retrovirus-derived reverse transcriptase domain is not derived from PERV.

[0032] In a particular embodiment, the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

[0033] In a particular embodiment, the genetically modified polypeptide comprises the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

[0034] In a particular embodiment, the template RNA further comprises a reverse transcriptase (RT) terminator sequence located between a heterogeneous object sequence and either (i) or (ii).

[0035] In a particular embodiment, the heterogeneous object sequence includes a sequence that encodes the target polypeptide or a portion thereof, or a sequence that is the inverse complementary chain of the sequence encoding the target polypeptide or a portion thereof.

[0036] In a particular embodiment, the polypeptide comprises a reverse transcriptase domain and an endonuclease domain, the endonuclease domain being a Cas9 domain, and the template RNA comprises (i) a gRNA spacer complementary to a first portion of the target gene, optionally comprising one or more consecutive nucleotides beginning at the 3' end of an adjacent nucleotide of the gRNA spacer; (ii) a gRNA scaffold that binds to the Cas9 domain; (iii) a heterologous object sequence comprising a mutation region for introducing a mutation into a second portion of the target gene (e.g., to correct a mutation therein) (optionally, the heterologous sequence comprises a 5'-to-3' post-edit homology region, a mutation region and a pre-edit homology region); and (iv) a primer-binding site (PBS) sequence comprising at least 5, 6, 7 or 8 bases having 100% identity with a third portion of the target gene.

[0037] In a particular embodiment, the target gene is a human PAH gene, and the template RNA is (i) a gRNA spacer complementary to the first portion of the human PAH gene, preferably having a sequence containing the core nucleotide of a gRNA spacer sequence from Table 1A, Table 1B, Table 1C or Table 1D of WO2023039435 incorporated herein by reference in whole, and optionally containing one or more consecutive nucleotides starting at the 3' end of an adjacent nucleotide of the gRNA spacer, or Table 5A-5F, 8A-8D, E3, E3A, BB, E5, E5A, E (ii) a gRNA spacer having a spacer sequence selected from 6 or E6A; (ii) a gRNA scaffold that binds to the Cas9 domain; (iii) a heterologous object sequence containing a mutation region for introducing a mutation into the second portion of the human PAH gene (for example, to correct a mutation therein) (the heterologous object sequence may include a 5' to 3' post-edit homologous region, a mutation region and a pre-edit homologous region); and (iv) a primer-binding site (PBS) sequence containing at least 5, 6, 7 or 8 bases having 100% identity with the third portion of the human PAH gene.

[0038] In a particular embodiment, the template RNA consists of the sequence of SEQ ID NO: 8258 (RNACS7570), SEQ ID NO: 8320 (RNACS229), SEQ ID NO: 8321 (RNACS1515), or SEQ ID NO: 8372.

[0039] In a particular embodiment, the reverse transcriptase domain and the endonuclease domain are linked by a peptide linker.

[0040] In a particular embodiment, the target site is a target site within the human genome.

[0041] Furthermore, a reaction mixture comprising cells and the system of the present invention is provided. In a particular embodiment, the cells are T cells (e.g., primary T cells).

[0042] Furthermore, a reaction mixture is provided that includes DNA containing a target site and the system of the present invention.

[0043] In a particular embodiment, the artificial RNA molecule comprises one or more chemically modified nucleotides.

[0044] Furthermore, a deoxyribonucleic acid (DNA) molecule encoding the artificial RNA molecule of the present invention is provided.

[0045] Furthermore, a pharmaceutical composition is provided comprising the artificial RNA molecule of the present invention, the system of the present invention, or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier. The pharmaceutically acceptable excipient or carrier may be selected from the group consisting of, for example, plasmid vectors, viral vectors, vesicles, and lipid nanoparticles. In a particular embodiment, the viral vector is an adeno-associated virus.

[0046] Furthermore, the present invention provides a host cell (e.g., a mammalian cell, e.g., a human cell) containing an artificial RNA molecule or system or DNA. The host cell may be, for example, a T cell (e.g., a primary T cell).

[0047] Furthermore, a method for producing the artificial RNA molecule of the present invention is provided. The method comprises synthesizing template RNA by in vitro transcription (e.g., solid-state synthesis) or by introducing DNA encoding the artificial RNA into a host cell under conditions that enable the production of template RNA.

[0048] A kit is also provided. The kit may include, for example, (a) the system, reaction mixture, DNA molecule or pharmaceutical composition of the present invention, and (b) instructions for use of the system, reaction mixture, DNA molecule or pharmaceutical composition.

[0049] Furthermore, lipid nanoparticles (LNPs) comprising the artificial RNA molecule or system of the present invention are provided.

[0050] Furthermore, a method for modifying a target site in the genomic DNA of a cell is provided. The method involves contacting a cell with the system of the present invention or one or more RNAs encoding the system of the present invention, thereby modifying a target site in the genomic DNA of the cell.

[0051] Furthermore, a method is provided for treating subjects having a disease or condition associated with a genetic defect. The method includes administering the system of the present invention to a subject, thereby treating the subject having a disease or condition associated with a genetic defect.

[0052] The above and other purposes, aspects, features and advantages of the exemplary embodiments can be made clearer and better understood by referring to the following description together with the accompanying drawings. [Brief explanation of the drawing]

[0053] [Figure 1] This graph shows the percentage of GFP-positive cells obtained 4 days after nucleofecting U2OS-BFP cells with mRNA encoding a genetically modified polypeptide without a HiBiT protein tag (mRNA encoding a genetically modified polypeptide without a HiBiT tag). [Figure 2] This graph shows the percentage of GFP-positive U2OS-BFP cells obtained 4 days after nucleofecting with mRNA encoding a genetically modified peptide fused with a HiBiT protein tag (mRNA encoding a HiBiT-tagged genetically modified polypeptide). [Figure 3] This graph shows the expression of mRNA encoding a HiBiT-tagged genetically modified polypeptide from mRNA with a reference 5'UTR, 6 hours after nucleofection in U2OS naive cells (which do not express BFP). [Figure 4] This graph shows the expression levels of HiBiT-tagged gene-modified polypeptides in Lonza's cryopreserved mouse hepatocytes (catalog number MCCP01) 24 hours after nucleofection. [Figure 5]This graph shows the expression levels of HiBiT-tagged genetically modified polypeptides from mRNA in fresh hepatocytes recovered from wild-type C57BL / 6 mice 24 hours after nucleofection. [Figure 6] This graph shows the percentage of GFP-positive cells obtained 4 days after nucleofecting U2OS-BFP cells with mRNA encoding a genetically modified polypeptide without a HiBiT protein tag (mRNA encoding a genetically modified polypeptide without a HiBiT tag). [Figure 7] This graph shows the percentage of GFP-positive cells 4 days after nucleofecting with mRNA encoding a genetically modified peptide fused with a HiBiT protein tag (mRNA encoding a HiBiT-tagged genetically modified polypeptide). [Figure 8] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from mRNA with a reference 3'UTR 6 hours after nucleofection in U2OS naive cells (which do not express BFP). [Figure 9] This graph shows the expression levels of HiBiT-tagged genetically modified polypeptides from mRNA in Lonza's cryopreserved mouse hepatocytes 24 hours after nucleofection. The dotted line indicates expression from mRNA with 3'UTR of reference 1. [Figure 10] This graph shows the expression levels of HiBiT-tagged genetically modified polypeptides from mRNA in hepatocytes freshly recovered from wild-type C57BL / 6 mice 24 hours after nucleofection. The dotted line indicates expression from mRNA with 3'UTR of reference 1. [Figure 11] This graph shows the percentage of GFP-positive cells obtained on day 4 after nucleofecting U2OS-BFP cells with mRNA encoding untagged HiBiT protein-modified polypeptides containing various 5'UTRs. The dotted line indicates the percentage of GFP produced by using a genetic modification system containing mRNA with reference 5'UTR and reference 1. [Figure 12]This graph shows the percentage of GFP-positive cells obtained on day 4 after nucleofecting U2OS-BFP cells with mRNA encoding HiBiT-tagged genetically modified polypeptides. The dotted line indicates the percentage of GFP produced by using a genetic modification system that includes mRNA containing reference 5'UTR and reference 1. [Figure 13] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from various 5'UTR-containing mRNAs 6 hours after nucleofection in U2OS naive cells (which do not express BFP). [Figure 14] This graph plots the expression of HiBiT-tagged genetically modified polypeptides from various 5'UTR-containing mRNAs in U2OS naive cells (non-BFP-expressing) against the percentage of GFP obtained by nucleofecting a genetically modified system containing the same HiBiT-tagged genetically modified polypeptide (normalized to expression from criterion 1 5'UTR-containing mRNA). Expression levels were obtained 6 hours after nucleofection. [Figure 15] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from various 5'UTR-containing mRNAs 24 hours after nucleofection in cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01). The dotted line indicates the expression levels of mRNA control 1 and genetically modified polypeptides from the 5'UTR. [Figure 16] This graph shows the expression levels of HiBiT-tagged genetically modified polypeptides from various 5'UTR-containing mRNAs 24 hours after nucleofection of hepatocytes freshly recovered from wild-type C57BL / 6 mice, normalized to the expression level of genetically modified polypeptides from mRNA containing the 5'UTR of Reference 1 (dotted line). [Figure 17]This graph shows the percentage of GFP-positive cells obtained on day 4 after nucleofecting U2OS-BFP cells with mRNA encoding untagged HiBiT protein-modified polypeptides containing various 3'UTRs. The dotted line indicates the percentage of GFP produced by using a genetic modification system containing mRNA with reference 3'UTR and reference 1. [Figure 18] This graph shows the percentage of GFP-positive cells obtained on day 4 after nucleofecting U2OS-BFP cells with mRNA encoding a HiBiT-tagged genetically modified peptide (mRNA encoding a HiBiT-tagged genetically modified polypeptide). The dotted line shows the percentage of GFP produced by using a genetic modification system that includes mRNA containing reference 3'UTR and reference 1. [Figure 19] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from various 3'UTR-containing mRNAs 6 hours after nucleofection in U2OS naive cells (which do not express BFP). The dotted line indicates expression from mRNA containing the reference 3'UTR and reference 1. [Figure 20] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from various 3'UTR-containing mRNAs 24 hours after nucleofection in cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01). The dotted line indicates the expression level of genetically modified polypeptides from the 3'UTR of mRNA reference 1. [Figure 21] This graph shows the expression of HiBiT-tagged genetically modified polypeptides from mRNAs comparing various 3'UTRs, 24 hours after nucleofection of hepatocytes freshly recovered from wild-type C57BL / 6 mice, normalized to the expression level of genetically modified polypeptides from mRNA containing 3'UTR of Reference 1 (dotted line). [Figure 22]This graph shows the expression of an exemplary genetically modified polypeptide (amino acid sequence shown by SEQ ID NO: 8257) from mRNA containing a modified poly(A) tail in U2OS-BFP cells 6 hours after nucleofection, normalized to the expression of the modified polypeptide expressed from mRNA containing a control-80A poly(A) tail. [Figure 23] This graph shows the percentage of GFP-positive cells obtained after nucleofecting USOS-BFP cells with modified poly-A tail-containing mRNA, normalized to the percentage of GFP-positive cells with control-80A poly-A tail-containing mRNA. [Figure 24] Figure 24A is a graph plotting the expression levels of genetically modified polypeptides from mRNAs containing various modified poly(A) tails (data from Figure 22) against the percentage of GFP-positive cells obtained (data from Figure 23). Correlations were calculated using Pearson's two-tailed test, yielding an R² of 0.5896. The Pearson analysis in Figure 24A includes four data points representing mRNAs exhibiting abnormally low genetically modified polypeptide expression, resulting in an R² of 0.7112. Figure 24B is a graph of the data from Figure 24A and the two-tailed test of the Pearson analysis, excluding the four data points representing mRNAs exhibiting abnormally low genetically modified polypeptide expression, which yielded an R² of 0.7112. [Figure 25] This graph shows the expression of genetically modified polypeptides from mRNA 18 hours after nucleofection of freshly recovered primary mouse liver cells. [Figure 26] Figure 26A is a flowchart of an experiment evaluating the expression of genetically modified polypeptides from mRNA over time, from 18 to 48 hours after nucleofection of primary mouse hepatocytes. Figure 26B is a graph of eight representative expression profiles of genetically modified polypeptides at 18, 24, 42, and 48 hours from eight mRNAs (six with modified polyA tails, one with a control-80A tail, and one with an external-80A tail) administered to primary mouse hepatocytes. [Figure 27-1]The expression profiles for all mRNAs evaluated in primary mouse hepatocytes are shown. [Figure 27-2] This is a continuation of Figure 27-1. [Figure 27-3] This is a continuation of Figure 27-2. [Figure 27-4] This is a continuation of Figure 27-3. [Figure 28] This is a ranking graph of the AUC (area under the curve) of the expression profiles of genetically modified polypeptides from mRNA containing different modified poly(A) tails, from 18 to 48 hours. [Figure 29] This graph ranks the percentage of modification in primary mouse liver cells that were nucleofected with gene modification systems containing mRNA with various modified poly(A) tails, according to the percentage of modification. [Figure 30] Graphs showing the time course of protein expression at 2, 4, 6, 8, and 24 hours from mRNA with test or reference UTR, analyzed by the Hibit assay, are shown. [Figure 31] Figure 30 shows a graph of the area under the protein expression curve (AUC) for 2 to 24 hours, calculated from the data. [Modes for carrying out the invention]

[0054] Various publications, articles, patents, and patent applications are cited or referenced in the background and throughout this Spec. Each of these references is incorporated herein by reference in its entirety. The documents, acts, materials, devices, articles, and other descriptions contained herein are for illustrative purposes only. Such descriptions do not constitute an endorsement that any or all of these constitute part of the prior art with respect to any invention disclosed or claimed.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention pertains. Otherwise, certain terms used herein have the meanings set forth herein.

[0056] It should be noted that, as used herein and in the accompanying claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly indicates otherwise.

[0057] Unless otherwise indicated, any numerical value, such as concentrations or concentration ranges described herein, should be understood to be modified in all cases by the term “approximately.” Therefore, numerical values ​​typically include ±10% of the listed values. For example, a concentration of 1 mg / mL includes 0.9 mg / mL to 1.1 mg / mL. Similarly, a concentration range of 1% to 10% (w / v) includes 0.9% (w / v) to 11% (w / v). Where used herein, unless the context clearly indicates otherwise, the use of numerical ranges obviously includes all possible subranges, all individual numerical values ​​within that range, including integers and fractions of values ​​within such ranges.

[0058] Unless otherwise indicated, the term “at least” preceding a set of elements should be understood to refer to each of those elements. Those skilled in the art can, by mere routine experimentation, recognize or confirm many equivalents of the particular embodiments of the invention described herein. Such equivalents are intended to be encompassed by the invention.

[0059] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” “contains,” or “containing,” or any other variation thereof, are intended to be non-exclusive or unrestrictive, indicating the inclusion of an integer or group of integers shown, but not the exclusion of any other integer or group of integers. For example, a composition, mixture, process, method, article, or apparatus that includes an enumeration of elements is not necessarily limited to those elements alone, and may include other elements that are not explicitly enumerated or are not specific to such composition, mixture, process, method, article, or apparatus. Furthermore, unless explicitly shown to the contrary, “or” refers to an inclusive “or” and not an exclusive “or.” For example, condition 1 or 2 is satisfied by one of the following: "1 is true (or exists) and 2 is false (or does not exist)", "1 is false (or does not exist) and 2 is true (or exists)", and "both 1 and 2 are true (or exist)".

[0060] Furthermore, as can be understood by those skilled in the art, the terms “about,” “approximately,” “generally,” “substantially,” and similar terms used herein when referring to the dimensions or features of preferred components of an invention should be understood to indicate that the dimensions / features described are not strict boundaries or parameters and do not exclude minor variations from which they are functionally the same or similar. At a minimum, such references, including numerical parameters, may include variations where the least significant digit does not vary, using mathematical and industrial principles permissible in the art (e.g., rounding, measurement or other systematic errors, manufacturing tolerances, etc.).

[0061] For sequence comparison, typically one sequence acts as a reference sequence, against which the test sequences are compared. When using a sequence comparison algorithm, the test sequences and reference sequences are input into a computer, and, if necessary, subsequence coordination and sequence algorithm program parameters are specified. The sequence comparison algorithm then calculates the sequence identity percentage of the test sequence(s) compared to the reference sequence based on the specified program parameters.

[0062] The optimal alignment of sequences for comparison can be achieved, for example, by the local homology algorithm of Smith and Waterman, Adv. Appl. Math. 1981;2:482; by the homology alignment algorithm of Needleman and Wunsch, J. Mol. Biol. 1970;48:443; by the similarity search method of Pearson and Lipman, Proc. Nat'l. Acad. Sci. USA 1988;85:2444; by computer implementation of these algorithms (GAP, BESTFIT, FASTA, and TFASTA (Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, WI)); or by visual inspection (generally, Current Protocols in Molecular Biology, edited by FM Ausubel et al., Current Protocols, Greene Publishing Associates, Inc. and John Wiley & Sons, This may be carried out by a joint venture of Inc. (see 1995 Addendum).

[0063] Examples of suitable algorithms for determining sequence identity percentage and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., J. Mol. Biol. 1990;215:403~410 and Altschul et al., Nucleic Acids Res. 1997;25:3389~3402, respectively. Software for performing BLAST analysis is publicly available through the National Center for Biotechnology Information. This algorithm first involves identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that, when aligned with words of the same length in the database sequence, fit or satisfy a certain positive threshold score T. T is called the neighbor word score threshold (Altschul et al., above). These initial neighbor word hits act as seeds for an initial search to find longer HSPs that contain them. Subsequently, word hits are stretched bidirectionally along each sequence, as long as the cumulative sort score can increase.

[0064] The cumulative score is calculated for nucleotide sequences using parameters M (reward score for matching residue pairs; always greater than 0) and N (penalty score for mismatched residues; always less than 0). For amino acid sequences, the cumulative score is calculated using a scoring matrix. Word hit extension in each direction is stopped if the cumulative alignment score falls by X units from its maximum achieved value; if the cumulative score becomes 0 or less due to the accumulation of residue alignments resulting in one or more negative scores; or if the end of any sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of alignment. The BLASTN program (for nucleotide sequences) uses word length (W) 11, expected value (E) 10, M=5, N=-4, and comparison of both strands as defaults. For amino acid sequences, the BLASTP program uses a word length (W) of 3, an expected value (E) of 10, and a BLOSUM62 scoring matrix by default (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 1989;89:10915).

[0065] In addition to calculating the sequence identity percentage, the BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul, Proc. Nat'l. Acad. Sci. USA 1993;90:5873~5787). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indicator of the probability that a match between two nucleotide or amino acid sequences may occur by chance. For example, if the smallest sum probability in a comparison of a test nucleic acid to a reference nucleic acid is less than about 0.1, more preferably less than about 0.01, and most preferably less than about 0.001, the nucleic acid is considered similar to the reference sequence.

[0066] A further indicator that two nucleic acid sequences or polypeptides are substantially identical is that the polypeptide encoded by the first nucleic acid is immunologically cross-reactive with the polypeptide encoded by the second nucleic acid, as described below. Therefore, if, for example, the two peptides differ only by conservative substitutions, the polypeptide is typically substantially identical to the second polypeptide. Another indicator that two nucleic acid sequences are substantially identical is that the two molecules hybridize with each other under stringent conditions.

[0067] As used herein, the term "expression cassette" refers to a nucleic acid construct containing sufficient nucleic acid elements for the expression of the nucleic acid molecule of the present invention.

[0068] As used herein, "gRNA spacer" refers to a portion of nucleic acid that is complementary to the target nucleic acid and, together with the gRNA scaffold, can target the Cas protein to the target nucleic acid.

[0069] As used herein, "gRNA scaffold" refers to a portion of nucleic acid that can bind to a Cas protein and, together with a gRNA spacer, can target the Cas protein to a target nucleic acid. In some embodiments, the gRNA scaffold includes a crRNA sequence, a tetraloop, and a tracrRNA sequence.

[0070] As used herein, the terms “peptide,” “polypeptide,” or “protein” may refer to a molecule consisting of amino acids that may be recognized as a protein by those skilled in the art. Conventional one-letter or three-letter notations of amino acid residues are used herein. The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acids of any length. The polymer may be linear or branched, may contain modified amino acids, or may be interspersed with non-amino acids. The term also encompasses naturally occurring or intervened amino acid polymers, including, for example, disulfide bond formation, glycosylation, lipid addition, acetylation, phosphorylation, or any other operation or modification, such as conjugation with labeled components. Furthermore, polypeptides containing one or more analogs of amino acids (including, for example, non-natural amino acids) and other modifications known in the art are included in this definition.

[0071] The peptide sequences described herein are written according to common convention, where the N-terminal region of the peptide is on the left and the C-terminal region is on the right. Although amino acid isomers are known, the type represented is the L-type of the amino acid unless otherwise clearly indicated.

[0072] In certain embodiments, “polypeptide” may be “genetically modified polypeptide.” When used interchangeably herein, “genetically modified polypeptide” and “retrotransposon genetically modified polypeptide” refer to polypeptides comprising a retrotransposase reverse transcriptase domain and a retrotransposase endonuclease domain, or polypeptides comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with the said domains, and which have the ability to incorporate a nucleic acid sequence (e.g., a sequence provided in a template nucleic acid) into a target DNA molecule (e.g., a mammalian host cell, e.g., a genomic DNA molecule within a host cell). In some embodiments, the endonuclease domain is a catalytically inactive endonuclease domain. In some embodiments, the retrotransposase reverse transcriptase domain and the retrotransposase endonuclease domain are derived from the same retrotransposase. In some embodiments, the genetically modified polypeptide has the ability to incorporate sequences substantially independently of host mechanisms. In some embodiments, the genetically modified polypeptide incorporates a sequence into a random location within the genome, and in some embodiments, the genetically modified polypeptide incorporates a sequence into a specific target site. In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively 1) facilitate binding to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. The genetically modified polypeptide comprises both naturally occurring polypeptides and engineered variants thereof, for example, variants having one or more amino acid substitutions to a naturally occurring sequence. The genetically modified polypeptide also comprises heterogeneous constructs, for example, constructs in which one or more of the domains listed above are heterogeneous to each other, whether through heterogeneous fusion (or other conjugates) of domains that would otherwise be wild-type domains and fusion of modified domains, for example, by means of substitution or fusion of heterogeneous subdomains or other substituted domains.Exemplary genetically modified polypeptides, systems comprising them, and methods of using them, which may be used in the methods provided herein, are described, for example, in WO2021 / 178717A2, incorporated herein by reference, and in Tables 10, 11, X, 3A, 3B, and Z1 contained therein. In some embodiments, the genetically modified polypeptide incorporates a sequence into a gene. In some embodiments, the genetically modified polypeptide incorporates a sequence outside the gene. When used herein, “genetically modified system” refers to a system comprising a genetically modified polypeptide and a template nucleic acid.

[0073] As used herein, the term “heterogeneically modified polypeptide” refers to a polypeptide comprising a retroviral reverse transcriptase, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with a retroviral reverse transcriptase, and having the ability to incorporate a nucleic acid sequence (e.g., a sequence provided by a template nucleic acid) into a target DNA molecule (e.g., a mammalian host cell, e.g., a genomic DNA molecule within a host cell). In some embodiments, heterogeneously modified polypeptides have the ability to incorporate sequences substantially independently of host mechanisms. In some embodiments, heterogeneously modified polypeptides incorporate sequences at random locations within the genome, and in some embodiments, heterogeneously modified polypeptides incorporate sequences at specific target sites. In some embodiments, the incorporated sequence includes deletions, substitutions, or insertions compared to the target DNA molecule. In some embodiments, heterogeneously modified polypeptides include one or more domains that collectively 1) facilitate binding to a template nucleic acid, 2) facilitate binding to a target DNA molecule, and 3) facilitate the incorporation of at least a portion of the template nucleic acid into the target DNA. Heterogeneically modified polypeptides include both naturally occurring polypeptides and manipulated variants thereof, such as variants having one or more amino acid substitutions to a naturally occurring sequence. Heterogeneically modified polypeptides also include heterogeneous constructs, such as constructs in which one or more of the domains listed above are heterogeneous, whether through heterogeneous fusion (or other conjugates) of domains that would otherwise be wild-type domains and fusion of modified domains, for example, by means of substitution or fusion of heterogeneous subdomains or other substituted domains. Exemplary heterogeneously modified polypeptides that may be used in the methods provided herein, as well as systems containing them and methods using them, are described, for example, in WO2021178720, incorporated herein by reference, for heterogeneously modified polypeptides containing retroviral reverse transcriptase domains. In some embodiments, heterogeneously modified polypeptides incorporate sequences into genes.In some embodiments, a heterologous genetically modified polypeptide incorporates its sequence into the sequence outside the gene.

[0074] As used herein, the term "domain" refers to a structure of a biomolecule that contributes to a particular function of the biomolecule. A domain may include a continuous region (e.g., a continuous sequence) or a separate discontinuous region (e.g., a discontinuous sequence) of a biomolecule. Examples of protein domains include, but are not limited to, endonuclease domains, DNA-binding domains, and reverse transcriptase domains. An example of a nucleic acid domain is a regulatory domain, such as a transcription factor-binding domain. In some embodiments, a domain (e.g., a Cas domain) may include two or more smaller domains (e.g., a DNA-binding domain and an endonuclease domain).

[0075] As used herein, “first strand” and “second strand” describe the individual DNA strands of the target DNA and are used to distinguish between the two DNA strands based on which strand the reverse transcriptase domain begins polymerization on, for example, where the target-primed synthesis begins. The first strand refers to the strand of target DNA on which the reverse transcriptase domain begins polymerization, for example, where the target-primed synthesis begins. The second strand refers to the other strand of target DNA. The designation of the first and second strands does not otherwise describe the target site DNA strands, and for example, in some embodiments, the first and second strands are nicked by the polypeptides described herein, but the designation of “first” and “second” strands does not imply the order in which such nicks occur.

[0076] When the term “heterogeneous” is used to describe the first element in reference to the second element, it means that the first and second elements do not exist in nature in the configuration described. For example, heterogeneous polypeptides, nucleic acid molecules, constructs, or sequences refer to (a) a polypeptide, nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule sequence that is not native to the cell in which it is expressed, (b) a polypeptide or nucleic acid molecule, or portion of a polypeptide or nucleic acid molecule that is modified or mutated compared to its native state, or (c) a polypeptide or nucleic acid molecule having modified expression compared to its native expression level under similar conditions. For example, heterogeneous regulatory sequences (e.g., promoters, enhancers) may be used to regulate the expression of a gene or nucleic acid molecule in a way that differs from how the gene or nucleic acid molecule is normally expressed in nature. In another example, heterogeneous domains of a polypeptide or nucleic acid sequence (e.g., a DNA-binding domain of a polypeptide, or a nucleic acid encoding the DNA-binding domain of a polypeptide) may be positioned relative to other domains, or may be different sequences or from different sources relative to other domains or portions of the polypeptide, or its encoding nucleic acid. In certain embodiments, heterologous nucleic acid molecules may be present in the native host cell genome, but may have altered expression levels, different sequences, or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous to the host cell or host genome, but may be introduced into the host cell by transformation (e.g., transfection, electroporation), and the added molecules may be incorporated into the host genome, or may exist transiently (e.g., as mRNA) or semi-stable for two or more generations (e.g., as episomal viral vectors, plasmids, or other self-replicating vectors) as extrachromosomal genetic material.

[0077] The term “nucleic acid molecule,” as used herein, refers to both RNA and DNA molecules, including but not limited to cDNA, genomic DNA, and mRNA, and also includes synthetic nucleic acid molecules, such as chemically synthesized or recombinantly produced nucleic acid molecules, such as RNA templates, as described herein. Nucleic acid molecules may be double-stranded or single-stranded, circular or linear. If single-stranded, the nucleic acid molecule may be a sense strand or an antisense strand. As an example for all sequences described herein in the general form of “Sequence Number,” unless otherwise indicated, “nucleic acid containing Sequence Number 1” means at least a portion of a nucleic acid having either (i) the sequence of Sequence Number 1, or (ii) a sequence complementary to Sequence Number 1. The choice between these two is indicated by the context in which Sequence Number 1 is used. For example, when a nucleic acid is used as a probe, the choice between the two is indicated by the requirement that the probe is complementary to the desired target. As will be readily apparent to those skilled in the art, the nucleic acid sequences of this disclosure may be chemically or biochemically modified, or may contain unnatural or derived nucleotide bases. Such modifications include, for example, labeling, methylation, substitution of one or more naturally occurring nucleotides by analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridines, psoralens, etc.), chelators, alkylators, and modified linkages (e.g., alpha-anomeric nucleic acids, etc.). It also includes synthetic molecules that mimic polynucleotides in terms of their ability to bind to a specified sequence via hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, molecules in which peptide bonds substitute for phosphate bonds within the molecular backbone. Other modifications may include, for example, analogs in which a ribose ring is a crosslinking moiety or other structure, such as “locked” nucleic acids.In various embodiments, nucleic acids are operable with additional gene elements, such as tissue-specific expression regulatory sequences (e.g., tissue-specific promoters and tissue-specific microRNA recognition sequences), as well as additional elements, such as reverse repeat sequences (e.g., reverse terminal repeat sequences, e.g., elements from or derived from viruses, e.g., AAV ITRs) and tandem repeat sequences, reverse repeat / direct repeat sequences (e.g., transposon reverse repeat sequences, e.g., transposon reverse repeat sequences also containing direct repeat sequences, e.g., reverse repeat sequences also containing direct repeat sequences), homologous regions (segments having varying degrees of homology to target DNA), UTRs (5', 3', or both 5' and 3' UTRs), and various combinations thereof. The nucleic acid elements of the system provided by the present invention can be provided in various topologies, including single-stranded, double-stranded, circular, linear, linear with open ends, linear with closed ends, and specific versions thereof, e.g., doggybone DNA (dbDNA), closed-end DNA (ceDNA).

[0078] As used herein, “insertion” of a sequence into a target site refers to the net addition of a DNA sequence at the target site, where, for example, a new nucleotide is present in the heterologous object sequence and no cognitive position is present in the unedited target site. In some embodiments, nucleotide alignment of the primer-binding site (PBS) sequence and the heterologous object sequence with respect to the target nucleic acid sequence may result in alignment gaps in the target nucleic acid sequence.

[0079] As used herein, a “deletion” generated by a heterologous object sequence at a target site refers to a net deletion of the DNA sequence at the target site, for example, where a nucleotide is present at the unedited target site and a cognitive position is absent in the heterologous object sequence. In some embodiments, nucleotide alignment of the PBS sequence and the heterologous object sequence to the target nucleic acid sequence may result in an alignment gap in the molecule containing the PBS sequence and the heterologous object sequence.

[0080] As used herein, the term "mutant region" refers to a region in template RNA that has one or more sequence differences compared to the corresponding sequence of the target nucleic acid. Sequence differences may include, for example, substitutions, insertions, frameshifts, or deletions.

[0081] When applied to nucleic acid sequences, the term "mutated" means that nucleotides in the nucleic acid sequence have been inserted, deleted, or altered compared to a reference (e.g., native) nucleic acid sequence. A single alteration may be produced at a single locus (point mutation), or multiple nucleotides may be inserted, deleted, or altered at a single locus. In addition, one or more alterations may be produced at any number of loci in a nucleic acid sequence. Nucleic acid sequences can be mutated by any method known in the art.

[0082] As used herein, “gene expression unit” is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably ligated to at least one effector sequence. The first nucleic acid sequence is operably ligated to the second nucleic acid sequence when the first nucleic acid sequence is functionally related to the second nucleic acid sequence. For example, if a promoter or enhancer affects the transcription or expression of a coding sequence, the promoter or enhancer is operably ligated to the coding sequence. The operably ligated DNA sequences may be contiguous or discontinuous. If it is necessary to join two protein-coding regions, the operably ligated sequences may be within the same reading frame.

[0083] The terms “host genome” or “host cell,” as used herein, refer to a cell and / or its genome into which protein and / or genetic material has been introduced. It should be understood that such terms are intended to refer not only to a specific target cell and / or genome, but also to the progeny of such cells and / or the genomes of such progeny. Such progeny may not be identical to the parent cell in practice, because certain modifications may occur in subsequent generations due to mutation or environmental influences, but they are still included within the scope of the term “host cell” as used herein. A host genome or host cell may be an isolated cell, a cell line grown in culture, or genomic material isolated from such a cell or cell line, or a host cell or host genome constituting a living tissue or organism. In some cases, a host cell may be, for example, an animal cell or a plant cell as described herein. In certain cases, a host cell may be a mammalian cell, a human cell, a bird cell, a reptile cell, a bovine cell, a horse cell, a pig cell, a goat cell, a sheep cell, a chicken cell, or a turkey cell. In certain cases, the host cells may be maize cells, soybean cells, wheat cells, or rice cells.

[0084] As used herein, “operable relationship” describes a functional relationship between two nucleic acid sequences, for example, 1) a promoter and 2) a heterologous object sequence, in which the promoter and the heterologous object sequence (e.g., the gene of interest) are directed so that, under favorable conditions, the promoter drives the expression of the heterologous object sequence. For example, a template nucleic acid having a promoter and a heterologous object sequence may be single-stranded and may be, for example, in either the (+) or (-) direction. The “operable relationship” between the promoter and the heterologous object sequence in this template means that, regardless of whether the template nucleic acid is transcribed into a particular state, if it is present in a favorable state (e.g., in the (+) direction, in the presence of the required catalyst and NTP), it will be transcribed accurately. The operable relationship applies similarly to other pairs of nucleic acids, including other tissue-specific expression regulatory sequences (e.g., enhancers, repressors, and microRNA recognition sequences), IR / DR, ITR, UTR, or homologous regions and sequences encoding heterologous object sequences or retroviral RT domains.

[0085] The terms “primer-binding site sequence” or “PBS sequence,” as used herein, refer to a portion of template RNA capable of binding to a region contained in a target nucleic acid sequence. In some cases, the PBS sequence is a nucleic acid sequence containing at least 3, 4, 5, 6, 7, or 8 nucleotides having 100% identity with a region contained in the target nucleic acid sequence. In some embodiments, the primer region contains at least 5, 6, 7, or 8 nucleotides having 100% identity with a region contained in the target nucleic acid sequence. While not theoretically bound, in some embodiments, if the template RNA contains both a PBS sequence and a heterologous object sequence, the PBS sequence binds to a region contained in the target nucleic acid sequence, enabling the reverse transcriptase domain to use that region as a primer for reverse transcription and the heterologous object sequence as a template for reverse transcription.

[0086] Genetically modified RNA molecules and systems containing them Genome manipulation promises tremendous therapeutic potential, including the ability to permanently address genetic diseases. However, existing genome manipulation methods are limited, for example, by the expression levels and / or stability of gene-editing polypeptides in the host cells to be edited. Therefore, there is a need for improved genome manipulation methods to address the need for improved systems and expression constructs for gene-editing polypeptides.

[0087] The present invention provides, in particular, artificial ribonucleic acid (RNA) molecules for improving the expression of polypeptides, preferably genetically modified polypeptides. Polypeptides may include, for example, a reverse transcriptase (RT) domain and optionally an endonuclease domain. Artificial RNA molecules may include, for example, 5' and / or 3' untranslated region (UTR) elements and / or poly(A) tails that are specifically designed to improve the expression of the polypeptide of interest.

[0088] Accordingly, artificial ribonucleic acid (RNA) molecules for enhancing the expression of polypeptides comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain are provided herein. The artificial RNA molecule may, for example, include (1) a nucleotide sequence encoding a polypeptide, (2) at least one of a 3'-untranslated region (3'UTR) element for enhancing the expression of (a) polypeptide, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or a 5'-untranslated region (5'UTR) element for enhancing the expression of (b) polypeptide, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263, and / or a poly(A) tail for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282-8314.

[0089] In a particular embodiment, the RNA molecule may include, for example, (1) a nucleotide sequence encoding a polypeptide, (2) (a) a 3'-untranslated region (3'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (b) a 5'-untranslated region (5'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263.

[0090] In a particular embodiment, the RNA molecule may include, for example, (1) a nucleotide sequence encoding a polypeptide, and (2) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314.

[0091] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises (1) at least one of (a) a 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201 to 8211, and / or (b) a 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212 to 8247, and / or (2) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314.

[0092] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises at least one of the following: (a) a 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269 and 8281; and / or (b) a 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212-8247 and 8259-8263.

[0093] In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314.

[0094] In certain embodiments, the 3'UTR element comprises a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269, and 8281. In certain embodiments, the 3'UTR element comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269, and 8281. In certain embodiments, the 3'UTR element comprises a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8256 and 8264-8269, wherein the nucleic acid sequence does not contain a 3'CUAG nucleotide.

[0095] In certain embodiments, the 3'UTR element of the RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281. In certain embodiments, the 3'UTR element of the RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281. In certain embodiments, the 3'UTR element of the artificial RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, or SEQ ID NO: 8211. In certain embodiments, the 3'UTR element of the artificial RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, or SEQ ID NO: 8211.

[0096] In a particular embodiment, the 5'UTR element comprises a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with a nucleic acid sequence selected from the group consisting of SEQ ID NOs.

[0097] In certain embodiments, the 5'UTR element of the RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263. In certain embodiments, the 5'UTR element of the RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263. In certain embodiments, the 5'UTR element of the artificial RNA molecule includes the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243. In certain embodiments, the 5'UTR element of the artificial RNA molecule consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243.

[0098] In a particular embodiment, the poly(A) tail for improving polypeptide expression comprises a nucleic acid sequence having at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 to 8233. In a particular embodiment, the poly(A) tail for improving polypeptide expression comprises SEQ ID NOs: 8201 to 8233.

[0099] In a particular embodiment, the poly(A)tail comprises the nucleic acid sequence of SEQ ID NOs: 8213, 8218, 8225, 8226, 8228, 8230, 8231, or 8233. In a particular embodiment, the poly(A)tail consists of the nucleic acid sequence of SEQ ID NOs: 8213, 8218, 8225, 8226, 8228, 8230, 8231, or 8233.

[0100] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8236; (2) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236; (3) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214; (4) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8243; (5) the 3'UTR element comprises SEQ ID NO: 8281 and the 5'UTR element comprises SEQ ID NO: 8259; or (6) the 3'UTR element comprises SEQ ID NO: 8281 and the 5'UTR element comprises SEQ ID NO: 8263. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element and a 5'UTR element, where (1) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8236; (2) the 3'UTR element is sequence number 8201 and the 5'UTR element is sequence number 8236; (3) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8214; (4) the 3'UTR element is sequence number 8209 and the 5'UTR element is sequence number 8243; (5) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8259; or (6) the 3'UTR element is sequence number 8281 and the 5'UTR element is sequence number 8263.

[0101] In a particular embodiment, the artificial RNA molecule includes a poly(A)tail, where the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0102] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8236, (2) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236, (3) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214, or (4) the 3'UTR element comprises (5) The 3'UTR element contains sequence number 8209, and the 5'UTR element contains sequence number 8243, or (6) the 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8259, or (7) the 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263; the poly(A)tail contains the nucleic acid sequence of sequence number 8294, sequence number 8299, sequence number 8306, sequence number 8307, sequence number 8309, sequence number 8311, sequence number 8312, or sequence number 8314. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8236, (2) the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236, (3) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214, or (4) the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8243; and the poly(A) tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0103] In certain embodiments, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (3) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail (4) The 3'UTR element includes sequence number 8209, the 5'UTR element includes sequence number 8243, and the poly(A) tail includes sequence number 8306 or sequence number 8307, (5) The 3'UTR element includes sequence number 8281, the 5'UTR element includes sequence number 8259, and the poly(A) tail includes sequence number 8306 or sequence number 8307, or (6) The 3'UTR element includes sequence number 8281, the 5'UTR element includes sequence number 8263, and the poly(A) tail includes sequence number 8306 or sequence number 8307. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; (3) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307; or (4) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8243, and the poly(A) tail comprises SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0104] In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element comprises SEQ ID NO: 8201, the 5'UTR element comprises SEQ ID NO: 8236, and the poly(A) tail comprises SEQ ID NO: 8307; (2) the 3'UTR element comprises SEQ ID NO: 8209, the 5'UTR element comprises SEQ ID NO: 8214, and the poly(A) tail comprises SEQ ID NO: 8306. In a particular embodiment, the artificial RNA molecule comprises a 3'UTR element, a 5'UTR element and a poly(A) tail, where (1) the 3'UTR element consists of SEQ ID NO: 8201, the 5'UTR element consists of SEQ ID NO: 8236, and the poly(A) tail consists of SEQ ID NO: 8307; (2) the 3'UTR element consists of SEQ ID NO: 8209, the 5'UTR element consists of SEQ ID NO: 8214, and the poly(A) tail consists of SEQ ID NO: 8306.

[0105] Furthermore, an artificial ribonucleic acid (RNA) molecule is provided for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, wherein the RNA molecule comprises (a) a nucleotide sequence encoding a polypeptide and (b) a poly(A) tail for improving polypeptide expression, the poly(A) tail comprising a nucleic acid sequence having a pattern of 12 to 18 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, 22 to 28 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, 32 to 38 adenosine nucleotides followed by at least 2 non-adenosine nucleotides, and 42 to 48 adenosine nucleotides.

[0106] In a particular embodiment, the artificial RNA molecule comprises a nucleic acid sequence comprising a pattern of 13-17 adenosine nucleotides followed by 1-4 non-adenosine nucleotides, 23-27 adenosine nucleotides followed by 1-5 non-adenosine nucleotides, 33-37 adenosine nucleotides followed by 2-6 non-adenosine nucleotides, and 43-47 adenosine nucleotides. In a particular embodiment, the artificial RNA molecule comprises a nucleic acid sequence comprising a pattern of 14-16 adenosine nucleotides followed by at least 2 (e.g., 2-6) non-adenosine nucleotides, 24-26 adenosine nucleotides followed by at least 2 (e.g., 2-6) non-adenosine nucleotides, 34-36 adenosine nucleotides followed by at least 2 (e.g., 2-6) non-adenosine nucleotides, and 44-46 adenosine nucleotides. In a particular embodiment, the artificial ribonucleic acid (RNA) molecule comprises a nucleic acid sequence comprising a pattern of 15 adenosine nucleotides followed by at least two (e.g., 2 to 4) non-adenosine nucleotides, 25 adenosine nucleotides followed by at least two (e.g., 2 to 4) non-adenosine nucleotides, 35 adenosine nucleotides followed by at least two (e.g., 2 to 4) non-adenosine nucleotides, and 45 adenosine nucleotides. In a particular embodiment, the artificial RNA molecule comprises a nucleic acid sequence comprising a pattern of 15 adenosine nucleotides followed by 1 to 4 non-adenosine nucleotides, 25 adenosine nucleotides followed by 1 to 5 non-adenosine nucleotides, 35 adenosine nucleotides followed by 2 to 6 non-adenosine nucleotides, and 45 adenosine nucleotides.

[0107] In certain embodiments, there are 2 to 10 non-adenosine nucleotides. In certain embodiments, the number of non-adenosine nucleotides within each group of non-adenosine nucleotides increases from 5' to 3' of the artificial RNA molecule. In certain embodiments, the number of non-adenosine nucleotides within each group of non-adenosine nucleotides from 5' to 3' of the artificial RNA molecule is 2, 3, and 4, respectively.

[0108] In certain embodiments, the non-adenosine nucleotide is selected from cytosine nucleotides or uridine nucleotides.

[0109] In a particular embodiment, the artificial RNA molecule comprises one or more chemically modified nucleotides.

[0110] Furthermore, a deoxyribonucleic acid (DNA) molecule encoding the artificial RNA molecule of the present invention is provided.

[0111] Genetically modified polypeptides In some embodiments, genetically modified polypeptides act as substantially autonomous protein machines capable of incorporating template nucleic acid sequences into target DNA molecules (e.g., genomic DNA molecules within mammalian host cells) substantially independently of host mechanisms. For example, genetically modified polypeptides may include a DNA-binding domain, a reverse transcriptase domain, and an endonuclease domain. In some embodiments, the DNA-binding function may include an RNA component, such as a gRNA spacer, that directs the protein to the DNA sequence. In other embodiments, genetically modified polypeptides may include a reverse transcriptase domain and an endonuclease domain. In some embodiments, the RNA template element is provided by the genetic modification system, where the RNA template element is typically heterologous to the genetically modified polypeptide element and provides an object sequence to be inserted (reverse transcribed) into the host genome. In some embodiments, the genetically modified polypeptide is capable of reverse transcription primed by the target. In some embodiments, the genetically modified polypeptide is capable of synthesizing a second strand.

[0112] Genetically modified polypeptides suitable for use in the compositions and methods described herein include, for example, polypeptides comprising reverse transcriptases (e.g., retroviral or retrotransposon reverse transcriptases), retrotransposes, DNA transposes, and recombinases (e.g., serine recombinases and tyrosine recombinases). Exemplary Gene Writer polypeptides, i.e., genetically modified polypeptides, as well as systems comprising them and methods of using them, are described, for example, in WO2020 / 047124 and WO2021 / 178720, which are incorporated herein by reference in their entirety, including amino acid and nucleic acid sequences.

[0113] For example, Table 3 of WO2020 / 047124 is incorporated herein by reference in its entirety. In some embodiments, the genetically modified polypeptide includes the amino acid sequence in column 8 of Table 3 of WO2020 / 047124, or any domain thereof (e.g., a DNA-binding domain, an RNA-binding domain, an endonuclease domain, or an RT domain), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with them. In some embodiments, the template RNA includes the sequence in Table 3 of WO2020 / 047124 (e.g., one or both of the 5' untranslated region in column 6 and the 3' untranslated region in column 7), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it.

[0114] Furthermore, exemplary genetically modified polypeptides, systems comprising genetically modified polypeptides, and methods using genetically modified polypeptides are described in WO2021 / 178720, which is incorporated herein by reference, for example, retroviral RT domains comprising amino acids and nucleic acid sequences. Exemplary genetically modified polypeptides and retroviral RT domain sequences are described in Tables 30, 31, and 44 of WO2021 / 178720. Accordingly, the genetically modified polypeptides described herein may include amino acid sequences or their domains (e.g., retroviral RT domains) listed in any of the tables referred to, or any of the functional fragments or variants described above, or amino acid sequences having at least 70%, 80%, 85%, 90%, 95%, or 99% identity with them.

[0115] Furthermore, exemplary retrotransposon gene-modified polypeptides, template nucleic acids, and systems comprising them are described, for example, in Tables 3A, 3B, 10, and 11 of WO2021178717A2, which is incorporated herein by reference in its entirety.

[0116] In some embodiments, the first genetically modified polypeptide is combined with a second polypeptide in the genetic modification system. In some embodiments, the second polypeptide may include an endonuclease domain. In some embodiments, the second polypeptide may include a polymerase domain, such as a reverse transcriptase domain. In some embodiments, the second polypeptide may include a DNA-dependent DNA polymerase domain. In some embodiments, the second polypeptide assists in the completion of genome editing, for example, by contributing to the synthesis of a second strand or a DNA repair mechanism.

[0117] In some embodiments, the genetically modified polypeptide comprises one or more domains that collectively promote 1) binding to a template nucleic acid, 2) binding to a target DNA molecule, and 3) integration of at least a portion of the template nucleic acid into the target DNA. In some embodiments, the genetically modified polypeptide is an engineered polypeptide comprising one or more amino acid substitutions to the corresponding naturally occurring sequence. In some embodiments, the genetically modified polypeptide comprises two or more domains that are heterogeneous to each other, for example, through heterogeneous fusion (or other conjugates) of domains that would otherwise be wild-type domains and fusion of modified domains, by means of substitution or fusion of heterogeneous subdomains or other substituted domains. For example, in some embodiments, one or more of the following may be said: the RT domain is heterogeneous to the DNA-binding domain (DBD), the DBD is heterogeneous to the endonuclease domain, or the RT domain is heterogeneous to the endonuclease domain.

[0118] A functionally modified polypeptide may consist of an unrelated DNA-binding domain, a reverse transcription domain, and an endonuclease domain. This modular structure allows for combinations of functional domains, e.g., dCas9 (DNA binding), MMLV reverse transcriptase (reverse transcription), and FokI (endonuclease). In some embodiments, multiple functional domains may arise from a single protein, e.g., Cas9 or Cas9 nickase (DNA binding, endonuclease).

[0119] In some embodiments, the genetically modified polypeptides described herein include a reverse transcriptase or RT domain (e.g., as described herein) containing a MoMLV RT sequence or a variant thereof. In embodiments, the MoMLV RT sequence includes one or more mutations selected from D200N, L603W, T330P, T306K, W313F, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, L435G, N454K, H594Q, D653N, R110S, and K103L. In embodiments, the MoMLV RT sequence includes a combination of mutations, e.g., D200N, L603W, and T330P, and optionally further including T306K and / or W313F.

[0120] In some embodiments, the genetically modified polypeptide comprises an endonuclease domain of nCas9 (as described herein, for example), including, for example, an N863A mutation (e.g., in spCas9) or an H840A mutation.

[0121] In certain embodiments, the polypeptide comprises a reverse transcriptase domain and an endonuclease domain. The endonuclease domain may be a Cas9 domain selected from, for example, a nickasase domain, such as a SpCas9 domain, a BlatCas9 domain, an Nme2 Cas9 domain, a PnpCas9 domain, a SauCas9 domain, a SauCas9-KKH domain, a SauriCas9 domain, a SauriCas9-KKH domain, a ScaCas9-Sc++ domain, a SpyCas9 domain, a SpyCas9-NG domain, a SpyCas9-SpRY domain, or a St1Cas9 domain. In certain embodiments, the Cas9 domain comprises an N670A mutation, an N611A mutation, an N605A mutation, an N580A mutation, an N588A mutation, an N872A mutation, an N863A mutation, an N622A mutation, or an H840A mutation.

[0122] The reverse transcriptase domain may be selected from, for example, retroviral transcriptase domains. In certain embodiments, the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus. The gamma retrovirus-derived reverse transcriptase domain may include amino acid sequences of reverse transcriptase domain sequences from a family selected from, for example, AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In certain embodiments, the gamma retrovirus-derived reverse transcriptase domain is not derived from PERV.

[0123] In a particular embodiment, the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

[0124] Preferably, the polypeptide contains the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

[0125] In some embodiments, the RT and endonuclease domains are joined by a flexible linker, for example, containing the amino acid sequence AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEAAAKEAAAKA (SEQ ID NO: 8275).

[0126] In some embodiments, the endonuclease domain is located at the N-terminus relative to the RT domain. In some embodiments, the endonuclease domain is located at the C-terminus relative to the RT domain.

[0127] In some embodiments, the genetically modified polypeptide has the ability to produce substitutions at target sites of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 or more nucleotides. In some embodiments, the substitutions are transition mutations. In some embodiments, the substitutions are transversion mutations. In some embodiments, the substitutions convert adenine to thymine, adenine to guanine, adenine to cytosine, guanine to adenine, guanine to cytosine, guanine to adenine, thymine to cytosine, thymine to adenine, thymine to guanine, cytosine to adenine, cytosine to guanine, or cytosine to thymine.

[0128] In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by altering, adding, or deleting sequences in promoters or enhancers, such as transcription factors. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene translation (e.g., alter amino acid sequences), insert or delete start or stop codons, or alter or modify the gene translation frame. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene splicing, for example, by inserting, deleting, or altering splice acceptors or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter the half-life of transcripts or proteins. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in cells (e.g., from cytoplasm to mitochondria, from cytoplasm to extracellular space (e.g., by adding secretory tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter (e.g., improve) protein folding (e.g., to prevent the accumulation of incorrectly folded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease gene activity, for example, the protein encoded by the gene.

[0129] Retargeting (e.g., retargeting of genetically modified polypeptides or nucleic acid molecules, or systems, as described herein) generally involves (i) orienting the polypeptide to bind and cleave at a target site, and / or (ii) designing the template RNA to have complementarity to the target sequence. In some embodiments, for example for target-primed reverse transcription (TPRT), the template RNA has complementarity to the 5' end of the nick on the first strand of the target sequence, for example, so that the 3' end of the template RNA anneals and the 5' end of the target site acts as a primer. In some embodiments, the endonuclease domain of the polypeptide and the 5' end of the RNA template are also modified as described.

[0130] In some embodiments, the genetically modified polypeptide includes modifications to the DNA-binding domain compared to, for example, the wild-type polypeptide. In some embodiments, the DNA-binding domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the original DNA-binding domain. In some embodiments, the DNA-binding domain is modified to include a heterogeneous functional domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the functional domain replaces at least a portion (e.g., all) of the polypeptide's previous DNA-binding domain. In some embodiments, the functional domain includes a zinc finger (e.g., a zinc finger that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest). In some embodiments, the functional domain includes a Cas domain (e.g., a Cas domain that specifically binds to a target nucleic acid (e.g., DNA) sequence of interest). In embodiments, the Cas domain includes Cas9 or its mutant or variant (e.g., as described herein, see, for example, SEQ ID NOs 8331-8369). In embodiments, the Cas domain is associated with a guide RNA (gRNA), as described herein, for example. In one embodiment, the Cas domain is directed by the gRNA to a target nucleic acid (e.g., DNA) sequence. In another embodiment, the Cas domain is encoded in the same nucleic acid (e.g., RNA) molecule as the gRNA. In yet another embodiment, the Cas domain is encoded in a different nucleic acid (e.g., RNA) molecule than the gRNA.

[0131] In some embodiments, the genetically modified polypeptide includes modifications to the endonuclease domain compared to, for example, the wild-type polypeptide. In some embodiments, the endonuclease domain includes additions, deletions, substitutions, or modifications to the amino acid sequence of the original endonuclease domain. In some embodiments, the endonuclease domain is modified to include heterologous functional domains that specifically bind to a target nucleic acid (e.g., DNA) sequence of interest and / or heterologous functional domains that induce endonuclease cleavage of the target nucleic acid (e.g., DNA) sequence of interest. In some embodiments, the endonuclease domain includes zinc fingers. In some embodiments, the endonuclease domain includes a Cas domain (e.g., Cas9 or its mutant or variant). In embodiments, the endonuclease domain including the Cas domain is associated with a guide RNA (gRNA), as described herein, for example. In some embodiments, the endonuclease domain is modified to include functional domains that do not target a specific target nucleic acid (e.g., DNA) sequence. In embodiments, the endonuclease domain includes a Fokl domain.

[0132] In some embodiments, the reverse transcriptase (RT) domain exhibits improved stringency for target-primed reverse transcription (TPRT) initiation compared to, for example, an endogenous RT domain. In some embodiments, the RT domain initiates TPRT if the target site immediately upstream of the nick in the first strand, for example, three nucleotides (nt) of the genomic DNA priming the RNA template, has at least 66% or 100% complementarity to the 3nt homology of the RNA template. In some embodiments, the RT domain initiates TPRT if the mismatch between the template RNA homology and the target DNA priming reverse transcription is less than five nt (e.g., mismatch is less than 1, 2, 3, 4, or 5 nt). In some embodiments, the RT domain is modified to increase stringency with respect to mismatch in priming the TPRT reaction, for example, here the RT domain does not tolerate any mismatch in the priming region, or tolerates fewer mismatches, compared to a wild-type (e.g., unmodified) RT domain.

[0133] In some embodiments, the RT domain includes an HIV-1 RT domain. In some embodiments, the HIV-1 RT domain initiates lower levels of synthesis, even with just three nucleotide mismatches, compared to alternative RT domains (as described, for example, by Jamburthugoda and Eickbush J Mol Biol 407(5):661-672 (2011), which are incorporated herein by reference in their entirety). In some embodiments, the RT domain forms a dimer (e.g., a heterodimer or homodimer). In some embodiments, the RT domain is a monomer. In some embodiments, the RT domain functions naturally as a monomer or dimer (e.g., a heterodimer or homodimer). In some embodiments, the RT domain functions naturally as a monomer and originates, for example, from a virus in which it functions as a monomer. In this embodiment, the RT domain is mouse leukemia virus (MLV; sometimes called MoMLV) (e.g., P03355), porcine endogenous retrovirus (PERV) (e.g., UniProt Q4VFZ2), mouse mammary tumor virus (MMTV) (e.g., UniProt P03365), avian reticuloendotheliopathy virus (AVIRE) (e.g., UniProtKB accession number P03360), feline leukemia virus (FLV or FeLV) (e.g., UniProtKB accession number P10273), Mason-Pfizer monkey virus (MPMV) (e.g., UniProt P07572), bovine leukemia virus (BLV) (e.g., UniProt P03361), human T-cell leukemia virus-1 (HTLV-1) (e.g., UniProt P03362), human foam virus (HFV) (e.g., UniProt The RT domain is selected from P14350), monkey foam virus (SFV) (e.g., SFV3L) (e.g., UniProt P23074 or P27401), or bovine foam virus / syntiocytotic virus (BFV / BSV) (e.g., UniProt O41894), or from functional fragments or variants thereof (e.g., amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with them). In some embodiments, the RT domain is dimer in its native functionality.In some embodiments, the RT domain is derived from a virus that functions as a dimer. In embodiments, the RT domain is derived from avens sarcoma / leukemia virus (ASLV) (e.g., UniProt A0A142BKH1), Roussarcoma virus (RSV) (e.g., UniProt P03354), avens myeloblastoma virus (AMV) (e.g., UniProt Q83133), human immunodeficiency virus type I (HIV-1) (e.g., UniProt P03369), human immunodeficiency virus type II (HIV-2) (e.g., UniProt P15833), simian immunodeficiency virus (SIV) (e.g., UniProt P05896), bovine immunodeficiency virus (BIV) (e.g., UniProt P19560), equine infectious anemia virus (EIAV) (e.g., UniProt P03371), or feline immunodeficiency virus (FIV) (e.g., UniProt P16088) (Herschhorn and Hizi Cell Mol Life Sci RT domains from 67(16):2717~2747(2010)) or functional fragments or variants thereof (e.g., amino acid sequences having at least 70%, 80%, 90%, 95%, or 99% identity with them) are selected. RT domains that are naturally heterodimers may also function as homodimers in some embodiments. In some embodiments, dimeric RT domains are expressed as fusion proteins, for example, as homodimeric fusion proteins or heterodimeric fusion proteins. In some embodiments, the RT function of the system is performed by a number of RT domains (e.g., as described herein). In further embodiments, the number of RT domains may be fused or separate, and may be, for example, on the same polypeptide or on different polypeptides.

[0134] In some embodiments, the genetically modified polypeptide has the function of cleaving a target DNA site via an endonuclease domain. In some embodiments, the genetically modified polypeptide includes a DNA-binding domain for, for example, binding to a target nucleic acid. In some embodiments, the domain of the genetically modified polypeptide (e.g., the Cas domain) includes two or more smaller domains, e.g., a DNA-binding domain and an endonuclease domain. When it is said that the DNA-binding domain (e.g., the Cas domain) binds to a target nucleic acid sequence, in some embodiments, it is understood that the binding is mediated by gRNA.

[0135] In some embodiments, the domain has two functions. For example, in some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA)-binding domain. For example, in some embodiments, the polypeptide includes a CRISPR-associated endonuclease domain that binds to a template RNA containing gRNA, binds to a target DNA sequence (e.g., complementary to a portion of the gRNA), and cleaves the target DNA sequence. In some embodiments, an endonuclease domain or endonuclease / DNA-binding domain from a heterogeneous source may be used or modified (e.g., by insertion, deletion, or substitution of one or more residues) in the gene modification system described herein.

[0136] In some embodiments, the endonuclease domain has nickase activity to nick a target site DNA on a first strand, for example, in some embodiments, the endonuclease domain cleaves genomic DNA at a target site near a site of modification on the strand extended by the write domain. In some embodiments, the endonuclease domain has nickase activity to nick a target site DNA on a first strand but not on a target site DNA on a second strand. For example, if the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain nicks a target site DNA strand containing a PAM site (for example, and does not nick a target site DNA strand that does not contain a PAM site). As a further example, if the polypeptide contains a CRISPR-related endonuclease domain having nickase activity, in some embodiments, the CRISPR-related endonuclease domain nicks a target site DNA strand that does not contain a PAM site (for example, and does not nick a target site DNA strand that contains a PAM site).

[0137] In some embodiments, the polypeptide comprises a single domain having endonuclease activity (e.g., a single endonuclease domain) which nicks both a first and a second strand. For example, in such an embodiment, the endonuclease domain may be a CRISPR-associated endonuclease domain, and the template nucleic acid (e.g., template RNA) comprises a gRNA spacer that directs nick formation on the first strand and an additional gRNA spacer that directs nick formation on the second strand. In some embodiments, the polypeptide comprises multiple domains having endonuclease activity, where the first endonuclease domain nicks on the first strand and the second endonuclease domain nicks on the second strand (or, in some cases, the first endonuclease domain does not nick (e.g., cannot nick) on the second strand, and the second endonuclease domain does not nick (e.g., cannot nick) on the first strand).

[0138] In some embodiments, the genetically modified polypeptides described herein include a Cas domain. In some embodiments, the Cas domain directs the genetically modified polypeptide to a target site identified by a gRNA spacer, thereby modifying the target nucleic acid sequence with "cis". In some embodiments, the genetically modified polypeptide is fused to the Cas domain. In some embodiments, the genetically modified polypeptide includes a CRISPR / Cas domain (also referred to herein as a CRISPR-related protein). In some embodiments, the CRISPR / Cas domain includes a protein involved in the CRISPR (clustered regulatory interspaced short palindromic repeat) system, such as a Cas protein, and may bind to a guide RNA, such as a single guide RNA (sgRNA).

[0139] Various CRISPR-related (Cas) genes or proteins can be used in the techniques provided by this disclosure, and the selection of a Cas protein depends on the specific conditions of the method. Specific examples of Cas proteins include a Class II system comprising Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cpf1, C2C1, or C2C3. In some embodiments, a Cas protein, e.g., a Cas9 protein, may be a Cas9 protein from any of various prokaryotic species. In some embodiments, a specific Cas protein, e.g., a specific Cas9 protein, is selected to recognize a specific protospacer-adjacent motif (PAM) sequence. In some embodiments, the DNA-binding domain or endonuclease domain comprises a sequence-targeted polypeptide, e.g., a Cas protein, e.g., Cas9. In certain embodiments, a Cas protein, e.g., a Cas9 protein, may be obtained from bacteria or archaea or synthesized using known methods. Further explanation of the CRISPR system can be found in WO2021 / 178898, which is incorporated herein by reference in its entirety.

[0140] Example arrangement of Cas9-linker-RT fusion In some embodiments, the genetically modified polypeptide (for example, a genetically modified polypeptide that is part of the system described herein) includes one amino acid sequence from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 80% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 90% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 95% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 99% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from sequence numbers 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the genetically modified polypeptide includes one amino acid sequence from sequence numbers 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0141] In some embodiments, the genetically modified polypeptide includes an amino acid sequence listed in Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes a linker sequence listed in Table T1, or a linker containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes an RT domain sequence listed in Table T1, or an RT domain containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes (i) a linker comprising a linker sequence listed in the row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (ii) an RT domain comprising an RT domain sequence listed in the same row of Table T1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0142] [Table 1]

[0143] In some embodiments, the genetically modified polypeptide includes an amino acid sequence listed in Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes a linker sequence listed in Table T2, or a linker containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes an RT domain sequence listed in Table T2, or an RT domain containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide includes (i) a linker comprising a linker sequence listed in the row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto, and (ii) an RT domain comprising an RT domain sequence listed in the same row of Table T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0144] [Table 2-1] [Table 2-2]

[0145] Exemplary example: Subsequences of genetically modified polypeptides In some embodiments, the genetically modified polypeptide comprises, in N-terminal to C-terminal order, one or more of the following (e.g., all 1, 2, 3, 4, 5, or 6): an N-terminal methionine residue, a first nuclear localization signal (NLS), a DNA-binding domain, a linker, an RT domain, and / or a second NLS. In some embodiments, the genetically modified polypeptide comprises, in N-terminal to C-terminal order, an NLS (e.g., a first NLS), a DNA-binding domain, a linker, and an RT domain, wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide comprises, in N-terminal to C-terminal order, a DNA-binding domain, a linker, an RT domain, and an NLS (e.g., a second NLS), wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide comprises, in N-terminal to C-terminal order, a first NLS, a DNA-binding domain, a linker, an RT domain, and a second NLS, wherein the linker and RT domains are amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker and RT domains of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.

[0146] In some embodiments, the genetically modified polypeptide comprises, in N-terminus to C-terminus order, an N-terminal methionine residue, a first nuclear localization signal (NLS) (e.g., one of the NLSs of SEQ ID NOs. 1 to 7743 and / or the NLSs of the genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto), and a DNA-binding domain (e.g., a Cas domain, e.g., a SpyCas domain). 9 domains, for example, the SpyCas9 domains listed in SEQ ID NOs. 8331-8369, or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto; or any one of SEQ ID NOs. 1-7743 and / or the DNA-binding domain of a genetically modified polypeptide listed in either Table T1 or T2, or having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. (An amino acid sequence), linker (for example, one of sequence numbers 1 to 7743 and / or a linker of a genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it), RT domain (for example, one of sequence numbers 1 to 7743 and / or the RT domain of a genetically modified polypeptide listed in either Table T1 or T2, or it It comprises an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the first one, and one or more second NLSs (e.g., one of any of sequence numbers 1 to 7743, and / or the second NLSs of any of the genetically modified polypeptides listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the first one) (e.g., all 1, 2, 3, 4, 5, or 6).In some embodiments, the genetically modified polypeptide further comprises a T2A sequence and / or a puromycin sequence (e.g., at the C-terminus of a second NLS) (e.g., the T2A sequence and / or puromycin sequence of any one of SEQ ID NOs: 1 to 7743 and / or any of the genetically modified polypeptides listed in Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In some embodiments, the nucleic acid encoding the genetically modified polypeptide (e.g., as described herein) encodes a T2A sequence, for example, where the T2A sequence is located between the region encoding the genetically modified polypeptide and a second region, where the second region optionally encodes a selectable marker, such as puromycin.

[0147] In certain embodiments, the first NLS includes a first NLS sequence of a genetically modified polypeptide having one of the amino acid sequences of SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the first NLS includes a first NLS sequence of a genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the first NLS sequence includes a C-myc NLS. In certain embodiments, the first NLS includes the amino acid sequence PAAKRVKLD (SEQ ID NO: 8277), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0148] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the first NLS and the DNA-binding domain. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the first NLS and the DNA-binding domain comprises the amino acid sequence GG.

[0149] In certain embodiments, the DNA-binding domain includes the DNA-binding domain of any one of the genetically modified polypeptides listed in SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the DNA-binding domain includes the DNA-binding domain of any genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the DNA-binding domain includes a Cas domain (e.g., listed in SEQ ID NOs: 8331 to 8369). In certain embodiments, the DNA-binding domain includes the amino acid sequence of a SpyCas9 polypeptide (e.g., the SpyCas9 polypeptides listed in SEQ ID NOs. 8331-8369, e.g., Cas9 N863A polypeptide), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the DNA-binding domain includes the amino acid sequence:

[0150] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the DNA-binding domain and the linker. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the DNA-binding domain and the linker comprises the amino acid sequence GG.

[0151] In certain embodiments, the linker includes the linker sequence of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the linker includes the linker sequence of a genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the linker includes the amino acid sequences listed in Table 10, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0152] [Table 3-1] [Table 3-2] [Table 3-3]

[0153] In some embodiments, the linker of the genetically modified polypeptide is (SGGS) n (Sequence ID 5025), (GGGS) n (Sequence ID 5026), (GGGGS) n (Sequence ID 5027), (G) n (EAAAK) n(Sequence ID 5028), (GGS) n or (XP) n Includes motifs selected from.

[0154] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the linker and the RT domain. In certain embodiments, the spacer sequence between the linker and the RT domain comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the linker and the RT domain comprises the amino acid sequence GG.

[0155] In certain embodiments, the RT domain includes the RT domain sequence of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the RT domain includes the RT domain sequence of any of the genetically modified polypeptides listed in Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the RT domain includes the amino acid sequence of any one of the SEQ ID NOs: 8001 to 8136, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In some embodiments, the RT domain has a length of approximately 400-500, 500-600, 600-700, 700-800, 800-900, or 900-1000 amino acids.

[0156] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the RT domain and the second NLS. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the RT domain and the second NLS comprises the amino acid sequence AG.

[0157] In a particular embodiment, the second NLS includes a second NLS sequence of any one genetically modified polypeptide from sequence numbers 1 to 7743. In a particular embodiment, the second NLS includes a second NLS sequence of a genetically modified polypeptide listed in either Table T1 or T2. In a particular embodiment, the second NLS sequence includes a plurality of partial NLS sequences. In an embodiment, the NLS sequence, for example, the second NLS sequence, includes a first partial NLS sequence, for example, an NLS sequence containing the amino acid sequence KRTADGSEFE (sequence number 8198), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In an embodiment, the NLS sequence, for example, the second NLS sequence, includes a second partial NLS sequence. In one embodiment, the NLS sequence, for example, the second NLS sequence, includes an SV40A5 NLS, for example, a double SV40A5 NLS, for example, an NLS containing the amino acid sequence KRTADGSEFESPKKKAKVE (SEQ ID NO: 8199), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In one particular embodiment, the NLS sequence, for example, the second NLS sequence, includes the amino acid sequence KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 8200), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it.

[0158] In certain embodiments, the genetically modified polypeptide further comprises a spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acids. In certain embodiments, the spacer sequence between the second NLS and the T2A sequence and / or the puromycin sequence comprises the amino acid sequence GSG.

[0159] Linker and RT domain In some embodiments, the genetically modified polypeptide includes a linker (as described herein, for example) and an RT domain (as described herein, for example). In certain embodiments, the genetically modified polypeptide includes a linker (as described herein, for example) and an RT domain (as described herein, for example) in the order of N-terminus to C-terminus.

[0160] In a particular embodiment, the linker includes an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker sequences listed in Table 10. In a particular embodiment, the linker includes one linker sequence from SEQ ID NOs. 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the linker includes one linker sequence from SEQ ID NOs. 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the linker includes one linker sequence from SEQ ID NOs. 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the linker includes a linker sequence of an exemplary genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes an RT domain sequence having an amino acid sequence selected from SEQ ID NOs. 8001 to 8136, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In certain embodiments, the RT domain includes an RT domain sequence of an exemplary genetically modified polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0161] In some embodiments, the genetically modified polypeptide comprises a portion of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743, wherein the portion comprises a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the portion.

[0162] In some embodiments, the genetically modified polypeptide includes a linker of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the genetically modified polypeptide includes a linker of any one of the genetically modified polypeptides SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the genetically modified polypeptide includes a linker of any one of the genetically modified polypeptides SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the linker. In some embodiments, the genetically modified polypeptide includes a linker of a genetically modified polypeptide listed in either Table T1 or T2, or a linker containing an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with such a linker.

[0163] In some embodiments, the genetically modified polypeptide includes the RT domain of any one of the genetically modified polypeptides SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the genetically modified polypeptide includes the RT domain of any one of the genetically modified polypeptides SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the genetically modified polypeptide includes the RT domain of any one of the genetically modified polypeptides SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the RT domain. In some embodiments, the genetically modified polypeptide includes an RT domain of a genetically modified polypeptide listed in either Table T1 or T2, or an RT domain having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with such an amino acid sequence.

[0164] In certain embodiments, the linker and RT domains of the genetically modified polypeptide include an amino acid sequence of the linker and RT domain of a genetically modified polypeptide having one amino acid sequence of any one of SEQ ID NOs: 1 to 7743 (or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In certain embodiments, the linker and RT domains of the genetically modified polypeptide include an amino acid sequence of the linker and RT domain having at least 80% identity with one of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include an amino acid sequence of the linker and RT domain having at least 90% identity with one of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include an amino acid sequence of the linker and RT domain having at least 95% identity with one of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains having at least 99% identity with any one of the linker and RT domains of SEQ ID NOs: 1 to 7743. In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains of genetically modified polypeptides having any one of the amino acid sequences of SEQ ID NOs: 6001 to 7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them). In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of linker and RT domains of genetically modified polypeptides having any one of the amino acid sequences of SEQ ID NOs: 4501 to 4541 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with them).In a particular embodiment, the linker and RT domain of the genetically modified polypeptide include amino acid sequences of the linker and RT domain from a single row in either Table T1 or T2 (for example, amino acid sequences from a single exemplary genetically modified polypeptide listed in either Table T1 or T2) (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).

[0165] In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains from two different amino acid sequences selected from SEQ ID NOs: 1 to 7743 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto). In certain embodiments, the linker and RT domains of the genetically modified polypeptide include amino acid sequences of the linker and RT domains from a different row in either Table T1 or T2 (or amino acid sequences having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto).

[0166] In certain embodiments, the genetically modified polypeptide further comprises a first NLS (e.g., 5'NLS) as described herein, for example. In certain embodiments, the genetically modified polypeptide further comprises a second NLS (e.g., 3'NLS) as described herein, for example. In certain embodiments, the genetically modified polypeptide further comprises an N-terminal methionine residue.

[0167] RT family and mutants In certain embodiments, the genetically modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, XMRV6, BLVAU, BLVJ, HTL1A, HTL1C, HTL1L, HTL32, HTL3P, ​​HTLV2, JSRV, MLVF5, MLVRD, MMTVB, MPMV, SSFVCP, SMRVH, SRV1, SRV2, and WDSV. In certain embodiments, the genetically modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.

[0168] In certain embodiments, the genetically modified polypeptide includes the amino acid sequence of the RT domain sequence from the MLVMS RT domain. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations listed in column 1 of Table M1, or corresponding point mutations. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations (Gen1 MLVMS) listed in column 3 of Table M1, or corresponding point mutations. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations at the amino acid positions of the RT domain listed in columns 1 and 2 of Table M2, or corresponding amino acid positions.

[0169] In certain embodiments, the genetically modified polypeptide includes the amino acid sequence of the RT domain sequence from the AVIRE RT domain. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations listed in column 2 of Table M1, or corresponding point mutations. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations (Gen2 AVIRE) listed in column 4 of Table M1, or corresponding point mutations. In embodiments, the amino acid sequence of the RT domain sequence includes one or more point mutations at the amino acid positions of the RT domain listed in columns 3 and 4 of Table M2, or corresponding amino acid positions. In certain embodiments, the RT domain includes IENSSP (e.g., at the C-terminus).

[0170] [Table 4]

[0171] [Table 5]

[0172] In certain embodiments, the genetically modified polypeptide includes an RT domain derived from a gamma retrovirus. In certain embodiments, the gamma retrovirus-derived RT domain of the genetically modified polypeptide includes an amino acid sequence of an RT domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6. In some embodiments, the gamma retrovirus-derived RT domain of the genetically modified polypeptide does not originate from PERV. In some embodiments, the RT includes one, two, three, four, five, or six or more mutations corresponding to the D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, E67K, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, H8Y, T306K, or D653N mutations in the RT domain of mouse leukemia virus reverse transcriptase. In some embodiments, the genetically modified polypeptide further includes a linker having at least 99% identity with one of the linker domains of SEQ ID NOs: 1 to 7743. In some embodiments, the genetically modified polypeptide further includes a linker having at least 99% or 100% identity with SEQ ID NO: 5217.

[0173] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of AVIRE RT (e.g., AVIRE_P03360 sequence, e.g., SEQ ID NO: 8001), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of AVIRE RT, further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, G330P, L605W, T306K, and W313F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of AVIRE RT, further comprising one, two, or three mutations selected from the group consisting of D200N, G330P, and L605W or corresponding positions in the same RT domain.

[0174] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of BAEVM RT (e.g., BAEVM_P10272 sequence, e.g., SEQ ID NO: 8004), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of BAEVM RT further comprising one, two, three, four, or five mutations selected from the group consisting of D198N, E328P, L602W, T304K, and W311F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of BAEVM RT further comprising one, two, or three mutations selected from the group consisting of D198N, E328P, and L602W or corresponding positions in the same RT domain.

[0175] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FFV RT (e.g., the FFV_O93209 sequence, e.g., SEQ ID NO: 8012), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further comprising one, two, three, or four mutations selected from the group consisting of D21N, T293N, T419P, and L393K or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further comprising one, two, or three mutations selected from the group consisting of D21N, T293N, and T419P or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of FFV RT further comprising the D21N mutation. In some embodiments, the RT domain comprises an amino acid sequence of FFV RT further comprising one, two, or three mutations selected from the group consisting of T207N, T333P, and L307K or corresponding positions in the same RT domain.

[0176] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FLV RT (e.g., FLV_P10273 sequence, e.g., SEQ ID NO: 8019), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FLV RT further comprising one, two, three, or four mutations selected from the group consisting of D199N, L602W, T305K, and W312F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of FLV RT further comprising one or two mutations selected from the group consisting of D199N and L602W or corresponding positions in the same RT domain.

[0177] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of FOAMV RT (e.g., the FOAMV_P14350 sequence, e.g., SEQ ID NO: 8021), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further comprising one, two, three, or four mutations selected from the group consisting of D24N, T296N, S420P, and L396K or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further comprising one, two, or three mutations selected from the group consisting of D24N, T296N, and S420P or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of FOAMV RT further comprising the mutant D24N or a corresponding position in the same RT domain. In some embodiments, the RT domain comprises an amino acid sequence of FOAMV RT further comprising one, two, or three mutations selected from the group consisting of T207N, S331P, and L307K or corresponding positions in the same RT domain.

[0178] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of GALV RT (e.g., the GALV_P21414 sequence, e.g., SEQ ID NO: 8027), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further comprising one, two, three, four, or five mutations selected from the group consisting of D198N, E328P, L600W, T304K, and W311F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further comprising one, two, or three mutations selected from the group consisting of D198N, E328P, and L600W or corresponding positions in the same RT domain.

[0179] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of KORV RT (e.g., KORV_Q9TTC1 sequence, e.g., SEQ ID NO: 8047), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further comprising one, two, three, four, five, or six mutations selected from the group consisting of D32N, D322N, E452P, L274W, T428K, and W435F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further comprising one, two, three, or four mutations selected from the group consisting of D32N, D322N, E452P, and L274W or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of GALV RT further comprising the D32N mutation. In some embodiments, the RT domain comprises an amino acid sequence of KORV RT further comprising one, two, three, four, or five mutations selected from the group consisting of D231N, E361P, L633W, T337K, and W344F or corresponding positions in the same RT domain. In some embodiments, the RT domain comprises an amino acid sequence of KORV RT further comprising one, two, or three mutations selected from the group consisting of D231N, E361P, and L633W or corresponding positions in the same RT domain.

[0180] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVAV RT (e.g., the MLVAV_P03356 sequence, e.g., SEQ ID NO: 8053), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVAV RT, further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of MLVAV RT, further comprising one, two, or three mutations selected from the group consisting of D200N, T330P, and L603W or corresponding positions in the same RT domain.

[0181] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVBM RT (e.g., the MLVBM_Q7SVK7 sequence, e.g., SEQ ID NO: 8056), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVBM RT further comprising one, two, three, four, or five mutations selected from the group consisting of D199N, T329P, L602W, T305K, and W312F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of MLVBM RT further comprising one, two, and three mutations selected from the group consisting of D200N, T330P, and L603W or corresponding positions in the same RT domain.

[0182] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVCB RT (e.g., MLVCB_P08361 sequence, e.g., SEQ ID NO: 8062), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVCB RT further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of MLVCB RT further comprising one, two, and three mutations selected from the group consisting of D200N, T330P, and L603W or corresponding positions in the same RT domain.

[0183] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVFF RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVFF RT, further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of MLVFF RT, further comprising one, two, and three mutations selected from the group consisting of D200N, T330P, and L603W or corresponding positions in the same RT domain.

[0184] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of MLVMS RT (e.g., MLVMS_reference sequence, e.g., SEQ ID NO: 8370; or MLVMS_P03355 sequence, e.g., SEQ ID NO: 8070), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of MLVMS RT, further comprising one, two, three, four, five, or six mutations selected from the group consisting of D200N, T330P, L603W, T306K, W313F, and H8Y, or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of MLVMS RT, further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F, or corresponding positions in the same RT domain. In some embodiments, the RT domain comprises an amino acid sequence of MLVMS RT further comprising one, two, or three mutations selected from the group consisting of D200N, T330P, and L603W or from corresponding positions in the same type of RT domain.

[0185] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of PERV RT (e.g., the PERV_Q4VFZ2 sequence, e.g., SEQ ID NO: 8099), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of PERV RT further comprising one, two, three, four, or five mutations selected from the group consisting of D196N, E326P, L599W, T302K, and W309F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of PERV RT further comprising one, two, or three mutations selected from the group consisting of D196N, E326P, and L599W or corresponding positions in the same RT domain.

[0186] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of SFV1 RT (e.g., SFV1_P23074 sequence, e.g., SEQ ID NO: 8105), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further comprising one, two, three, or four mutations selected from the group consisting of D24N, T296N, N420P, and L396K or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further comprising one, two, or three mutations selected from the group consisting of D24N, T296N, and N420P or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of SFV1 RT further comprising D24N or corresponding positions in the same RT domain.

[0187] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of SFV3L RT (e.g., the SFV3L_P27401 sequence, e.g., SEQ ID NO: 8111), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further comprising one, two, three, or four mutations selected from the group consisting of D24N, T296N, N422P, and L396K or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further comprising one, two, or three mutations selected from the group consisting of D24N, T296N, and N422P or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of SFV3L RT further comprising the mutant D24N or a corresponding position in the same RT domain. In some embodiments, the RT domain comprises an amino acid sequence of SFV3L RT further comprising one, two, or three mutations selected from the group consisting of T307N, N333P, and L307K or corresponding positions in the same RT domain.

[0188] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of WMSV RT (e.g., WMSV_P03359 sequence, e.g., SEQ ID NO: 8131), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of WMSV RT, further comprising one, two, three, four, or five mutations selected from the group consisting of D198N, E328P, L600W, T304K, and W311F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of WMSV RT, further comprising one, two, or three mutations selected from the group consisting of D198N, E328P, and L600W or corresponding positions in the same RT domain.

[0189] In some embodiments, the RT domain includes the amino acid sequence of the RT domain of XMRV6 RT (e.g., XMRV6_A1Z651 sequence, e.g., SEQ ID NO: 8134), or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the RT domain includes the amino acid sequence of XMRV6 RT further comprising one, two, three, four, or five mutations selected from the group consisting of D200N, T330P, L603W, T306K, and W313F or corresponding positions in the same RT domain. In some embodiments, the RT domain includes the amino acid sequence of XMRV6 RT further comprising one, two, or three mutations selected from the group consisting of D200N, T330P, and L603W or corresponding positions in the same RT domain.

[0190] In certain embodiments, the RT domain of the genetically modified polypeptide includes the amino acid sequence of the RT domain of AVIRE RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the RT domain includes the amino acid sequence of the RT domain contained in the sequences listed in column 1 of Table A5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide further includes a linker having at least 99% or 100% identity with SEQ ID NO: 5217.

[0191] In certain embodiments, the RT domain of the genetically modified polypeptide includes the amino acid sequence of the RT domain of MLVMS RT, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In embodiments, the RT domain includes the amino acid sequence of the RT domain contained in any of the sequences listed in columns 2 to 6 of Table A5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto. In some embodiments, the genetically modified polypeptide further includes a linker having at least 99% or 100% identity with SEQ ID NO: 5217.

[0192] [Table 6-1] [Table 6-2] [Table 6-3] [Table 6-4] [Table 6-5] [Table 6-6]

[0193] system In one embodiment, the disclosure relates to a system comprising a nucleic acid molecule encoding a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., a template RNA, e.g., as described herein). In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide contains one or more silent mutations in its coding region (e.g., in the sequence encoding the RT domain) compared to the nucleic acid molecules described herein. In a particular embodiment, the system further comprises a gRNA (e.g., a gRNA that binds to the inducing polypeptide, e.g., on the opposite strand of the target DNA to which the genetically modified polypeptide binds).

[0194] In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide encodes a polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0195] In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, which includes a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the aforementioned portion. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, which includes a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the aforementioned portion. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, which includes a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the aforementioned portion. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide includes a portion of the polypeptide listed in either Table T1 or T2, which includes a linker and an RT domain, or a sequence encoding an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identical to the aforementioned portion.

[0196] In certain embodiments, the nucleic acid molecule encoding a genetically modified polypeptide includes a linker of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the nucleic acid molecule encoding a genetically modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the nucleic acid molecule encoding a genetically modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the nucleic acid molecule encoding a genetically modified polypeptide includes a linker of a polypeptide listed in either Table T1 or T2, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with such a linker.

[0197] In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide includes an RT domain of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the nucleic acid molecule encoding the genetically modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the nucleic acid molecule encoding the genetically modified polypeptide includes the RT domain of a polypeptide listed in either Table T1 or T2, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0198] In one embodiment, the disclosure relates to a system comprising a genetically modified polypeptide (e.g., as described herein) and a template nucleic acid (e.g., template RNA, e.g., as described herein).

[0199] In certain embodiments, the genetically modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the genetically modified polypeptide comprises a polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0200] In certain embodiments, the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, which includes a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the aforementioned portion. In certain embodiments, the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, which includes a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the aforementioned portion. In certain embodiments, the genetically modified polypeptide includes a portion of an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, which includes a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the aforementioned portion. In a particular embodiment, the genetically modified polypeptide is a part of a polypeptide listed in either Table T1 or T2, which includes a linker and an RT domain, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with the aforementioned part.

[0201] In certain embodiments, the genetically modified polypeptide includes a linker of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes a linker of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the genetically modified polypeptide includes a linker of a polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with such a polypeptide.

[0202] In certain embodiments, the genetically modified polypeptide includes an RT domain of an amino acid sequence selected from SEQ ID NOs: 1 to 7743, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 6001 to 7743, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In certain embodiments, the genetically modified polypeptide includes an RT domain of a polypeptide having an amino acid sequence selected from SEQ ID NOs: 4501 to 4541, or a sequence encoding an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity with it. In a particular embodiment, the genetically modified polypeptide includes the RT domain of a polypeptide listed in either Table T1 or T2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, or 99% identity thereto.

[0203] DNA modification system Furthermore, a system for modifying DNA is provided herein. The gene modification system may include, for example, (a) a gene-modified polypeptide or a nucleic acid molecule encoding a gene-modified polypeptide, wherein the gene-modified polypeptide comprises (i) a reverse transcriptase (RT) domain and an endonuclease domain having DNA-binding functionality, or an endonuclease domain and a separate DNA-binding domain, and (b) a template RNA.

[0204] Accordingly, a system for modifying DNA is provided herein. The system includes, for example, (a) an artificial ribonucleic acid (RNA) molecule for enhancing the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising (1) a nucleotide sequence encoding a polypeptide, (2)(i) a 3'-untranslated region (3'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) a 5'-untranslated region (5'UTR) element for enhancing polypeptide expression, comprising SEQ ID NOs: 8212-8247 and An artificial RNA molecule comprising at least one 5'UTR element containing a nucleic acid sequence selected from the group consisting of 8259-8263, and / or (3) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282-8314; and (b) (e.g., 5' to 3') (i) optionally a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) a sequence that binds to a polypeptide, (iii) a heterologous object sequence, and (iv) optionally a template RNA (or DNA encoding the template RNA) containing a 3' target homologous domain. The heterologous object sequence may include, for example, modifications compared to the corresponding original sequence (e.g., wild-type sequence), where the modifications improve the speed, fidelity, or speed and fidelity of reverse transcription primed by the target by reverse transcriptase.

[0205] In a particular embodiment, the system is, for example, (a) an artificial ribonucleic acid (RNA) molecule for enhancing the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising (1) a nucleotide sequence encoding the polypeptide, (2)(i) a 3'-untranslated region (3'UTR) element for enhancing polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) a 5'-untranslated region (5'UTR) element for enhancing polypeptide expression, comprising SEQ ID NO. 821 An artificial nucleic acid molecule comprising at least one 5'UTR element containing a nucleic acid sequence selected from the group consisting of 2-8247 and 8259-8263, and / or (3) a poly(A) tail for enhancing polypeptide expression, the poly(A) tail containing a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282-8314; and (b) (e.g., 5' to 3') (i) optionally a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) a sequence that binds to a polypeptide, (iii) a heterogeneous object sequence, and (iv) optionally a template RNA (or DNA encoding the template RNA) containing a 3' target homologous domain.Preferably, heterogeneous object sequences have one, two, or all of the following characteristics: i) they do not contain self-complementary sequences, e.g., self-complementary sequences that form a hairpin structure under stringent conditions, for example; or, if self-complementary sequences are present, they have the following characteristics: (1) each self-complementary sequence is no longer than 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides; (2) the self-complementary sequences form a hairpin including an arm of no longer than 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides; or (3) the self-complementary sequences include at least 1, 2, 3, 4, or 5 positions that are incomplementary (e.g., mismatched or bulge) with their partner sequence; and (4) they do not contain repetitive sequences (e.g., single, di, or trinucleotide repetitive sequences); or, if repetitive sequences are present, they are sequences of no longer than 12, 11, 10, 9, 8, 7, or 6 nucleotides.

[0206] Template nucleic acid or RNA The gene modification systems described herein can modify host target DNA sites by using template nucleic acids. In some embodiments, the gene modification systems described herein transcribe an RNA sequence template into a host target DNA site by target-primed reverse transcription (TPRT). By modifying DNA sequences(s) through direct reverse transcription of RNA sequence templates into the host genome, the gene modification systems can insert object sequences into the target genome without requiring the introduction of exogenous DNA sequences into host cells (unlike, for example, CRISPR systems), thus eliminating the exogenous DNA insertion step. The gene modification systems can also use object sequences to delete or introduce substitutions from the target genome. Therefore, the gene modification systems provide a platform for the use of customized RNA sequence templates containing object sequences, such as sequences containing heterologous genetic codes and / or functional information.

[0207] In some embodiments, the template nucleic acid includes one or more sequences (e.g., two sequences) that bind to the genetically modified polypeptide.

[0208] In some embodiments, the template nucleic acid includes RNA. In some embodiments, the template nucleic acid includes DNA (e.g., single-stranded or double-stranded DNA).

[0209] In some embodiments, the template nucleic acid includes one or more (e.g., two) homologous domains having homology to the target sequence. In some embodiments, the homologous domains are approximately 10–20, 20–50, or 50–100 nucleotides in length.

[0210] In some embodiments, the template RNA may include a gRNA sequence for directing, for example, a genetically modified polypeptide to a target site of interest. In some embodiments, the template RNA may include (i) a gRNA spacer (e.g., 5' to 3') optionally binding to a target site (e.g., the second strand of a site in the target genome), (ii) a gRNA scaffold (e.g., a genetically modified polypeptide or Cas polypeptide) optionally binding to a polypeptide as described herein, (iii) a heterologous object sequence containing a mutant region (optionally, the heterologous object sequence includes a first homologous region, a mutant region, and a second homologous region from 5' to 3'), and (iv) a primer-binding site (PBS) sequence containing a 3' target homologous domain.

[0211] The template nucleic acid (e.g., template RNA) component of the genome editing systems described herein is typically capable of binding to the system's modified polypeptide. In some embodiments, the template nucleic acid (e.g., template RNA) has a 3' region capable of binding to the modified polypeptide. The binding region, e.g., the 3' region, may be a structured RNA region capable of binding to the system's modified polypeptide, e.g., an RNA region having at least one, two, or three hairpin loops. The binding region may associate the template nucleic acid (e.g., template RNA) with any of the polypeptide modules. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with an RNA-binding domain within the polypeptide. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with the reverse transcription domain of the genome-modified polypeptide (e.g., specifically binding to the RT domain). In some embodiments, the template nucleic acid (e.g., template RNA) may associate with a DNA-binding domain of the polypeptide, e.g., gRNA associates with a Cas9-derived DNA-binding domain. In some embodiments, the binding region may also provide DNA target recognition, for example, by hybridizing the gRNA to a target DNA sequence and binding to a polypeptide, such as a Cas9 domain. In some embodiments, a template nucleic acid (e.g., template RNA) may associate with multiple components of a polypeptide, such as a DNA-binding domain and a reverse transcription domain.

[0212] In some embodiments, the template nucleic acid is template RNA. In some embodiments, the template RNA contains one or more modified nucleotides. For example, in some embodiments, the template RNA contains one or more deoxyribonucleotides. In some embodiments, a region of the template RNA is replaced with DNA nucleotides, for example, to improve molecular stability. For example, the 3' end of the template may contain DNA nucleotides, while the remainder of the template contains RNA nucleotides that can be reverse transcribed. For example, in some embodiments, the heterologous object sequence consists mainly or entirely of RNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% RNA nucleotides). In some embodiments, the PBS sequence consists mainly or entirely of DNA nucleotides (e.g., at least 90%, 95%, 98%, or 99% DNA nucleotides). In other embodiments, the heterologous object sequence for writing into the genome may contain DNA nucleotides. In some embodiments, the DNA nucleotides in the template are copied into the genome by a domain capable of performing DNA-dependent DNA polymerase activity. In some embodiments, DNA-dependent DNA polymerase activity is provided by a DNA polymerase domain within the polypeptide. In some embodiments, DNA-dependent DNA polymerase activity is provided by a reverse transcriptase domain that also has the ability to perform DNA-dependent DNA polymerization, such as the synthesis of a second strand. In some embodiments, the template molecule consists solely of DNA nucleotides.

[0213] In some embodiments, the system described herein comprises two nucleic acids, each containing together the sequence of the template RNA described herein. In some embodiments, the two nucleic acids are non-covalently linked to one another, for example, directly linked to one another (e.g., via base pairing), or indirectly linked as part of a complex containing one or more additional molecules.

[0214] The template RNA described herein may include (1) a gRNA spacer, (2) a gRNA scaffold, (3) a heterogeneous object sequence, and (4) a primer-binding site (PBS) sequence from 5' to 3'.

[0215] As described herein, a gRNA spacer can direct a genetic modification system to a target nucleic acid, and a gRNA scaffold can facilitate the association of a template RNA with the Cas domain of a genetically modified polypeptide, thereby enabling editing of the target sequence. In certain embodiments, a gRNA comprising a gRNA spacer and a gRNA scaffold, but without a heterogeneous object sequence or PBS sequence, may be used, for example, to induce nick formation of a second strand.

[0216] As described herein, heterologous object sequences can be used as templates for genetically modified polypeptides for reverse transcription to write a desired sequence into a target nucleic acid. In some embodiments, the heterologous object sequence includes a 5' to 3' post-edit homologous region, a mutation region, and a pre-edit homologous region. Though not bound by theory, a reverse transcription (RT) performing reverse transcription on the template RNA would first reverse transcribe the pre-edit homologous region, then the mutation region, and then the post-edit homologous region, thereby creating a DNA strand containing the desired mutation with homologous regions on both sides.

[0217] As described herein, the template nucleic acid may include, for example, a PBS sequence. In some embodiments, the PBS sequence is located at 3' of a heterogeneous object sequence and contains only one, two, three, four, or five or fewer mismatches with sequences adjacent to the site modified by the system described herein, or sequences complementary to sequences adjacent to the site modified by the system / genetic modification polypeptide. In some embodiments, the PBS sequence binds within one, two, three, four, five, six, seven, eight, nine, or ten nucleotides of the nic site in the target nucleic acid molecule. In some embodiments, the binding of the PBS sequence to the target nucleic acid molecule enables the initiation of TPRT, for example, by having the 3' homologous domain act as a primer for TPRT primed by the target. In some embodiments, the PBS sequence has lengths of 3-5, 5-10, 10-30, 10-25, 10-20, 10-19, 10-18, 10-17, 10-16, 10-15, 10-14, 10-13, 10-12, 10-11, 11-30, 11-25, 11-20, 11-19, 11-18, 11 ~17, 11~16, 11~15, 11~14, 11~13, 11~12, 12~30, 12~25, 12~20, 12~19, 12~18, 12~17, 12~16, 12~15, 12~14, 12~13, 13~30, 13~25, 13~20, 13~19, 13~18, 13~17, 13~16 , 13-15, 13-14, 14-30, 14-25, 14-20, 14-19, 14-18, 14-17, 14-16, 14-15, 15-30, 15-25, 15-20, 15-19, 15-18, 15-17, 15-16, 16-30, 16-25, 16-20, 16-19, 16-18, 16 The PBS sequence has lengths of ~17, 17~30, 17~25, 17~20, 17~19, 17~18, 18~30, 18~25, 18~20, 18~19, 19~30, 19~25, 19~20, 20~30, 20~25, or 25~30 nucleotides, for example, lengths of 10~17, 12~16, or 12~14 nucleotides. In some embodiments, the PBS sequence has lengths of 5~20, 8~16, 8~14, 8~13, 9~13, 9~12, or 10~12 nucleotides, for example, lengths of 9~12 nucleotides.

[0218] Template nucleic acids and template RNAs are described in WO2021 / 248102, which is incorporated herein by reference in its entirety.

[0219] In a particular embodiment, the template RNA further comprises a reverse transcriptase (RT) terminator sequence located between a heterogeneous object sequence and either (i) or (ii).

[0220] In a particular embodiment, the heterogeneous object sequence includes a sequence that encodes the target polypeptide or a portion thereof, or a sequence that is the inverse complementary chain of the sequence encoding the target polypeptide or a portion thereof.

[0221] In a particular embodiment, the polypeptide comprises a reverse transcriptase domain and an endonuclease domain, the endonuclease domain being a Cas9 domain, and the template RNA comprises (i) a gRNA spacer complementary to a first portion of the target gene, optionally comprising one or more consecutive nucleotides beginning at the 3' end of an adjacent nucleotide of the gRNA spacer, (ii) a gRNA scaffold that binds to the Cas9 domain, (iii) a heterologous object sequence comprising a mutation region for introducing a mutation into a second portion of the target gene (e.g., to correct a mutation therein) (optionally, the heterologous sequence comprises a 5' to 3' post-edit homologous region, a mutation region and a pre-edit homologous region), and (iv) a primer-binding site (PBS) sequence comprising at least 5, 6, 7 or 8 bases having 100% identity with a third portion of the target gene.

[0222] Furthermore, template RNA sequences are provided in Tables 1A-1D, 5A-5F, 8A-8D, E3, E3A, BB, E5, E5A, E6 and E6A of WO2023039435, which are incorporated herein by reference in their entirety.

[0223] In a particular embodiment, the target gene is a human PAH gene, and the template RNA is (i) a gRNA spacer complementary to the first portion of the human PAH gene, preferably having a sequence containing the core nucleotide of a gRNA spacer sequence from Table 1A, Table 1B, Table 1C or Table 1D of WO2023039435 incorporated herein by reference in whole, and optionally containing one or more consecutive nucleotides starting at the 3' end of an adjacent nucleotide of the gRNA spacer, or Table 5A-5F, 8A-8D, E3, E3A, BB, E5, E5A, E (ii) a gRNA spacer having a spacer sequence selected from 6 or E6A; (ii) a gRNA scaffold that binds to the Cas9 domain; (iii) a heterologous object sequence containing a mutation region for introducing a mutation into the second portion of the human PAH gene (for example, to correct a mutation therein) (the heterologous object sequence may include a 5' to 3' post-edit homologous region, a mutation region and a pre-edit homologous region); and (iv) a primer-binding site (PBS) sequence containing at least 5, 6, 7 or 8 bases having 100% identity with the third portion of the human PAH gene.

[0224] In a particular embodiment, the template RNA consists of the sequence of SEQ ID NO: 8258 (RNACS7570), SEQ ID NO: 8320 (RNACS229), SEQ ID NO: 8321 (RNACS1515), or SEQ ID NO: 8372.

[0225] In a particular embodiment, the target site is a target site within the human genome.

[0226] therapeutic application By incorporating coding genes into RNA sequence templates, the system can address therapeutic needs, for example, by providing the expression of therapeutic transgenes in individuals with loss-of-function mutations, by replacing gain-of-function mutations with normal transgenes, by providing regulatory sequences for deleting gain-of-function mutation expression, and / or by controlling the expression of operably linked genes, transgenes, and the system. In certain embodiments, the RNA sequence template encodes a promoter region specific to the therapeutic needs of the host cell, such as a tissue-specific promoter or enhancer. In yet other embodiments, the promoter may be operably linked to the coding sequence.

[0227] In some embodiments, the systems described herein may be used to create insertions, deletions, substitutions, or combinations thereof in cells, tissues, or subjects. In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation). In some embodiments, insertions, deletions, substitutions, or combinations thereof increase or decrease gene expression (e.g., transcription or translation) by altering, adding, or deleting sequences in sequences that bind to promoters or enhancers, such as transcription factors. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene translation (e.g., alter amino acid sequences), insert or delete start or stop codons, or alter or modify gene translation frames.

[0228] In some embodiments, insertions, deletions, substitutions, or combinations thereof alter gene splicing, for example, by inserting, deleting, or modifying splice acceptors or donor sites. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter the half-life of a transcript or protein. In some embodiments, insertions, deletions, substitutions, or combinations thereof alter protein localization in cells (e.g., from cytoplasm to mitochondria, from cytoplasm to extracellular space (e.g., by adding secretory tags)). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter (e.g., improve) protein folding (e.g., to prevent the accumulation of misfolded proteins). In some embodiments, insertions, deletions, substitutions, or combinations thereof alter, increase, or decrease gene activity, for example, the protein encoded by the gene.

[0229] This disclosure, in part, relates to methods for modifying target sites in the genomic DNA of cells. In some embodiments, the method involves contacting cells with the systems described herein, template RNA, viruses, virus-like particles or viromosomes, or LNPs, or DNA encoding them, thereby modifying target sites in the genomic DNA of cells.

[0230] Therefore, a method is provided for modifying a target site in the genomic DNA of a cell. The method involves contacting a cell with the system of the present invention or one or more RNAs encoding the system of the present invention, thereby modifying a target site in the genomic DNA of the cell. In a particular embodiment, the cell is a T cell (e.g., a primary T cell).

[0231] This disclosure relates in part to methods for treating subjects having a disease or condition associated with a genetic defect. In some embodiments, the method includes administering to a subject a system, template RNA, virus, virus-like particle or virosome, or LNP, or DNA encoding them, thereby treating the subject having a disease or condition associated with a genetic defect. In some embodiments, the disease or condition associated with a genetic defect is an indication listed in any of Tables 9 to 12 of International Patent Application Publication No. 2021 / 178720, which is incorporated herein by reference in its entirety, including the aforementioned table, and / or an indication in which the genetic defect is a defect in a gene listed in any of Tables 9 to 12. In some embodiments, the subject is a human subject.

[0232] Therefore, a method is provided for treating subjects having a disease or condition associated with a genetic defect. The method comprises administering the system of the present invention to a subject, thereby treating the subject having a disease or condition associated with a genetic defect.

[0233] Accordingly, methods for treating phenylketonuria (PKU) or hyperphenylalaninemia (e.g., mild or severe hyperphenylalaninemia) in subjects requiring treatment are provided herein. In some embodiments, the treatment results in improvement of one or more symptoms associated with PKU or hyperphenylalaninemia, as shown in WO2023039435, which is incorporated herein by reference in its entirety.

[0234] In some embodiments, treatment with the gene modification system described herein results in one or more of the following compared to subjects having a PKU not treated with the gene modification system described herein: (a) increased phenylalanine hydroxylase (PAH) activity, efficiency and / or function; (b) decreased phenylalanine concentration in blood and / or cerebrospinal fluid; (c) increased tyrosine concentration in blood; (d) restoration of normal synthesis of dopamine, norepinephrine and / or melanin; (e) reduced urea production; and / or (f) improved protein retention and / or Phe utilization.

[0235] Administration and Delivery The compositions and systems described herein may be used in vitro or in vivo. In some embodiments, the system or components of the system are delivered to cells (e.g., mammalian cells, e.g., human cells) in vitro or in vivo. In some embodiments, the cells are eukaryotic cells, e.g., multicellular organisms, e.g., animals, e.g., mammals (e.g., humans, pigs, cattle), birds (e.g., poultry, e.g., chickens, turkeys, or ducks), or fish cells. In some embodiments, the cells are non-human animal cells (e.g., research animals, livestock, or companion animals). In some embodiments, the cells are stem cells (e.g., hematopoietic stem cells), fibroblasts, or T cells. In some embodiments, the cells are immune cells, e.g., T cells (e.g., Treg, CD4, CD8, γδ, or memory T cells), B cells (e.g., memory B cells or plasma cells), or NK cells. In some embodiments, the cells are non-dividing cells, e.g., non-dividing fibroblasts or non-dividing T cells.

[0236] In one embodiment, the system and / or components of the system are delivered as nucleic acids. For example, a genetically modified polypeptide may be delivered in the form of DNA or RNA encoding the polypeptide, and the template RNA may be delivered in the form of RNA or its complementary DNA transcribed into RNA. In some embodiments, the system or components of the system are delivered in one, two, three, four or more separate nucleic acid molecules. In some embodiments, the system or components of the system are delivered as a combination of DNA and RNA. In some embodiments, the system or components of the system are delivered as a combination of DNA and protein. In some embodiments, the system or components of the system are delivered as a combination of RNA and protein. In some embodiments, the genetically modified polypeptide is delivered as a protein.

[0237] In some embodiments, the system or its components are delivered to cells, such as mammalian or human cells, using a vector. The vector may be, for example, a plasmid or a virus. In some embodiments, delivery is performed in vivo, in vitro, ex vivo, or in situ. In some embodiments, the virus is an adeno-associated virus (AAV), lentivirus, or adenovirus. In some embodiments, the system or its components are delivered to cells by virus-like particles or viromosomes. In some embodiments, delivery uses two or more viruses, virus-like particles, or viromosomes.

[0238] In one embodiment, the compositions and systems described herein may be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicle structures consisting of a monolayer or multilayer lipid bilayer surrounding an internal aqueous compartment and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes may be anionic, neutral, or cationic. Liposomes are biocompatible, non-toxic, capable of delivering both hydrophilic and lipophilic drug molecules, protecting their cargoes from degradation by plasma enzymes, and transporting their cargoes across biological membranes and the blood-brain barrier (BBB) ​​(see, for example, Spuch and Navarro, Journal of Drug Delivery, Vol. 2011, Article No. 469679, p. 12, 2011. doi:10.1155 / 2011 / 469679 for an overview).

[0239] Vesicles may be produced from several different types of lipids, however, phospholipids are most commonly used to generate liposomes as drug carriers. Methods for the preparation of multilayer vesicle lipids are known in the art (see, for example, U.S. Patent No. 6,693,086, the teaching relating to the preparation of multilayer vesicle lipids is incorporated herein by reference). Vesicle formation may be spontaneous when a lipid film is mixed with an aqueous solution, but it can also be facilitated by applying force in the form of shaking using a homogenizer, sonicator, or extruder (see, for example, Spuch and Navarro, Journal of Drug Delivery, Vol. 2011, Article No. 469679, p. 12, 2011. doi:10.1155 / 2011 / 469679 for an overview). The extruded lipids can be prepared by extruding them through a size-reducing filter, as described in Templeton et al., Nature Biotech, 15:647-652, 1997, the teachings relating to the preparation of the extruded lipids are incorporated herein by reference.

[0240] Various nanoparticles, such as liposomes, lipid nanoparticles, cationic lipid nanoparticles, ionizable lipid nanoparticles, polymer nanoparticles, gold nanoparticles, dendrimers, cyclodextrin nanoparticles, micelles, or combinations thereof, can be used for delivery.

[0241] The methods and systems provided by the present invention may use any suitable carrier or delivery mode, which in certain embodiments include lipid nanoparticles (LNPs). Any LNP known in the art may be used. LNPs are described in WO2021 / 178898, which is incorporated herein by reference in whole.

[0242] kit Furthermore, this disclosure covers, in part, (a) a system, template nucleic acid (e.g., template RNA), reaction mixture, DNA molecule, RNA molecule, or pharmaceutical composition as described herein, and (b) a kit comprising instructions for use of the system, template nucleic acid (e.g., template RNA), reaction mixture, DNA molecule, RNA molecule, or pharmaceutical composition as described herein. In some embodiments, the kit further comprises a cell (e.g., a cell from a cell line, or a cell from a subject, e.g., a human cell) or DNA (e.g., genomic DNA or a vector) containing a target site (e.g., a target site that is a target of the system or template RNA).

[0243] Embodiment This application includes, but is not limited to, the following numbered embodiments.

[0244] Embodiment 1 is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(a) A 3'-untranslated region (3'UTR) element for improving the expression of a polypeptide, the 3'UTR element comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 - 8211, 8256, 8264 - 8269, and 8281, and / or (b) A 5'-untranslated region (5'UTR) element for improving the expression of a polypeptide, the 5'UTR element comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212 - 8247 and 8259 - 8263 At least one of the above, and / or (3) A poly(A) tail for improving the expression of a polypeptide, the poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 - 8314 An RNA molecule comprising the above.

[0245] Embodiment 1a is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) A nucleotide sequence encoding a polypeptide, (2)(a) A 3'-untranslated region (3'UTR) element for improving the expression of a polypeptide, the 3'UTR element consisting of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 - 8211, 8256, 8264 - 8269, and 8281, and / or (b) A 5'-untranslated region (5'UTR) element for improving the expression of a polypeptide, the 5'UTR element consisting of a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212 - 8247 and 8259 - 8263 At least one of the above, and / or (3) A poly(A) tail for improving the expression of a polypeptide, the poly(A) tail comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 - 8314 An RNA molecule comprising the above.

[0246] Embodiment 1b is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a nucleotide sequence encoding the polypeptide, and (a) 3'-untranslated region (3'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (b) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8212-8247 and 8259-8263. It is an RNA molecule containing at least one of the following.

[0247] Embodiment 1c is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a nucleotide sequence encoding the polypeptide, and (a) 3'-untranslated region (3'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (b) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247 and 8259-8263. It is an RNA molecule containing at least one of the following.

[0248] Embodiment 1d is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, (a) a nucleotide sequence encoding a polypeptide, and (b) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 to 8233. It is an RNA molecule that contains [something].

[0249] Embodiment 1e is an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, (a) a nucleotide sequence encoding a polypeptide, and (b) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314. It is an RNA molecule that contains [something].

[0250] Embodiment 2 is an artificial RNA molecule according to Embodiment 1 or 1b, wherein the 3'UTR element contains the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

[0251] Embodiment 2a is an artificial RNA molecule according to Embodiment 1a or 1c, wherein the 3'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

[0252] Embodiment 3 is an artificial RNA molecule according to any one of Embodiments 1, 1b, or 2, wherein the 5'UTR element contains the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

[0253] Embodiment 3a is an artificial RNA molecule according to any one of Embodiments 1a, 1c, or 2a, wherein the 5'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

[0254] Embodiment 4 includes a 3'UTR element and a 5'UTR element, (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, (5) The 3'UTR element contains sequence number 8281 and the 5'UTR element contains sequence number 8259, or (6) The 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263, The artificial RNA molecule is described in any one of Embodiments 1, 1b, 2, or 3.

[0255] Embodiment 4a includes a 3'UTR element and a 5'UTR element, (1) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8236, (2) The 3'UTR element consists of sequence number 8201 and the 5'UTR element consists of sequence number 8236, (3) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8214, (4) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8243, (5) The 3'UTR element consists of sequence number 8281 and the 5'UTR element consists of sequence number 8259, or (6) The 3'UTR element consists of sequence number 8281, and the 5'UTR element consists of sequence number 8263. An artificial RNA molecule according to any one of Embodiment 1a, 1c, 2a or 3a.

[0256] Embodiment 4b is an artificial RNA molecule according to Embodiment 4, comprising a 3'UTR element and a 5'UTR element, wherein the 3'UTR element contains SEQ ID NO: 8201 and the 5'UTR element contains SEQ ID NO: 8236.

[0257] Embodiment 4c is an artificial RNA molecule according to Embodiment 4, comprising a 3'UTR element and a 5'UTR element, wherein the 3'UTR element contains SEQ ID NO: 8209 and the 5'UTR element contains SEQ ID NO: 8214.

[0258] Embodiment 5 is an artificial RNA molecule according to any one of Embodiments 1, 1d and 2 - 4c, wherein the poly(A) tail contains a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0259] Embodiment 5a is an artificial RNA molecule according to any one of Embodiments 1, 1e and 2 - 4c, wherein the poly(A) tail consists of a nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0260] Embodiment 5b is an artificial RNA molecule according to Embodiment 5, wherein the poly(A) tail contains a nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0261] Embodiment 5c is an artificial RNA molecule according to Embodiment 5a, wherein the poly(A) tail consists of a nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0262] Embodiment 5d comprises a 3'UTR element, a 5'UTR element and a poly(A) tail. (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (5) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8259, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (6) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8263, and the poly(A) tail contains sequence number 8306 or sequence number 8307, The artificial RNA molecule is described in any one of Embodiments 1 to 3a.

[0263] Embodiment 6 is an artificial RNA molecule according to any one of Embodiments 1 to 5c, wherein the polypeptide comprises a reverse transcriptase domain and an endonuclease domain.

[0264] Embodiment 6a is an artificial RNA molecule according to any one of Embodiments 1 to 5c, wherein the polypeptide is a heterogeneously modified polypeptide or a retrotransposon gene-modified polypeptide.

[0265] Embodiment 7 is an artificial RNA molecule according to Embodiment 6 or 6a, wherein the endonuclease domain is a niccasse domain.

[0266] Embodiment 7a is an artificial RNA molecule according to Embodiment 7, wherein the nickase domain is a Cas9 domain, and optionally the Cas9 domain is selected from the SpCas9 domain, BlatCas9 domain, Nme2 Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain.

[0267] Embodiment 7b is the artificial RNA molecule described in Embodiment 7a, wherein the Cas9 domain includes the N670A mutation, N611A mutation, N605A mutation, N580A mutation, N588A mutation, N872A mutation, N863A mutation, N622A mutation, or H840A mutation.

[0268] Embodiment 7c is an artificial RNA molecule according to any one of Embodiments 7 to 7b, wherein the reverse transcriptase domain is selected from retroviral reverse transcriptase domains.

[0269] Embodiment 7d is the artificial RNA molecule described in Embodiment 7c, wherein the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus.

[0270] Embodiment 7e is an artificial RNA molecule according to Embodiment 7d, wherein the gamma retrovirus-derived reverse transcriptase domain comprises the amino acid sequence of a reverse transcriptase domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.

[0271] Embodiment 7f is an artificial RNA molecule according to Embodiment 7c or 7d, wherein the reverse transcriptase domain derived from gamma retrovirus is not derived from PERV.

[0272] Embodiment 7g is an artificial RNA molecule according to any one of Embodiments 7c to 7f, wherein the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

[0273] Embodiment 7h is an artificial RNA molecule according to any one of Embodiments 7 to 7g, wherein the polypeptide comprises the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

[0274] Embodiment 8 is an artificial ribonucleic acid (RNA) molecule, (1)(a) A 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8201 to 8211, and / or (b) A 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247. At least one of the following, and / or (2) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314. It is an artificial ribonucleic acid (RNA) molecule containing [the specified element].

[0275] Embodiment 8a is an artificial ribonucleic acid (RNA) molecule, (1)(a) A 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8201 to 8211, and / or (b) A 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247. At least one of the following, and / or (2) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314. It is an artificial ribonucleic acid (RNA) molecule composed of [the specified components].

[0276] Embodiment 8b is, (a) A 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8201-8211, and / or (b) A 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247. It is an artificial ribonucleic acid (RNA) molecule containing [the specified element].

[0277] Embodiment 8c is, (a) A 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8201-8211, and / or (b) A 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247. It is an artificial ribonucleic acid (RNA) molecule composed of [the specified components].

[0278] Embodiment 8d is an artificial ribonucleic acid (RNA) molecule containing a poly(A)tail that includes a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8282 to 8314.

[0279] Embodiment 8e is an artificial ribonucleic acid (RNA) molecule containing a poly(A)tail consisting of a nucleic acid sequence selected from the group comprising SEQ ID NOs: 8282 to 8314.

[0280] Embodiment 9 is an artificial RNA molecule according to Embodiment 8 or 8b, wherein the 3'UTR element contains the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, or SEQ ID NO: 8211.

[0281] Embodiment 9a is the artificial RNA molecule described in Embodiment 8a, wherein the 3'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, or SEQ ID NO: 8211.

[0282] Embodiment 10 is an artificial RNA molecule according to any one of Embodiments 8, 8b, or 9, wherein the 5'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243.

[0283] Embodiment 10a is an artificial RNA molecule according to any one of Embodiments 8a, 8c, or 9a, wherein the 5'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243.

[0284] Embodiment 11 includes a 3'UTR element and a 5'UTR element, (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, or (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, The artificial RNA molecule is as described in any one of claims 8, 8b, 9, or 10.

[0285] Embodiment 11a includes a 3'UTR element and a 5'UTR element, (1) The 3'UTR element consists of sequence number 8209, and the 5'UTR element consists of sequence number 8236. (2) The 3'UTR element consists of sequence number 8201, and the 5'UTR element consists of sequence number 8236, (3) The 3'UTR element consists of sequence number 8209, and the 5'UTR element consists of sequence number 8214. (4) The 3'UTR element consists of sequence number 8209, and the 5'UTR element consists of sequence number 8243. The artificial RNA molecule is as described in any one of claims 8a, 8c, 9a, or 10a.

[0286] Embodiment 12 is an artificial RNA molecule according to Embodiments 8, 8d, 9, 10, or 11, wherein the poly(A)tail contains the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0287] Embodiment 12a is the artificial RNA molecule according to Embodiment 12, wherein the poly(A)tail contains the nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0288] Embodiment 12b is an artificial RNA molecule according to any one of Embodiments 8, 8e, 9a, 10a, or 11a, wherein the poly(A)tail consists of the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0289] Embodiment 12c is the artificial RNA molecule described in Embodiment 12b, wherein the poly(A)tail consists of the nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0290] Embodiment 12d includes a 3'UTR element, a 5'UTR element, and a poly(A) tail. (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307. This is an artificial RNA molecule described in any one of Embodiments 8 to 11a.

[0291] Embodiment 13 is a system for modifying DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide containing a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(i) A 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. At least one of the following, and / or (3) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial RNA molecules containing; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) The system includes a heterogeneous object sequence, which includes modifications compared to the corresponding original sequence (e.g., wild-type sequence), where the modifications improve the speed, fidelity, or speed and fidelity of the reverse transcription primed by the target by the reverse transcriptase.

[0292] Embodiment 13a is a system for modifying DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a nucleotide sequence encoding a polypeptide, and (i) 3'-untranslated region (3'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. An artificial RNA molecule containing at least one of the following; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) The system includes a heterogeneous object sequence, which includes modifications compared to the corresponding original sequence (e.g., wild-type sequence), where the modifications improve the speed, fidelity, or speed and fidelity of the reverse transcription primed by the target by the reverse transcriptase.

[0293] Embodiment 13b is a system for modifying DNA, (a) an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a poly(A)tail containing (i) a nucleotide sequence encoding a polypeptide and (ii) a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 to 8233; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) The system includes a heterogeneous object sequence, which includes modifications compared to the corresponding original sequence (e.g., wild-type sequence), where the modifications improve the speed, fidelity, or speed and fidelity of the reverse transcription primed by the target by the reverse transcriptase.

[0294] Embodiment 14 is a system for modifying DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide containing a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(i) A 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. At least one of the following, and / or (3) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial nucleic acid molecules including; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) Preferably, the heterogeneous object array has the following characteristics: i) it does not contain self-complementary arrays, such as self-complementary arrays that form a hairpin structure under stringent conditions, or if self-complementary arrays are present, it has the following characteristics: (1) Each self-complementary sequence has a length of 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides or less. (2) The self-complementary sequence forms a hairpin containing an arm of length 10, 9, 8, 7, 6, 5, 4 or 3 nucleotides or less, or (3) The self-complementary sequence includes at least 1, 2, 3, 4, or 5 positions that are complementary to its partner sequence (e.g., mismatch or bulge). Having one, two, or all of the following, (4) It does not contain any repeating sequences (e.g., single, di, or trinucleotide repeating sequences), or if it does contain repeating sequences, they must be sequences of 12, 11, 10, 9, 8, 7, or 6 nucleotides or less in length. It is a system that has one or both of the following:

[0295] Embodiment 14a is a system for modifying DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a nucleotide sequence encoding a polypeptide, and (i) 3'-untranslated region (3'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and / or (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. An artificial nucleic acid molecule comprising at least one of the following; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) Preferably, the heterogeneous object array has the following characteristics: i) it does not contain self-complementary arrays, such as self-complementary arrays that form a hairpin structure under stringent conditions, or if self-complementary arrays are present, it has the following characteristics: (1) Each self-complementary sequence has a length of 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides or less. (2) The self-complementary sequence forms a hairpin containing an arm of length 10, 9, 8, 7, 6, 5, 4 or 3 nucleotides or less, or (3) The self-complementary sequence includes at least 1, 2, 3, 4, or 5 positions that are complementary to its partner sequence (e.g., mismatch or bulge). Having one, two, or all of the following, (4) It does not contain any repeating sequences (e.g., single, di, or trinucleotide repeating sequences), or if it does contain repeating sequences, they must be sequences of 12, 11, 10, 9, 8, 7, or 6 nucleotides or less in length. It is a system that has one or both of the following:

[0296] Embodiment 14b is a system for modifying DNA, (a) an artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide comprising a reverse transcriptase (RT) domain and optionally an endonuclease domain, comprising a poly(A)tail containing (i) a nucleotide sequence encoding a polypeptide and (ii) a nucleic acid sequence selected from the group consisting of SEQ ID NOs: 8201 to 8233; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) Preferably, the heterogeneous object array has the following characteristics: i) it does not contain self-complementary arrays, such as self-complementary arrays that form a hairpin structure under stringent conditions, or if self-complementary arrays are present, it has the following characteristics: (1) Each self-complementary sequence has a length of 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides or less. (2) The self-complementary sequence forms a hairpin containing an arm of length 10, 9, 8, 7, 6, 5, 4 or 3 nucleotides or less, or (3) The self-complementary sequence includes at least 1, 2, 3, 4, or 5 positions that are complementary to its partner sequence (e.g., mismatch or bulge). Having one, two, or all of the following, (4) It does not contain any repeating sequences (e.g., single, di, or trinucleotide repeating sequences), or if it does contain repeating sequences, they must be sequences of 12, 11, 10, 9, 8, 7, or 6 nucleotides or less in length. It is a system that has one or both of the following:

[0297] Embodiment 15 is the system according to Embodiments 13, 13a, 14, or 14a, wherein the 3'UTR element includes the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

[0298] Embodiment 15a is the system according to any one of Embodiments 13, 13a, 14, or 14a, wherein the 3'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

[0299] Embodiment 16 is the system according to any one of Embodiments 13, 13a, 14, 14a, or 15, wherein the 5'UTR element includes the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

[0300] Embodiment 16a is the system according to any one of Embodiments 13, 13a, 14, 14a, 15, or 15a, wherein the 5'UTR element consists of the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

[0301] Embodiment 17 comprises an artificial RNA molecule containing a polypeptide-encoding nucleotide sequence, a 3'UTR element, and a 5'UTR element. (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, (5) The 3'UTR element contains sequence number 8281 and the 5'UTR element contains sequence number 8259, or (6) The 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263, The system is described in any one of embodiments 13, 13a, 14, 14a, and 15-16a.

[0302] Embodiment 17a comprises an artificial RNA molecule containing a polypeptide-encoding nucleotide sequence, 3'UTR and 5'UTR, (1) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8236, (2) The 3'UTR element consists of sequence number 8201 and the 5'UTR element consists of sequence number 8236, (3) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8214, (4) The 3'UTR element consists of sequence number 8209 and the 5'UTR element consists of sequence number 8243, (5) The 3'UTR element consists of sequence number 8281 and the 5'UTR element consists of sequence number 8259, or (6) The 3'UTR element consists of sequence number 8281, and the 5'UTR element consists of sequence number 8263. The system is described in any one of embodiments 13, 13a, 14, 14a, and 15-16a.

[0303] Embodiment 17b is the system according to Embodiment 17, wherein the artificial RNA molecule comprises a nucleotide sequence encoding a polypeptide, a 3'UTR element and a 5'UTR element, where the 3'UTR element comprises SEQ ID NO: 8201 and the 5'UTR element comprises SEQ ID NO: 8236.

[0304] Embodiment 17c is the system according to Embodiment 17, wherein the artificial RNA molecule comprises a nucleotide sequence encoding a polypeptide, a 3'UTR element and a 5'UTR element, where the 3'UTR element comprises SEQ ID NO: 8209 and the 5'UTR element comprises SEQ ID NO: 8214.

[0305] Embodiment 18 is a system according to any one of Embodiments 13, 13b, 14, 14b and 15-17c, wherein the poly(A)tail includes the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312 or SEQ ID NO: 8314.

[0306] Embodiment 18a is a system according to any one of Embodiments 13, 13b, 14, 14b and 15-18, wherein the poly(A)tail consists of the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

[0307] Embodiment 18b is the system according to Embodiment 18, wherein the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0308] Embodiment 18c is the system according to Embodiment 18a, wherein the poly(A)tail consists of the nucleic acid sequence of SEQ ID NO: 8306 or SEQ ID NO: 8307.

[0309] Embodiment 18d includes a 3'UTR element, a 5'UTR element, and a poly(A) tail. (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (5) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8259, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (6) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8263, and the poly(A) tail contains sequence number 8306 or sequence number 8307, The system is as described in any one of embodiments 13 to 16a.

[0310] Embodiment 19 is a system according to any one of Embodiments 13 to 18c, wherein the polypeptide comprises a reverse transcriptase domain and an endonuclease domain.

[0311] Embodiment 19a is the system according to any one of Embodiments 13 to 18c, wherein the polypeptide is a heterogeneously modified polypeptide or a retrotransposon gene-modified polypeptide.

[0312] Embodiment 20 is the system described in Embodiment 19, wherein the endonuclease domain is a nickasase domain.

[0313] Embodiment 20a is the system described in Embodiment 20, wherein the nickarse domain is a Cas9 domain.

[0314] Embodiment 20b is the system described in Embodiment 20a, wherein the Cas9 domain is selected from the SpCas9 domain, BlatCas9 domain, Nme2 Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain.

[0315] Embodiment 20c is a system according to any one of Embodiments 20 to 20b, wherein the Cas9 domain includes the N670A mutation, N611A mutation, N605A mutation, N580A mutation, N588A mutation, N872A mutation, N863A mutation, N622A mutation, or H840A mutation.

[0316] Embodiment 20d is the system according to any one of Embodiments 20 to 20c, wherein the reverse transcriptase domain is selected from retroviral reverse transcriptase domains.

[0317] Embodiment 20e is the system described in Embodiment 20d, wherein the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus.

[0318] Embodiment 20f is the system according to Embodiment 20e, wherein the reverse transcriptase domain derived from a gamma retrovirus comprises the amino acid sequence of a reverse transcriptase domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6.

[0319] Embodiment 20g is the system described in Embodiment 20e or 20f, wherein the reverse transcriptase domain derived from gamma retrovirus is not derived from PERV.

[0320] Embodiment 20h is a system according to any one of Embodiments 20c to 20f, wherein the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

[0321] Embodiment 20i is the system according to any one of Embodiments 20 to 20h, wherein the polypeptide comprises the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

[0322] Embodiment 21 is the system according to any one of Embodiments 13 to 20i, wherein the template RNA further comprises an RT terminator sequence located between a heterogeneous object sequence and either (i) or (ii).

[0323] Embodiment 22 is a system according to any one of Embodiments 13 to 21, wherein the heterogeneous object sequence includes a sequence that encodes the target polypeptide or a portion thereof, or a sequence that is the inverse complementary chain of the sequence encoding the target polypeptide or a portion thereof.

[0324] Embodiment 22a is the system according to any one of Embodiments 13 to 22, further comprising a second polypeptide containing an endonuclease domain or a nucleic acid encoding a second polypeptide.

[0325] Embodiment 22b is the system according to any one of Embodiments 13 to 22a, wherein (a) further comprises a DNA-binding domain.

[0326] Embodiment 22c is the system according to any one of Embodiments 13 to 22b, wherein the second polypeptide further comprises a DNA-binding domain.

[0327] Embodiment 23 is a polypeptide comprising a reverse transcriptase domain and an endonuclease domain, wherein the endonuclease domain is a Cas9 domain, and the template RNA is (i) A gRNA spacer that is complementary to the first portion of the target gene and optionally contains one or more consecutive nucleotides starting at the 3' end of an adjacent nucleotide of the gRNA spacer. (ii) gRNA scaffold that binds to the Cas9 domain, (iii) heterologous object sequences containing mutation regions for introducing mutations into a second portion of the target gene (for example, to correct mutations within it) (where the heterologous sequence may include a 5' to 3' post-edit homologous region, a mutation region and a pre-edit homologous region), and (iv) A primer binding site (PBS) sequence containing at least 5, 6, 7, or 8 bases that is 100% identical to the third portion of the target gene. The system is one of the embodiments 13 to 22c, including the above.

[0328] Embodiment 24 is a human PAH gene, and the template RNA is (i) A gRNA spacer that is complementary to the first portion of the human PAH gene, preferably having a sequence containing the core nucleotide of a gRNA spacer sequence from Table 1A, Table 1B, Table 1C or Table 1D of WO2023039435 incorporated herein by reference in its entirety, and optionally containing one or more consecutive nucleotides starting at the 3' end of an adjacent nucleotide of the gRNA spacer, or having a spacer sequence selected from Tables 5A-5F, 8A-8D, E3, E3A, BB, E5, E5A, E6 or E6A of WO2023039435 incorporated herein by reference in its entirety. (ii) gRNA scaffold that binds to the Cas9 domain, (iii) a heterologous object sequence containing a mutation region for introducing a mutation into the second portion of the human PAH gene (for example, to correct a mutation within it) (where the heterologous object sequence may include a post-edit homologous region from 5' to 3', a mutation region and a pre-edit homologous region), and (iv) A primer-binding site (PBS) sequence containing at least 5, 6, 7, or 8 bases that is 100% identical to the third portion of the human PAH gene. This is the system described in Embodiment 23, which includes the following:

[0329] Embodiment 25 is a system according to any one of Embodiments 13 to 24, wherein the template RNA includes SEQ ID NO: 8258 (RNACS7570), SEQ ID NO: 8320 (RNACS229), SEQ ID NO: 8321 (RNACS1515), or SEQ ID NO: 8372.

[0330] Embodiment 26 is a system according to any one of Embodiments 13 to 25, wherein the reverse transcriptase domain and the endonuclease domain are linked by a peptide linker.

[0331] Embodiment 27 is the system described in any one of Embodiments 13 to 26, wherein the target site is a target site within the human genome.

[0332] Embodiment 28 is a reaction mixture comprising cells and the system described in any one of Embodiments 13 to 27.

[0333] Embodiment 28a is the reaction mixture according to Embodiment 28, wherein the cells are T cells (e.g., primary T cells).

[0334] Embodiment 29 is a reaction mixture comprising DNA containing a target site and the system described in any one of Embodiments 13 to 27.

[0335] Embodiment 30 is an artificial RNA molecule according to any one of Embodiments 1 to 12c, comprising one or more chemically modified nucleotides.

[0336] Embodiment 31 is a deoxyribonucleic acid (DNA) molecule that encodes an artificial RNA molecule described in any one of Embodiments 1 to 12c.

[0337] Embodiment 32 is a pharmaceutical composition comprising an artificial RNA molecule described in any one of Embodiments 1 to 12c and 30, a system described in any one of Embodiments 13 to 27, or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier.

[0338] Embodiment 33 is the pharmaceutical composition according to Embodiment 32, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

[0339] Embodiment 34 is the pharmaceutical composition according to Embodiment 33, wherein the viral vector is an adeno-associated virus.

[0340] Embodiment 35 is a host cell containing an artificial RNA molecule or system or DNA molecule as described in any one of the embodiments above.

[0341] Embodiment 35a is the host cell according to Embodiment 35, wherein the host cell is a T cell (e.g., a primary T cell).

[0342] Embodiment 35b is the host cell according to Embodiment 35 or 35a, wherein the host cell is a mammalian cell.

[0343] Embodiment 35c is the host cell described in Embodiment 35b, wherein the mammalian cell is a human cell.

[0344] Embodiment 36 is a method for producing an artificial RNA molecule according to any one of Embodiments 1 to 12c, comprising synthesizing template RNA by in vitro transcription (e.g., solid-state synthesis) or by introducing DNA encoding the artificial RNA molecule into a host cell under conditions that enable the production of template RNA.

[0345] Embodiment 37 is, (a) A system according to any one of Embodiments 13 to 27, a reaction mixture according to Embodiment 28 or 29, a DNA molecule according to Embodiment 31, or a pharmaceutical composition according to any one of Embodiments 32 to 34, and (b) Instructions for use of the system, reaction mixture, DNA molecule, or pharmaceutical composition. This is a kit that includes [the following items].

[0346] Embodiment 38 is a lipid nanoparticle (LNP) comprising an artificial RNA molecule described in any one of Embodiments 1 to 12c and 30, or a system described in any one of Embodiments 13 to 27.

[0347] Embodiment 39 is a method for modifying a target site in the genomic DNA of a cell, comprising contacting the cell with the system described in any one of Embodiments 13 to 27, or one or more RNAs encoding the system, thereby modifying the target site in the genomic DNA of the cell.

[0348] Embodiment 40 is a method for treating a subject having a disease or condition associated with a genetic defect, comprising administering to the subject the system described in any one of Embodiments 13 to 27, thereby treating the subject having a disease or condition associated with a genetic defect.

[0349] [Examples] The following embodiments of the present invention are for illustrative purposes only, to further illustrate the properties of the present invention. It should be understood that the following embodiments are not limiting to the present invention, and that the scope of the present invention should be determined by the accompanying claims.

[0350] [Example 1] Evaluation of the expression and activity of exemplary genetically modified polypeptides from mRNA with reference 5'UTR in U2OS cells and primary mouse hepatocytes. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding a genetically modified polypeptide to quantify the activity of mRNAs with various criterion 5'UTRs for targeted genetic modification function in the U2OS cell line, as well as the expression of genetically modified polypeptides in the U2OS cell line and primary mouse hepatocytes.

[0351] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) One of the 5'UTRs listed in Table 1, (3) A coding sequence (CDS) encoding a genetically modified polypeptide, and in some embodiments, a HiBiT protein tag having the nucleotide sequence of SEQ ID NO: 8255 is attached to the carboxyl terminus of the genetically modified polypeptide, (4) The 3'UTR having the nucleotide sequence of sequence number 8256, and (5) A poly A tail containing 80 adenine (A) residues It contained [something].

[0352] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0353] For example, the genetically modified polypeptide was encoded by mRNA RNAV209 and contained the amino acid sequence of sequence number 8257.

[0354] In this example, the template RNA co-delivered with the mRNA has the following nucleotide sequence: This is RNACS7570 (SEQ ID NO: 8258) containing mG*mC*mU*mU. Nucleotide modifications are indicated as follows: phosphorothioate bonds are indicated by an asterisk (*), 2'-O-methyl groups are indicated by (m) preceding the nucleotide, and ribonucleotides are indicated by "r" preceding the nucleotide.

[0355] [Table 7]

[0356] To compare the efficiency of mRNAs with different criterion 5'UTRs in inducing targeted gene modification, a landing pad cell line called U2OS-BFP was used. U2OS-BFP cells were engineered to stably express blue fluorescent protein (BFP). Template RNA was designed to convert the genomic DNA sequence encoding BFP to the DNA sequence encoding GFP. The gene modification activity of the system was evaluated using the percentage of GFP-positive cells (GFP%) from the recovered U2OS-BFP cell samples. Since the 5'UTR of the mRNA was varied in each reaction, the template RNA and gene modification polypeptide remained unchanged.

[0357] 0.2 μg of mRNA encoding a genetically modified polypeptide and 2 μg of template RNA were co-delivered to 250,000 U2OS-BFP cells on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle® (Lonza, Basel, Switzerland). On day 4, the cells were harvested and BFP and GFP expression were examined by flow cytometry (Figure 1). Data were analyzed using FlowJo to determine GFP%. Correlations were examined using GraphPad Prism.

[0358] Using mRNA encoding HiBiT-tagged genetically modified polypeptides, the GFP% induced by reference 5'UTR ranged from 69.0% to 57.3%, demonstrating that all reference 5'UTRs promote the expression of genetically modified polypeptides in a manner sufficient to facilitate BFP-to-GFP editing.

[0359] The GFP% induced by the gene modification system containing mRNA encoding a HiBiT-tagged genetically modified polypeptide with a reference 5'UTR ranged from 58.7% to 34.8%, demonstrating that the reference 5'UTR also promotes the expression of the genetically modified polypeptide in a manner sufficient to facilitate editing from BFP to GFP (Figure 2).

[0360] To directly compare the expression levels of genetically modified polypeptides from mRNA with a reference 5'UTR, 2 μg of HiBiT set mRNA and 2 μg of template RNA were co-delivered to 250,000 BFP-nonexpressing U2OS cells (U2OS-naive) by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Six hours after nucleofection, the cells were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. The results showed that the expression levels of genetically modified polypeptides varied by approximately twice the normalized reference 5'UTR (reference 1) across the tested reference 5'UTR (Figure 3).

[0361] Furthermore, to evaluate the effect of the reference 5'UTR on the expression of exemplary genetically modified polypeptides in primary cells, primary mouse hepatocytes from two sources were used: cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01) and fresh hepatocytes recovered from wild-type C57BL / 6 mice. In each reaction, all other segments of mRNA remained identical except for the template RNA and the 5'UTR. 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 4 μg of template RNA were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. 24 hours after nucleofection, the hepatocytes were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system.

[0362] The results show that all criterion 5'UTRs successfully promoted the expression of exemplary genetically modified polypeptides in cryopreserved mouse hepatocytes (Figure 4). Two criterion 5'UTRs, Criterion 5 and Criterion 1, induced higher expression of genetically modified polypeptides in cryopreserved primary mouse hepatocytes than the other criterion 5'UTRs.

[0363] Furthermore, the expression levels of HiBiT-tagged genetically modified polypeptides from mRNA were evaluated 24 hours after nucleofection in hepatocytes freshly recovered from wild-type C57BL / 6 mice (Figure 5). The results indicate that all reference 5'UTRs successfully promoted the expression of exemplary genetically modified polypeptides in freshly recovered primary mouse hepatocytes. Two reference 5'UTRs, reference 5 and reference 1, induced higher genetically modified polypeptide expression in freshly recovered primary mouse hepatocytes than the other reference 5'UTRs.

[0364] Taken together, these results demonstrate that the criterion 5'UTRs tested have the ability to induce genetically modified polypeptide expression in several different cell types, including primary hepatocytes from different sources, and that some criterion 5'UTRs induce higher expression levels than others tested.

[0365] [Example 2] Evaluation of the expression and activity of exemplary genetically modified polypeptides from mRNA containing a 5'UTR in U2OS cells. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding a genetically modified polypeptide to quantify the activity of mRNAs with various 5'UTRs regarding the targeted genetic modification function and expression of genetically modified polypeptides in the U2OS cell line.

[0366] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) One of the 5'UTRs listed in Table 2, or 5'UTR Criterion 1, (3) A coding sequence (CDS) encoding a genetically modified polypeptide, and in some embodiments, a HiBiT protein tag whose coding nucleotide sequence is sequence number 8255 is attached to the carboxyl terminus of the genetically modified polypeptide, (4) The 3'UTR having the nucleotide sequence of sequence number 8256, and (5) Poly-A tail containing 80 A residues It contained [something].

[0367] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0368] The genetically modified polypeptide was encoded by mRNA RNAV209 (SEQ ID NO: 8257) and contained the amino acid sequence provided in Example 1. In this example, the template RNA co-delivered with the mRNA was RNACS7570 (SEQ ID NO: 8258) and contained the nucleotide sequence shown in Example 1. The 5'UTR tested in this example was compared to the 5'UTR of Reference 1 from Example 1.

[0369] [Table 8-1] [Table 8-2]

[0370] To compare the efficiency of mRNAs with different 5'UTRs in inducing targeted gene modification, we used a landing pad cell line called U2OS-BFP. U2OS-BFP cells were engineered to stably express blue fluorescent protein (BFP). Template RNA was designed to convert the genomic DNA sequence encoding BFP to the DNA sequence encoding GFP. The gene modification activity of the system was evaluated using the percentage of GFP-positive cells (GFP%) from the recovered U2OS-BFP cell samples. Since the 5'UTR of the mRNA was varied in each reaction, the template RNA and gene modification polypeptide remained unchanged.

[0371] 0.2 μg of mRNA and 2 μg of the template RNA RNACS7570 (SEQ ID NO: 8258) were co-delivered to 250,000 U2OS-BFP cells on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. On day 4, the cells were harvested and BFP and GFP expression was examined by flow cytometry. Data were analyzed using FlowJo to determine the GFP% of the cells. Significance and correlation were calculated using GraphPad Prism.

[0372] Using mRNA encoding HiBiT-tagged genetically modified polypeptides, GFP% ranged from 69.7% to 53.2%, indicating that many of the tested 5'UTRs promoted the expression of genetically modified polypeptides in a manner sufficient to facilitate editing from BFP to GFP (Figure 11). The levels of editing using genetic modification systems containing mRNA with several 5'UTRs (e.g., 70-4c (69.7%), 70-2c (69.6%), 70-5c (66.9%), 50-2c (65.9%), and 50-3c (65.8%)) were higher than or equivalent to the levels obtained using genetic modification systems containing mRNA with Reference 5'UTR Reference 1 in Example 1.

[0373] Furthermore, the percentage of GFP-positive cells obtained 4 days after nucleofecting U2OS-BFP cells with mRNA encoding HiBiT-tagged genetically modified polypeptides was evaluated. The percentage of GFP induced by HiBiT-tagged genetically modified systems containing mRNA with a 5'UTR ranged from 58.0% to 39.4%, indicating that many of the tested 5'UTRs promoted the expression of genetically modified polypeptides in a manner sufficient to facilitate editing from BFP to GFP (Figure 12). The levels of editing using genetically modified systems containing several mRNAs with 5'UTRs (e.g., 70-2a (58.0%), 70-3c (57.7%), 50-1c (55.2%), 50-7 (54.9%), 70-1d (54.0%), 70-3a (54.0%), 70-2b (53.8%)) were higher than or equivalent to the levels obtained when using genetically modified systems containing mRNA with a 5'UTR of criterion 1.

[0374] To directly compare the expression levels of genetically modified polypeptides from mRNAs with different 5'UTRs, 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 2 μg of template RNA were co-delivered to 250,000 BFP-nonexpressing U2OS-naive cells by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Six hours after nucleofection, the cells were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. The results showed that several 5'UTRs, e.g., 50-7, 70-3c, 50-1c, 70-2a, 70-2b, and 70-3a, enabled 15%–30% higher expression of the genetically modified polypeptide compared to mRNA containing criterion 1 5'UTR (Figure 13).

[0375] Furthermore, in U2OS-naive cells (not expressing BFP), the expression of genetically modified polypeptides from mRNAs encoding HiBiT-tagged genetically modified polypeptides containing various 5'UTRs (normalized to expression from mRNA containing 5'UTR of criterion 1) was evaluated against GFP% obtained by nucleofecting a HiBiT-tagged genetically modified system containing the same mRNA. Expression levels were obtained 6 hours after nucleofection (Figure 14). The data allowed for the analysis of the correlation between the expression levels of genetically modified polypeptides shown in Figure 13 and the genetic modification activity shown in Figure 12. Correlations were calculated using Pearson's two-tailed test for correlation, and R 2 A value of 0.7554 was obtained. The results suggest that when expressed from mRNA containing a 5'UTR, the expression level of the exemplary genetically modified polypeptide correlates well with genetically modified function. Therefore, improvements in genetically modified function may be due to stronger expression of genetically modified polypeptides from mRNAs having these novel 5'UTRs. The results suggest that the editing activity of a genetically modified system containing mRNA encoding a genetically modified polypeptide can be improved using 5'UTRs that increase the expression of the genetically modified polypeptide described herein.

[0376] [Example 3] Evaluation of the expression of exemplary genetically modified polypeptides from mRNA with 5'UTR in primary mouse hepatocytes. This example illustrates the quantification of the expression of exemplary genetically modified polypeptides in primary mouse hepatocytes using an exemplary genetic modification system that includes genetically modified polypeptides and mRNAs encoding various reference 5'UTRs.

[0377] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) One of the 5'UTRs listed in Table 2, or 5'UTR Criterion 1, (3) A coding sequence (CDS) is provided, which encodes a genetically modified polypeptide and has a HiBiT protein tag whose coding sequence is sequence number 8255, attached to the carboxyl terminus of the genetically modified polypeptide. (4) The 3'UTR whose nucleotide sequence is the sequence of sequence number 8256, (5) Poly-A tail containing 80 A residues It contained [something].

[0378] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0379] The genetically modified polypeptide was the same as the polypeptide described in Examples 1 and 2.

[0380] In this example, the template RNA co-delivered with the mRNA is RNACS7570 (SEQ ID NO: 8258), which includes the nucleic acid sequence and modifications shown in Example 1. The 5'UTR tested in this example was compared to the 5'UTR of Reference 1 from Example 1 (see Table 2).

[0381] Furthermore, to evaluate the effectiveness of 5'UTR in promoting the expression of exemplary genetically modified polypeptides in primary cells, primary mouse hepatocytes from two sources were used: cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01) and fresh hepatocytes recovered from wild-type C57BL / 6 mice. In each reaction, the template RNA and genetically modified polypeptides remained unchanged while the mRNA 5'UTR was altered.

[0382] 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 4 μg of template RNA were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. 24 hours after nucleofection, the hepatocytes were lysed, and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. Absolute HiBiT protein expression (pmol / cell) was calculated using a standard HiBiT quantity ladder prepared with Promega's HiBiT reference protein (catalog number N3010) and a cell number ladder prepared with ThermoFisher's PrestoBlue® cell viability reagent (catalog number A13261). The results show that several 5'UTRs, such as 70-2b, 50-7, 70-3c, and 50-1c, enabled the expression of genetically modified polypeptides that were higher than or equivalent to the 5'UTR of criterion 1 (Figure 15).

[0383] Furthermore, the expression of HiBiT-tagged genetically modified polypeptides from mRNAs containing various 5'UTRs was evaluated 24 hours after nucleofection in hepatocytes freshly recovered from wild-type C57BL / 6 mice and normalized to the expression levels of genetically modified polypeptides from mRNAs containing 5'UTR 1. The results show that the 50-1c 5'UTR enabled comparable expression of genetically modified polypeptides compared to mRNAs containing 5'UTR 1 (Figure 16). In addition, the tested 5'UTRs had the ability to enable the expression of exemplary genetically modified polypeptides in primary hepatocytes from two different sources, and some 5'UTRs enabled expression equivalent to or higher than that of 5'UTR 1.

[0384] [Example 4] Evaluation of the expression and activity of exemplary genetically modified polypeptides from mRNA containing 3'UTR in U2OS cells and primary mouse hepatocytes. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding a genetically modified polypeptide to quantify the activity of mRNAs with various criterion 3'UTRs for targeted gene modification function in the U2OS cell line, as well as the expression of genetically modified polypeptides in the U2OS cell line and primary mouse hepatocytes.

[0385] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) 5'UTR having the nucleotide sequence of sequence number 8260, (3) A coding sequence (CDS) encoding a genetically modified polypeptide, and in some embodiments, a HiBiT protein tag whose coding nucleotide sequence is sequence number 8255 is attached to the carboxyl terminus of the genetically modified polypeptide, (4) One of the 3'UTRs listed in Table 3, and (5) Poly-A tail containing 80 A residues It contained [something].

[0386] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0387] The 3'UTRs listed in Table 3 contain terminal CUAGs as a result of the nucleic acid synthesis method used. The experiments described herein may also be performed using 3'UTRs lacking terminal CUAGs, and similar results are expected.

[0388] The genetically modified polypeptides were the same as those described in Examples 1-3.

[0389] In this example, the template RNA co-delivered with the mRNA was RNACS7570 (SEQ ID NO: 8258), as described in Example 1 above.

[0390] [Table 9]

[0391] To compare the efficiency of mRNAs with different criterion 3'UTRs in inducing targeted gene modification, a landing pad cell line called U2OS-BFP was used. U2OS-BFP cells were engineered to stably express blue fluorescent protein (BFP). Template RNA was designed to convert the genomic DNA sequence encoding BFP to the DNA sequence encoding GFP. The gene modification activity of the system was evaluated using the percentage of GFP-positive cells (GFP%) from the recovered U2OS-BFP cell samples. Since the 3'UTR of the mRNA was varied in each reaction, the template RNA and gene modification polypeptide remained unchanged.

[0392] 0.2 μg of mRNA encoding a genetically modified polypeptide and 2 μg of template RNA RNACS7570 (SEQ ID NO: 8258) were co-delivered to 250,000 U2OS-BFP cells on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. On day 4, the cells were harvested and BFP and GFP expression were examined by flow cytometry. Data were analyzed using FlowJo to determine the GFP% of the cells. Correlations were examined using GraphPad Prism.

[0393] The results show the percentage of GFP-positive U2OS-BFP cells detected by flow cytometry on day 4 (GFP%) (Figure 6). Using mRNA encoding a HiBiT-tagged genetically modified polypeptide, the GFP% induced by mRNA with a reference 3'UTR ranged from 69.0% to 54.8%, indicating that all reference 3'UTRs promote the expression of the genetically modified polypeptide in a manner sufficient to facilitate editing from BFP to GFP.

[0394] Furthermore, the percentage of GFP-positive cells was examined using flow cytometry four days after nucleofecting U2OS-BFP cells with mRNA encoding HiBiT-tagged genetically modified polypeptides (Figure 7). The percentage of GFP induced by the HiBiT-tagged genetically modified system containing mRNA with a reference 3'UTR ranged from 47.8% to 34.8%, indicating that the reference 3'UTR also promotes the expression of genetically modified polypeptides in a manner sufficient to facilitate editing from BFP to GFP.

[0395] To directly compare the expression levels of genetically modified polypeptides from mRNA with a reference 3'UTR, 2 μg of mRNA and 2 μg of template RNA from the HiBiT set were co-delivered to 250,000 BFP-nonexpressing U2OS cells (U2OS-naive) by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Six hours after nucleofection, the cells were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. The results showed that the expression levels of genetically modified polypeptides varied across the tested reference 3'UTR, ranging from approximately 0.5 times or less of the normalized reference 3'UTR (reference 1) (Figure 8).

[0396] Furthermore, to evaluate the effect of the reference 3'UTR on the expression of exemplary genetically modified polypeptides in primary cells, primary mouse hepatocytes from two sources were used: cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01) and fresh hepatocytes recovered from wild-type C57BL / 6 mice. In each reaction, all other segments of mRNA remained identical, apart from the template RNA and the 3'UTR.

[0397] 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 4 μg of template RNA were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. 24 hours after nucleofection, the hepatocytes were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. The results showed that all criterion 3'UTRs successfully promoted the expression of exemplary genetically modified polypeptides in cryopreserved mouse hepatocytes (Figure 9). Two criterion 3'UTRs, Criterion 2 and Criterion 3, induced higher genetically modified polypeptide expression in cryopreserved primary mouse hepatocytes than the other criterion 3'UTRs.

[0398] Furthermore, the expression levels of HiBiT-tagged genetically modified polypeptides from RNA were evaluated 24 hours after nucleofection in hepatocytes from freshly recovered wild-type C57BL / 6 mice. The results showed that all criterion 3'UTRs successfully promoted the expression of exemplary genetically modified polypeptides in freshly recovered primary mouse hepatocytes (Figure 10). Criterion 1, Criterion 2, and Criterion 4, the three criterion 3'UTRs, induced higher genetically modified polypeptide expression in freshly recovered primary mouse hepatocytes than the other criterion 3'UTRs.

[0399] Taken together, these results demonstrate that the tested criterion 3'UTRs have the ability to induce genetically modified polypeptide expression in several different cell types, including primary hepatocytes from different sources, and that some criterion 3'UTRs induce higher expression levels than others tested.

[0400] [Example 5] Evaluation of the expression and activity of exemplary genetically modified polypeptides from mRNA containing a 3'UTR in U2OS cells. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding a genetically modified polypeptide to quantify the activity of mRNAs with various 3'UTRs regarding the targeted genetic modification function and expression of genetically modified polypeptides in the U2OS cell line.

[0401] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) 5'UTR having the nucleotide sequence of sequence number 8260, (3) A coding sequence (CDS) encoding a genetically modified polypeptide, and in some embodiments, a HiBiT protein tag whose coding nucleotide sequence is sequence number 8255 is attached to the carboxyl terminus of the genetically modified polypeptide, (4) One of the 3'UTRs listed in Table 3 or Table 4, and (5) Poly-A tail containing 80 A residues It contained [something].

[0402] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0403] The genetically modified polypeptides were the same as those described in Examples 1-4.

[0404] In this example, the template RNA co-delivered with the mRNA was RNACS7570 (SEQ ID NO: 8258), as described in Example 1. The 3'UTR tested in this example was compared to the 3'UTR of Reference 1 from Example 4.

[0405] [Table 10]

[0406] To compare the efficiency of mRNAs with different 3'UTRs in inducing targeted gene modification, a landing pad cell line called U2OS-BFP was used. U2OS-BFP cells were engineered to stably express blue fluorescent protein (BFP). Template RNA was designed to convert the genomic DNA sequence encoding BFP to the DNA sequence encoding GFP. The gene modification activity of the system was evaluated using the percentage of GFP-positive cells (GFP%) from the recovered U2OS-BFP cell samples. Since the 3'UTR of the mRNA was varied in each reaction, the template RNA and gene modification polypeptide remained unchanged.

[0407] 0.2 μg of mRNA and 2 μg of template RNA were co-delivered to 250,000 U2OS-BFP cells on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. On day 4, the cells were harvested and BFP and GFP expression were examined by flow cytometry. Data were analyzed using FlowJo to determine the GFP% of the cells. Significance and correlation were calculated using GraphPad Prism.

[0408] Using mRNA encoding HiBiT-tagged genetically modified polypeptides, GFP percentages ranged from 63.9% to 48.8%, indicating that many of the tested 3'UTRs promoted the expression of the genetically modified polypeptide in a manner sufficient to facilitate editing from BFP to GFP (Figure 17). The levels of editing using genetic modification systems containing mRNA with several 3'UTRs (e.g., 14-UUCG (63.9%), 3WJ-1 (62.9%), 24-GCAA (59.5%), and 3WJ-3 (58.9%)) were higher than or equivalent to the levels obtained using genetic modification systems containing mRNA with 3'UTRs according to Criterion 1 of Example 4.

[0409] Furthermore, the percentage of GFP-positive cells obtained on day 4 after nucleofecting U2OS-BFP cells with mRNA encoding HiBiT-tagged genetically modified peptides was evaluated. The percentage of GFP induced by gene modification systems containing mRNA encoding HiBiT-tagged genetically modified polypeptides with a 3'UTR ranged from 41.6% to 25.0%, indicating that many of the tested 3'UTRs promoted the expression of the genetically modified polypeptide in a manner sufficient to facilitate editing from BFP to GFP (Figure 18). The levels of editing using gene modification systems containing mRNA with several 3'UTRs (e.g., 3WJ-3 (41.6%), 3WJ-1 (38.1%), G4-2 (35.9%), 14-GCAA (35.0%), 3WJ-2 (34.5%), and 14-UUCG (35.0%)) were higher than or equivalent to the levels obtained when using gene modification systems containing mRNA with 3'UTRs of criterion 1.

[0410] To directly compare the expression levels of genetically modified polypeptides from mRNAs with different 3'UTRs, 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 2 μg of template RNA were co-delivered to 250,000 BFP-nonexpressing U2OS-naive cells by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Six hours after nucleofection, the cells were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. The results show that several 3'UTRs, e.g., G4-2, 3UTR14-UUCG, 3WJ-3, 3UTR14-GCAA, 3UTR24-UUCG, and G4-1, enabled higher or similar expression of genetically modified polypeptides compared to mRNA containing criterion 1's 3'UTR (Figure 19). The dotted line indicates expression from mRNA containing the criterion 3'UTR.

[0411] [Example 6] Evaluation of the expression of exemplary genetically modified polypeptides from mRNA containing 3'UTR in primary mouse hepatocytes. This example illustrates the quantification of the expression of exemplary genetically modified polypeptides in primary mouse hepatocytes using an exemplary genetic modification system that includes genetically modified polypeptides and mRNAs encoding various 3'UTRs.

[0412] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) The nucleotide sequence is the sequence of sequence number 8260, 5'UTR, (3) A coding sequence (CDS) is provided, which encodes a genetically modified polypeptide and has a HiBiT protein tag whose coding nucleotide sequence is sequence number 8255, attached to the carboxyl terminus of the genetically modified polypeptide. (4) One of the 3'UTRs listed in Table 4, or 3'UTR criterion 1, and (5) Poly-A tail containing 80 A residues It contained [something].

[0413] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0414] The genetically modified polypeptide was the same as the polypeptide shown in Example 1.

[0415] In this example, the template RNA co-delivered with the mRNA was RNACS7570 (SEQ ID NO: 8258), which included the nucleic acid sequence and modifications shown in Example 1. The 3'UTR tested in this example was compared to the 3'UTR of Reference 1 from Example 4.

[0416] Furthermore, to evaluate the effectiveness of 3'UTR in promoting the expression of exemplary genetically modified polypeptides in primary cells, primary mouse hepatocytes from two sources were used: cryopreserved mouse hepatocytes from Lonza (catalog number MCCP01) and fresh hepatocytes recovered from wild-type C57BL / 6 mice. In each reaction, the template RNA and genetically modified polypeptides remained unchanged while the mRNA 3'UTR was altered.

[0417] 2 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 4 μg of template RNA were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. 24 hours after nucleofection, the hepatocytes were lysed, and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. Absolute HiBiT protein expression (pmol / cell) was calculated using a standard HiBiT quantity ladder prepared with Promega's HiBiT reference protein (catalog number N3010) and a cell number ladder prepared with ThermoFisher's PrestoBlue® cell viability reagent (catalog number A13261). The results show that several 3'UTRs, such as 24-UUCG, G4-2, and 3WJ-1, enabled the expression of genetically modified polypeptides at a higher or equivalent level than the 3'UTR of criterion 1 (Figure 20).

[0418] Furthermore, the expression of HiBiT-tagged genetically modified polypeptides from mRNA was examined 24 hours after nucleofection in hepatocytes from freshly recovered wild-type C57BL / 6 mice. The results showed that some 3'UTRs, such as G4-2, enabled the expression of comparable genetically modified polypeptides compared to mRNA containing the 3'UTR of reference 1 (Figure 21).

[0419] Taken together, these results indicate that the tested 3'UTRs were capable of enabling the expression of exemplary genetically modified polypeptides in primary hepatocytes from two different sources, with some 3'UTRs enabling expression equivalent to or higher than that of criterion 1 3'UTRs.

[0420] [Example 7] Expression of exemplary genetically modified polypeptides from mRNA with a modified polyA tail in U2OS cells, and evaluation of genetic modification activity. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding the genetically modified polypeptide and including a modified poly(A) tail, for quantifying the genetic modification activity and expression of the genetically modified polypeptide in the U2OS cell line.

[0421] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) 5'UTR having the nucleotide sequence of sequence number 8259, (3) A coding sequence (CDS) that encodes the genetically modified polypeptide fused to a HiBiT protein tag having the coding nucleotide sequence of SEQ ID NO: 8255, which is attached to the carboxyl terminus of the genetically modified polypeptide. (4) The 3'UTR having the nucleotide sequence of sequence number 8317, (5) Poly A tails listed in Table 5 or Table 6 It contained [something].

[0422] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0423] The genetically modified polypeptide was encoded by mRNA RNAV209 and contained the amino acid sequence of sequence number 8257.

[0424] In this example, the template RNA co-delivered with the mRNA is RNACS7570 (SEQ ID NO: 8258) containing the following nucleotide sequence: mG*mC*mC*rGrArArGrCrArCrUrGrCrCrArCrGrCrCrGrUrGrUrUrUrArGrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrArGrGrCrCrUrArGrGrUrArUrCrAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrArCrCrCrUrGrArCrGrGrCrUrArCrGrGrCrUrGrCrArGrUrG*mC*mU*mU. Nucleotide modifications are indicated as follows: phosphorothioate bonds are indicated by asterisks, 2'-O-methyl groups are indicated by an "m" preceding the nucleotide, and ribonucleotides are indicated by an "r" preceding the nucleotide.

[0425] [Table 11]

[0426] [Table 12-1] [Table 12-2]

[0427] To evaluate the expression levels of genetically modified polypeptides from mRNA with a modified poly(A) tail, we used a landing pad cell line called U2OS-BFP. U2OS-BFP was an engineered U2OS cell line that stably expressed blue fluorescent protein (BFP) from its genome.

[0428] 0.25 μg of mRNA encoding a HiBiT-tagged genetically modified polypeptide and 2 μg of template RNA were co-delivered to 250,000 U2OS-BFP cells by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Six hours after nucleofection, the cells were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system.

[0429] As shown in Figure 1, 18 of the 31 modified polyA tails enabled the expression of genetically modified polypeptides at or above the level of mRNA with the control 80A tail. Some of the modified polyA tails, such as T26, T28, T25, T30, T33, T27, T24, T16, and T18, dramatically increased the expression of genetically modified polypeptides by up to four times compared to the control 80A tail. While not theoretically bound, the results suggest that intervening non-adenosine nucleotides in the center of the polyA tail can enhance the protein expression activity of mRNA. The results further suggest that intervening cytosine or uridine nucleotides may result in a superior enhancement of protein expression compared to intervening guanosine nucleotides. The results further suggest that modified poly(A) following the pattern 15A-25A-35A-45A (hyphens indicate intervening non-adenosine nucleotides) or similar results in superior protein expression compared to modified poly(A) following a more uniform pattern, e.g., 30A-30A-30A. While not theoretically bound, higher levels of expression of genetically modified polypeptides are thought to increase the targeted gene modification activity of the system.

[0430] The effect of poly(A) tails on targeted gene modification (e.g., through changes in the expression of the modified polypeptide) was evaluated using U2OS-BFP cells. Template RNA was designed to convert the DNA sequence encoding BFP in the genome to the DNA sequence encoding GFP. The percentage of GFP-positive cells (GFP%) from the recovered U2OS-BFP cell samples was used to evaluate the system's gene modification activity. Since the poly(A) tail of mRNA was altered in each reaction, the template RNA and modified polypeptide remained unchanged.

[0431] 0.25 μg of mRNA and 2 μg of template RNA RNACS7570 (SEQ ID NO: 8258) were co-delivered to 250,000 U2OS-BFP cells on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. On day 4, the cells were harvested and BFP and GFP expression was examined by flow cytometry. Data were analyzed using FlowJo to determine the percentage of GFP-positive cells. Significance and correlation were calculated using GraphPad Prism.

[0432] Before normalization, U2OS-BFP cells nucleofected with mRNA containing the control-80A polytail showed a 68% GFP positivity rate on day 4. Many cell samples nucleofected with mRNA containing different modified polyA tails exhibited a significant amount of GFP-positive cells. Samples nucleofected with 22 different modified polyA tails showed GFP% at or above the control level, with T26 (SEQ ID NO: 8307), the most superior polyA tail, showing an 88% GFP+ rate. The results indicate that many of the tested modified polyA tails successfully promoted the expression of genetically modified polypeptides sufficient for gene editing in U2OS cells. The results further demonstrate that the 22 modified polyA tails were particularly high performers, promoting higher gene editing activity in U2OS-BFP cells than the control-80A polyA tail. In addition, the expression enhancement provided by intervening cytosine or uridine nucleotides may promote a superior improvement in gene modification activity compared to the expression enhancement provided by intervening guanosine nucleotides. The results further suggest that the expression enhancement provided by modified polyA following pattern 15A-25A-35A-45A (hyphens indicate intervening non-adenosine nucleotides) or similar may promote a superior improvement in gene modification activity compared to the expression enhancement provided by modified polyA following a more uniform pattern, e.g., 30A-30A-30A.

[0433] The data in Figures 22 and 23 are analyzed by correlation analysis (Figures 24A-24B). The correlation analysis in Figures 24A-B suggests that for most mRNAs with modified polyA tails, the expression level of the genetically modified polypeptide correlated well with genetic modification function. Therefore, improvements in genetic modification function may be due to stronger expression of the genetically modified polypeptide from these novel mRNAs with modified polyA tails. The results suggest that the editing activity of a genetically modified system containing mRNA encoding a genetically modified polypeptide can be improved using polyA tails that increase the expression of the genetically modified polypeptide described herein.

[0434] [Example 8] Evaluation of the expression and genetic modification activity of exemplary genetically modified polypeptides from mRNA with a modified polyA tail in primary mouse hepatocytes. This example illustrates the use of an exemplary genetic modification system containing template RNA and mRNA encoding the genetically modified polypeptide and containing a modified poly(A) tail, for quantifying the genetic modification activity and expression of genetically modified polypeptides in primary mouse hepatocytes.

[0435] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) 5'UTR having the nucleotide sequence of sequence number 8259, (3) A coding sequence (CDS) encoding the genetically modified polypeptide, fused to a HiBiT protein tag having the coding nucleotide sequence of SEQ ID NO: 8255, which is attached to the carboxyl terminus of the genetically modified polypeptide. (4) The 3'UTR having the nucleotide sequence of sequence number 8317, (5) Poly A tails listed in Table 5 or Table 6 It contained [something].

[0436] In this example, the genetically modified polypeptide encoded by mRNA is (1) Endonucleases and / or DNA-binding domains, (2) Peptide linker, and (3) Reverse transcriptase (RT) domain It contained [something].

[0437] The genetically modified polypeptide was the same as the polypeptide described in Example 1.

[0438] In this example, the gene modification system included mRNA encoding a gene modification polypeptide and one of two template RNAs. Template RNA A was RNACS229 (SEQ ID NO: 8320): mU*mC*mA*rGrArGrArArGrCrUrGrGrGrCrCrArCrCrGrUrUrUrUrArGrAmGmCmUmAmGmAmAmUmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUrArGrUrCrCrGrU rUrArUrCrAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUrGrGrArGrCrArGrUrArArUrGrGrCrUrGrGrUrGrGrCrCrCrArGrC*mU*mU*mC That was the case.

[0439] The template RNA B has the following nucleotide sequence: It was RNACS1515 (SEQ ID NO: 8321) containing mU*mU*mA*rCrCrArArCrUrUrUrCrUrCrCrArUrGrGrCrCrGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrArGrGrGrCrUrArGrUrUrCrCrGrUrUrUrArUrCrAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU.

[0440] For expression analysis, the template RNA RNACS7570 (shown in Example 1) was co-delivered with the mRNA.

[0441] Furthermore, to evaluate the expression and gene modification activity associated with the modified poly(A) tail in primary cells, fresh primary mouse hepatocytes recovered from wild-type C57BL / 6 mice were used. In each gene modification system used for primary cells, all segments of the template RNA and mRNA encoding the modified polypeptide remained unchanged, except for the poly(A) tail.

[0442] To evaluate the expression of genetically modified polypeptides 18 hours after nucleofection, 2 μg of mRNA and 4 μg of template RNA RNACS7570 (SEQ ID NO: 58) were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. 18 hours after nucleofection, the hepatocytes were lysed and HiBiT expression was analyzed using Promega's Nano-Glo® HiBiT lysis detection system. Of the 31 modified polyA tails tested, 24 enabled 2-fold or higher expression of genetically modified polypeptides compared to the control-80A tail (Figure 25). The most effective modified polyA tail, T25, dramatically increased genetically modified polypeptide expression by approximately 15-fold compared to the control-80A tail. The results suggest that many of the modified polyA tails, such as T25, T18, T31, T13, T28, T32, T24, T01, T14, T33, T21, T17, T04, T26, T08, T27, T02, T07, T30, T05, T10, T29, T16, and T22, successfully enabled the expression of genetically modified polypeptides in primary mouse hepatocytes. The results highlight that the central, intervening non-adenosine nucleotides of the polyA tail greatly enhance mRNA expression in primary cells, consistent with the improved expression described in U2OS-BFP cells described above.

[0443] Experiments were conducted to evaluate the expression of genetically modified polypeptides from mRNA over a time course of 18 to 48 hours after nucleofection (Figure 26A). 2 μg of mRNA and 4 μg of template RNA RNACS7570 (SEQ ID NO: 8240) were co-delivered to 100,000 primary mouse hepatocytes by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. Two pairs of reactions for each reaction were prepared in two 96-well plates, nucleofected, and then the same reactions were mixed in one 96-well deep-well plate to obtain 200,000 nucleofected hepatocytes for each reaction condition. These hepatocytes were then divided into four 96-well plates and cultured at 37°C for 18, 24, 42, and 48 hours, respectively. At each time point, one of the four plates was transferred to -80°C for storage. At the time of analysis, hepatocytes in all four plates were thawed and lysed, and HiBiT expression was examined using Promega's Nano-Glo® HiBiT lysis detection system.

[0444] The data showed that nucleofection of primary cells with mRNA containing many modified poly(A) tails resulted in a sharp increase in the expression of the modified polypeptide compared to the expression provided by control mRNA, and this increase was maintained throughout the entire study period, with expression starting at its peak 18 hours after nucleofection and generally decreasing over time (Figure 26B). In particular, T25, the most performing modified poly(A) tail, was associated with an approximately 13.5-fold increase in modified polypeptide expression 18 hours after nucleofection compared to control-80A, with further increases exceeding control-80A being 13.7-fold at 24 hours after nucleofection, 10.6-fold at 42 hours after nucleofection, and 15.0-fold at 48 hours after nucleofection (Figure 26B). Expression profiles for all mRNAs used in this example are shown in Figures 27A–27G.

[0445] The area under the curve (AUC) was calculated in GraphPad Prism using the 18-48 hour expression profiles obtained above from each tested mRNA. The results show that many nucleofections of mRNA with modified polyA tails enabled a dramatic increase in the AUC of the modified polypeptide compared to the AUC from mRNA containing the control-80A polyA tail (Figure 28). The results showed that modified polyA tail T25 enabled a 13-fold increase in AUC, modified polyA tail T18 enabled a 9-fold increase, modified polyA tail T31 enabled a 7-fold increase, four modified polyA tails (T13, T14, T28 and T01) enabled a 6-fold increase, three modified polyA tails (T24, T33 and T32) enabled a 5-fold increase, and four modified polyA tails (T08, T26, T04 and T21) enabled a 4-fold increase, four modified polyA tails (T02, T27, 05, and T17) enabled a 3-fold increase, six modified polyA tails (T30, T07, T16, T10, T29, and T22) enabled a 2-fold increase, and three modified polyA tails (T03, T23, and T15) enabled comparable AUC (all comparisons are against the AUC of control-80A polyA tail mRNA). The up to 13-fold increase in expression AUC, superior to the control polyA tail, demonstrates the ability of these modified polyA tails to increase the expression of exemplary genetically modified peptides from mRNA.

[0446] Combined, these results further suggest that intervening cytosine or uridine nucleotides may result in a superior improvement in protein expression compared to intervening guanosine nucleotides. Combined, the results further suggest that modified polyA following pattern 15A-25A-35A-45A (hyphens indicate intervening non-adenosine nucleotides) or similar results in a superior improvement in protein expression compared to modified polyA following a more uniform pattern, e.g., 30A-30A-30A.

[0447] To evaluate the ability of mRNA encoding a genetically modified polypeptide and containing a modified poly(A) tail to induce targeted gene modification in primary cells, a gene modification system consisting of one of two exemplary template RNAs and one mRNA encoding a genetically modified polypeptide and containing a modified poly(A) tail was co-delivered to primary mouse hepatocytes freshly recovered from wild-type C57B / L6 mice. The two template RNAs targeted gene modification in the wild-type mouse Fah gene and wild-type sequence (GGAGC G GTAATG C CTGGTGG (Sequence ID 8318) is a mutated disease sequence (GGAGC) observed in a mouse model of tyrosinemia. A GTAATG G The nucleotides were converted to CTGGTGG (sequence number 8319). Underlined nucleotides indicate the modified nucleotides. In each reaction, the two template RNAs and all segments of the mRNA remained unchanged, except for the poly(A) tail.

[0448] 1 μg of mRNA, 4 μg of template RNA A, and 5 μg of template RNA B were co-delivered to 100,000 primary mouse hepatocytes on day 0 by nucleofection using Lonza's Amaxa® Nucleofector 96-well Shuttle®. The hepatocytes were harvested on day 5, and genomic DNA (gDNA) was subsequently extracted. The gDNA samples were submitted to Targeted Amplicon Sequencing, where the copy number of wild-type unmodified sequences and the number of sequences with intended modifications in the target region of the mFah gene were counted. The percentage of mutational modifications was calculated by dividing the copy number of gDNA sequences with intended modifications by the total number of gDNA copies in the sample and multiplying by 100.

[0449] The results show that many mRNAs with modified polyA tails enabled dramatically higher modification percentages in primary mouse hepatocytes compared to control mRNAs containing 80A polyA (which resulted in 2% modification) (Figure 29). For example, the T26 tail enabled 11% modification, the T16 tail enabled 10% modification, the T25 tail enabled 9% modification, the T33 tail enabled 7% modification, the T14 and T13 tails enabled 6% modification, the T24, T03, T15 and T28 tails enabled 5% modification, the T32, T01, T04, T31 and T17 tails enabled 4% modification, the T18, T07, T30 and T02 tails enabled 3% modification, and the T23 and T10 tails enabled 2% modification. The results show that the use of modified polyA in mRNA encoding exemplary genetically modified polypeptides resulted in up to approximately 5-fold improvement in modification percentage (up to 11% modification) compared to control mRNA containing 80A, demonstrating that modified polyA can lead to improved effectiveness of genetic modification systems in primary cells. The results further suggest that increased expression mediated by intervening cytosine or uridine nucleotides may promote improved genetic modification activity more effectively than increased expression mediated by intervening guanosine nucleotides. Furthermore, the results suggest that increased expression mediated by modified polyA following a pattern of 15A-25A-35A-45A (hyphens indicate intervening non-adenosine nucleotides) or similar may promote improved genetic modification activity more effectively than increased expression mediated by modified polyA following a more uniform pattern, e.g., 30A-30A-30A.

[0450] [Example 9] Evaluation of the expression of exemplary genetically modified polypeptides from mRNA with UTR and polyA tail in primary human T cells. This example evaluates the expression of genetically modified polypeptides in primary human T cells from mRNA possessing exemplary UTRs and exemplary poly-A tails.

[0451] In this embodiment, mRNA is in the following segments. (1) 5' cap, (2) One of the 5'UTRs listed in Table 7, (3) A coding sequence (CDS) encoding a genetically modified polypeptide fused to a HiBiT protein tag having the coding nucleotide sequence of sequence number 8255, (4) One of the 3'UTRs listed in Table 8, and (5) Poly A Tail (T26): AAAAAAAAAAAAAAAUUAAAAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAUUUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (Sequence number 8307) It contained [something].

[0452] In this example, the genetically modified polypeptide encoded by all mRNA is RNAIVT3689. It contained the following amino acid sequence.

[0453] In this example, the template RNA co-delivered with all mRNA is It contains the nucleotide sequence.

[0454] [Table 13]

[0455] [Table 14]

[0456] The performance of selected 5' and 3' UTRs for mRNA was compared by evaluating their efficiency in promoting RNA writer protein expression in primary human T cells. For rapid and reliable protein expression analysis, a C-terminal Hibit tag was added to all constructs to enable Promega's Nano-Glo® HiBiT lysis detection system. The number of cells in each well was determined using PrestoBlue® cell viability reagent to normalize the Hibit readouts. In this example, the 5' and 3' UTRs of these mRNAs were varied, so the coding region and poly(A) tail of all mRNAs remained unchanged. When comparing protein expression, the template RNA was included in all reactions and remained unchanged.

[0457] Cryopreserved primary human T cells were thawed and cultured in activation medium for 3 days. On the day of the experiment, 0.2 μg of mRNA and 1 μg of template RNA were co-delivered to 500,000 activated human T cells by nucleofection using Lonza's Amaxa Nucleofector 96-well Shuttle®. Cells were harvested and lysed 2, 4, 6, 8, and 24 hours after nucleofection and maintained at -80°C. After collecting all samples from the six time points, the frozen cell lysates were tested according to the manufacturer's protocol using Promega's Nano-Glo® HiBiT lysis detection system.

[0458] Figure 30 shows graphs of the time course of protein expression at 2, 4, 6, 8, and 24 hours from mRNAs with test or reference UTRs, analyzed by the HiBiT assay. The results show that one mRNA with 70-2b (SEQ ID NO: 8236) and 3WJ-3 (SEQ ID NO: 8209) UTRs performed better than the control mRNA with reference UTR (5'Ref (SEQ ID NO: 8259) + 3'Ref (SEQ ID NO: 8373)), exhibiting higher protein expression at all time points. Another mRNA with 70-2b (SEQ ID NO: 8236) and 14-UUCG (SEQ ID NO: 8201) UTRs expressed more protein at earlier time points (2, 4, and 6 hours) than the control with reference UTR (5'Ref + 3'Ref). The mRNAs containing UTR 50-1c (SEQ ID NO: 8214) and 3WJ-3 (SEQ ID NO: 8209), as well as two other mRNAs containing 70-4c (SEQ ID NO: 8243) and 3WJ-3 (SEQ ID NO: 8209), come very close to the performance of the control mRNA (5'Ref + 3'Ref).

[0459] Figure 31 shows a graph of the area under the protein expression curve (AUC) for 2–24 hours calculated from Figure 30. The results show that one mRNA with UTR 70-2b and 3WJ-3 is superior to the control mRNA with reference UTR (5'Ref + 3'Ref) by having a higher overall protein expression AUC. Another mRNA with UTR 70-2b and 14-UUCG allows for a similar protein expression AUC to the control mRNA (5'Ref + 3'Ref).

[0460] Combined, these results demonstrate the following: (1) UTR 70-2b and 3WJ-3 perform better than the reference UTR (5'Ref + 3'Ref) and enable high protein expression in activated primary human T cells. (2) UTR 70-2b and 14-UUCG have similar efficiencies to reference UTR (5'Ref + 3'Ref) and enable high protein expression in activated primary human T cells. (3) These UTRs and combinations of UTRs, as well as heterogeneously modified polypeptides (e.g., shown in Example 1) and retrotransposon gene-modified polypeptides (e.g., RNAIVT3689 used in this example), are generally effective in promoting high protein expression of the target protein.

Claims

1. An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide containing a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(a) A 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and (b) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. At least one of the following, and / or (3) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial RNA molecules containing this substance.

2. The artificial RNA molecule according to claim 1, wherein the 3'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

3. The artificial RNA molecule according to claim 1 or 2, wherein the 5'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

4. Including 3'UTR elements and 5'UTR elements, (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, (5) The 3'UTR element contains sequence number 8281 and the 5'UTR element contains sequence number 8259, or (6) The 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263, An artificial RNA molecule according to any one of claims 1 to 3.

5. The artificial RNA molecule according to claim 4, wherein the 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236.

6. The artificial RNA molecule according to claim 4, wherein the 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214.

7. An artificial RNA molecule according to any one of claims 1 to 6, wherein the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

8. The artificial RNA molecule according to claim 7, wherein the poly(A)tail comprises the nucleic acid sequence of sequence number 8306 or sequence number 8307.

9. Includes 3'UTR elements, 5'UTR elements and poly(A) tails, (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (5) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8259, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (6) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8263, and the poly(A) tail contains sequence number 8306 or sequence number 8307, An artificial RNA molecule according to any one of claims 1 to 3.

10. The artificial RNA molecule according to any one of claims 1 to 9, wherein the polypeptide is a heterogeneously modified polypeptide or a retrotransposon gene-modified polypeptide.

11. An artificial RNA molecule according to any one of claims 1 to 9, wherein the polypeptide comprises a reverse transcriptase domain and an endonuclease domain.

12. The artificial RNA molecule according to claim 11, wherein the endonuclease domain is a Cas9 domain selected from nickase domains, such as the SpCas9 domain, BlatCas9 domain, Nme2 Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain, and the reverse transcriptase domain is selected from retroviral reverse transcriptase domains.

13. The artificial RNA molecule according to claim 12, wherein the Cas9 domain comprises an N670A mutation, an N611A mutation, an N605A mutation, an N580A mutation, an N588A mutation, an N872A mutation, an N863A mutation, an N622A mutation, or an H840A mutation.

14. The artificial RNA molecule according to claim 12 or 13, wherein the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus, preferably comprising an amino acid sequence of a reverse transcriptase domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6, and preferably the gamma retrovirus-derived reverse transcriptase domain is not derived from PERV.

15. An artificial RNA molecule according to any one of claims 11 to 14, wherein the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

16. The artificial RNA molecule according to claim 11, wherein the polypeptide comprises the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

17. An artificial ribonucleic acid (RNA) molecule, (1)(a) A 3'-untranslated region (3'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8201 to 8211, and (b) A 5'-untranslated region (5'UTR) element containing a nucleic acid sequence selected from the group consisting of sequence numbers 8212-8247. At least one of the following, and / or (2) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial ribonucleic acid (RNA) molecules containing [the specified substance].

18. The artificial RNA molecule according to claim 17, wherein the 3'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, or SEQ ID NO: 8211.

19. The artificial RNA molecule according to claim 17 or 18, wherein the 5'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, or SEQ ID NO: 8243.

20. Including 3'UTR elements and 5'UTR elements, (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, or (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, An artificial RNA molecule according to any one of claims 17 to 19.

21. An artificial RNA molecule according to any one of claims 17 to 20, wherein the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

22. The artificial RNA molecule according to claim 21, wherein the poly(A)tail comprises the nucleic acid sequence of sequence number 8306 or sequence number 8307.

23. Includes 3'UTR elements, 5'UTR elements and poly(A) tails, (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307. An artificial RNA molecule according to any one of claims 17 to 19.

24. It is a system that modifies DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide containing a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(i) A 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. At least one of the following, and / or (3) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial RNA molecules containing; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) A system comprising, where a heterogeneous object sequence is modified compared to the corresponding original sequence (e.g., wild-type sequence), where the modification improves the speed, fidelity, or speed and fidelity of reverse transcription primed by the target by reverse transcriptase.

25. It is a system that modifies DNA, (a) An artificial ribonucleic acid (RNA) molecule for improving the expression of a polypeptide containing a reverse transcriptase (RT) domain and optionally an endonuclease domain, (1) Nucleotide sequence encoding polypeptide, (2)(i) A 3'-untranslated region (3'UTR) element for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8201-8211, 8256, 8264-8269 and 8281, and (ii) 5'-untranslated region (5'UTR) elements for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8212-8247 and 8259-8263. At least one of the following, and / or (3) A poly(A) tail for improving polypeptide expression, comprising a nucleic acid sequence selected from the group consisting of SEQ ID NOs. 8282 to 8314. Artificial nucleic acid molecules including; and (b) (For example, from 5' to 3') (i) Depending on the case, a sequence that binds to a target site in DNA (e.g., the second strand of a site in the target genome), (ii) Sequences that bind to polypeptides, (iii) arrays of heterogeneous objects, and (iv) In some cases, the 3' target homologous domain template RNA (or DNA encoding template RNA) Preferably, the heterogeneous object array has the following characteristics: i) it does not contain self-complementary arrays, such as self-complementary arrays that form a hairpin structure under stringent conditions, or if self-complementary arrays are present, it has the following characteristics: (1) Each self-complementary sequence has a length of 10, 9, 8, 7, 6, 5, 4, or 3 nucleotides or less. (2) The self-complementary sequence forms a hairpin containing an arm of length 10, 9, 8, 7, 6, 5, 4 or 3 nucleotides or less, or (3) The self-complementary sequence includes at least 1, 2, 3, 4, or 5 positions that are complementary to its partner sequence (e.g., mismatch or bulge). Having one, two, or all of the following, (4) It does not contain any repeating sequences (e.g., single, di, or trinucleotide repeating sequences), or if it does contain repeating sequences, they must be sequences of 12, 11, 10, 9, 8, 7, or 6 nucleotides or less in length. A system having one or both of the following.

26. The system according to claim 24 or 25, wherein the 3'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8201, SEQ ID NO: 8204, SEQ ID NO: 8209, SEQ ID NO: 8211, SEQ ID NO: 8256, or SEQ ID NO: 8281.

27. The system according to any one of claims 24 to 26, wherein the 5'UTR element comprises the nucleic acid sequence of SEQ ID NO: 8214, SEQ ID NO: 8230, SEQ ID NO: 8235, SEQ ID NO: 8236, SEQ ID NO: 8243, SEQ ID NO: 8259, or SEQ ID NO: 8263.

28. The artificial RNA molecule contains a nucleotide sequence encoding a polypeptide, a 3'UTR element, and a 5'UTR element. (1) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8236, (2) The 3'UTR element contains sequence number 8201 and the 5'UTR element contains sequence number 8236, (3) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8214, (4) The 3'UTR element contains sequence number 8209 and the 5'UTR element contains sequence number 8243, (5) The 3'UTR element contains sequence number 8281 and the 5'UTR element contains sequence number 8259, or (6) The 3'UTR element contains sequence number 8281, and the 5'UTR element contains sequence number 8263, The system according to any one of claims 24 to 27.

29. The system according to claim 28, wherein the 3'UTR element includes sequence number 8201 and the 5'UTR element includes sequence number 8236.

30. The system according to claim 28, wherein the 3'UTR element includes sequence number 8209 and the 5'UTR element includes sequence number 8214.

31. The system according to any one of claims 24 to 30, wherein the poly(A)tail comprises the nucleic acid sequence of SEQ ID NO: 8294, SEQ ID NO: 8299, SEQ ID NO: 8306, SEQ ID NO: 8307, SEQ ID NO: 8309, SEQ ID NO: 8311, SEQ ID NO: 8312, or SEQ ID NO: 8314.

32. The system according to claim 31, wherein the poly(A)tail comprises the nucleic acid sequence of sequence number 8306 or sequence number 8307.

33. Includes 3'UTR elements, 5'UTR elements and poly(A) tails, (1) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (2) The 3'UTR element contains sequence number 8201, the 5'UTR element contains sequence number 8236, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (3) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8214, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (4) The 3'UTR element contains sequence number 8209, the 5'UTR element contains sequence number 8243, and the poly(A) tail contains sequence number 8306 or sequence number 8307, (5) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8259, and the poly(A) tail contains sequence number 8306 or sequence number 8307, or (6) The 3'UTR element contains sequence number 8281, the 5'UTR element contains sequence number 8263, and the poly(A) tail contains sequence number 8306 or sequence number 8307, The system according to any one of claims 24 to 27.

34. The system according to any one of claims 24 to 33, wherein the polypeptide is a heterogeneously modified polypeptide or a retrotransposon gene-modified polypeptide.

35. The system according to any one of claims 24 to 33, wherein the polypeptide comprises a reverse transcriptase domain and an endonuclease domain.

36. The system according to claim 35, wherein the endonuclease domain is a Cas9 domain selected from nickerse domains, such as the SpCas9 domain, BlatCas9 domain, Nme2 Cas9 domain, PnpCas9 domain, SauCas9 domain, SauCas9-KKH domain, SauriCas9 domain, SauriCas9-KKH domain, ScaCas9-Sc++ domain, SpyCas9 domain, SpyCas9-NG domain, SpyCas9-SpRY domain, or St1Cas9 domain, and the reverse transcriptase domain is selected from retroviral reverse transcriptase domains.

37. The system according to claim 36, wherein the Cas9 domain comprises an N670A mutation, an N611A mutation, an N605A mutation, an N580A mutation, an N588A mutation, an N872A mutation, an N863A mutation, an N622A mutation, or an H840A mutation.

38. The system according to any one of claims 35 to 37, wherein the retroviral reverse transcriptase domain is a reverse transcriptase domain derived from a gamma retrovirus, preferably comprising an amino acid sequence of a reverse transcriptase domain sequence from a family selected from AVIRE, BAEVM, FFV, FLV, FOAMV, GALV, KORV, MLVAV, MLVBM, MLVCB, MLVFF, MLVMS, PERV, SFV1, SFV3L, WMSV, and XMRV6, and preferably the gamma retrovirus-derived reverse transcriptase domain is not derived from PERV.

39. The system according to any one of claims 35 to 38, wherein the reverse transcriptase domain contains one, two, three, four, five, or six or more mutations corresponding to the following mutations in the reverse transcriptase domain of mouse leukemia virus reverse transcriptase: D200N, L603W, T330P, D524G, E562Q, D583N, P51L, S67R, T197A, H204R, E302K, F309N, W313F, L435G, N454K, H594Q, L671P, E69K, or D653N.

40. The system according to any one of claims 35 to 39, wherein the polypeptide comprises the amino acid sequence of SEQ ID NO: 8257 or SEQ ID NO: 8371.

41. The system according to any one of claims 24 to 40, wherein the template RNA further comprises an RT terminator sequence located between a heterogeneous object sequence and either (i) or (ii).

42. The system according to any one of claims 24 to 41, wherein the heterogeneous object sequence includes a sequence that codes for the target polypeptide or a portion thereof, or a sequence that is the inverse complementary chain of the sequence that codes for the target polypeptide or a portion thereof.

43. The polypeptide contains a reverse transcriptase domain and an endonuclease domain, the endonuclease domain is a Cas9 domain, and the template RNA is (i) A gRNA spacer that is complementary to the first portion of the target gene and optionally contains one or more consecutive nucleotides starting at the 3' end of an adjacent nucleotide of the gRNA spacer. (ii) gRNA scaffold that binds to the Cas9 domain, (iii) heterologous object sequences containing mutation regions for introducing mutations into a second portion of the target gene (for example, to correct mutations within it) (where the heterologous sequence may include a 5' to 3' post-edit homologous region, a mutation region and a pre-edit homologous region), and (iv) A primer binding site (PBS) sequence containing at least 5, 6, 7, or 8 bases that is 100% identical to the third portion of the target gene. A system according to any one of claims 24 to 41, including the system described in any one of claims 24 to 41.

44. The target gene is the human PAH gene, and the template RNA is, (i) A gRNA spacer that is complementary to the first portion of the human PAH gene, has a sequence containing the core nucleotide of the gRNA spacer sequence, and optionally contains one or more consecutive nucleotides starting at the 3' end of the nucleotide adjacent to the gRNA spacer. (ii) gRNA scaffold that binds to the Cas9 domain, (iii) a heterologous object sequence containing a mutation region for introducing a mutation into the second portion of the human PAH gene (for example, to correct a mutation within it) (where the heterologous object sequence may include a post-edit homologous region from 5' to 3', a mutation region and a pre-edit homologous region), and (iv) A primer-binding site (PBS) sequence containing at least 5, 6, 7, or 8 bases that is 100% identical to the third portion of the human PAH gene. The system according to claim 43, including the system described in claim 43.

45. The system according to any one of claims 24 to 44, wherein the reverse transcriptase domain and the endonuclease domain are linked by a peptide linker.

46. The system according to any one of claims 24 to 45, wherein the target site is a target site within the human genome.

47. A reaction mixture comprising cells and the system according to any one of claims 24 to 46.

48. The reaction mixture according to claim 47, wherein the cells are T cells (e.g., primary T cells).

49. A reaction mixture comprising DNA containing a target site and the system according to any one of claims 24 to 47.

50. An artificial RNA molecule according to any one of claims 1 to 23, comprising one or more chemically modified nucleotides.

51. A deoxyribonucleic acid (DNA) molecule encoding an artificial RNA molecule according to any one of claims 1 to 23.

52. A pharmaceutical composition comprising an artificial RNA molecule according to any one of claims 1 to 23 and 50, a system according to any one of claims 24 to 47, or one or more nucleic acids encoding it, and a pharmaceutically acceptable excipient or carrier.

53. The pharmaceutical composition according to claim 52, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of plasmid vectors, viral vectors, vesicles, and lipid nanoparticles.

54. The pharmaceutical composition according to claim 53, wherein the viral vector is an adeno-associated virus.

55. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising the artificial RNA or system or DNA molecule described in any one of the above claims.

56. The host cell according to claim 55, wherein the host cell is a T cell (e.g., a primary T cell).

57. A method for producing an artificial RNA molecule according to any one of claims 1 to 23, comprising synthesizing template RNA by in vitro transcription (e.g., solid-state synthesis) or by introducing DNA encoding the artificial RNA molecule into a host cell under conditions that enable the production of template RNA.

58. (a) a system according to any one of claims 24 to 47, a reaction mixture according to any one of claims 47 to 49, a DNA molecule according to claim 51, or a pharmaceutical composition according to any one of claims 52 to 54, and (b) Instructions for use of the system, reaction mixture, DNA molecule, or pharmaceutical composition. A kit that includes this.

59. Lipid nanoparticles (LNPs) comprising an artificial RNA molecule according to any one of claims 1 to 23 and 50, or a system according to any one of claims 24 to 47.

60. A method for modifying a target site in the genomic DNA of a cell, comprising contacting the cell with the system described in any one of claims 24 to 47, or one or more nucleic acids encoding the system, thereby modifying the target site in the genomic DNA of the cell.

61. A method for treating a subject having a disease or condition associated with a genetic defect, comprising administering to the subject the system described in any one of claims 24 to 47, thereby treating the subject having a disease or condition associated with a genetic defect.