Multi-element system for site-specific genome modification

The genome editing system employing RTC and GIC derived from non-LTR retrotransposons addresses the challenges of immune responses and non-site-specific integration by enabling efficient, site-specific transgene insertion into the genome, enhancing research and clinical applications.

JP2025517630APending Publication Date: 2025-06-10RGT UNIV OF CALIFORNIA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024564803
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-02
Filing Date
2023-05-02
Publication Date
2025-06-10

Smart Images

  • Figure 2025517630000001_ABST
    Figure 2025517630000001_ABST
Patent Text Reader

Abstract

The present invention includes systems, compositions, and methods for performing modular gene editing by processes related to reverse transcriptase. More specifically, the present invention provides systems and methods using modified nucleotides and modified peptides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to related applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 337,564, filed May 2, 2022, the disclosure of which is hereby incorporated by reference in its entirety for all purposes.

[0002] Reference to the Sequence Listing This application is filed with a sequence listing in electronic form. This sequence listing is created in [[XX]] with the file name [[XXXX.xml]] and has a size of [[XX]] bytes. The information recorded in this electronic sequence listing is hereby incorporated by reference in its entirety into this specification.

[0003] Statement regarding government support This invention was made with government support under grant numbers GM139306 and HL156819 awarded by the National Institutes of Health. The United States government has certain rights in this invention.

Background Art

[0004] The introduction of transgenes into the genomes of eukaryotes enables the improvement, correction, and / or modification of gene expression, and at the same time, helps to treat and relieve the symptoms of diseases. If a transgene can actually be inserted, it is possible to restore loss-of-function mutations, inhibit gain-of-function mutations, externally control RNA expression and / or protein expression, introduce specificity of isoform expression, express recombinant genes and recombinant proteins, and obtain other useful results.

[0005] However, there are still significant hurdles to overcome in current methods of introducing genetic material into cells and inserting it into the genome. For example, in methods of delivering DNA to target cells, the DNA needs to pass through the cytoplasm, which often induces destructive or harmful immune responses. Also, in site-specific integration methods where DNA is introduced into the genome by homologous recombination (HR), mutagenic double-strand DNA breaks may be introduced, and there is a risk of disrupting the target genome or epigenome at the integration site. In higher eukaryotes, especially in post-mitotic cells, DNA integration is often non-site-specific. This is because homologous recombination is suppressed during most of the cell cycle, and non-homologous end joining is favored.

[0006] Any method that can flexibly accommodate the length of the DNA to be introduced and effectively and site-specifically insert a transgene into the genome of living cells without introducing the DNA into the cytoplasm would make a great contribution to human biology, animal biology, microbial biology, and plant biology, leading to a strong advancement in research and clinical applications.

[0007] As one such method, a method of introducing a transgene sequence as RNA and using it as a template for complementary DNA (cDNA) synthesis by reverse transcriptase (RT) can be considered. However, at present, molecular signals that can induce RNA introduced into mammalian cells to be replicated as a template for inserting a transgene into the genome have not been identified.

[0008] In addressing such problems, a group of genes known as non-long terminal repeat (non-LTR) retroelements (REs) or non-LTR retrotransposons, which are synonymous with the former, can be a promising solution. These genes are capable of self-amplification within the host genome and exert their function by expressing a non-LTR retrotransposon reverse transcriptase (RT) protein. This RT protein binds to its own retroelement transcript RNA, uses it as a template, and synthesizes cDNA (RT primer extension) by using the nick introduced into genomic DNA (by the catalytic action of the endonuclease (EN) domain of this RT protein) as a primer for the initiation of cDNA synthesis. This process is known as target-primed reverse transcription (TPRT), by which a copy of the double-stranded DNA retroelement is added into the genome.

[0009] WO2022 / 155055 describes a system consisting of two components for site-specific transgene insertion into a safe harbor site of the human genome. As these two components, a non-LTR retroelement reverse transcriptase (RT) and a template RNA are used. This non-LTR retroelement reverse transcriptase (RT) has been recombined to enable the insertion of the full length of the transgene instead of the natural retroelement that tends to insert and cleave at the 5' side, and the template RNA is adapted to this non-LTR retroelement reverse transcriptase (RT). The synthesis mechanism of the first DNA strand to be inserted is target-primed reverse transcription (TPRT) induced by the 3'-side module of the template RNA, and this reverse transcription is enhanced by the non-natural 3'-tail that is part of the 3'-side module. The 5'-side module of the template RNA imparts in vivo stability to the template RNA, increases the bioavailability of the template RNA when binding to the RT protein, and induces the synthesis of the second strand.

[0010] The present disclosure provides compositions and methods for inserting and expressing a transgene into the genome of a eukaryotic cell, particularly the genome of a human cell, by creating a biopolymer construct from a portion of a retroelement array. SUMMARY OF THE INVENTION Means for Solving the Problems

[0011] The present invention provides compositions, methods, and / or uses of proteins and nucleotides, and modified proteins and modified polynucleotides, for inserting a transgene into a target genome by target primed reverse transcription (TPRT) using components derived from non-long terminal repeat (non-LTR) retrotransposons.

[0012] The present invention is a genome editing system comprising (i) at least one reverse transcriptase construct (RTC) comprising a polynucleotide encoding a polypeptide having an enzyme activity for reverse transcribing a polynucleotide template, and (ii) at least one gene insertion construct (GIC) comprising at least one polynucleotide template suitable for reverse transcription by the polypeptide encoded by the at least one RTC and provides a system.

[0013] In some embodiments, the genome editing system comprises (i) at least one reverse transcriptase construct (RTC), and (ii) at least one gene insertion construct (GIC) wherein the RTC comprises at least one reverse transcriptase module (RTC:RT module), and the RTC:RT module comprises mRNA encoding a reverse transcriptase (RT), at least one 5'-side module (RTC:5'-side module), and / or at least one 3'-side module (RTC:3'-side module), and The GIC includes at least one RNA template suitable for reverse transcription by a polypeptide encoded by the at least one RTC, and the GIC includes at least one GIC:5'-side module, at least one GIC:payload module, and / or at least one GIC:3'-side module.

[0014] In some embodiments, the RT module includes mRNA encoding a reverse transcriptase derived from an organism selected from birds, arthropods, fish, urochordates, and other animals (including mammals and humans).

[0015] In some embodiments, the genome editing system i) an RTC 5'-side module including a 5' untranslated region (5'UTR), a Kozak sequence, an unnatural translation start codon, and / or a 5' cap; ii) an RT module including mRNA encoding a reverse transcriptase derived from an organism selected from the group consisting of Zonotrichia albicollis (ZoAl), Taeniopygia guttata (TaGu), Tinamus guttatus (TiGu), Oryzias latipes (OrLa), and Tribolium castaneum (strain B) (TriCasB); iii) an RTC 3'-side module including a translation stop codon of the reverse transcriptase, a 3' untranslated region (3'UTR), and a polyA tail; iv) a GIC:5'-side module including a sequence derived from the 5' region of a natural retroelement, an rRNA sequence, a ribozyme sequence, a folding motif sequence, and / or an RNA polymerase terminator sequence; (v) A GIC: payload module comprising at least one transgene ORF or non-coding RNA (ncRNA) sequence, a promoter sequence of the transgene, an internal ribosome entry site (IRES), a 5' untranslated sequence of the transgene, a 3' untranslated sequence of the transgene, a polyadenylation signal sequence of the transgene, and / or an ncRNA processing sequence of the transgene; and (iv) A GIC: 3'-side module comprising a reverse transcriptase recognition sequence, an rRNA sequence, and / or an A tract sequence comprising.

[0016] In some embodiments, the at least one reverse transcriptase construct comprises at least one biopolymer, which biopolymer comprises at least one nucleic acid, at least one amino acid, and any combination thereof. In some embodiments, the polynucleotide of the reverse transcriptase construct described in (i) above comprises mRNA encoding reverse transcriptase. In some embodiments, the template consisting of the polynucleotide of the gene insertion construct described in (ii) above comprises RNA. In some embodiments, the polynucleotide of the reverse transcriptase construct described in (i) above comprises mRNA encoding reverse transcriptase, and the template consisting of the polynucleotide of the gene insertion construct described in (ii) above comprises a different (distinct) RNA. In some embodiments, the gene insertion construct comprises an RNA template different from the mRNA encoding the reverse transcriptase described in (i) above.

[0017] In some embodiments, the at least one reverse transcriptase construct comprises at least one reverse transcriptase open reading frame (ORF) module (RTC: RT module), and may optionally comprise at least one 5' untranslated region (UTR) module (RTC: 5'-side module), at least one 3' UTR module (RTC: 3'-side module), and any combination thereof.

[0018] In some embodiments, the at least one reverse transcriptase module comprises at least one reverse transcriptase or encodes at least one reverse transcriptase.

[0019] In some embodiments, the at least one reverse transcriptase module comprises or encodes at least one reverse transcriptase derived from a non-long terminal repeat (non-LTR) retroelement.

[0020] In some embodiments, the at least one reverse transcriptase module comprises or encodes a non-native translation start codon.

[0021] In some embodiments, the at least one reverse transcriptase comprises at least one DNA binding domain, at least one RNA binding domain, at least one cDNA synthesis domain, at least one endonuclease domain, and any combination thereof.

[0022] In some embodiments, at least one of the at least one reverse transcriptase domain, the at least one target DNA binding domain, the at least one template RNA binding domain, and the at least one endonuclease domain, and any combination thereof, is derived from a different species than the species from which at least one of the remaining domains is derived.

[0023] In some embodiments, at least one 5'-side module of the reverse transcriptase construct comprises or encodes at least one RNA polymerase promoter, at least one 5' untranslated region (5'-UTR), at least one Kozak sequence, at least one 5' cap, and any combination thereof.

[0024] In some embodiments, at least one 3'-side module of the reverse transcriptase construct comprises, encodes, or consists of at least one reverse transcriptase translation termination codon, at least one 3' untranslated region (3'UTR), at least one polyA tract and / or polyA tail, and any combination thereof.

[0025] In some embodiments, the at least one reverse transcriptase module comprises, encodes, or consists of at least one structure described in FIGS. 2-5 or any combination thereof.

[0026] In some embodiments, the at least one reverse transcriptase construct comprises, encodes, or is encoded by at least one of SEQ ID NOs: 1-57. In some embodiments, the at least one reverse transcriptase construct comprises mRNA encoding a reverse transcriptase protein derived from a species selected from the group consisting of TriCasB, NaViB, OrLa, ZoAl, TiGu, TaGu, GeFo, DroSi, BoMo, DrMerc, DrMe, GaAc, PuPu, AdVa, HyMaA, CiIn, LiPo, TriCan, LeCo, and any combination thereof.

[0027] In some embodiments, the at least one gene insertion construct comprises, encodes, or consists of at least one nucleic acid biopolymer. In some embodiments, the gene insertion construct comprises template RNA.

[0028] In some embodiments, the at least one gene insertion construct comprises, encodes, or consists of at least one GIC: 5'-side module, which is an arbitrary component, at least one GIC: payload module, at least one GIC: 3'-side module, which is an arbitrary component, and any combination thereof.

[0029] In some embodiments, the at least one GIC:5' side module comprises or encodes at least one sequence derived from the 5' region of a natural retroelement and may comprise or encode at least one rRNA sequence, at least one ribozyme (RZ) sequence, at least one folding motif sequence, or any combination thereof.

[0030] In some embodiments, at least one rRNA sequence, which is any component of the GIC:5' side module, comprises or encodes a 1 - 30 nt rRNA of interest.

[0031] In some embodiments, at least one ribozyme sequence, which is any component of the GIC:5' side module, comprises or encodes at least one self - cleaving ribozyme, and the self - cleaving ribozyme may comprise the ribozyme of hepatitis delta virus (HDV).

[0032] In some embodiments, at least one ribozyme sequence, which is any component of the GIC:5' side module, comprises or encodes a ribozyme derived from the 5' region of at least one non - long terminal repeat retroelement. In some embodiments, at least one folding motif sequence, which is any component of the GIC:5' side module, comprises or encodes at least one self - folding RNA sequence motif, and the self - folding RNA sequence motif may comprise at least one hairpin motif, at least one stem - loop motif, at least one paired stem motif, or any combination thereof within the RZ.

[0033] In some embodiments, the GIC:5' side module comprises, or encodes, at least one of SEQ ID NOs: 60 to 154, or a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) to at least one of SEQ ID NOs: 60 to 154. In some embodiments, the GIC:5' side module comprises a sequence derived from a species selected from the group consisting of OrLa, TriCasB, TriCasA, ZoAl, TiGu, DroSi, LeCo, CiIn, FoRa, TriCan, HDV-28, HDV-24, HDV-21, HDV-13, HDV-36, and any combination thereof.

[0034] In some embodiments, the at least one GIC:3' side module comprises, or encodes, at least one reverse transcriptase recognition sequence, and may comprise, or encode, at least one rRNA sequence, at least one A tract sequence, or any combination thereof.

[0035] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module comprises, or encodes, at least one sequence that interacts with at least one reverse transcriptase. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module comprises a sequence selected from the group consisting of SEQ ID NOs: 200 to 224.

[0036] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is derived from the 3' region of a natural retroelement.

[0037] In some embodiments, at least one rRNA sequence, which is any component of the GIC:3' side module, comprises, or encodes, 1 to 30 nt of rRNA.

[0038] In some embodiments, at least one A tract sequence, which is any component of the GIC: 3' side module, comprises or encodes a sequence consisting of about 1 to 50 adenine bases.

[0039] In some embodiments, the at least one GIC: 3' side module comprises or encodes at least one of SEQ ID NOs: 200 to 224 or at least one of SEQ ID NOs: 300 to 329. In some embodiments, the GIC: 3' side module comprises a sequence derived from a species selected from the group consisting of OrLa, TriCasB, TaGu, GeFo, ZoAl, NaViB, DroSi, PuPu, LiPo, BoMo, GaAc, LeCo, CiIn, DrMe, DrNa, DrMer, TriCan, AdVa, HyMaA, and any combination thereof.

[0040] In some embodiments, the at least one GIC: payload module comprises or encodes at least one transgene ORF sequence, and may also comprise or encode at least one promoter sequence of the transgene, at least one 5' untranslated sequence of the transgene, at least one 3' untranslated sequence of the transgene, at least one polyadenylation signal sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, at least one ncRNA processing sequence, and / or other 3' end processing sequences or stabilization signals, or any combination thereof.

[0041] In some embodiments, the at least one transgene sequence comprises or encodes at least one sequence of interest for insertion into the genome of a subject.

[0042] In some embodiments, the at least one promoter sequence of the transgene comprises or encodes at least one sequence that promotes the expression of the transgene in the genome of a subject.

[0043] In some embodiments, the at least one GIC:payload module comprises at least one 5' untranslated sequence of the transgene, and the at least one 5' untranslated sequence comprises or encodes at least one 5' untranslated region of the mRNA of the transgene.

[0044] In some embodiments, the at least one 3' untranslated sequence of the transgene comprises or encodes at least one 3' untranslated region of the mRNA of the transgene.

[0045] In some embodiments, the at least one polyadenylation signal sequence of the transgene comprises or encodes at least one polyadenylation signal of the transgene.

[0046] In some embodiments, the at least one non-coding RNA (ncRNA) processing sequence and / or other 3' end processing sequence or stabilization signal of the transgene comprises or encodes at least one termination signal, at least one 3' processing signal, and any combination thereof of at least one ncRNA expressed from the transgene.

[0047] In some embodiments, the at least one GIC:payload module comprises or encodes a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) to at least one of SEQ ID NOs: 411-422 and SEQ ID NOs: 499-536 and any combination thereof.

[0048] In some embodiments, at least one of the at least one GIC:5'-side module and the at least one GIC:3'-side module comprises or encodes at least one sequence derived from a species different from the species from which the non-long terminal repeat retroelement from which the other is derived is derived.

[0049] In some embodiments, the at least one gene insertion construct comprises or encodes at least one structure shown in at least one of the drawings attached hereto, for example, at least one structure described in FIGS. 6-9 and any combination thereof.

[0050] In some embodiments, the genome editing system (i) at least one reverse transcriptase construct, and (ii) at least one gene insertion construct comprising, the at least one reverse transcriptase construct comprises or encodes or is encoded by at least one sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 1-57, the at least one gene insertion construct comprises at least one sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 60-154, 250-276, 278-279, 280-289, 300-329, 400-404, 405-407, 411-422 and 499-536. In some embodiments, the mRNA sequence transfected to produce the reverse transcriptase protein is split and expressed from a plasmid and encodes the amino acid sequences of multiple proteins.

[0051] In some embodiments, the genome editing system is (i) at least one reverse transcriptase construct, and (ii) at least one gene insertion construct comprising wherein the at least one reverse transcriptase construct comprises or is encoded by at least one sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 1 to 57, wherein the at least one gene insertion construct a GIC: 5'-side module comprising a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 60 to 154, an rRNA sequence which is an optional component comprising a sequence selected from the group consisting of SEQ ID NOs: 250 to 276, or a sequence having one, two or three nucleotide changes as compared with a sequence selected from the group consisting of SEQ ID NOs: 250 to 276, a GIC: payload module comprising at least one transgene sequence, a GIC: 3'-side module comprising a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 300 to 329, a reverse transcriptase recognition sequence of the GIC: 3'-side module, comprising a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 200 to 224, Selected from the group consisting of SEQ ID NOs: 280 to 289 and sequences containing one, two, or three nucleotide substitutions in SEQ ID NOs: 280 to 289, the rRNA sequence of the GIC: 3'-side module, and The A-tract sequence of the GIC: 3'-side module, containing 1 to 100 adenine bases comprises.

[0052] In some embodiments, the 5'UTR of the RTC 5'-side module comprises a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with SEQ ID NO: 58.

[0053] In some embodiments, the 3'UTR of the RTC 3'-side module comprises a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with SEQ ID NO: 59.

[0054] In some embodiments, the genome editing system comprises at least one of the gene insertion constructs described herein or a construct for synthesizing a gene insertion construct (GIC: synthesis construct) encoding the same.

[0055] In some embodiments, at least one of the at least one reverse transcriptase construct and the at least one gene insertion construct comprises or encodes at least one sequence derived from a species different from the species from which the retroelement from which the other is derived.

[0056] In some embodiments, the genome editing system comprises at least one of the combinations of (i) at least one reverse transcriptase construct described herein and (ii) at least one gene insertion construct described herein.

[0057] Furthermore, provided is a method of inserting at least one transgene into a target genome, the method comprising administering to the target, in an effective amount, at least one of the gene insertion systems (GIS) of the present disclosure.

[0058] In some embodiments, the transgene is inserted into one or more target sites of the target genome, and the one or more target sites may comprise at least one safe harbor site.

[0059] In some embodiments, the at least one safe harbor site, which is an optional component, comprises at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may comprise at least one 28S rDNA sequence.

[0060] In some embodiments, the at least one method comprises administering at least one of the gene insertion systems formulated with at least one delivery agent.

[0061] In some embodiments, the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may comprise at least one lipid nanoparticle.

[0062] Furthermore, provided is a pharmaceutical composition comprising at least one of the gene insertion systems according to the claims, and optionally comprising at least one additive, at least one delivery agent, at least one adjuvant, and any combination thereof.

[0063] Furthermore, provided is a method of treating a therapeutic indication in a subject in need of treatment of the therapeutic indication, the method comprising administering to the subject, in an effective amount, at least one of the gene insertion systems of the present disclosure or the pharmaceutical composition of the present disclosure.

[0064] In some embodiments, the therapeutic indication is caused by a deletion of telomerase activity.

[0065] In some embodiments, the at least one gene insertion system comprises at least one TERT transgene.

[0066] Furthermore, a kit for producing the gene insertion system of the present disclosure is provided. In some embodiments, the kit comprises the pharmaceutical composition of the present disclosure. In some embodiments, the kit may further comprise a buffer, a DNA plasmid, or a protocol for producing the gene insertion system or the pharmaceutical composition.

[0067] Furthermore, a method is provided that includes de novo design of a 5'-side module that recruits a host mechanism for introducing a nick into the second strand to synthesize the second strand. In some embodiments, in the method, (a) contains rRNA of a predetermined length (described herein) at a predetermined position, (b) promotes the folding of ribozyme (RZ), and / or (c) the de novo design of the 5'-side module configured to recruit host cell mechanisms increases the insertion efficiency.

[0068] In another aspect, the present disclosure provides a method of inserting at least one transgene into the genome of a cell, the method comprising contacting the cell with at least one of the gene insertion systems (GIS) of the present disclosure.

[0069] In some embodiments, the transgene is inserted into one or more target sites of the genome of the subject, and the one or more target sites may include at least one safe harbor site. In some embodiments, the at least one safe harbor site, which is an optional component, includes at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may include at least one 28S rDNA sequence.

[0070] In some embodiments, the method includes administering at least one of the gene insertion systems formulated with at least one delivery agent. In some embodiments, the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may include at least one lipid nanoparticle.

[0071] In some embodiments, the transgene is inserted with a target site specificity of greater than 90% (e.g., a target site specificity of greater than 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99%).

[0072] In some embodiments, the reverse transcriptase construct includes RNA encoding a reverse transcriptase derived from Zoanthus sp. (ZoA1), Taeniopygia guttata (TaGu) or Thamnophis gigas (TiGU), or includes an amino acid sequence having at least 90% identity with SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 29 or SEQ ID NO: 25.

[0073] In some embodiments, the transgene is expressed at the target site over a period of 3 months or more.

[0074] In some embodiments, the cells are contacted with the gene insertion system, and the molar ratio of the reverse transcriptase construct to the gene insertion construct is about 10:1 to 1:20.

[0075] In some embodiments, the method is an in vitro method, an ex vivo method or an in vivo method.

[0076] In some embodiments, the cells are selected from the group consisting of primary cells, transformed cells, epithelial cells, fibroblasts, human cells, monkey cells and mouse cells.

[0077] In some embodiments, the cells are allogeneic cells or autologous cells. In some embodiments, the autologous cells are cells with a matched HLA.

[0078] The present invention includes these combinations as if all combinations of the specific embodiments described herein were described in detail.

Brief Description of the Drawings

[0079]

Figure 1

[0080]

Figure 2

[0081]

Figure 3

[0082]

Figure 4

[0083]

Figure 5

[0084]

Figure 6

[0085]

Figure 7

[0086]

Figure 8

[0087]

Figure 9

[0088]

Figure 10

[0089]

Figure 11

[0090]

Figure 12A - 12B

[0091]

Figure 13

[0092]

Figure 14A - 14B

[0093]

Figure 15

[0094]

Figure 16

[0095]

Figure 17

[0096]

Figure 18A - 18B

[0097]

Figure 19

[0098]

Figure 20

[0099]

Figure 21

Mode for Carrying Out the Invention

[0100] I. Introduction Throughout the following detailed description and the entire specification, unless inappropriate or otherwise stated, the terms "a" and "an" mean one or more, and the term "or" means "and / or". The examples and embodiments described in this specification are for illustrative purposes only. From this, those skilled in the art are suggested that various improvements and changes can be made, and these improvements and changes are included in the gist and scope of this application and the scope of the appended claims. All publications, patents, and patent applications cited in this specification (including the cited references in these documents) are hereby incorporated by reference in their entirety for all purposes.

[0101] The present invention provides a system and method for genome editing and / or gene modification, including inserting a transgene into a target genome. As used herein, the system of the present invention, referred to as "gene insertion system (GIS)", may include (a) at least one reverse transcriptase or at least one reverse transcriptase (RT) construct (RTC) encoding the same, and (b) at least two components including an RNA construct used as a template for reverse transcription or encoding the same and expressed separately from the RTC, i.e., a gene insertion construct (GIC) (i.e., a GIS consisting of two components). As used herein, the term "construct" means any biopolymer artificially designed or synthesized. This biopolymer may be composed of, for example, nucleic acids (e.g., DNA or RNA), amino acids, or any combination thereof. In some embodiments, both (a) and (b) are RNA constructs. In some embodiments, (a) is an amino acid construct (i.e., a protein) and (b) is an RNA construct.

[0102] Also provided is a recombinant reverse transcriptase construct (RTC) capable of performing target primed reverse transcription (TPRT). As used herein, the term "target primed reverse transcription (TPRT)" means a process in which reverse transcriptase uses the available 3'-end of the DNA at the target site as a primer to initiate cDNA synthesis.

[0103] Furthermore, the systems and methods provided herein may be capable of inserting a transgene at a sequence-specific position of a target DNA (referred to herein as a "target site"), such as a safe harbor site. As used herein, the terms "safe harbor" and "safe harbor site" mean a site in the genome of a subject that, for example, does not adversely affect the function of the subject's cells even if the target DNA sequence is disrupted by the insertion of a heterologous sequence. Exemplary safe harbor sites utilized in the present invention are present in a portion of the subject's genome that encodes ribosomal RNA (rRNA). An example of a safe harbor site is an rRNA precursor transcribed by RNA polymerase I encoded by a locus referred to herein as the "ribosomal DNA (rDNA) locus", which rRNA precursor contains sequences encoding 5.8S rRNA, 18S rRNA, or 28S rRNA.

[0104] The present disclosure demonstrates that it can be programmed to insert a DNA transgene into a safe harbor site of the genome of a cell (e.g., a human cell) by delivering only RNA. In some embodiments, an RNA template encoding the transgene to be inserted and a messenger RNA encoding the reverse transcriptase necessary to convert the RNA template into genomic DNA are delivered to the cell. The method of delivering only RNA of the present invention is expected to facilitate the transition to human gene therapy by utilizing an RNA delivery mechanism that targets specific types of cells in a non-toxic and highly efficient manner, which is currently undergoing technological innovation.

[0105] In some embodiments, the expression of reverse transcriptase (RT) from a plasmid is combined with the transfection of an RNA template. In some embodiments, by heterologously combining a 5'-side module of a template of a transgene, which contains a natural R2 retroelement sequence or a part thereof, with reverse transcriptase, there is an advantage that a full-length sequence can be site-specifically inserted instead of inserting a truncated retroelement sequence. In some embodiments, the template RNA includes a 3'-side module having a 3'UTR sequence of a retroelement derived from the same species as the reverse transcriptase. In some embodiments, this 3'UTR further includes a 3' polyA tract that increases the specific insertion efficiency into the target site.

[0106] The present disclosure provides the following improvements and advantages compared to prior art systems and methods.

[0107] The inventors were able to demonstrate the following. (i) The reverse transcriptase (RT) protein derived from birds showed significant activity for the insertion of a transgene, and the transgene was functionally expressed in more than 20% of the transfected cells. The reverse transcriptase derived from birds showed extremely high selectivity for the replication of template RNA containing a 3'UTR and a 3' polyA tract derived from birds.

[0108] (ii) By heterologously combining the 3'UTR of the R2 retroelement derived from birds with the reverse transcriptase protein, it can be more effective than the natural combination.

[0109] (iii) The de novo generated and optimized non-natural 5'-side module is more effective, resulting in a more than one-digit increase in site-specific insertion efficiency.

[0110] (iv) The natural 5'-side modules (such as TCA, TCA5, TCARZ, etc.) derived from TriCasA are derived from R2 retroelements of a clade that is completely different from avian reverse transcriptase proteins, and such natural 5'-side modules derived from TriCasA can be even more effective.

[0111] (v) Instead of transfecting template RNA after plasmid expression of reverse transcriptase, the transgene is inserted by co-transfecting and delivering two RNA systems.

[0112] (vi) Since multiple transgenes can be inserted into a single cell by transfecting two RNA systems, multiplexing of gene delivery becomes possible with a single RNA administration. By such a method, it becomes possible to insert multiple therapeutic transgenes into the genome of a single cell, and such multiple therapeutic transgenes include multiple transgenes encoding therapeutic proteins respectively, multiple transgenes encoding separate subunits of multiple therapeutic proteins respectively, combinations of therapeutic proteins, and combinations of therapeutic RNAs.

[0113] (vii) By delivering two RNA systems, transgenes can be expressed in a wide variety of cells, and such cells include primary cell lines, non-dividing cells, and cells with a slow division rate, and include mouse cells, monkey cells, and human cells.

[0114] (viii) Genome sequencing has shown that the insertion of transgenes is site-specific.

[0115] (ix) The inserted transgene expression cassette stably expresses over several months.

[0116] Components derived from retro - elements

[0117] The reverse transcriptase construct (RTC) and / or gene insertion construct (GIC) of the present invention may be derived from a part of at least one non-long terminal repeat (non-LTR) type retroelement and / or may contain components (also referred to as "modules" in the same sense) that do not exist in nature. Without wishing to be bound by any theory, FIG. 1 (upper figure) shows a target genome containing a natural retroelement 100, in this case, a non-long terminal repeat (non-LTR) type retroelement. As can be seen from this figure, the target DNA 110 may contain at least one target insertion site 120, and the natural retroelement 130 may be inserted into this target insertion site. In the enlarged view (lower figure), the structure of a natural retroelement as an example is further examined in more detail. In this figure, the 5' region 131 of the natural retroelement is located before the translation start site 132. The 5' region of the natural retroelement is usually not translated into an amino acid biopolymer. The 5' region of the natural retroelement may contain nucleic acid sequences that are recognized by the reverse transcriptase (RT) of this natural retroelement itself in later insertions and / or affect the synthesis of the second strand of the natural retroelement. The translation start site 132 is the first nucleotide that is translated into an amino acid. The open reading frame 133 of the reverse transcriptase of the natural retroelement encodes a reverse transcriptase that can recognize, bind to, and use the RNA transcript of this natural retroelement itself as a template for reverse transcription. The open reading frame of the reverse transcriptase of the natural retroelement extends to the translation termination site 134, but this open reading frame and the translation termination site are distinguished. The 3' region 135 of the natural retroelement is usually not translated into an amino acid biopolymer and may contain nucleic acid sequences that are recognized by the reverse transcriptase of the natural retroelement itself. The 5' region 131 and the 3' region 135 may or may not exist. When the 5' region 131 and the 3' region 135 exist, they may replicate the surrounding target site sequences and / or contain sequences not encoded by the RNA template of the natural retroelement.

[0118] Suitable retroelements from which the components of the gene insertion system (GIS) may be derived include, for example, but are not limited to, non-LTR retroelements of the RLE type, APE type or Penelope type. The non-LTR retrotransposon of the RLE type may be derived from any one of a number of clades including, but not limited to, R2, R4, CRE, Genie, HERO, NeSL. The non-LTR retrotransposon of the APE type may be derived from any one of a number of clades including, but not limited to, I, R1, L1, Tx1, CR1, Rex1, Jockey, L2, Tad, RTE, RTEX, Ingi, Vingi, TRAS, SART or any combination thereof. In some embodiments, the components of the gene insertion system may be derived from a retroelement inserted into rDNA, i.e., a so-called R factor, for example, from a retroelement of the R1 clade or R2 clade. In some embodiments, the retroelement of the R2 clade may have specificity for the insertion site of the standard R2 retroelement, or may be derived from the R8 retroelement and / or R9 retroelement derived from the upper clade of the R2 clade by changing the target sequence of the standard R2 retroelement, or may be derived from the R2NS retroelement derived by losing specificity for the target site.

[0119] Each component of the gene insertion system (GIS) may be derived from a part or domain of a retroelement obtained from some biological species, including species phylogenetically distant from the target. For example, suitable retroelements from which each component of the gene insertion system may be derived include birds (e.g., Zonotrichia albicollis, Taeniopygia guttata, Tinamus guttatus, and Geospiza fortis), fish (e.g., Pungitis pungitis, Oryzias latipes, Danio rerio, Oryzias melastigma, Petromyzon marinus, Salmo trutta, Salmo salar, or Gasterosteus aculeatus), insects (e.g., Drosophila mercatorum, Drosophila melanogaster, Nasonia vitripennis, Tribolium castaneum, Drosophila simulans, Apis cerana, and Bombyx mori), crustaceans (e.g., Lepidurus couesii and Triops cancriformis), other invertebrates (e.g., Limulus polyphemus, Hydra magnipapillata, or Adineta vaga), chordates (e.g., Ciona intestinalis), mammals, and retroelements found in any combination of these.

[0120] In some embodiments, each component of the gene insertion system may be derived from a part or domain of any sequence disclosed herein.

[0121] II. Composition of the Gene Insertion System (GIS) Throughout the description of the present disclosure, the system of the present invention for inserting genetic material (e.g., a transgene) into a target genome is referred to as a "gene insertion system (GIS)". The gene insertion system of the present disclosure may be composed of a plurality of biopolymer constructs, and these biopolymer constructs are co-administered to insert at least one transgene via target primed reverse transcription (TPRT). These biopolymer constructs may be biopolymers consisting of amino acids, biopolymers consisting of nucleic acids, hybrid biopolymers containing both amino acids and nucleic acids, or any combination thereof. In some examples, the gene insertion system of the present disclosure consists of at least two biopolymers, namely, at least one reverse transcriptase construct (RTC) and at least one gene insertion construct (GIC). In such an example, the reverse transcriptase construct (RTC) includes, for example, a reverse transcriptase or means for performing reverse transcription by encoding a reverse transcriptase, and the gene insertion construct (GIC) includes or encodes at least one RNA sequence that may be used as a template for synthesizing cDNA by the reverse transcriptase construct (RTC).

[0122] Since the biopolymer constructs of the present invention are themselves composed of a plurality of modules, the gene insertion system of the present disclosure may be modified by combining these modules as needed so as to exhibit a desired function. As used herein, the term "module" means a part of a construct defined by its function (e.g., a functional domain of a protein) or a part of a construct defined by its sequence (e.g., an amino acid sequence or a nucleic acid sequence).

[0123] Reverse Transcriptase Construct (RTC) The gene insertion system of the present invention includes an active reverse transcriptase protein such as a reverse transcriptase derived from a non-LTR retroelement, or at least one reverse transcriptase construct (RTC) encoding the same. As used herein, the term "RTC" means a biopolymer construct that includes at least one reverse transcriptase (RT) or encodes at least one reverse transcriptase (RT). In some embodiments, at least one reverse transcriptase construct (RTC) used in the gene insertion system of the present invention may include, but is not limited to, an amino acid biopolymer such as a polypeptide, protein, proprotein, or any combination thereof. In some embodiments, at least one reverse transcriptase construct (RTC) used in the gene insertion system of the present invention may include, but is not limited to, a nucleic acid biopolymer such as RNA, DNA, or any combination thereof. In some embodiments, at least one reverse transcriptase construct (RTC) may include at least one mRNA construct.

[0124] Structure of the Reverse Transcriptase Construct (RTC) The reverse transcriptase construct (RTC) of the present invention may include at least one RTC: reverse transcriptase module (RTC:RT module), at least one 5'-side module (RTC:5'-side module) which is an optional component, at least one 3'-side module (RTC:3'-side module) which is an optional component, and any combination thereof. In some examples of RTC, the RTC:5'-side module and the RTC:3'-side module may be optional components, and one or both of them may not be present. In some embodiments, at least one RTC may include a linear RNA biopolymer or may be delivered as a linear RNA biopolymer to a target. In some embodiments, at least one RTC may include an mRNA biopolymer or may be delivered as an mRNA biopolymer to a target.

[0125] Referring to FIG. 2, an exemplary structure of a reverse transcriptase construct (RTC) 200 that is a linear RNA biopolymer (e.g., mRNA) is shown. As shown in this figure, in the RTC that is an mRNA biopolymer, the RTC:5′-side module 210 is an optional component of the RTC and, when present within the RTC, may include sequences that modify the immunogenicity of the RTC and / or sequences that control the expression of the RTC:RT module 220. For example, the RTC:5′-side module may include at least one 5′ cap (e.g., Clean Cap AG from TriLink, m7(3′OMeG)(5′)ppp(5′)(2′OMeA)pG), at least one 5′ untranslated region (5′-UTR), at least one Kozak sequence, at least one promoter, and any combination thereof, and may encode these. The start codon is a three-base nucleic acid sequence known to initiate translation and defines the 5′ end of the RTC:RT module. The RTC:RT module (detailed later) includes the region extending from the start codon to the stop codon, but does not include the stop codon. The RTC:3′-side module 230, which is an optional component, when present within the RTC, includes the region extending from the stop codon to the 3′ end of the RTC. The RTC:3′-side module, when present within the RTC, may include sequences that modify the immunogenicity of the RTC and / or sequences that control the expression of the RTC:RT module. For example, the RTC:3′-side module may include a translation stop codon, 3′UTR, polyadenosine sequence, polyadenylation signal, or any combination thereof, and may encode these.

[0126] In some embodiments, at least one RTC may contain a plasmid or may be delivered as a plasmid to a subject. In some embodiments, at least one RTC may contain mRNA or pro-mRNA or may be delivered as mRNA or pro-mRNA to a subject. In some embodiments, at least one RTC may contain a protein or may be delivered as a protein to a subject. In some embodiments, at least one RTC may contain a proprotein or may be delivered as a proprotein to a subject.

[0127] RTC: RT module The RT module of the RTC contains or encodes at least one compound or composition having reverse transcription activity. Specific examples of such compounds or compositions include, but are not limited to, a group of enzyme proteins known as reverse transcriptase (RT). In some embodiments, the RT module may contain or encode a biopolymer derived from at least one reverse transcriptase found in retroelement genes (i.e., retroelement reverse transcriptase). In some embodiments, the RTC:RT module contains or encodes at least one reverse transcriptase derived from a non-long terminal repeat (non-LTR) type retroelement.

[0128] Reverse transcriptase As used herein, the term "reverse transcriptase (RT)" is used in the broadest sense and means any biopolymer having reverse transcription activity. In some embodiments, the RT used in the present invention is a non-LTR type RT or its genome derived from Zonotrichia albicollis, Taeniopygia guttata, Tinamus guttatus, Geospiza fortis, Pungitis pungitis, Oryzias latipes, Danio rerio, Oryzias melastigma, Petromyzon marinus, Salmo trutta, Salmo salar, Gasterosteus aculeatus, Drosophila mercatorum, Drosophila melanogaster, Nasonia vitripennis, Tribolium castaneum, Drosophila simulans, Apis cerana, Bombyx mori, Lepidurus couesii, Triops cancriformis, Limulus polyphemus, Hydra magnipapillata, Adineta vaga, Ciona intestinalis, other birds, other arthropods, other fish, other urochordates, other animals (including mammals and humans), or may be derived from these non-LTR type RTs or their genomes.

[0129] In some embodiments, at least one RTC:RT module used in the gene insertion system (GIS) of the present disclosure may include at least one of SEQ ID NOs: 1-57, may encode at least one of SEQ ID NOs: 1-57, or may be encoded by at least one of SEQ ID NOs: 1-57. In some embodiments, the at least one RTC:RT module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 1-57, may encode this sequence, or may be encoded by this sequence. In some embodiments, the RTC:RT module includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity with SEQ ID NOs: 1-57.

[0130] In some embodiments, the at least one RTC:RT module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 17-21 (ZoA1 RT sequence), may encode this sequence, or may be encoded by this sequence.

[0131] In some embodiments, the at least one RTC:RT module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 26-29 (TaGu RT sequence), may encode this sequence, or may be encoded by this sequence.

[0132] In some embodiments, at least one RTC:RT module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 1-5 (TriCasB RT sequences).

[0133] In some embodiments, the RTC:RT module may include, or may encode, a protein shown to have TPRT activity by a suitable TPRT assay. A suitable TPRT assay, by way of example, includes the steps of: (i) transfecting a cell population with an expression plasmid encoding an RT protein having a suitable tag (e.g., a FLAG tag) for affinity purification; (ii) lysing the cell population and recovering and purifying the expressed protein product by any suitable method known in the art; (iii) preparing recombinant template RNA by any method known in the art (e.g., T7 RNA polymerase); (iv) combining the purified RT protein, the recombinant template, and a nucleotide solution containing a target site oligonucleotide double-stranded DNA having a radiolabeled lower strand in a medium that promotes reverse transcription by RT; and (v) recovering and analyzing the product by any suitable method known in the art (e.g., denaturing PAGE), but is not limited thereto.

[0134] Reverse transcriptases (RTs) suitable for use in the present invention may be composed of multiple functional domains. In some embodiments, for example, as shown in FIG. 3, at least one reverse transcriptase 300 includes at least one DNA binding domain 310, at least one RNA binding domain 320, at least one cDNA synthesis domain 330, at least one endonuclease domain 340, and any combination thereof. It should be noted that in this figure, only one configuration that can be arranged for each domain is presented. In some embodiments, each of the illustrated domains may be included in this reverse transcriptase (RT) at various frequencies, and / or these domains may be included in any order. In some embodiments, the DNA binding domain or the RNA binding domain may be derived from a different type of polypeptide than the reverse transcriptase (RT), and may be a sequence not known in the eukaryotic genome (for example, a de novo generated DNA binding domain or RNA binding domain).

[0135] Start codon Using various methods known in the art, at least one non-natural translation initiation codon may be added to the nucleic acid sequence encoding reverse transcriptase. This non-natural translation initiation codon may be added at any position in the sequence derived from the non-LTR retroelement that can induce the production of functional reverse transcriptase. For example, in a wild-type non-LTR retroelement, from a known reference point (e.g., from the amino acid sequence motif of the ORF of the native retroelement reverse transcriptase), about 1 base, about 10 bases, about 50 bases, about 100 bases, about 150 bases, about 200 bases, about 250 bases, about 300 bases, about 350 bases, about 400 bases, about 450 bases, about 500 bases, about 550 bases, about 600 bases, about 650 bases, about 700 bases, about 750 bases, about 800 bases, about 850 bases, about 900 bases, about 950 bases, or at least one non-natural start codon may be added at a position more than about 1000 bases away. The position of the translation initiation codon may be selected according to various considerations obvious to those skilled in the art engaged in optimizing or regulating protein expression in the target cells of interest using recombinant techniques, depending on the length, sequence composition, activity, biological stability, avoidance of aggregation or localization of the polypeptide, and / or may be selected so as to improve the biological stability of the mRNA encoding the protein.

[0136] The translation initiation codon may be any three nucleotides known to be translated by ribosomes, depending on or independent of another sequence or structure in the mRNA. In some embodiments, the non-natural translation initiation codon is AUG.

[0137] RTC: 5'-side module The reverse transcriptase construct (RTC) of the present invention may include at least one RTC:5'-side module. Usually, the RTC:5'-side module includes a biopolymer component that is not translated, and this non-translated biopolymer component can, for example, modify the immunogenicity of the gene insertion construct, assist in localizing the gene insertion construct to the target intracellular region, control or modify the expression of the RTC:RT module contained in the gene insertion construct, label the gene insertion construct for identification, assist in purifying the gene insertion construct, control the degradation of the gene insertion construct, exogenously or endogenously regulate the activity and / or function of the gene insertion construct, and any combination thereof, but is not limited thereto.

[0138] In some embodiments, at least one RTC:5'-side module may include at least one 5'UTR and may encode the same. In some embodiments, at least one RTC:5'-side module may include at least one 5' cap and may encode the same. In some embodiments, at least one RTC:5'-side module may include at least one microRNA binding sequence and may encode the same. In some embodiments, at least one RTC:5'-side module may include at least one RNA polymerase promoter and may encode the same.

[0139] In some embodiments, at least one RTC:5'-side module used in the gene insertion system of the present disclosure includes the 5'UTR of SEQ ID NO: 58.

[0140] In multiple embodiments, the inventors used one 5’UTR and one 3’UTR for the transfected mRNA, obtained from the vaccine sequences of BioNTech reported by the WHO. Further, instead of using polyA polymerase after transcription, the polyA region encoded by its template was used. This polyA region was composed of 30 adenosines, a 10nt linker, and 70 adenosines. Further, a Type IIS restriction site was introduced to cleave the template for mRNA transcription without adding extra 3’ terminal bases. All mRNAs were capped with TriLink’s AG clean cap, i.e., m7(3’OMeG)(5’)ppp(5’)(2’OMeA)pG). The UTRs were selected to express reverse transcriptase tissue-specifically, for example, to perform specific translational regulation depending on the cell type.

[0141] In some embodiments, the RTC:5’ side module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 58.

[0142] RTC: 3'-side module The reverse transcriptase construct (RTC) of the present invention may comprise at least one RTC:3' side module. Usually, the RTC:3' side module contains a non-translated biopolymer component, and this non-translated biopolymer component may, for example, modify the immunogenicity of a gene insertion construct, assist in localizing the gene insertion construct to a target intracellular region, control or modify the expression of the RTC:RT module contained in the gene insertion construct, label the gene insertion construct for identification, assist in purifying the gene insertion construct, control the degradation of the gene insertion construct, exogenously or endogenously regulate the activity and / or function of the gene insertion construct, and any combination thereof, but is not limited thereto.

[0143] In some embodiments, at least one RTC:3' side module may comprise at least one 3'UTR. In some embodiments, at least one RTC:3' side module may comprise at least one polyA tract, i.e., polyA tail, and may encode this. In some embodiments, at least one RTC:3' side module may comprise at least one microRNA binding sequence, and may encode this.

[0144] In some embodiments, at least one RTC:3' side module used in the gene insertion system of the present disclosure comprises a 3'UTR and a polyA tail shown in SEQ ID NO: 59.

[0145] In some embodiments, the RTC:3' side module comprises a 3'UTR having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 59.

[0146] Modularity of the Reverse Transcriptase Construct (RTC) The reverse transcriptase construct (RTC) of the present invention may be designed by combining at least one RTC:RT module, at least one RTC:5'-side module as an optional component, and / or at least one RTC:3'-side module as an optional component so as to obtain a desired function or activity. In some embodiments, the reverse transcriptase construct includes at least one RTC:5'-side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:3'-side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:RT module. In some embodiments, the reverse transcriptase construct includes at least one RTC:5'-side module, at least one RTC:RT module, and at least one RTC:3'-side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:5'-side module and at least one RTC:RT module. In some embodiments, the reverse transcriptase construct includes at least one RTC:RT module and at least one RTC:3'-side module.

[0147] In some embodiments, the reverse transcriptase construct of the present invention may not include both at least one RTC:5'-side module and at least one RTC:3'-side module. In some embodiments, the reverse transcriptase construct of the present invention may not include either at least one RTC:5'-side module or at least one RTC:3'-side module. In some embodiments, the reverse transcriptase construct of the present invention may not include at least one RTC:5'-side module. In some embodiments, the reverse transcriptase construct of the present invention may not include at least one RTC:3'-side module.

[0148] In some embodiments, at least one reverse transcriptase construct may comprise any combination of (a) at least one RTC:5'-side module selected from SEQ ID NO: 58, encoding SEQ ID NO: 58, or encoded by SEQ ID NO: 58; (b) at least one RTC:RT module selected from SEQ ID NOs: 1-57, encoding any of SEQ ID NOs: 1-57, or encoded by any of SEQ ID NOs: 1-57; and / or (c) at least one RTC:3'-side module selected from SEQ ID NO: 59, encoding SEQ ID NO: 59, or encoded by SEQ ID NO: 59.

[0149] Exemplary Reverse Transcriptase Construct (RTC) The reverse transcriptase construct (RTC) used in the present invention may comprise at least one of SEQ ID NOs: 1-57, may encode at least one of SEQ ID NOs: 1-57, or may be encoded by at least one of SEQ ID NOs: 1-57. In some embodiments, the reverse transcriptase construct (RTC) may comprise, encode, or be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 17-21.

[0150] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, encode, or be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 17-21.

[0151] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 26-29.

[0152] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 24 or 25.

[0153] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 1-5.

[0154] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 35-37.

[0155] In some embodiments, at least one reverse transcriptase construct (RTC) may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 32-34.

[0156] In some embodiments, at least one reverse transcriptase construct (RTC) comprises the structure shown in FIG. 5.

[0157] Regulatory factors of the Reverse Transcriptase Construct (RTC) The reverse transcriptase construct (RTC) of the present invention may further comprise any number of regulatory factors, which may be included in any of the RTC modules. As used herein, the term "regulatory factor" means a sequence, region, or domain that enables control of the expression or activity of a biopolymer that is part of the construct.

[0158] For example, an RNA-based RTC may comprise any number of microRNA (miRNA) binding sites or small interfering RNA (siRNA) binding sites. Without wishing to be bound by any theory, the presence of an RNA interference (RNAi) binding site may prevent the expression of RT protein in certain types of cells based on the presence of the RNAi transcriptome. In this way, the gene insertion system of the present invention can exclude the target type of cells. As used herein, the term "miRNA binding site or siRNA binding site" means an RNA sequence complementary to at least one miRNA or siRNA.

[0159] In some embodiments, the RTC may comprise at least one miRNA and / or siRNA contained in the transgene inserted by the gene insertion system, or at least one miRNA binding site and / or siRNA binding site complementary to at least one miRNA and / or siRNA encoded by the transgene. By including such miRNA binding sites or siRNA binding sites, the gene insertion system of the present invention may generally be capable of regulating by itself the number of transgene insertions achieved by administration of the gene insertion system once, and / or may be capable of preventing repeated insertion of the transgene after the first administration. Thus, the gene insertion system may have an improved ability to re-administer or co-administer to a subject.

[0160] Gene Insertion Construct (GIC) The gene insertion system of the present invention comprises at least one gene insertion construct (GIC), which usually contains at least one target sequence (i.e., "payload sequence") intended for insertion into the genome of a subject or encodes such a sequence. As used herein, the term "gene insertion construct (GIC)" refers to a biopolymer construct that contains at least one RNA sequence or encodes at least one RNA sequence, and is characterized in that the RNA sequence can function as a template for reverse transcription when recognized by at least one reverse transcriptase contained in at least one RTC:RT module or encoded by at least one RTC:RT module. In some embodiments, the at least one gene insertion construct used in the gene insertion system of the present invention may contain, but is not limited to, nucleic acid biopolymers such as RNA, DNA, or any combination thereof.

[0161] Structure of the Gene Insertion Construct (GIC) The gene insertion construct (GIC) of the present invention may include at least one GIC: 5'-side module, at least one GIC: payload module, at least one GIC: 3'-side module, and any combination thereof, and may encode these. In some embodiments, at least one GIC may include a plasmid or may be delivered to a subject as a plasmid. In some embodiments, at least one GIC may include linear RNA or may be delivered to a subject as linear RNA.

[0162] In some embodiments, the at least one GIC: 5'-side module is an arbitrary component. In some embodiments, the at least one GIC: 3'-side module is an arbitrary component. In some embodiments, the gene insertion construct (GIC) of the present invention includes at least one GIC: payload module or encodes at least one GIC: payload module, but does not include and does not encode at least one GIC: 5'-side module and / or at least one GIC: 3'-side module.

[0163] Referring to FIG. 6, an exemplary linear RNA gene insertion construct (GIC) 400 is shown, and the GIC: 5'-side module 410, which is an arbitrary component, extends from the 5'-end of the GIC sequence to the end 420 of the GIC: 5'-side module. The GIC: payload module 430 extends towards the 3'-side of the GIC: 5'-side module (if present) to the end 440 of the GIC: payload module. Further, the GIC: 3'-side module 450 extends to the 3'-end of the GIC. Details of each of these features are described below.

[0164] GIC: 5'-side module In the gene insertion construct (GIC) of the present disclosure, the GIC:5' side module may include at least one sequence derived from the 5' region of a natural retroelement, and may encode such a sequence. Without wishing to be bound by any theory, this 5' side module may interact with at least one RNA binding domain of reverse transcriptase, an RNA sequence capable of synthesizing a second strand upon insertion of a transgene, an RNA sequence that reduces the immunogenicity of the gene insertion construct, an RNA sequence that provides useful features for the stability and / or purification of the gene insertion construct, and any combination thereof, and may encode such an RNA sequence.

[0165] Structure of the GIC: 5'-side module In a plurality of embodiments, the 5' side module includes a 5' rRNA sequence and a ribozyme (RZ) sequence. In some embodiments, the 5' rRNA sequence and the RZ sequence do not necessarily have to be completely separated. In some embodiments, the 5' side module includes a "folding sequence" that may be separated from the RZ sequence. In some embodiments, the GIC:5' side module may include at least one rRNA sequence (or other target site sequence), at least one ribozyme (RZ) sequence, at least one folding sequence, and any combination thereof, and may encode these.

[0166] Referring back to FIG. 6, the enlarged view of the GIC:5' side module 410 (lower left) shows the structure of an exemplary GIC:5' side module. When the GIC:5' rRNA sequence 411 is present at the 5' end of the 5' side module, this GIC:5' rRNA sequence 411 may contain an RNA sequence complementary to a target DNA sequence located on the 5' side of the target insertion site or in the vicinity of the target insertion site, and may encode such an RNA sequence. When the ribozyme (RZ) sequence 412 of the GIC:5' side module is present, this RZ sequence of the GIC:5' side module may contain at least one RNA sequence having a folded structure of a self-cleaving ribozyme, and this RNA sequence may release a functional GIC from the transcribed 5' leader sequence by self-cleaving, or such self-cleavage may not occur. When the RZ sequence of the GIC:5' side module is folded and active, it self-cleaves and incorporates the GIC:5' rRNA sequence as part of the RZ at or near the 5' end of the GIC. The folding motif sequence 413, which is an optional component of the GIC:5' side module, may contain at least one RNA sequence that is predicted or demonstrated to fold autonomously, and this RNA sequence is considered useful for physically and / or kinetically separating the folding of the RZ of the GIC:5' side module from the folding of the payload sequence. Further, within region 414 or at position 420 between the GIC 5' side module 410 and the payload module 430, a GIC sequence may be added to stop transcription initiated from an endogenous promoter sequence adjacent to the target site in the cell or for other control. In some embodiments, the endogenous promoter sequence adjacent to the target site in the cell may be utilized for the expression of the payload, and this aspect is an example of a case where the expression of the payload is regulated by adding a GIC sequence at position 420 and / or position 440 (e.g., initiating or terminating the translation of an RNA transcript containing the payload sequence by a host promoter).Furthermore, region 414 may include an RNA polymerase (RNAP) termination sequence to prevent read-through of the gene at the target insertion site by RNA polymerase. In some embodiments, the RNAP is RNAP I (Pol I), and when the GIC payload module is integrated into the gene target site of ribosomal DNA, the termination sequence blocks the read-through transcription of Pol I. In some embodiments, the RNAP transcription termination sequence includes the sequence shown by 5’-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3’ (SEQ ID NO: 537).

[0167] rRNA sequence of the GIC: 5'-side module GIC: At least one rRNA sequence of the 5’-side module is any component of the GIC: 5’-side module. When an rRNA sequence is present in the GIC: 5’-side module, this rRNA sequence may include a human ribosomal RNA (rRNA) sequence or other sequence that is homologous and / or complementary to at least one target DNA sequence located on the 5’-side of the target insertion site, and may encode such a sequence. Without wishing to be bound by any theory, this rRNA sequence may induce the synthesis of the second strand of the inserted cDNA transgene by mobilizing at least one endogenous DNA repair mechanism. In some embodiments, the rRNA sequence of the GIC: 5’-side module is located on the 5’-side of the RZ sequence of the GIC: 5’-side module. In some embodiments, the GIC: 5’-side module does not include a sequence containing an rRNA genomic sequence.

[0168] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 36 nt and may encode rRNA of about 1 to 36 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 30 nt and may encode rRNA of about 1 to 30 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 28 nt and may encode rRNA of about 1 to 28 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 26 nt and may encode rRNA of about 1 to 26 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 13 nt and may encode rRNA of about 1 to 13 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 to 11 nt and may encode rRNA of about 1 to 11 nt.

[0169] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise rRNA of about 1 nt, about 2 nt, about 3 nt, about 4 nt, about 5 nt, about 6 nt, about 7 nt, about 8 nt, about 9 nt, about 10 nt, about 11 nt, about 12 nt, about 13 nt, about 14 nt, about 15 nt, about 16 nt, about 17 nt, about 18 nt, about 19 nt, about 20 nt, about 21 nt, about 22 nt, about 23 nt, about 24 nt, about 25 nt, about 26 nt, about 27 nt, about 28 nt, about 29 nt, about 30 nt, about 31 nt, about 32 nt, about 33 nt, about 34 nt, about 35 nt, or about 36 nt and may encode rRNA of such lengths.

[0170] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 30 nt and may encode an rRNA of about 30 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 36 nt and may encode an rRNA of about 36 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 28 nt and may encode an rRNA of about 28 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 26 nt and may encode an rRNA of about 26 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 13 nt and may encode an rRNA of about 13 nt. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise an rRNA of about 11 nt and may encode an rRNA of about 11 nt. In some embodiments, the rRNA sequence of the GIC:5'-side module comprises a 5' G nucleotide.

[0171] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise at least one of SEQ ID NOs: 250 to 276, may encode at least one of these sequences, and may be encoded by at least one of these sequences. In some embodiments, at least one rRNA sequence of the GIC:5'-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75% or 70% homology with at least one of SEQ ID NOs: 250 to 276, may encode this sequence, and may be encoded by this sequence. In some embodiments, at least one rRNA sequence of the GIC:5'-side module comprises a sequence having one, two or three nucleotide changes or nucleotide substitutions as compared to a sequence selected from the group consisting of SEQ ID NOs: 250 to 276.

[0172] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75% or 70% homology with SEQ ID NO: 252. In some embodiments, at least one rRNA sequence of the GIC:5'-side module includes a sequence having one, two or three nucleotide changes compared to SEQ ID NO: 252.

[0173] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75% or 70% homology with SEQ ID NO: 254. In some embodiments, at least one rRNA sequence of the GIC:5'-side module includes a sequence having one, two or three nucleotide changes compared to SEQ ID NO: 254.

[0174] In some embodiments, at least one rRNA sequence of the GIC:5'-side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75% or 70% homology with SEQ ID NO: 255. In some embodiments, at least one rRNA sequence of the GIC:5'-side module includes a sequence having one, two or three nucleotide changes compared to SEQ ID NO: 255.

[0175] RZ sequence of the GIC: 5'-side module GIC: The RZ sequence of the 5'-side module is an arbitrary component of the GIC: 5'-side module. When present in the GIC: 5'-side module, it contains at least one self-cleaving ribozyme or at least one sequence having a folded structure of a self-cleaving ribozyme, or encodes these (collectively referred to as "RZ"). Without wishing to be bound by any theory, such an RZ motif generates a 5'OH terminus in the GIC, and this 5'OH terminus may be, for example, a 5' terminus generated by self-cleavage. In a stable tertiary structure, the generated 5'OH terminus may reduce the innate immune response to exogenous RNA, and the degradation of the GIC by a 5'-3' exonuclease that initiates cleavage depending on 5'-monophosphate may be reduced. The possibility that the GIC is recognized by the target cell as mRNA or other undesirable types of RNA rather than as template RNA may also be reduced.

[0176] In some embodiments, at least one RZ sequence of the GIC: 5'-side module contains a ribozyme derived from the 5'-region of at least one non-LTR retroelement or encodes such a ribozyme. In some embodiments, at least one RZ sequence of the GIC: 5'-side module contains a ribozyme derived from the 5'-region of a non-LTR retroelement obtained from Itoyo, American lobster, Kitano Tomiyo, Kyouso Yadorikobachi, Galapagos finch, medaka, nodojiro shidotod, goldfish, koknustomodoki (e.g., R2 derived from strain A or B), nodojiro shigidachou, other birds, other arthropods, other fish, other tunicates, other animals or a similar genome, or encodes such a ribozyme.

[0177] In some embodiments, the RZ sequence of the GIC:5’-side module comprises or encodes an RZ capable of forming the secondary and tertiary structures of the RZ of hepatitis delta virus (HDV), which RZ may be modified from sequences found in nature and / or may be de novo designed without using known genomic sequences. In some embodiments, the folded RZ sequence that crosslinks stem P2 paired with stem P1 paired in HDV is also called junction (J)1 / 2, and part or all of it is included in a target site sequence of a desired length (e.g., 5’ rRNA) or in a desired target site sequence further protected by the formation of a stem loop. In some embodiments, by incorporating the paired stem 4 (P4) of the folded RZ sequence of HDV into the design, it may be possible to purify the gene insertion construct (GIC) without denaturation, for example, by binding to the native or modified sequence of the coat protein of PP7 phage or MS2 phage. In some embodiments, the RZ sequence is designed and optimized to minimize or eliminate non-productive folding. In some embodiments, the RZ sequence is designed and optimized to minimize the number of uridine nucleotides. In some embodiments, the RZ sequence is designed and optimized such that all or some of the standard ribonucleotides can be replaced with nucleotide analogs incorporated during the synthesis process of the template RNA.

[0178] In some embodiments, at least one RZ sequence of the GIC:5'-side module may include at least one of SEQ ID NOs: 60 to 154, may encode at least one of SEQ ID NOs: 60 to 154, or may be encoded by at least one of SEQ ID NOs: 60 to 154. In some embodiments, the RZ sequence folds autonomously to form an active ribozyme. In some embodiments, the RZ sequence includes an internal rRNA sequence at its 5'-end. In some embodiments, the RZ sequence has an extended sequence at the 5'-end or 3'-end. In some embodiments, the RZ sequence is a catalytically inactive RZ sequence. In some embodiments, at least one RZ sequence of the GIC:5'-side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 60 to 154, may encode this sequence, or may be encoded by this sequence. In some embodiments, the RZ sequence of the GIC:5'-side module includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NOs: 60 to 154.

[0179] In some embodiments, at least one RZ sequence of the GIC:5'-side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 60, may encode this sequence, or may be encoded by this sequence.

[0180] In some embodiments, at least one RZ sequence of the GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 64.

[0181] In some embodiments, at least one RZ sequence of the GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 67.

[0182] In some embodiments, at least one RZ sequence of the GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 101.

[0183] In some embodiments, at least one RZ sequence of the GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 121.

[0184] In some embodiments, at least one RZ sequence of the GIC:5'-side module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 122.

[0185] In some embodiments, at least one RZ sequence of the GIC:5'-side module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 137.

[0186] Folding sequence of the GIC: 5'-side module GIC: The folding array of the 5'-side module is an arbitrary component of the 5'-side module. When this folding array exists, this folding array contains at least one RNA sequence motif having a specially designed structure. In some embodiments, the self-regulating folding RNA sequence motif contains at least one hairpin motif, and this hairpin motif may exist, for example, after blocking the misfolding of the RZ sequence by forming a base pair with the payload region to be transcribed later. In some embodiments, the 5'-side module region designed to improve the folding of the productive template RNA forms a base pair or interacts directly or indirectly with another template RNA region of the payload module or the 3'-side module. In some embodiments, at least one RNA sequence motif that induces the folding of the template RNA may contain at least one stem-loop motif that binds a protein bridge to another stem-loop motif. In some embodiments, the folding array of the 5'-side module may promote the pairing of the mRNA encoding reverse transcriptase and the template RNA, for example, by promoting the packaging of the mRNA encoding reverse transcriptase and the template RNA together in an individual delivery medium at a 1:1 stoichiometric ratio. In some embodiments, the folding array of the 5'-side module may promote the pairing of the template RNA and the endogenous RNA of the target cell, for example, for the purpose of achieving stabilization, localization, and / or other useful results of the template RNA.

[0187] In some embodiments, at least one folding array of the GIC:5'-side module may include at least one of SEQ ID NOs: 278 to 279, may encode at least one of these sequences, or may be encoded by at least one of these sequences. In some embodiments, at least one folding array of the GIC:5'-side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 278 to 279, may encode this sequence, or may be encoded by this sequence. In some embodiments, the folding array of the GIC:5'-side module includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NOs: 278 to 279.

[0188] In some embodiments, at least one folding array of the GIC:5'-side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 278, may encode this sequence, or may be encoded by this sequence.

[0189] In some embodiments, at least one folding array of the GIC:5'-side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 279, may encode this sequence, or may be encoded by this sequence.

[0190] Modularity of the GIC: 5'-side module Each component of the 5'-side module disclosed in this specification may be used interchangeably with each other in a combinatorial manner to design a 5'-side module having the functionality required for a specific gene insertion system or the desired functionality.

[0191] In some embodiments, at least one GIC: 5'-side module includes at least one rRNA sequence of the GIC: 5'-side module. In some embodiments, at least one GIC: 5'-side module includes at least one RZ sequence of the GIC: 5'-side module. In some embodiments, at least one GIC: 5'-side module includes at least one folding sequence of the GIC: 5'-side module. In some embodiments, at least one GIC: 5'-side module includes at least one rRNA sequence of the GIC: 5'-side module and at least one RZ sequence of the GIC: 5'-side module. In some embodiments, at least one GIC: 5'-side module includes at least one rRNA sequence of the GIC: 5'-side module, at least one RZ sequence of the GIC: 5'-side module, and at least one folding sequence of the GIC: 5'-side module.

[0192] In some embodiments, at least one GIC: 5'-side module (a) at least one rRNA sequence selected from any one of SEQ ID NOs: 250 to 276, encoding any one of these, or encoded by any one of these; (c) at least one RZ sequence selected from any one of SEQ ID NOs: 60 to 154, encoding any one of these, or encoded by any one of these; and / or (d) at least one folding sequence selected from any one of SEQ ID NOs: 278 to 279, encoding any one of these, or encoded by any one of these may be included in any combination.

[0193] Exemplary GIC: 5'-side module In some embodiments, at least one GIC:5'-side module may comprise at least one of SEQ ID NOs: 60-154, may encode at least one of SEQ ID NOs: 60-154, or may be encoded by at least one of SEQ ID NOs: 60-154. In some embodiments, at least one GIC:5'-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 60-154, may encode this sequence, or may be encoded by this sequence. In some embodiments, the GIC:5'-side module comprises a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NOs: 60-154.

[0194] In some embodiments, at least one GIC:5'-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 60, 61, 77, and 79-83, may encode this sequence, or may be encoded by this sequence.

[0195] In some embodiments, at least one GIC:5'-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 62 or 63, may encode this sequence, or may be encoded by this sequence.

[0196] In some embodiments, at least one GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 121.

[0197] In some embodiments, at least one GIC:5' side module may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 117 - 119.

[0198] GIC: 3'-side module

[0199] The 3' side module used in the gene insertion construct (GIC) of the present disclosure may include, or may encode, at least one sequence derived from the 3' UTR of a natural retroelement. Generally, the 3' side module includes components that facilitate recognition and binding of the gene insertion construct by reverse transcriptase, components that arrange the payload module for reverse transcription, and components that stabilize the RNA of the gene insertion construct.

[0200] Structure of the GIC: 3'-side module In some embodiments, the GIC:3' side module may include, or may encode, at least one reverse transcriptase (RT) recognition sequence, at least one rRNA sequence which is an optional component, at least one A - tract sequence which is an optional component, and any combination thereof.

[0201] Referring back to FIG. 6, the enlarged view (lower right) shows an example of the structure of the GIC:3'-side module 450. The 5'-end of the GIC:3'-side module is the reverse transcriptase recognition sequence 451 of the GIC:3'-side module. This reverse transcriptase recognition sequence 451 may contain a sequence recognized or bound by at least one reverse transcriptase, or may encode such a sequence. When the rRNA sequence 452 of the GIC:3'-side module is present, this rRNA sequence 452 may be located on the 3'-side of the reverse transcriptase recognition sequence of the GIC:3'-side module, may contain a sequence homologous to the target site region, or may encode this, for example, it may be a 28S rRNA nucleotide that can form a base pair with the 3'-end of the TPRT primer. Further, when the A-tract sequence 453 of the GIC:3'-side module is present, this A-tract sequence 453 may contain a sequence rich in adenosine or a sequence in which adenosines are arranged in series, the length of which may be limited, for example, the length of this sequence may be 10 to 60 nt, and the A-tract sequence 453 may be arranged at the 3'-end of the GIC:3'-side module.

[0202] Reverse transcriptase recognition sequence of the GIC: 3'-side module GIC: The reverse transcriptase recognition sequence of the 3'-side module may include at least one sequence that interacts with at least one reverse transcriptase, or at least one sequence recognized by at least one reverse transcriptase, and may encode such a sequence. Without wishing to be bound by any theory, at least one RNA sequence included in the reverse transcriptase recognition sequence of the GIC: 3'-side module may bind at least transiently to at least one template RNA binding domain of a reverse transcriptase such as a retroelement reverse transcriptase. The length and sequence identity of the reverse transcriptase recognition sequence of the GIC: 3'-side module may be configured to place a reverse transcriptase in the gene insertion construct (GIC), such that the first nucleotide reverse transcribed by the reverse transcriptase can be placed at the intended 3'-end of the transgene to be inserted. As used herein, the reverse transcriptase recognition sequence of the GIC: 3'-side module may also be referred to as the "3'-UTR of the GIC: 3'-side module".

[0203] In some embodiments, the reverse transcriptase recognition sequence of at least one GIC:3’-side module is derived from the 3’ region of a natural retroelement or includes the 3’ region of a natural retroelement. In some embodiments, the reverse transcriptase recognition sequence of at least one GIC:3’-side module is derived from the 3’ region of a non-LTR retroelement obtained from the genome of Itoyo, Drosophila melanogaster, American lobster, Kitano Tomiyo, Chrysomya megacephala, Galapagos finch, medaka, Nodularia spumigena, golden butterfly, Koknus tomokio, Nodularia shigedai, Drosophila simulans, silkworm moth, A. vaga, other birds, other arthropods, other fish, other tunicates, other animals, or organisms similar thereto. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3’-side module is modified from the 3’ region of a natural retroelement by increasing the folding stability or uniformity. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3’-side module is designed and / or selected to have a desired affinity and / or specificity in the interaction with reverse transcriptase, or is designed and / or selected to have another mechanism that confers a desired function as a template for reverse transcription. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3’-side module is designed and / or selected so as not to interact with or affect the endogenous components of the target cell, and / or is designed and / or selected so as not to have a harmful effect on the host cell.

[0204] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module (i.e., the 3'UTR sequence of the GIC:3' side module) may include at least one of SEQ ID NOs: 200 to 224, may encode at least one of SEQ ID NOs: 200 to 224, or may be encoded by at least one of SEQ ID NOs: 200 to 224. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 200 to 221, may encode this sequence, or may be encoded by this sequence. In some embodiments, the reverse transcriptase recognition sequence of the GIC:3' side module is a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity with SEQ ID NOs: 200 to 224.

[0205] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 202, may encode this sequence, or may be encoded by this sequence.

[0206] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NOs: 204, 222, 223 or 224, may encode this sequence, or may be encoded by this sequence.

[0207] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 203, may encode this sequence, or may be encoded by this sequence.

[0208] In some embodiments, the GIC:3’-side module contains a reverse transcriptase recognition sequence derived from a species different from the species from which the reverse transcriptase encoded by the reverse transcriptase construct (RTC) is derived. For example, in some embodiments, the reverse transcriptase recognition sequence may be derived from one species of bird, and the reverse transcriptase may be derived from another species of bird. In some embodiments, the reverse transcriptase recognition sequence is derived from a bird selected from any one of the common starling, goldfinch, pied starling, and Galapagos finch, and the reverse transcriptase is selected from a bird different from the selected bird (e.g., the common starling, goldfinch, pied starling, or Galapagos finch). In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from a bird selected from any one of the common starling, goldfinch, pied starling, and Galapagos finch, and the reverse transcriptase recognition sequence is selected from a bird different from the selected bird (e.g., the common starling, goldfinch, pied starling, or Galapagos finch). In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 18 or 20, and the reverse transcriptase recognition sequence is selected from an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 203, 204, 205, or 222-224. In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 27 or 29, and the reverse transcriptase recognition sequence is selected from an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 202, 204, 205, or 222-224.In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from amino acid sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with SEQ ID NO: 25, and the reverse transcriptase recognition sequence is selected from amino acid sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with SEQ ID NO: 202, 203, 204 or 222-224. In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from amino acid sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with SEQ ID NO: 31, and the reverse transcriptase recognition sequence is selected from amino acid sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with SEQ ID NO: 202, 203 or 205.

[0209] rRNA sequence of the GIC: 3'-side module GIC: The rRNA sequence of the 3'-side module, i.e., the sequence that forms base pairs with the TPRT primer immediately downstream of the nick introduced into the non-rDNA target site, is an arbitrary component of the 3'-side module. When present in the 3'-side module, this rRNA sequence may contain a human ribosomal RNA (rRNA) sequence. Without wishing to be bound by any theory, the length and sequence identity of the rRNA sequence of the GIC: 3'-side module affect how accurately and efficiently the gene insertion system disclosed herein can insert a transgene into the target genome. For example, depending on the length of the selected rRNA sequence of the GIC: 3'-side module, reverse transcription may be initiated from an internal sequence, potentially resulting in efficient shortening of the inserted transgene or enabling insertion into off-target sites. In either case, the insertion efficiency and specificity of the transgene at the target site of interest are reduced. RTC and GIC are configured such that the base pairs formed between the primer sequence immediately downstream of the nick introduced into the target site and the rRNA sequence of the GIC: 3'-side module need to be of a specific length. This further improves the fidelity with respect to the utilization of the target site and enables more efficiently obtaining the exact ligation site where the transgene is inserted. The optimal length of the GIC: 3' rRNA is less than 20 nt, and particularly if it is 4 nt, it can receive a strong stimulus from the entire 4 bp base pairs formed at the nick of the target site. Therefore, if RTC randomly introduces nicks, using 4 nt of GIC: 3' rRNA is considered to achieve optimal transgene insertion for only 1 out of 256 nicks.

[0210] In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 1 to 30 nt. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 1 to 20 nt. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 1 to 10 nt. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 1 to 5 nt.

[0211] In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain a portion of an rRNA of about 1 nt, about 2 nt, about 3 nt, about 4 nt, about 5 nt, about 6 nt, about 7 nt, about 8 nt, about 9 nt, about 10 nt, about 11 nt, about 12 nt, about 13 nt, about 14 nt, about 15 nt, about 16 nt, about 17 nt, about 18 nt, about 19 nt, about 20 nt, about 21 nt, about 22 nt, about 23 nt, about 24 nt, about 25 nt, about 26 nt, about 27 nt, about 28 nt, about 29 nt, or about 30 nt, and may encode a portion of an rRNA of such a length.

[0212] In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 20 nt. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 4 nt. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain or encode an rRNA of about 10 nt.

[0213] In some embodiments, at least one rRNA sequence of the GIC:3' side module may include at least one of SEQ ID NOs: 280 to 285. In some embodiments, at least one rRNA sequence of the GIC:3' side module is selected from the group consisting of SEQ ID NOs: 280 to 289 and sequences containing one, two or three nucleotide substitutions in SEQ ID NOs: 280 to 289.

[0214] A - tract sequence of the GIC: 3'-side module The A-tract sequence of the GIC:3' side module is an arbitrary component of the 3' side module. When present in the 3' side module, this A-tract sequence includes a terminal poly sequence containing a plurality of adenosines (A) arranged in series. Without wishing to be bound by any theory, the A-tract sequence of the GIC:3' side module may stabilize or protect the GIC from further processing on the 3' side. Furthermore, in cells, the GIC is recognized as mRNA, making it less likely for the GIC to undergo degradation related to the assembly, transport, and translation of ribonucleoproteins. Additionally, at least one A-tract sequence of the GIC:3' side module may protect the GIC from binding by common single-stranded RNA-binding proteins and may assist in arranging the GIC:3' rRNA sequence so that it can form base pairs with the target site primer. For clarity, this A-tract sequence is different from the natural mRNA poly-A tail sequence in which usually more than about 100 to 200 nt of adenosines are arranged in series.

[0215] In some embodiments, at least one A-tract sequence, which is any component of the GIC:3’-side module, comprises a sequence consisting of about 1 to 50 adenosines or encodes a sequence consisting of about 1 to 50 adenosines. For example, the A-tract sequence, which is any component of the GIC:3’-side module, may comprise or encode a sequence consisting of about 1 to 50 adenosines, about 5 to 50 adenosines, about 10 to 50 adenosines, about 15 to 50 adenosines, about 20 to 50 adenosines, about 25 to 50 adenosines, about 30 to 50 adenosines, about 35 to 50 adenosines, about 40 to 50 adenosines, about 45 to 50 adenosines, about 1 to 45 adenosines, about 5 to 45 adenosines, about 10 to 45 adenosines, about 15 to 45 adenosines, about 20 to 45 adenosines, about 25 to 45 adenosines, about 30 to 45 adenosines, about 35 to 45 adenosines, about 40 to 45 adenosines, about 1 to 40 adenosines, about 5 to 40 adenosines, about 10 to 40 adenosines, about 15 to 40 adenosines, about 20 to 40 adenosines, about 25 to 40 adenosines, about 30 to 40 adenosines, about 35 to 40 adenosines, about 1 to 35 adenosines, about 5 to 35 adenosines, about 10 to 35 adenosines, about 15 to 35 adenosines, about 20 to 35 adenosines, about 25 to 35 adenosines, about 30 to 35 adenosines, about 1 to 30 adenosines, about 5 to 30 adenosines, about 10 to 30 adenosines, about 15 to 30 adenosines, about 20 to 30 adenosines, about 25 to 30 adenosines, about 1 to 25 adenosines, about 5 to 25 adenosines, about 10 to 25 adenosines, about 15 to 25 adenosines, about 20 to 25 adenosines, about 1 to 20 adenosines, about 5 to 20 adenosines, about 10 to 20 adenosines, about 15 to 20 adenosines, about 1 to 15 adenosines, about 5 to 15 adenosines, about 10 to 15 adenosines, about 1 to 10 adenosines, about 5 to 10 adenosines, or about 1 to 5 adenosines.In some embodiments, the A tract sequence of the GIC:3’-side module comprises from about 1 to about 100, from about 1 to about 90, from about 1 to about 80, from about 1 to about 70, or from about 1 to about 60 adenosines.

[0216] In some embodiments, at least one A tract sequence, which is any component of the GIC:3’-side module, comprises a sequence consisting of about 20 to about 25 adenosines or encodes a sequence consisting of about 20 to about 25 adenosines.

[0217] In some embodiments, at least one A tract sequence, which is any component of the GIC:3’-side module, comprises a sequence consisting of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, or about 50 adenosines or encodes a sequence consisting of such a number of adenosines. In some embodiments, the A tract sequence of the GIC:3’-side module comprises 22 adenosines.

[0218] Modularity of the GIC: 3'-side module Each component of the 3’-side module disclosed herein may be used interchangeably with each other in a combinatorial manner to design a 3’-side module having the functionality required for a particular gene insertion system or the desired functionality.

[0219] In some embodiments, at least one GIC:3' side module comprises at least one reverse transcriptase recognition sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module comprises at least one rRNA sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module comprises at least one A-tract sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module comprises at least one reverse transcriptase recognition sequence of the GIC:3' side module and at least one rRNA sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module comprises at least one reverse transcriptase recognition sequence of the GIC:3' side module and at least one A-tract sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module comprises at least one reverse transcriptase recognition sequence of the GIC:3' side module, at least one rRNA sequence of the GIC:3' side module, and at least one A-tract sequence of the GIC:3' side module.

[0220] In some embodiments, at least one GIC:3' side module (a) at least one reverse transcriptase recognition sequence selected from any one of SEQ ID NOs: 200 to 221, encoding any one of these, or encoded by any one of these; (b) at least one rRNA sequence selected from any one of SEQ ID NOs: 280 to 289, encoding any one of these, or encoded by any one of these; and / or (c) at least one A-tract sequence may be included in any combination.

[0221] Exemplary GIC: 3'-side module In some embodiments, at least one GIC:3’-side module may comprise at least one of SEQ ID NOs: 300 to 329, may encode at least one of SEQ ID NOs: 300 to 329, or may be encoded by at least one of SEQ ID NOs: 300 to 329. In some embodiments, at least one GIC:3’-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one sequence selected from the group consisting of SEQ ID NOs: 300 to 329, may encode this sequence, or may be encoded by this sequence. In some embodiments, at least one GIC:3’-side module comprises a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity) with a sequence selected from the group consisting of SEQ ID NOs: 300 to 329 or any combination thereof. In some embodiments, the GIC:3’-side module comprises a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with SEQ ID NOs: 300 to 329.

[0222] In some embodiments, at least one GIC:3’-side module may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 313 to 320, may encode this sequence, or may be encoded by this sequence.

[0223] In some embodiments, at least one GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with a sequence selected from the group consisting of "GACGGTAGC TAGGTTCGCA AGGCAGCCAC AAGCCAAAGA TAGGTAGGGT GCTCATAGTG AGTAGGGACA GTGCCTTTTG ATTCACAACG CGTCAATACC ATCTGACACG GATACCCTTA CCGGACTTGT CATGATCTCC CAGACTTGTC CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO: 222), "CCGGACTTGT CATGATCTCC CAGACTTGTC CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO: 223), and "CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO: 224). In some embodiments, such a sequence may include a 3' side sequence represented by TAGCaaaaaaaaaaaaaaaaaaaaaa (SEQ ID NO: 538).

[0224] In some embodiments, at least one GIC:3' side module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 315, may encode this sequence, or may be encoded by this sequence.

[0225] In some embodiments, at least one GIC: 3'-side module may comprise, may encode, or may be encoded by a sequence having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 307.

[0226] In some embodiments, at least one GIC: 3'-side module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 316.

[0227] GIC: Payload module The GIC: payload module used in the gene insertion construct (GIC) of the present invention comprises, or encodes, at least one payload sequence that functions as part of a template for reverse transcription and is inserted into the genome of interest by the gene insertion system disclosed herein. As used herein, the term "payload sequence" or simply "payload" means a biopolymer sequence intended to be inserted into a target genome by at least one gene insertion system of the present invention. The payload sequence of the present invention may comprise at least one transgene.

[0228] As used herein, the term "transgene" is used in the broadest sense and means any gene sequence inserted into the genome of interest by the gene insertion system of the present invention. For example, transgenes include sequences not normally found in the genome of interest, or sequences that are normally found in the genome of interest but not normally found at the target insertion site. Transgenes include, but are not limited to, sequences encoding a desired expression product (e.g., at least one of mRNA, microRNA, siRNA, rRNA, tRNA, long non-coding RNA, cytoplasmic small RNA, nuclear small RNA, nucleolar small RNA, Cajal body small RNA, circular RNA, peptide, polypeptide, and / or protein), and / or sequences controlling the expression of at least one transgene. In some embodiments, the transgene encodes a protein selected from telomerase reverse transcriptase (TERT, e.g., human TERT), phenylalanine hydroxylase (PAH, e.g., human PAH), factor VIII (e.g., human factor VIII), mutant factor VIII with various lengths of the B domain (e.g., hFactor VIII N6 and hFactor VIII N6 variants), and factor IX (e.g., human factor IX). In some embodiments, the transgene encodes a regulatory RNA. In some embodiments, the transgene encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody. In some embodiments, the transgene encodes a protein that can be used for the treatment of diseases selected from the genes listed in Table X below.

[0229]

Table 1

[0230] Structure of the payload module

[0231] The GIC: Payload module may contain at least one (e.g., one, two, three, or more) transgene sequences, and may further contain at least one promoter sequence of the transgene, at least one 5' untranslated sequence of the transgene, at least one 3' untranslated sequence of the transgene, at least one polyadenylation signal sequence or polyA tail sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, and any combination thereof.

[0232] Referring back to FIG. 6, the structure of an exemplary payload module 430 is shown in the upper enlarged view. If there is a promoter sequence 431, which is any component of the transgene, this promoter sequence 431 may contain at least one promoter that may control the expression of the transgene inserted in the target cell, and may encode such a promoter. The 5'UTR sequence 432, which is any component of the transgene, may contain a sequence encoding the 5'UTR of the transgene mRNA when the inserted transgene is expressed, and may encode such a sequence. The transgene sequence 433 of the payload module may contain at least one transgene sequence that is reverse transcribed and inserted by the gene insertion system of the present disclosure. For example, this sequence may contain the ORF of the gene of interest and may encode the ORF of the gene of interest. The 3'UTR sequence 434, which is any component of the transgene, may contain at least one 3'UTR of the expressed transgene mRNA and may encode this 3'UTR. Similarly, the polyadenylation signal sequence 435, which is any component of the transgene, may contain the polyadenylation signal of the expressed transgene mRNA and may encode this polyadenylation signal. Furthermore, the non-coding RNA (ncRNA) processing sequence 436, which is any component of the transgene, may contain the termination signal and / or 3' processing signal of the nrRNA expressed by the transgene and may encode such a signal.

[0233] Promoter sequence of the transgene and RNAP II 5' UTR sequence When a promoter sequence of the introduced gene is present, this promoter sequence may include at least one promoter sequence including means for promoting the expression of the introduced gene in the target genome, and may encode such a promoter sequence. Such means for promoting the expression of the gene and / or the introduced gene are well known in the art, and examples of such means include the insertion of a known promoter sequence to the 5'-side of the gene of interest. A person skilled in the art may select the type of the promoter sequence based on the introduced gene and other specific factors used, and thus, it can be understood that any suitable promoter may be used in practicing the present disclosure.

[0234] Exemplary promoters used in the present disclosure may be constitutive promoters or inducible promoters. In some embodiments, the promoter sequence of the introduced gene may include at least one promoter of RNA polymerase I to III (RNAP I, RNAP II or III), and may encode such a promoter. In some embodiments, instead of or in addition to the promoter, at least one ribozyme or other motif that enables the release of the RNA transcript of the introduced gene transcribed by the rDNA RNAP I of the host cell may be included in at least one introduced gene in the same region as the promoter, and may encode such a ribozyme or other motif.

[0235] In some embodiments, at least one promoter sequence of the transgene comprises or encodes at least one human U1 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene comprises or encodes at least one human U3 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene comprises or encodes at least one human U6 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene comprises or encodes at least one tRNA promoter.

[0236] When the 5’ UTR sequence of the transgene is present, this 5’ UTR sequence comprises or encodes at least one mRNA 5’ UTR of the inserted transgene. Typically, this 5’ UTR sequence comprises or encodes a sequence that is not translated into an amino acid biopolymer by ribosomes in the cell when the inserted transgene is expressed by the cell. Such sequences include, for example, 5’UTRs that are naturally associated with the transgene, 5’UTRs that are non-natural to the transgene (including sequences derived from the 5’-flanking sequences of retroelements), “synthetic” 5’UTRs that may not be associated with a known wild-type gene, and any combination thereof.

[0237] One of ordinary skill in the art will understand that the choice of the 5’ UTR sequence of the transgene depends on the identity of the transgene and other specific factors used, so any known 5’UTR sequence or newly discovered 5’UTR sequence may be suitable for use in the 5’ sequence of the payload module of the transgene.

[0238] In some embodiments, at least one promoter sequence of the transgene may comprise at least one of SEQ ID NOs: 400-404 and 408-409, may encode at least one of SEQ ID NOs: 400-404 and 408-409, or may be encoded by at least one of SEQ ID NOs: 400-404 and 408-409. In some embodiments, at least one promoter sequence of the transgene may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 400-404 and 408-409, may encode this sequence, or may be encoded by this sequence.

[0239] In some embodiments, at least one promoter sequence of the transgene may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 400, may encode this sequence, or may be encoded by this sequence.

[0240] In some embodiments, at least one promoter sequence of the transgene may comprise a sequence having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 401, may encode this sequence, or may be encoded by this sequence.

[0241] In some embodiments, at least one promoter sequence of the transgene may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 403.

[0242] In some embodiments, at least one promoter sequence of the transgene comprises a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 404.

[0243] In some embodiments, at least one promoter sequence of the transgene comprises a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 408.

[0244] In some embodiments, at least one promoter sequence of the transgene comprises a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 409.

[0245] In some embodiments, the GIC:payload module includes an RNA polymerase (RNAP) terminator sequence located 5' to the promoter sequence of the transgene. In some embodiments, the RNAP is RNAP I (Pol I), and when the GIC payload module is integrated into the target site of the ribosomal DNA gene, the termination sequence blocks the read-through transcription of Pol I. In some embodiments, the RNAP terminator sequence includes the sequence shown by 5'-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3' (SEQ ID NO: 537).

[0246] Transgene sequence The transgene sequence of the payload module includes or encodes at least one sequence of interest for insertion into the genome of the subject. As used herein, "sequence of interest" means a biopolymer sequence that includes or encodes at least one desired expression product. In some embodiments, the transgene encodes a protein selected from hTERT, hPAH, hFactor VIII, variant human factor VIII with various lengths of the B domain (e.g., hFactor VIII N6 and hFactor VIII N6 variants), and factor IX (e.g., human factor IX). In some embodiments, the transgene encodes a regulatory RNA. In some embodiments, the transgene encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody. In some embodiments, the transgene encodes a protein that can be used for the treatment of a disease selected from the genes listed in Table X above.

[0247] Any sequence of interest may be suitable for the practice of the present disclosure, and is not limited by the source of the sequence (i.e., regardless of the biological species of the source, whether it is a natural or artificial sequence), nor is it limited by the length of the sequence.

[0248] In some embodiments, at least one transgene sequence may comprise at least one of SEQ ID NOs: 411 to 422, may encode at least one of SEQ ID NOs: 411 to 422, or may be encoded by at least one of SEQ ID NOs: 411 to 422. In some embodiments, at least one transgene sequence may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 411 to 422, may encode this sequence, or may be encoded by this sequence.

[0249] In some embodiments, at least one transgene sequence may comprise a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with SEQ ID NO: 419 or 420, may encode this sequence, or may be encoded by this sequence.

[0250] In some embodiments, at least one transgene sequence may comprise a sequence having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of SEQ ID NOs: 421 to 422, may encode this sequence, or may be encoded by this sequence.

[0251] In some embodiments, at least one transgene sequence may include, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 518 - 536.

[0252] 3'UTR sequence of the transgene and polyadenylation signal When a 3'UTR sequence of the transgene is present, this 3'UTR sequence includes or encodes at least one mRNA 3'UTR of the inserted transgene. Usually, when the inserted transgene is expressed by the cell, this 3'UTR sequence includes or encodes a sequence that is not translated into an amino acid biopolymer by ribosomes in the cell. Such sequences include, for example, 3'UTRs naturally associated with the transgene, 3'UTRs non - natural to the transgene (including sequences derived from the 3' - flanking sequences of retroelements), "synthetic" 3'UTRs not associated with known wild - type genes, and any combination thereof.

[0253] One of ordinary skill in the art will understand that the choice of the 3'UTR sequence of a transgene depends on the identity of the transgene and other specific factors used, so any known 3'UTR sequence or newly discovered 3'UTR sequence may be suitable for use in the 3' sequence of the payload module of the transgene.

[0254] When the introduced gene has a polyadenylation signal sequence, this polyadenylation signal sequence contains or encodes at least one polyadenylation signal of the mRNA of the introduced gene. Any suitable known or novel polyadenylation signal may be used in the template module of the present disclosure. For the sake of clarity, at least one polyadenylation signal present in the inserted introduced gene or encoded within the inserted introduced gene provides RNAP II that adds a polyA tail to the mRNA expression product or ncRNA expression product of the introduced gene.

[0255] In some embodiments, at least one 3' UTR sequence of the introduced gene may contain a sequence selected from at least one of SEQ ID NOs: 405 to 407. In some embodiments, at least one 3' UTR sequence of the introduced gene may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one of SEQ ID NOs: 405 to 407.

[0256] In some embodiments, at least one 3' UTR sequence of the introduced gene may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with SEQ ID NO: 405, may encode this sequence, or may be encoded by this sequence.

[0257] In some embodiments, at least one 3’ UTR sequence of the transgene may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 406.

[0258] In some embodiments, at least one 3’ UTR sequence of the transgene may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 407.

[0259] Non - coding RNA (ncRNA) processing sequence of the transgene When an ncRNA processing sequence of the transgene is present, the ncRNA processing sequence comprises or encodes a sequence that controls the expression or processing of an ncRNA (e.g., transfer RNA (tRNA), rRNA, microRNA, siRNA, snRNA, etc.) expressed from the transgene. In some embodiments, at least one non-coding RNA (ncRNA) processing sequence comprises or encodes at least one termination signal, at least one 3’ processing signal, and any combination thereof, of at least one ncRNA expressed from the transgene.

[0260] In some embodiments, at least one ncRNA processing sequence of the transgene comprises or encodes at least one MALAT1 3′-end processing signal and / or MALAT1 3′-end protection signal. In some embodiments, at least one ncRNA processing sequence of the transgene comprises or encodes at least one RNA triple helix-forming end protection structure. In some embodiments, at least one ncRNA processing sequence of the transgene comprises or encodes at least one endonuclease recruitment structure, endonuclease recruitment site or endonuclease recruitment motif. In some embodiments, at least one ncRNA processing sequence of the transgene comprises or encodes at least one poly-thymidine tract. In some embodiments, at least one RNA 3′-end termination sequence and / or RNA 3′-end processing sequence of the transgene comprises the SalI termination box of RNAP I.

[0261] Modularity of the payload module Each component of the GIC:payload module disclosed herein may be used interchangeably with each other in a combinatorial manner to design a 3′-side module having the functionality required for a particular gene insertion system or the desired functionality.

[0262] In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one transgene sequence. In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one promoter sequence of the transgene. In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one 5’ UTR sequence of the transgene. In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one 3’ UTR sequence of the transgene. In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one polyadenylation signal sequence of the transgene. In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one ncRNA processing sequence of the transgene.

[0263] In some embodiments, at least one GIC:payload module may comprise, or may encode, at least one transgene sequence, at least one promoter sequence of the transgene, at least one 5’ UTR sequence of the transgene, at least one 3’ UTR sequence of the transgene, at least one polyadenylation signal sequence of the transgene, and / or at least one ncRNA processing sequence of the transgene.

[0264] In some embodiments, at least one GIC:payload module (a) at least one promoter sequence and 5’UTR sequence of the transgene selected from any one of SEQ ID NOs: 400-404 and a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with any one of SEQ ID NOs: 400-404; (b) At least one transgene sequence selected from any one of SEQ ID NOs: 411 to 422 and SEQ ID NOs: 499 to 536, and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with SEQ ID NOs: 411 to 422 or 499 to 536, or encoding a sequence selected from these sequences, or encoded by a sequence selected from these sequences; and (c) At least one 3' UTR sequence and polyadenylation signal of a transgene selected from SEQ ID NOs: 405 to 407 and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with SEQ ID NOs: 405 to 407 may be included in any combination.

[0265] Exemplary GIC: Payload module In some embodiments, at least one GIC:payload module may include at least one sequence selected from SEQ ID NOs: 499 to 536, may encode at least one sequence selected from SEQ ID NOs: 499 to 536, or may be encoded by at least one sequence selected from SEQ ID NOs: 499 to 536. In some embodiments, at least one GIC:payload module may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with at least one sequence selected from SEQ ID NOs: 499 to 536, may encode this sequence, or may be encoded by this sequence.

[0266] In some embodiments, at least one GIC:payload module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 419, 420, 518, or 519.

[0267] In some embodiments, at least one GIC:payload module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 421, 422, 520, or 521.

[0268] In some embodiments, at least one GIC:payload module may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 522, 523, 524, or 525.

[0269] Modularity of the Gene Insertion Construct (GIC) Each component of the gene insertion construct (GIC) disclosed herein (i.e., the GIC:5' module, the GIC:3' module, and the GIC:payload module) may be used interchangeably with each other in a combinatorial manner to design a GIC having the functionality required for a particular gene insertion system or the desired functionality.

[0270] In some embodiments, at least one gene insertion construct (GIC) comprises at least one GIC:5'-side module. In some embodiments, at least one gene insertion construct comprises at least one GIC:payload module. In some embodiments, at least one gene insertion construct comprises at least one GIC:3'-side module. In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module and at least one GIC:payload module. In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module and at least one GIC:3'-side module. In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module, at least one GIC:payload module, and at least one GIC:3'-side module.

[0271] In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module comprising a GIC:5'-side module retroelement sequence derived from a retroelement of the same species as the source of the reverse transcriptase recognition sequence of the GIC:3'-side module. In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module comprising a GIC:5'-side module retroelement sequence derived from a retroelement of a species different from the source of the reverse transcriptase recognition sequence of the GIC:3'-side module. In some embodiments, at least one gene insertion construct comprises at least one GIC:5'-side module comprising a GIC:5'-side module sequence not naturally found in the eukaryotic ecosystem that is generally useful for at least one gene insertion construct comprising the reverse transcriptase recognition sequence of the GIC:3'-side module.

[0272] In some embodiments, the gene insertion construct comprises a combination of a GIC:5' side module sequence derived from the source shown in FIG. 7 and a GIC:3' side module sequence. In FIG. 7, A1 is Zonotrichia albicollis, A2 is Taeniopygia guttata, A3 is Tinamus guttatus, A4 is Geospiza fortis, B1 is Pungitis pungitis, B2 is Oryzias latipes, B3 is Gasterosteus aculeatus, C1 is Nasonia vitripennis, C2 is Drosophila melanogaster, C3 is Tribolium castaneum, C4 is Bombyx mori, C5 is Drosophila simulans, C6 is Drosophila mercatorum, D1 is Lepidurus couseii, D2 is Triops cancriformis, E1 is Hydra magnipapillata, E2 is Limulus polyphemus, E3 is Adineta vaga, and E4 is Ciona intestinalis.

[0273] In some embodiments, at least one gene insertion construct is (a) At least one GIC: 5'-side module, which is selected from the sequences of SEQ ID NO: 250-276, sequences having one, two or three nucleotide changes or nucleotide substitutions compared to SEQ ID NO: 250-276, SEQ ID NO: 60-154, sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) compared to SEQ ID NO: 60-154, SEQ ID NO: 278-279, and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) compared to SEQ ID NO: 278-279, or encodes a sequence selected from these sequences, or is encoded by a sequence selected from these sequences; at least one GIC: 5'-side module; (b) At least one GIC: payload module, which is selected from any one of SEQ ID NO: 411-422 and 499-525, and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) compared to SEQ ID NO: 411-422 or 499-522, or encodes a sequence selected from these sequences, or is encoded by a sequence selected from these sequences; at least one GIC: payload module; and / or (c) At least one GIC: 3'-side module, which is selected from any one of SEQ ID NO: 300-329, and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) compared to SEQ ID NO: 300-329, or encodes a sequence selected from these sequences, or is encoded by a sequence selected from these sequences; at least one GIC: 3'-side module It may contain such combinations in any combination, may encode such combinations, or may be encoded by such combinations.

[0274] Exemplary Gene Insertion Construct (GIC) In some embodiments, at least one gene insertion construct (GIC) may comprise at least one of SEQ ID NOs: 411-422 and 499-525, may encode at least one of SEQ ID NOs: 411-422 and 499-525, or may be encoded by at least one of SEQ ID NOs: 411-422 and 499-525. In some embodiments, at least one gene insertion construct may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of SEQ ID NOs: 411-422 and 499-536.

[0275] In some embodiments, at least one gene insertion construct may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 419, 420, 518, or 519.

[0276] In some embodiments, at least one gene insertion construct may comprise, may encode, or may be encoded by a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 421, 422, 520, or 521.

[0277] In some embodiments, at least one gene insertion construct may comprise, may encode, or may be encoded by a sequence having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to SEQ ID NO: 522, 523, 524, or 525.

[0278] Design and modularity of the Gene Insertion System Each component of the gene insertion system (GIS) disclosed herein (i.e., the reverse transcriptase construct (RTC) and the gene insertion construct (GIC)) may be used interchangeably with each other in a combinatorial manner to design a gene insertion system (GIS) having the required functionality or desired functionality.

[0279] In some embodiments, at least one gene insertion system may comprise at least one reverse transcriptase construct (RTC). In some embodiments, at least one gene insertion system may comprise at least one gene insertion construct (GIC). In some embodiments, at least one gene insertion system may comprise at least one reverse transcriptase construct (RTC) and at least one gene insertion construct (GIC).

[0280] Composition of biopolymers containing the Gene Insertion System A composition of biopolymers comprising each component of the gene insertion system (GIS) disclosed herein may be selected in a combinatorial manner from those disclosed herein to design a gene insertion system (GIS) having the required functionality or desired functionality.

[0281] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an RNA biopolymer. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an mRNA biopolymer.

[0282] In some embodiments, at least one gene insertion construct (GIC) may be introduced into at least one subject as an RNA biopolymer. In some embodiments, at least one gene insertion construct (GIC) may be introduced into at least one subject as a linear RNA biopolymer.

[0283] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an RNA biopolymer, and at least one gene insertion construct (GIC) may be introduced into at least one subject as an RNA biopolymer.

[0284] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an mRNA biopolymer, and at least one gene insertion construct (GIC) may be introduced into at least one subject as an RNA biopolymer.

[0285] In some embodiments, at least one reverse transcriptase construct (RTC) and / or at least one gene insertion construct (GIC) may be introduced into at least one subject as a DNA biopolymer. In some embodiments, at least one reverse transcriptase construct (RTC) and / or at least one gene insertion construct (GIC) may be introduced into at least one subject as a plasmid.

[0286] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an amino acid biopolymer. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as a protein.

[0287] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as an amino acid biopolymer, and at least one gene insertion construct (GIC) may be introduced into at least one subject as an RNA biopolymer. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as a plasmid, and at least one gene insertion construct (GIC) may be introduced into at least one subject as an RNA biopolymer.

[0288] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as a plasmid, and at least one gene insertion construct (GIC) may be introduced into at least one subject as a plasmid. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one subject as RNA (e.g., mRNA), and at least one gene insertion construct (GIC) may be introduced into at least one subject as a plasmid.

[0289] Pair of reverse transcriptases The gene insertion system of the present invention may be optimized for a desired function by controlling the interaction between the gene insertion construct (GIC) and the reverse transcriptase construct (RTC) included in this gene insertion system, and by designing or selecting the composition of at least one of both of them. For example, by modifying the composition of GIC and / or RTC, as observed by detecting the insertion using PCR, sequencing, and / or the expression of the transgene which is the payload, the efficiency, speed, and / or fidelity of the insertion of the full-length payload can be changed; as observed by sequencing, hybridization, or other methods for visualizing the position of the inserted DNA on the gene, the sequence specificity and / or chromosomal position in the selection of the target site where the payload is inserted can be changed; the selectivity in which the RTC utilizes only the GIC administered as a template for reverse transcription can be changed, and the like. In the present specification, the term "paired reverse transcriptase" means a specific RTC:RT module sequence administered in combination with a specific GIC sequence.

[0290] Although not wishing to be bound by any theory, the modification of the interaction between the RTC and the GIC may be achieved by selecting the RTC:RT module and the GIC:5'-side module and / or the GIC:3'-side module. For example, the specificity of the RTC for the GIC may be modified by selecting components derived from retroelements of the same or different species. In the present specification, when two components of the gene insertion system are derived from retroelements of the same species, they are said to be homologous. In contrast, when two components of the gene insertion system are derived from retroelements of different species, they are said to be heterologous.

[0291] In some embodiments, at least one RTC:RT module comprises or encodes at least one sequence derived from a retroelement of a different species than at least one retroelement from which the GIC:5'-side module array and / or the GIC:3'-side module array is derived (referred to herein as "reverse transcriptase of a heterologous pair").

[0292] In some embodiments, the sequences derived from the retroelements contained in the RTC and the GIC are both derived from retroelements of the same species (referred to herein as "reverse transcriptase of a homologous pair").

[0293] In some embodiments, compared to the reverse transcriptase of a homologous pair, the reverse transcriptase of a heterologous pair may have increased specificity.

[0294] As used herein, the term "specificity" means the possibility that the paired reverse transcriptase efficiently and / or selectively utilizes the template RNA intended for the insertion of the transgene.

[0295] In some embodiments, at least one gene insertion system comprises a combination of at least one GIC and at least one reverse transcriptase paired therewith, and such a combination is shown in FIG. 7.

[0296] Exemplary Gene Insertion System In some embodiments, at least one gene insertion system is (a) at least one reverse transcriptase construct (RTC) selected from any one of SEQ ID NOs: 1 to 59, and having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity) with any one of SEQ ID NOs: 1 to 59, or encoding a sequence selected from these sequences, or being encoded by a sequence selected from these sequences; and (b) At least one gene insertion construct (GIC), which is selected from, encodes, or is encoded by a sequence selected from any one of the sequences of SEQ ID NO: 250 - 276, a sequence having one, two, or three nucleotide changes or nucleotide substitutions compared to SEQ ID NO: 250 - 276, SEQ ID NO: 60 - 154, a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) to SEQ ID NO: 60 - 154, SEQ ID NO: 278 - 279, a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) to SEQ ID NO: 278 - 279, SEQ ID NO: 411 - 422 or 499 - 536, a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) to SEQ ID NO: 411 - 422 or 499 - 536, SEQ ID NO: 300 - 329, and / or a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) to SEQ ID NO: 300 - 329 It may contain these in any combination, may encode these in any combination, and may be encoded by any combination of these.

[0297] III. Formulations and Delivery Mechanisms Nucleic acids In some embodiments, the reverse transcriptase construct (RTC) or gene insertion construct (GIC) may contain one or more modified nucleotides, and examples of modified nucleotides include, but are not limited to, nucleobase modifications, sugar-modified nucleotides, and / or backbone modifications. In some embodiments, the RTC construct or GIC construct may contain a combination of modifications, for example, a combination of nucleobase modification and backbone modification.

[0298] In some embodiments, the modified nucleotide may be a nucleotide with a modified nucleobase. "Modified base" means a nucleotide base such as adenine, cytosine, thymine, guanine, uracil, xanthine, inosine, queuosine, etc., which is modified by substitution or addition of one or more groups or atoms, but is not limited thereto. In some embodiments, the modified nucleotide may be a nucleotide with a modified backbone.

[0299] The RTC construct and / or GIC construct may include one or more substitutions, insertions and / or additions, deletions, and covalent modifications as compared to a reference sequence, particularly the sequence of interest, and such RTC constructs and / or GIC constructs are also included within the scope of the present invention.

[0300] In some embodiments, the RTC construct and / or GIC construct includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, polyA sequence, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, nitrosylation of thiol groups and tyrosine residues, etc.).

[0301] The RTC construct and / or GIC construct may include modifications useful for sugars, nucleobases, or internucleoside linkages (e.g., phosphate linkages, phosphodiester linkages, or phosphodiester backbones).

[0302] In some embodiments, the modification may include a chemical modification or a cell-guided modification. For example, some examples of intracellular RNA modifications are reported by Lewis and Pan in "RNA modifications and structures cooperate to guide RNA-protein interactions" published in Nat Reviews Mol Cell Biol, 2017, 18:202-210, but are not limited thereto.

[0303] In some embodiments, immune evasion may be enhanced by chemical modification of RNA. RNA may be synthesized and / or modified by methods established in the art.

[0304] In some embodiments, at least one RNA construct may include at least one modified uracil. Examples of uracil modifications include 5-methyluridine, 5-methoxyuridine, pseudouridine, N 1 -methylpseudouridine, and / or 2-thiouridine. In some embodiments, at least one RNA construct may include at least one modified adenosine. An example of an adenosine modification is 2,6-diaminopurine deoxyribonucleotide.

[0305] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions or backbone modifications in one or more RNAs include modifications or substitutions of phosphodiester bonds.

[0306] Delivery mechanism The gene insertion system (GIS) of the present invention may be introduced into a subject via any delivery mechanism known in the art. As used herein, "delivery mechanism" means a method or composition used to introduce a gene insertion system, a component of a gene insertion system, or a product of a gene insertion system into a subject. Examples of delivery mechanisms include, but are not limited to, delivery vehicles, direct transfection (using transfection reagents, etc.), transplantation of cells pre-transfected with the gene insertion system, and any combination thereof.

[0307] Delivery medium In some embodiments, the gene insertion system of the present invention may be formulated with a delivery vehicle. Delivery vehicles generally facilitate transfection of cells of a subject in vivo or in vitro by protecting each component of the gene insertion system from degradation in the extracellular environment, promoting uptake by the cells of the subject, promoting endosomal escape, and any combination thereof. Examples of delivery vehicles include, but are not limited to, nanoparticles such as lipid-based nanoparticles (e.g., lipid nanoparticles (LNP), liposomes, and micelles) and non-lipid-based nanoparticles (e.g., virus-like particles (VLP) and polymeric delivery particles).

[0308] Nanoparticles In some embodiments, the delivery medium may contain at least one type of nanoparticles. As used herein, the term "nanoparticles" usually may mean particles having a size of 10 to 1000 nm. For example, 10 nm, 15 nm, 20 nm, 25 nm, 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, 150 nm, 155 nm, 160 nm, 165 nm, 170 nm, 175 nm, 180 nm, 185 nm, 190 nm, 195 nm, 200 nm, 205 nm, 210 nm, 215 nm, 220 nm, 225 nm, 230 nm, 235 nm, 240 nm, 245 nm, 250 nm, 255 nm, 260 nm, 265 nm, 270 nm, 275 nm, 280 nm, 285 nm, 290 nm, 295 nm, 300 nm, 305 nm, 310 nm, 315 nm, 320 nm, 325 nm, 330 nm, 335 nm, 340 nm, 345 nm, 350 nm, 355 nm, 360 nm, 365 nm, 370 nm, 375 nm, 380 nm, 385 nm, 390 nm, 395 nm, 400 nm, 405 nm, 410 nm, 415 nm, 420 nm, 425 nm, 430 nm, 435 nm, 440 nm, 445 nm, 450 nm, 455 nm, 460 nm, 465 nm, 470 nm, 475 nm, 480 nm, 485 nm, 490 nm, 495 nm, 500 nm, 505 nm, 510 nm, 515 nm, 520 nm, 525 nm, 530 nm, 535 nm, 540 nm, 545 nm, 550 nm, 555 nm, 560 nm, 565 nm, 570 nm, 575 nm, 580 nm, 585 nm, 590 nm, 595 nm, 600 nm, 605 nm, 610 nm, 615 nm, 620 nm, 625 nm, 630 nm, 635 nm, 640 nm, 645 nm, 650 nm, 655 nm, 660 nm, 665 nm, 670 nm, 675 nm, 680 nm, 685 nm, 690 nm, 695 nm, 700 nm, 705 nm, 710 nm, 715 nm, 720 nm, 725 nm, 730 nm, 735 nm, 740 nm, 745 nm, 750 nm, 755 nm, 760 nm, 765 nm, 770 nm,Particles may be of sizes 775 nm, 780 nm, 785 nm, 790 nm, 795 nm, 800 nm, 805 nm, 810 nm, 815 nm, 820 nm, 825 nm, 830 nm, 835 nm, 840 nm, 845 nm, 850 nm, 855 nm, 860 nm, 865 nm, 870 nm, 875 nm, 880 nm, 885 nm, 890 nm, 895 nm, 900 nm, 905 nm, 910 nm, 915 nm, 920 nm, 925 nm, 930 nm, 935 nm, 940 nm, 945 nm, 950 nm, 955 nm, 960 nm, 965 nm, 970 nm, 975 nm, 980 nm, 985 nm, 990 nm, 995 nm, or 1000 nm.

[0309] Lipid - based particles In some embodiments, the delivery medium may comprise at least one lipid-based nanoparticle, examples of lipid-based nanoparticles include, but are not limited to, lipid nanoparticles (LNP), liposomes, micelles, and any combination thereof.

[0310] Lipid nanoparticles In some embodiments, the delivery medium may be a lipid nanoparticle (LNP). Generally, an LNP has an outer lipid layer containing a hydrophilic outer surface in contact with a non-LNP environment, a non-aqueous or aqueous internal space (i.e., micelle-like LNP and vesicle-like LNP, respectively), and at least one hydrophobic intermembrane space. The LNP membrane may be of a non-lamellar structure or a lamellar structure, and may be composed of one, two, three, four, five, or more than five layers. The LNP may be solid or semi-solid. In some embodiments, at least one cargo or payload (such as a gene insertion system (GIS)) may be included in the internal space, intermembrane space, outer surface of the LNP, or any combination thereof.

[0311] LNPs useful in the present invention are known in the art and generally include ionizable (cationic) lipids, phospholipids, cholesterol, and polymer-modified lipids. Without wishing to be bound by any theory, cholesterol may promote membrane fusion and assist in the stability of the LNP; phospholipids may promote endosomal escape and provide structure to the LNP bilayer; polymer-modified lipids may suppress aggregation of the LNP and "protect" the LNP from non-specific endocytosis by immune cells; and ionizable (cationic) lipids may promote endosomal escape and form a complex with negatively charged cargo (such as polynucleotides of the gene insertion system).

[0312] In some embodiments, the gene insertion system of the present invention may be incorporated into lipid nanoparticles (LNP). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid), at least one non-cationic lipid (e.g., phospholipid), at least one sterol (e.g., cholesterol), at least one polymer-modified lipid (e.g., PEG lipid), or any combination thereof. In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid), at least one non-cationic lipid (e.g., phospholipid), at least one sterol (e.g., cholesterol), and at least one polymer-modified lipid (e.g., PEG lipid). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid), at least one non-cationic lipid, and at least one sterol (e.g., cholesterol). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid), at least one non-cationic lipid (e.g., phospholipid), and at least a polymer-modified lipid (e.g., PEG lipid). In some embodiments, the LNP may be composed of at least one non-cationic lipid (e.g., phospholipid), at least one sterol (e.g., cholesterol), and at least one polymer-modified lipid (e.g., PEG lipid). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid) and at least one non-cationic lipid (e.g., phospholipid). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid) and at least one sterol. In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid) and at least one polymer-modified lipid (e.g., PEG lipid).In some embodiments, the LNP may be composed of at least one non-cationic lipid (e.g., phospholipid) and at least one sterol (e.g., cholesterol). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid) and at least one polymer-modified lipid (e.g., PEG lipid). In some embodiments, the LNP may be composed of at least one sterol (e.g., cholesterol) and at least one polymer-modified lipid (e.g., PEG lipid). In some embodiments, the LNP may be composed of at least one cationic lipid (e.g., ionizable cationic lipid). In some embodiments, the LNP may be composed of at least one non-cationic lipid (e.g., phospholipid). In some embodiments, the LNP may be composed of sterol (e.g., cholesterol). In some embodiments, the LNP may be composed of polymer-modified lipid (e.g., PEG lipid).

[0313] The LNP described herein may be formed using techniques known in the art. As an example, in a microchannel, a delivery medium carrying a gene insertion system can be formed by mixing an acidic aqueous solution containing the gene insertion system with an organic solution containing lipids, but is not limited to this method.

[0314] Micelles In some embodiments, the delivery medium contains at least one type of micelle. In some embodiments, the micelle may be composed of the same components as the lipid nanoparticles, but basically has a different manufacturing method. As used herein, "micelle" means small particles that do not have an aqueous intra-particle space. Without wishing to be bound by any theory, the intra-particle space of the micelle does not contain lipid head groups, but rather contains the hydrophobic tails of the lipids that make up the micelle membrane and a gene insertion system (GIS) that may bind thereto.

[0315] Liposomes In some embodiments, the delivery medium comprises at least one liposome. In some embodiments, the liposome may be composed of the same components and in the same amounts as the lipid nanoparticles, but basically has a different production method. As used herein, the term "liposome" means a vesicle composed of at least one lipid bilayer surrounding an aqueous internal nanoparticle space. Further, liposomes are different from extracellular vesicles and generally do not originate from precursor cells / host cells. Liposomes include those having a plurality of concentric bilayers separated by a narrow aqueous space and having a diameter of several hundred nanometers (i.e., (large) multilamellar vesicles (MLV)), those having a diameter of less than 50 nm (small unilamellar vesicles (SUV)), and those having a diameter of 50 to 500 nm (large unilamellar vesicles (LUV)).

[0316] Exosomes In some embodiments, the delivery medium comprises at least one exosome. Generally, the term "exosome" means an extracellular vesicle, which is a small membrane-bound body derived from endocytosis. The exosome membrane generally has a lamellar structure composed of a lipid bilayer and has an aqueous internal nanoparticle space. Exosomes tend to contain components of the host cell membrane / precursor cell membrane in addition to the preset components. Without wishing to be bound by any theory, exosomes are generally released from host cells / precursor cells into the extracellular environment after the multivesicular body fuses with the cell plasma membrane.

[0317] Virus - like particles In some embodiments, the delivery medium comprises at least one type of virus-like particle (VLP). Generally, virus-like particles are non-infectious vesicles mainly composed of a viral-derived protein capsid, coat, shell, or sheath (all of which are used interchangeably with the same meaning herein) that can carry a gene insertion system (GIS). In some embodiments, the VLP may be synthesized by the self-assembly of a viral capsid protein sequence expressed by an intracellular mechanism and the incorporation of a gene insertion system (GIS). In some embodiments, the VLP may be formed by preparing each component of the capsid and the gene insertion system (GIS) and self-assembling them without using the intracellular mechanisms related to expression.

[0318] Examples of the viral family and virus species from which the VLP is derived include, but are not limited to, Parvoviridae, Retroviridae, Flaviviridae, Paramyxoviridae, adeno-associated virus, HIV, hepatitis C virus, HPV, bacteriophage, or combinations thereof.

[0319] Polymeric delivery particles In some embodiments, the delivery medium may comprise at least one polymeric delivery particle. As used herein, "polymeric delivery particle" means a non-aggregating delivery particle composed of a soluble polymer bound to a portion constituting a gene insertion system via various linking groups. In some embodiments, the polymeric delivery particle may comprise any of the polymers described herein.

[0320] In some embodiments, the delivery medium may comprise nucleic acid nanoparticles (NANP). Generally, "nucleic acid nanoparticles" are small particles formed from non-coding nucleic acid sequences that can form a three-dimensional structure capable of carrying a cargo (e.g., each component of a gene insertion system) through interactions.

[0321] Encapsulation In some embodiments, the delivery vehicle may completely encapsulate the gene insertion system disclosed herein. In some embodiments, the delivery vehicle may partially encapsulate the gene insertion system disclosed herein. In some embodiments, in the final formulation, substantially 0% of the gene insertion system contained in the delivery vehicle is exposed to the external environment of the delivery vehicle (i.e., the gene insertion system is completely encapsulated). In some embodiments, the gene insertion system is bound to the delivery vehicle, but at least a portion thereof is exposed to the external environment of the delivery vehicle.

[0322] In some embodiments, the delivery medium may be characterized by encapsulation efficiency, i.e., by the proportion of gene insertion systems that are not exposed to the external environment of the delivery medium. For the sake of clarity, a delivery medium formulation with an encapsulation efficiency of approximately 100% means that substantially the entire gene insertion system is completely encapsulated by the delivery medium, and an encapsulation rate of approximately 0% means that there is substantially no gene insertion system encapsulated in the delivery medium, for example, a delivery medium to which the gene insertion system is bound to the external surface of the delivery medium. In some embodiments, the encapsulation efficiency of the delivery medium may be less than about 100%, less than about 95%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5%. In some embodiments, the encapsulation efficiency of the delivery medium may be about 90 - 100%, about 80 - 100%, about 70 - 100%, about 60 - 100%, about 50 - 100%, about 40 - 100%, about 30 - 100%, about 20 - 100%, about 10 - 100%, about 80 - 90%, about 70 - 90%, about 60 - 90%, about 50 - 90%, about 40 - 90%, about 30 - 90%, about 20 - 90%, about 10 - 90%, about 70 - 80%, about 60 - 80%, about 50 - 80%, about 40 - 80%, about 30 - 80%, about 20 - 80%, about 10 - 80%, about 60 - 70%, about 50 - 70%, about 40 - 70%, about 30 - 70%, about 20 - 70%, about 10 - 70%, about 40 - 50%, about 30 - 50%, about 20 - 50%, about 10 - 50%, about 30 - 40%, about 20 - 40%, about 10 - 40%, about 20 - 30%, about 10 - 30%, or about 10 - 20%.

[0323] Physical properties of the nanoparticles as the delivery medium In some embodiments, the delivery medium is characterized by its shape. In some embodiments, the delivery medium may be substantially spherical, substantially cylindrical (i.e., tubular), substantially disc-shaped, but is not limited thereto.

[0324] In some embodiments, the delivery medium is characterized by its size. In some embodiments, the size of the delivery medium can be defined as its diameter. As used herein in connection with the size of the delivery medium, "diameter" means the diameter of the largest circular cross-section of the delivery medium. In some embodiments, the diameter of the delivery medium may be from 30 nm to about 150 nm. For example, the diameter of the delivery medium may be about 40 - 150 nm, about 50 - 150 nm, about 60 - 150 nm, about 70 - 150 nm, about 80 - 150 nm, about 90 - 150 nm, about 100 nm or more, about 110 - 150 nm, about 120 - 150 nm, about 130 - 150 nm, about 140 - 150 nm, about 30 - 140 nm, about 40 - 140 nm, about 50 - 140 nm, about 60 - 140 nm, about 70 - 140 nm, about 80 - 140 nm, about 90 - 140 nm, about 100 - 140 nm, about 110 - 140 nm, about 120 - 140 nm, about 130 - 140 nm, about 140 - 140 nm, about 30 - 140 nm, about 40 - 130 nm, about 50 - 130 nm, about 60 - 130 nm, about 70 - 130 nm, about 80 - 130 nm, about 90 - 130 nm, about 100 - 130 nm, about 110 - 130 nm, about 120 - 130 nm, about 30 - 120 nm, about 40 - 120 nm, about 50 - 120 nm, about 60 - 120 nm, about 70 - 120 nm, about 80 - 120 nm, about 90 - 120 nm, about 100 - 120 nm, about 110 - 120 nm, about 30 - 110 nm, about 40 - 110 nm, about 50 - 110 nm, about 60 - 110 nm, about 70 - 110 nm, about 80 - 110 nm, about 90 - 110 nm, about 100 - 110 nm, about 30 - 100 nm, about 40 - 100 nm, about 50 - 100 nm, about 60 - 100 nm, about 70 - 100 nm, about 80 - 100 nm, about 90 - 100 nm, about 30 - 90 nm, about 40 - 90 nm, about 50 - 90 nm, about 60 - 90 nm, about 70 - 90 nm, about 80 - 90 nm, about 30 - 80 nm, about 40 - 80 nm, about 50 - 80 nm, about 60 - 80 nm, about 70 - 80 nm, about 30 - 70 nm, about 40 - 70 nm, about 50 - 70 nm, about 60 - 70 nm, about 30 - 60 nm, about 40 - 60 nm, about 50 - 60 nm, about 30 - 50 nm, about 40 - 50 nm, or about 30 - 40 nm.

[0325] In some embodiments, a population of the delivery vehicles, e.g., all delivery vehicles from one formulation, may be characterized by measuring the uniformity of the physical properties (e.g., size, shape, or mass) of the particles included in the population. In some embodiments, this uniformity may be expressed as the polydispersity index (PI) of the population. In some embodiments, this uniformity may be expressed as the heterogeneity (D) of the population. As used herein, the terms "polydispersity index" and "heterogeneity" may be used interchangeably with the same meaning.

[0326] In some embodiments, the PI of a population of delivery vehicles from a formulation is from about 0.1 to 1. In some embodiments, the PI of a population of delivery vehicles from a formulation is from about 0.1 to 1, from about 0.1 to 0.8, from about 0.1 to 0.6, from about 0.1 to 0.4, from about 0.1 to 0.2, from about 0.2 to 1, from about 0.2 to 0.8, from about 0.2 to 0.6, from about 0.2 to 0.4, from about 0.4 to 1, from about 0.4 to 0.8, from about 0.4 to 0.6, from about 0.6 to 1, from about 0.6 to 0.8, or from about 0.8 to 1. In some embodiments, the PI of a population of delivery vehicles from a formulation is less than about 1, less than about 0.5, less than about 0.4, less than about 0.3, less than about 0.2, or less than about 0.1.

[0327] Targeting of delivery In some embodiments, delivery vehicles formulated to include the gene insertion system of the invention may facilitate the localization of the gene insertion system of the invention to any of the targeted regions, tissues, cells, or physiological systems described herein (i.e., the delivery vehicle "targets" a particular location). In some embodiments, targeting may be achieved by a component of the delivery vehicle included in the formulation. In some embodiments, the delivery vehicle may include a targeting agent.

[0328] Targeting agent In some embodiments, the delivery medium may contain at least one targeting agent. As used herein, the term "targeting agent" may, in some embodiments, refer to a moiety, compound, antibody (e.g., a moiety that targets a specific cell or a specific type of cell) that specifically binds to a particular type or classification of cells and / or other particular types of compounds. In some embodiments, the targeting agent may have an affinity for (i.e., be specific for) the surface of a particular target cell, a target cell surface antigen, a target cell receptor, or a combination thereof.

[0329] In some embodiments, the "targeting agent" may refer to an agent that can target the delivery medium to a particular type or classification of cells by exerting a particular action (e.g., a cleavage action) when exposed to a particular type or classification of substance and / or cells.

[0330] In some embodiments, the term "targeting agent" may refer to an agent that is part of the delivery medium and may not itself have specificity for a particular type or classification of cells, but plays a role in the specificity of the delivery medium for its target.

[0331] In some embodiments, the inclusion of at least one targeting agent in the delivery medium may increase the efficiency (e.g., the total amount or rate) at which the gene insertion system delivered by the delivery medium is taken up by cells. In some embodiments, the inclusion of at least one targeting agent in the delivery medium may increase the specificity (e.g., the total amount or rate) at which the gene insertion system delivered by the delivery medium is taken up by cells. As used herein, "specificity" means that the uptake efficiency of the cell by the target cell is higher than that by non-target cells.

[0332] In some embodiments, suitable targeting agents include, but are not limited to, small molecule targeting agents (e.g., sugar moieties), antibodies, antibody-like molecules, peptides, vitamins (e.g., folic acid), saccharides (e.g., lactose and galactose), artificial affinity molecules (e.g., peptidomimetics or aptamers), antibody fragments, single-chain variable region fragments (scFv), cell surface receptors (e.g., T cell receptors (TCR), B cell receptors (BCR) or chimeric antigen receptors (CAR)), and one or more of any combination thereof.

[0333] In some embodiments, cell surface antigens that can be targeted by the targeting agent include cell surface molecules of the target cells. Examples of suitable cell surface molecules include, but are not limited to, proteins, saccharides, lipids, or other antigens on the cell surface. In some embodiments, the cell surface antigen translocates intracellularly.

[0334] In some specific embodiments, the delivery medium may contain two or more targeting agents.

[0335] In some embodiments, at least one targeting agent may be incorporated into the lipid membrane of the nanoparticles. In some embodiments, at least one targeting agent may be presented on the outer surface of the nanoparticles. In some embodiments, at least one targeting agent may be bound to the lipid component of the nanoparticles. In some embodiments, at least one targeting agent may be bound to the polymer component of the nanoparticles. In some embodiments, a polymer-modified lipid can be formed and a delivery medium can be formed by copolymerizing a monomer containing a portion of the targeting agent (e.g., a polymerizable derivative of the targeting agent, e.g., an (alkyl)acrylic acid derivative of a peptide). In some embodiments, at least one targeting agent may be linked to the membrane structure of the nanoparticles and the internal or external aqueous environment of the nanoparticles through hydrophobic or hydrophilic interactions, thereby being linked between at least one targeting agent, the membrane structure of the nanoparticles, and the internal or external aqueous environment of the nanoparticles. In some embodiments, at least one targeting agent is bound to the peptide / protein component of the membrane structure of the nanoparticles. In some embodiments, at least one targeting agent is bound to a suitable linker moiety bound to a component of the membrane structure of the nanoparticles. In some embodiments, any combination of forces and bonds can be used to bind the targeting agent to the nanoparticles.

[0336] In some embodiments, one or more targeting agents may be linked to at least one polymer of the delivery medium via a linking moiety. In some embodiments, this linking moiety may be a cleavable linking moiety (e.g., a linking moiety containing a cleavable bond). In some embodiments, the linking moiety may contain a bond that may be cleaved by a specific enzyme (e.g., phosphatase or protease). In some embodiments, the linking moiety may contain a bond that may be cleavable by a change in intracellular pH, redox potential, or other intracellular parameters. In some embodiments, the linking moiety may contain a bond that may be cleaved by exposure to matrix metalloproteinase (MMP).

[0337] Direct transfection In some embodiments, the gene insertion system (GIS) disclosed herein may directly transfect target cells without using the delivery medium. In some embodiments, the gene insertion system (GIS) disclosed herein may transfect target cells using any technique known in the art. Such techniques include, but are not limited to, chemical transfection methods (e.g., calcium phosphate contact method), physical transfection methods (e.g., electroporation method, microinjection method, gene gun delivery method), etc. In some embodiments, direct transfection may be performed using lipid-based transfection reagents such as, but not limited to, Lipofectamine, Lipofectamine 2000, and any combination thereof.

[0338] Transplantation of transfected cells In some embodiments, the gene insertion system of the present invention may be introduced into a cell population in vitro (e.g., via direct transfection as described herein) and then transplanted into a subject. In some embodiments, the cell population for transplantation may be stem cells. In some embodiments, the cell population for transplantation may be cells derived from the subject. In some embodiments, the transplantation may be performed by methods known in the art.

[0339] IV. Pharmaceutical Compositions and Routes of Administration The present invention provides a pharmaceutical composition for administering a gene insertion system (GIS) to a subject. In some embodiments, the present invention provides a pharmaceutical composition for use as a medicine in the treatment of a therapeutic indication. In some embodiments, the pharmaceutical composition comprises at least one active ingredient (e.g., the gene insertion system (GIS) of the present invention) and at least one pharmaceutically acceptable additive, adjuvant, carrier, diluent, or any combination thereof. In some embodiments, the pharmaceutical composition is prepared as a formulation suitable for at least one route of administration. In some embodiments, the pharmaceutical composition is prepared as a formulation such that a predetermined dose of at least one active ingredient (e.g., gene insertion system (GIS)) is delivered. It can also be prepared as a formulation to be delivered according to a predetermined schedule.

[0340] As used herein, the term "pharmaceutical composition" means a composition comprising at least one active ingredient and optionally further comprising one or more pharmaceutically acceptable additives. As used herein, the term "active ingredient" generally means the gene insertion system (GIS) described herein, the gene payload delivered by the gene insertion system (GIS) for insertion into the genome of a subject, or the expression product of the gene payload delivered by the gene insertion system (GIS) described herein.

[0341] Pharmaceutical Preparations and Pharmaceutical Compositions The gene insertion system of the present invention may achieve the following by being formulated using one or more additives. (1) Increased stability of the gene insertion system or the delivery mechanism comprising the gene insertion system, (2) Increased transfection or transduction into cells, (3) Sustained or delayed introduction of the gene insertion system into the cells of a subject, (4) Modification of biodistribution (e.g., targeting of the gene insertion system to a specific tissue or a specific type of cell), (5) Increased expression of the encoded gene, (6) Modification of the release characteristics of the encoded protein, and / or (7) Controllable expression of the gene insertion system and / or its payload.

[0342] The pharmaceutical preparation may include, but is not limited to, physiological saline, liposomes, lipid nanoparticles, polymers, peptides, proteins, cells transfected with a gene insertion system (e.g., for transfer or transplantation into a subject), and any combination thereof.

[0343] In some embodiments, the formulations of the pharmaceutical compositions described herein may be prepared by any method known in the pharmacological art or developed in the future. Generally, the methods for preparing such formulations include the step of combining the active ingredient with additives and / or one or more other auxiliary components.

[0344] The formulations of the gene insertion systems and pharmaceutical compositions described herein may be carried out by any method known in the pharmacological art or developed in the future. Generally, the methods for preparing such formulations include the step of mixing the active ingredient with additives and / or one or more other auxiliary components, and optionally, the steps of dividing, shaping, and / or packaging the resulting mixture into the desired single-dose or multi-dose dosage forms, as necessary and / or desired.

[0345] The pharmaceutical compositions described herein may be prepared, packaged, and / or sold as a single unit dose and / or multiple unit doses. As used herein, the term "unit dose" means an individual quantity of the pharmaceutical composition containing a predetermined amount of the active ingredient. Generally, the amount of the active ingredient is equal to the dose of the active ingredient administered to the subject and / or a convenient fraction of such a dose, such as half or one-third of such a dose.

[0346] In some embodiments, the additive is approved for human and veterinary use. In some embodiments, the additive may be approved by the US Food and Drug Administration. In some embodiments, the additive may comply with the standards of the United States Pharmacopeia (USP), European Pharmacopeia (EP), British Pharmacopoeia, and / or International Pharmacopoeia. In some embodiments, the purity of the pharmaceutically acceptable additive may be at least 100%, at least 99%, at least 98%, at least 97%, at least 96%, or 95%. In some embodiments, the additive may be of pharmaceutical grade.

[0347] In some embodiments, the relative amounts of the pharmaceutically acceptable additive, active ingredient, and / or additional ingredients may vary widely for each pharmaceutical composition of the present invention. In some embodiments, the relative amounts may vary widely depending on the body length, condition, and / or identity of the subject being treated. In some embodiments, the relative amounts may vary widely depending on the route of administration of the pharmaceutical composition of the present invention. For example, the pharmaceutical composition of the present invention may contain 0.1% to 100% (e.g., 0.1% to 99%, 0.5 to 50%, 1 to 30%, 5 to 80%, or at least 80% (w / w)) of the active ingredient.

[0348] Additives, Diluents, and Inactive Ingredients In some embodiments, the pharmaceutical composition of the present invention may contain additives known or novel in the art. Examples of suitable additives include, but are not limited to, all kinds of preservatives, isotonic agents, thickeners, emulsifiers, solvents, dispersion media, diluents or other liquid solvents, dispersion aids or suspension aids, surfactants, and combinations thereof. In some embodiments, the additive may be selected to be suitable for the particular dosage form desired.

[0349] In some embodiments, the formulations described herein may contain at least one inert ingredient. As used herein, the term "inert ingredient" means one or more substances contained in the formulation that do not contribute to the activity of the active ingredient of the pharmaceutical composition of the present invention. In some embodiments, all or some of the inert ingredients contained in the pharmaceutical composition of the present invention may be those approved by the US Food and Drug Administration (FDA), or all of the inert ingredients contained in the pharmaceutical composition of the present invention may be those not approved by the US Food and Drug Administration (FDA).

[0350] In some embodiments, the pharmaceutical formulations disclosed herein may contain cations or anions. In some embodiments, the pharmaceutical formulations of the present disclosure contain, but are not limited to, metal cations such as Ca 2+ , Zn 2+ , Mn 2+ , Cu 2+ , Mg + , and any combination thereof. In some embodiments, the pharmaceutical formulations of the present disclosure may contain a polymer complexed with a metal cation.

[0351] In some embodiments, the pharmaceutical compositions of the present disclosure may contain one or more pharmaceutically acceptable salts. As used herein, the term "pharmaceutically acceptable salt" means a derivative of a compound of the present disclosure modified by converting an acid moiety or a base moiety present in the parent compound into the salt form (e.g., by reacting the free base with a suitable organic acid). Examples of pharmaceutically acceptable salts of the present invention include conventional non-toxic salts of the parent compound formed using non-toxic inorganic or organic acids. Further, pharmaceutically acceptable salts include, but are not limited to, alkali or organic salts of acidic residues such as carboxylic acids, and inorganic or organic salts of basic residues such as amines.

[0352] In some embodiments, the pharmaceutical composition of the present disclosure may contain at least one solvent. In some embodiments, when the solvent is water, the solvate is generally referred to as a "hydrate".

[0353] Routes of Administration The gene insertion system (GIS) of the present invention, for example, a pharmaceutical composition containing the gene insertion system (GIS) described herein, can be administered by any route as long as it is a delivery route capable of incorporating the gene insertion system (GIS) into the cells of a subject. Acceptable administration routes include transauricular administration (intratympanic administration or transmastoid administration), biliary perfusion, buccal administration (intrabuccal administration), cardiac perfusion, sacral block, conjunctival administration, transdermal administration, transdental administration (administration to one or more teeth), intracoronal administration, diagnostic administration, otic instillation, electroosmosis, endocervical administration, intranasal administration, intratracheal administration, enema, enteral administration (administration into the intestinal tract), epidermal administration (puncture from above the skin), epidural administration (administration into the epidural space), extraamniotic administration, administration via an extracorporeal circulation device, ophthalmic instillation (administration onto the conjunctiva), gastrointestinal administration, hemodialysis, infiltration administration, insufflation administration (suction from the nose), interstitial administration, intraperitoneal administration, intraamniotic administration, intraarterial administration (administration into an artery), intraarticular administration, intrahepatic duct administration, intratracheobronchial administration, intracapsular administration, intracardiac administration (administration into the heart), intracartilaginous administration (administration into cartilage), intrasacral administration (administration into the cauda equina), intracavitary injection (administration into a pathological cavity), intracorporeal cavernous administration (administration to the base of the penis), intracerebral administration (administration to the cerebrum), intracerebroventricular administration (administration to the ventricle), intracisternal administration (administration into the cisterna magna / cisterna cerebellomedullaris), intracorneal administration (administration into the cornea), intracoronary administration (administration into the coronary artery), intracorporeal cavernous administration (administration into the expansible space of the corpus cavernosum of the penis), intradermal administration (administration to the skin itself), intradiscal administration (administration into the intervertebral disc), intraductal administration (administration into a duct), intraduodenal administration (administration into the duodenum), intradural administration (administration into the dura mater or subdural space), intraepidermal administration (administration to the epidermis), intraesophageal administration (administration to the esophagus), intragastric administration (administration into the stomach), intramuscular administration (administration to the muscle), intramyocardial administration (administration into the myocardium), intraocular administration (administration into the interior of the eye), intraosseous injection (administration into the bone marrow), intraovarian administration (administration into the ovary), parenchymal administration (administration to brain tissue), pericardial administration (administration into the pericardium), intraperitoneal administration (intraperitoneal infusion or injection),Intrapleural administration (administration into the pleura), intraprostatic administration (administration into the prostate), intralung administration (administration into the lung or its bronchi), intracavitary administration (administration into the nasal cavity or the periorbital cavity), intraspinal administration (administration into the spine), intra-synovial administration (administration into the synovial cavity of a joint), intratendinous administration (administration into a tendon), intratesticular administration (administration into the testis), intrathecal administration (administration into the spinal canal), intrathecal administration (administration into the cerebrospinal fluid at any site of the cerebrospinal axis), intrathoracic administration (administration into the thoracic cavity), intraluminal administration (administration into the lumen of an organ), intratumoral administration (administration into a tumor), intratympanic administration (administration into the middle ear), intrauterine administration, intravaginal administration, intravascular administration (administration into one or more blood vessels), intravenous administration (administration into a vein), intravenous bolus, intravenous drip infusion, intraventricular administration (administration into the ventricle), intravesical instillation, intravitreal administration (administration into the interior of the eye), iontophoresis (a method of introducing ions of a soluble salt into living tissue using an electric current), irrigation (immersion irrigation or flushing of an open wound or body cavity), laryngeal administration (direct administration to the larynx), nasal administration (administration via the nasal cavity), nasogastric administration (a method of administering into the stomach through the nasal cavity), nerve block, occlusive therapy (a method of covering a site with a covering material to close it after administration via an external route), ophthalmic administration (administration to the external eye), oral administration (administration via the oral cavity), oropharyngeal administration (direct administration to the oral cavity and pharynx), parenteral administration, transdermal administration, perijoint administration, peridural administration, perineural administration, periodontal administration, photopheresis, rectal administration, respiratory administration (administration into the respiratory tract by oral inhalation or nasal inhalation to obtain a local or systemic effect), retrobulbar administration (administration behind the pons or behind the eye), soft tissue administration, subarachnoid administration, subconjunctival administration, subcutaneous administration (administration under the skin), sublabial administration, sublingual administration, submucosal administration, topical administration, transdermal administration (diffusion through intact skin for systemic distribution), transmucosal administration (diffusion through a mucosa), transplacental administration (administration through or via the placenta), transtracheal administration (administration through the tracheal wall), transtympanic administration (administration through or via the tympanic cavity), transvaginal administration, ureteral administration (administration to the ureter), urethral administration (administration to the urethra), vaginal administration as well as spinal administration, but not limited thereto.,

[0354] In some embodiments, the pharmaceutical composition of the present disclosure may be administered in a manner that can pass through the vascular barrier, the blood-brain barrier, or other epithelial barriers. The gene insertion system may be administered in any suitable form, specifically, solutions, suspensions, solid forms, solid forms suitable for dissolution in solutions, solid forms capable of being suspended in solutions, and any combination thereof, but not limited thereto.

[0355] In some embodiments, the gene insertion system may be delivered to the subject via multiple administration routes. It may be administered to two, three, four, five, or six or more sites of the subject.

[0356] In some embodiments, the gene insertion system may be delivered to the subject via a single administration route.

[0357] In some embodiments, the gene insertion system may be administered to the subject using bolus injection.

[0358] In some embodiments, the gene insertion system may be administered to the subject using a sustained delivery method (i.e., infusion) over several minutes, hours, or days. The infusion rate may be varied according to delivery parameters such as the characteristics of the subject, the desired distribution, and the formulation used, but is not limited thereto.

[0359] In some embodiments, the gene insertion system may be delivered by intramuscular delivery routes such as subcutaneous injection or intravenous injection, but is not limited thereto.

[0360] In some embodiments, the gene insertion system may be delivered by oral administration such as gastrointestinal administration or buccal administration, but is not limited thereto.

[0361] In some embodiments, the gene insertion system may be delivered by intraocular delivery routes such as intravitreal injection or eye drops of eye drops, but is not limited thereto.

[0362] In some embodiments, the gene insertion system may be delivered by an intranasal delivery route such as, but not limited to, a nasal drop or a nasal spray.

[0363] In some embodiments, the gene insertion system may be administered to a subject by injection into the periphery such as, but not limited to, intramuscular injection, intraperitoneal injection, intravenous injection, conjunctival injection, or joint injection.

[0364] In some embodiments, the gene insertion system may be delivered by injection into the cerebrospinal fluid pathway such as, but not limited to, intrathecal administration or intraventricular administration.

[0365] In some embodiments, the gene insertion system may be delivered by a systemic delivery route such as, but not limited to, intravascular administration.

[0366] In some embodiments, the gene insertion system may be administered to a subject by parenchymal administration.

[0367] In some embodiments, the gene insertion system may be administered to a subject by topical administration.

[0368] In some embodiments, the gene insertion system may be administered to a subject by intracranial delivery.

[0369] In some embodiments, the gene insertion system may be administered to a subject by intramuscular administration.

[0370] In some embodiments, the gene insertion system may be administered to a subject by intravenous administration.

[0371] In some embodiments, the gene insertion system may be administered to a subject by subcutaneous administration.

[0372] In some embodiments, the gene insertion system may be delivered by two or more administration routes.

[0373] Injections and Parenteral Administration In some embodiments, the pharmaceutical compositions described herein may be administered parenterally. Liquid dosage forms for parenteral and oral administration include, but are not limited to, pharmaceutically acceptable solutions, emulsions, microemulsions, elixirs, suspensions, and / or syrups. The liquid dosage forms may include, in addition to the active ingredient, inert diluents commonly used in the art, such as solubilizing agents, water or other solvents, and emulsifying agents (e.g., polyethylene glycol, propylene glycol, 1,3 - butylene glycol, tetrahydrofurfuryl alcohol, isopropyl alcohol, ethyl alcohol, ethyl carbonate, ethyl acetate, benzyl alcohol, benzyl benzoate, dimethylformamide, oils, glycerol, and sorbitan fatty acid esters), and any combination thereof. Examples of oils include cottonseed oil, peanut oil, corn oil, germ oil, olive oil, castor oil, and sesame oil, and mixtures thereof. In some embodiments, the pharmaceutical compositions of the present disclosure include solubilizing agents such as alcohol, oil, glycol, CREMOPHOR®, modified oil, polysorbate, polymer, cyclodextrin, and / or combinations thereof. In some embodiments, surfactants such as hydroxypropylcellulose are included.

[0374] In some embodiments, the injectable preparation may comprise a sterile aqueous suspension or a sterile oily suspension for injection. The sterile injectable solution may be formulated by known techniques using suitable wetting agents, dispersing agents and / or suspending agents. The sterile injectable preparation may be a sterile injectable suspension, a sterile injectable solution and / or a sterile injectable emulsion in a non-toxic diluent and / or solvent acceptable for parenteral administration. In some embodiments, the sterile injectable preparation may be a 1,3-butanediol solution. In some embodiments, acceptable media and solvents include, but are not limited to, Ringer's solution (United States Pharmacopeia), water, isotonic saline, and sterile non-volatile oils. In some embodiments, non-irritating non-volatile oils (e.g., synthetic monoglycerides or synthetic diglycerides) are included as the non-volatile oil. In some embodiments, fatty acids such as oleic acid can also be used in the preparation of the injection.

[0375] In some embodiments, the injectable formulation may be sterilized by filtration through a bacteria-trapping filter and / or by addition of a sterilizing agent. In some embodiments, the sterilizing agent may be in the form of a sterile solid composition that is soluble or dispersible in a sterile injectable medium such as sterile water before use.

[0376] In order to sustain the effect of the active ingredient, it is often desirable to delay the absorption of the active ingredient from subcutaneous or intramuscular injection. In some embodiments, the absorption of the parenterally administered pharmaceutical composition can be delayed by dissolving or suspending the pharmaceutical composition of the present disclosure in an oily solvent. In some embodiments, the absorption of the active ingredient may be delayed by using a liquid suspension of a low water-soluble amorphous or crystalline component. The absorption rate of the active ingredient depends on the dissolution rate, which depends on the crystal size and crystal form.

[0377] Oral Administration In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be administered orally. Solid dosage forms for oral administration include tablets, capsules, powders, pills, and granules. Generally, in solid dosage forms, the active ingredient is mixed with at least one pharmaceutically acceptable inert additive, and such additives include dicalcium phosphate or sodium citrate; binders (e.g., carboxymethyl cellulose, alginates, gelatin, polyvinylpyrrolidone, sucrose, and gum arabic); fillers or extenders (e.g., starch, lactose, sucrose, glucose, mannitol, and silicic acid); disintegrants (e.g., agar, calcium carbonate, potato starch or tapioca starch, alginic acid, certain silicates, and sodium carbonate); absorption promoters (e.g., quaternary ammonium compounds); humectants (e.g., glycerol); liquid retardants (e.g., paraffin); absorbents (e.g., kaolin and bentonite clay); wetting agents (e.g., cetyl alcohol and glycerol monostearate); lubricants (e.g., talc, calcium stearate, magnesium stearate, solid polyethylene glycol, sodium lauryl sulfate); and any combination thereof, but not limited thereto. In the case of tablets, capsules, and pills, these dosage forms may contain buffering agents.

[0378] Liquid dosage forms for oral administration may contain the additives described above for parenteral administration. Oral compositions may also contain adjuvants such as emulsifying agents, wetting agents, suspending agents, flavoring agents, sweetening agents, and / or aromatic agents in addition to inert diluents.

[0379] Topical or Transdermal Administration In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be formulated for topical administration. The skin can be an ideal target site for delivery because it is easily accessible. In some embodiments, the delivery routes to the skin or through the skin of the pharmaceutical compositions described herein include, but are not limited to, topical application (e.g., for cosmetic use and / or local / regional treatment), intradermal injection (e.g., for cosmetic use and / or local / regional treatment), and systemic delivery (e.g., for the treatment of skin diseases affecting both the skin area and the extra-dermal area).

[0380] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be delivered using various coated dressings (e.g., band-aids or wound dressings) to effectively and / or conveniently implement the methods described herein. In some embodiments, the dressing or bandage may contain a sufficient amount of the pharmaceutical composition described herein so that the user can perform multiple treatments by themselves.

[0381] Examples of dosage forms for topical administration and / or transdermal administration include lotions, creams, ointments, gels, sprays, pastes, powders, solutions, inhalants, and / or patches. Generally, dosage forms for topical administration and / or transdermal administration may be formulated by mixing the active ingredient with pharmaceutically acceptable additives, buffers, and / or preservatives as needed under aseptic conditions.

[0382] In some embodiments, a transdermal patch may be used. The transdermal patch may have the additional advantage of being able to controllably deliver the pharmaceutical composition described herein to the body. Generally, the transdermal patch may be prepared by dissolving and / or dispersing the pharmaceutical composition described herein in a suitable medium. In some embodiments, the delivery rate may be controlled by dispersing the pharmaceutical composition of the present invention in a polymer matrix and / or gel, or by providing a membrane that controls the release rate, or by combining these.

[0383] In some embodiments, suitable formulations for topical administration include liquid preparations and / or semi-liquid preparations (e.g., liniments and lotions), water-in-oil emulsions and / or oil-in-water emulsions (e.g., ointments, creams and / or pastes), solutions and / or suspensions, and any combination thereof.

[0384] Ophthalmic or Otic Administration In some embodiments, the pharmaceutical compositions described herein may be in a formulation suitable for ophthalmic administration, otic administration, or both. Generally, such formulations may be in the form of eye drops and / or ear drops, specifically including, but not limited to, solutions and / or suspensions in which the active ingredient is added to an aqueous liquid additive and / or an oily liquid additive. In some embodiments, such eye drops and / or ear drops may contain salts, buffers, one or more other additional ingredients described herein, and combinations thereof. In some embodiments, suitable formulations for ophthalmic administration include liposome preparations and / or the active ingredient in microcrystalline form. In some embodiments, the pharmaceutical compositions described herein may be administered subretinally.

[0385] Pulmonary Administration In some embodiments, the pharmaceutical compositions described herein may be in a formulation suitable for pulmonary administration. In some embodiments, pulmonary administration is performed via the buccal cavity. In some embodiments, the pharmaceutical compositions described herein may contain dry particles containing the active ingredient. In some embodiments, the diameter of the dry particles for pulmonary administration may range from about 0.5 to 7 nm or from about 1 to 6 nm.

[0386] In some embodiments, the pharmaceutical compositions described herein may be administered using a self-injecting solvent / powder supply container. Generally, the active ingredient may be dissolved and / or suspended in a low-boiling propellant within a sealed container. In some embodiments, the pharmaceutical compositions described herein may be in the form of a dry powder administered using an apparatus comprising a dry powder reservoir in which such powder is dispersed by the flow of the propellant. In some embodiments in which dry powder is used, the powder may comprise particles having a diameter greater than 0.5 nm for at least 98% of the particles by weight and a diameter less than 7 nm for at least 95% of the particles by number. In some embodiments, at least 95% of the particles have a diameter greater than 1 nm by weight and at least 90% of the particles have a diameter less than 6 nm by number. In some embodiments, the dry pharmaceutical composition containing the powder may contain a diluent (e.g., a saccharide) in the form of solid fine powder and may be conveniently provided in unit dosage form.

[0387] In some embodiments, examples of the low-boiling propellant include liquid propellants having a boiling point below 65°F at atmospheric pressure. In some embodiments, the content of the propellant in the pharmaceutical composition of the present invention may be 50% to 99.9% (w / w), and the content of the active ingredient in the pharmaceutical composition of the present invention may be 0.1% to 20% (w / w). In some embodiments, the propellant may contain additional components, specifically, liquid nonionic surfactants, solid anionic surfactants, solid diluents (e.g., solid diluents having a particle size of the same order of magnitude as the particles containing the active ingredient), and any combinations thereof, but are not limited thereto.

[0388] In some embodiments, the pharmaceutical composition formulated for pulmonary delivery may be in the form of droplets of a solution, suspension, and combinations thereof. Such formulations may be administered using an atomizing device and / or a nebulizer device when prepared, packaged, and / or sold as a solution, suspension, or combination thereof. In some embodiments, these solutions and / or suspensions may be sterilized. Exemplary solutions and / or suspensions include aqueous alcohol compositions and / or diluted alcohol compositions. In some embodiments, the pharmaceutical composition formulated for pulmonary delivery may include a flavor (e.g., sodium saccharin), a volatile oil, a surfactant, a buffer, a preservative (e.g., methyl hydroxybenzoate), and any combination thereof. In some embodiments, the average diameter of the droplets provided by the pulmonary administration route may range from about 0.1 nm to about 200 nm.

[0389] Intranasal, Nasal, or Buccal Administration In some embodiments, the pharmaceutical composition described herein may be administered by intranasal administration, nasal administration, or both. In some embodiments, the pharmaceutical composition for intranasal delivery may include the additives described herein for pulmonary delivery. In some embodiments, the pharmaceutical composition for intranasal administration includes a coarse powder containing the active ingredient, and its average particle size is about 0.2 μm to 500 μm. In some embodiments, the pharmaceutical composition for intranasal administration may be administered by rapidly inhaling through the nasal cavity from a powder container held near the nose, i.e., by sniffing from the nose. Exemplary pharmaceutical formulations may include from about 0.1% (w / w) to 100% (w / w) of the active ingredient and may further include one or more additional ingredients described herein.

[0390] In some embodiments, the pharmaceutical compositions described herein may be formulations suitable for buccal administration, and such formulations include, but are not limited to, tablets, troches, and combinations thereof. Generally, such tablets or troches may be made using conventional methods, may contain (as non-limiting examples) from 0.1% to 20% (w / w) of the active ingredient, may include any combination of compositions soluble by oral ingestion and compositions disintegratable by oral ingestion, and may include one or more of the additional ingredients described herein as optional components. In some embodiments, a pharmaceutical composition suitable for buccal administration may include any combination of powders, aerosolized liquids and / or suspensions, or atomized liquids and / or suspensions, each containing the active ingredient, and these dosage forms may be in the form of dispersed particles having an average particle size of about 0.1 nm to 200 nm and / or dispersed droplets having a droplet diameter of about 0.1 nm to 200 nm. In some embodiments, the pharmaceutical composition for buccal administration may further include one or more of the additional ingredients described herein.

[0391] Depot Administration In some embodiments, the pharmaceutical compositions described herein may be formulated in the form of a depot for sustained release. In some embodiments, the pharmaceutical compositions described herein are spatially retained within or in the vicinity of the target tissue.

[0392] Injectable depot dosage forms are generally prepared by forming a microencapsulated matrix of the pharmaceutical composition encapsulated in a biodegradable polymer (e.g., polylactic acid - polyglycolide). Generally, the release rate of the pharmaceutical composition can be controlled by varying the ratio of the pharmaceutical composition to the polymer and by varying the nature of the particular polymer used. Suitable biodegradable polymers include, but are not limited to, poly(orthoesters) and poly(anhydrides). Depot injectable formulations are prepared by encapsulating the pharmaceutical composition within liposomes or microemulsions that are compatible with living tissue.

[0393] Rectal and Vaginal Administration In some embodiments, the pharmaceutical compositions described herein may be administered rectally, vaginally, or in combination thereof. Generally, compositions for rectal or vaginal administration are suppositories, which can be prepared by mixing the active ingredient with a suitable non-irritating additive (e.g., polyethylene glycol, cocoa butter, or suppository wax) that is solid at ambient temperature but liquid at body temperature. The active ingredient is released when the suppository dissolves in the rectal or vaginal cavity.

[0394] Dosage The gene insertion system (GIS) of the present invention and / or a pharmaceutical composition comprising the gene insertion system (GIS) may be administered in an amount (i.e., dosage) that produces a desired effect (e.g., a desired therapeutic effect or research result) in a subject. In some embodiments, the desired dosage may be determined based on parameters of the subject (e.g., the subject's body length, condition or nature), parameters of the effect (e.g., the degree of response required, the threshold of the therapeutic effect, the duration of the effect, or the side effects presented), or any combination thereof. In some embodiments, the appropriate dosage may be determined prior to the first administration and may be based on at least one assay testing at least one parameter of the subject. In some embodiments, the appropriate dosage may be determined after the first administration and may be based on at least one assay testing at least one parameter of the effect. In some embodiments, the dosage may be maintained without change throughout the course of administration. In some embodiments, the dosage may be changed one, two, or multiple times throughout the course of administration.

[0395] In some embodiments, the dosage may be expressed as the mass ratio of the active ingredient to the mass of the subject (e.g., mass in units of mg / kg). For example, the dosage may be 0.1 to 100 mg / kg, 1 to 100 mg / kg, 2 to 100 mg / kg, 3 to 100 mg / kg, 4 to 100 mg / kg, 5 to 100 mg / kg, 6 to 100 mg / kg, 7 to 100 mg / kg, 8 to 100 mg / kg, 9 to 100 mg / kg, 10 to 100 mg / kg, 15 to 100 mg / kg, 20 to 100 mg / kg, 25 to 100 mg / kg, 30 to 100 mg / kg, 35 to 100 mg / kg, 40 to 100 mg / kg, 45 to 100 mg / kg, 50 to 100 mg / kg, 55 to 100 mg / kg, 60 to 100 mg / kg, 65 to 100 mg / kg, 70 to 100 mg / kg, 75 to 100 mg / kg, 80 to 100 mg / kg, 85 to 100 mg / kg, 90 to 100 mg / kg, 95 to 100 mg / kg, 0.1 to 95 mg / kg, 1 to 95 mg / kg, 2 to 95 mg / kg, 3 to 95 mg / kg, 4 to 95 mg / kg, 5 to 95 mg / kg, 6 to 95 mg / kg, 7 to 95 mg / kg, 8 to 95 mg / kg, 9 to 95 mg / kg, 10 to 95 mg / kg, 15 to 95 mg / kg, 20 to 95 mg / kg, 25 to 95 mg / kg, 30 to 95 mg / kg, 35 to 95 mg / kg, 40 to 95 mg / kg, 45 to 95 mg / kg, 50 to 95 mg / kg, 55 to 95 mg / kg, 60 to 95 mg / kg, 65 to 95 mg / kg, 70 to 95 mg / kg, 75 to 95 mg / kg, 80 to 95 mg / kg, 85 to 95 mg / kg, 90 to 95 mg / kg, 0.1 to 90 mg / kg, 1 to 90 mg / kg, 2 to 90 mg / kg, 3 to 90 mg / kg, 4 to 90 mg / kg, 5 to 90 mg / kg, 6 to 90 mg / kg, 7 to 90 mg / kg, 8 to 90 mg / kg, 9 to 90 mg / kg, 10 to 90 mg / kg, 15 to 90 mg / kg, 20 to 90 mg / kg, 25 to 90 mg / kg, 30 to 90 mg / kg, 35 to 90 mg / kg, 40 to 90 mg / kg, 45 to 90 mg / kg, 50 to 90 mg / kg, 55 to 90 mg / kg, 60 to 90 mg / kg, 65 to 90 mg / kg, 70 to 90 mg / kg, 75 to 90 mg / kg, 80 to 90 mg / kg, 85 to 90 mg / kg, 0.1 - 85 mg / kg, 1 - 85 mg / kg, 2 - 85 mg / kg, 3 - 85 mg / kg, 4 - 85 mg / kg, 5 - 85 mg / kg, 6 - 85 mg / kg, 7 - 85 mg / kg, 8 - 85 mg / kg, 9 - 85 mg / kg, 10 - 85 mg / kg, 15 - 85 mg / kg, 20 - 85 mg / kg, 25 - 85 mg / kg, 30 - 85 mg / kg, 35 - 85 mg / kg, 40 - 85 mg / kg, 45 - 85 mg / kg, 50 - 85 mg / kg, 55 - 85 mg / kg, 60 - 85 mg / kg, 65 - 85 mg / kg, 70 - 85 mg / kg, 75 - 85 mg / kg, 80 - 85 mg / kg, 0.1 - 80 mg / kg, 1 - 80 mg / kg, 2 - 80 mg / kg, 3 - 80 mg / kg, 4 - 80 mg / kg, 5 - 80 mg / kg, 6 - 80 mg / kg, 7 - 80 mg / kg, 8 - 80 mg / kg, 9 - 80 mg / kg, 10 - 80 mg / kg, 15 - 80 mg / kg, 20 - 80 mg / kg, 25 - 80 mg / kg, 30 - 80 mg / kg, 35 - 80 mg / kg, 40 - 80 mg / kg, 45 - 80 mg / kg, 50 - 80 mg / kg, 55 - 80 mg / kg, 60 - 80 mg / kg, 65 - 80 mg / kg, 70 - 80 mg / kg, 75 - 80 mg / kg, 0.1 - 75 mg / kg, 1 - 75 mg / kg, 2 - 75 mg / kg, 3 - 75 mg / kg, 4 - 75 mg / kg, 5 - 75 mg / kg, 6 - 75 mg / kg, 7 - 75 mg / kg, 8 - 75 mg / kg, 9 - 75 mg / kg, 10 - 75 mg / kg, 15 - 75 mg / kg, 20 - 75 mg / kg, 25 - 75 mg / kg, 30 - 75 mg / kg, 35 - 75 mg / kg, 40 - 75 mg / kg, 45 - 75 mg / kg, 50 - 75 mg / kg, 55 - 75 mg / kg, 60 - 75 mg / kg, 65 - 75 mg / kg, 70 - 75 mg / kg, 0.1 - 70 mg / kg, 1 - 70 mg / kg, 2 - 70 mg / kg, 3 - 70 mg / kg, 4 - 70 mg / kg, 5 - 70 mg / kg, 6 - 70 mg / kg, 7 - 70 mg / kg, 8 - 70 mg / kg, 9 - 70 mg / kg, 10 - 70 mg / kg, 15 - 70 mg / kg, 20 - 70 mg / kg, 25 - 70 mg / kg, 30 - 70 mg / kg, 35 - 70 mg / kg, 40 - 70 mg / kg, 45 - 70 mg / kg, 50 - 70 mg / kg, 55 - 70 mg / kg, 60 - 70 mg / kg, 65 - 70 mg / kg, 0.1 - 65 mg / kg, 1 - 65 mg / kg, 2 - 65 mg / kg, 3 - 65 mg / kg, 4 - 65 mg / kg, 5 - 65 mg / kg, 6 - 65 mg / kg, 7 - 65 mg / kg, 8 - 65 mg / kg, 9 - 65 mg / kg, 10 - 65 mg / kg, 15 - 65 mg / kg, 20 - 65 mg / kg, 25 - 65 mg / kg, 30 - 65 mg / kg, 35 - 65 mg / kg, 40 - 65 mg / kg, 45 - 65 mg / kg, 50 - 65 mg / kg, 55 - 65 mg / kg, 60 - 65 mg / kg, 0.1 - 60 mg / kg, 1 - 60 mg / kg, 2 - 60 mg / kg, 3 - 60 mg / kg, 4 - 60 mg / kg, 5 - 60 mg / kg, 6 - 60 mg / kg, 7 - 60 mg / kg, 8 - 60 mg / kg, 9 - 60 mg / kg, 10 - 60 mg / kg, 15 - 60 mg / kg, 20 - 60 mg / kg, 25 - 60 mg / kg, 30 - 60 mg / kg, 35 - 60 mg / kg, 40 - 60 mg / kg, 45 - 60 mg / kg, 50 - 60 mg / kg, 55 - 60 mg / kg, 0.1 - 55 mg / kg, 1 - 55 mg / kg, 2 - 55 mg / kg, 3 - 55 mg / kg, 4 - 55 mg / kg, 5 - 55 mg / kg, 6 - 55 mg / kg, 7 - 55 mg / kg, 8 - 55 mg / kg, 9 - 55 mg / kg, 10 - 55 mg / kg, 15 - 55 mg / kg, 20 - 55 mg / kg, 25 - 55 mg / kg, 30 - 55 mg / kg, 35 - 55 mg / kg, 40 - 55 mg / kg, 45 - 55 mg / kg, 50 - 55 mg / kg, 0.1 - 50 mg / kg, 1 - 50 mg / kg, 2 - 50 mg / kg, 3 - 50 mg / kg, 4 - 50 mg / kg, 5 - 50 mg / kg, 6 - 50 mg / kg, 7 - 50 mg / kg, 8 - 50 mg / kg, 9 - 50 mg / kg, 10 - 50 mg / kg, 15 - 50 mg / kg, 20 - 50 mg / kg, 25 - 50 mg / kg, 30 - 50 mg / kg, 35 - 50 mg / kg, 40 - 50 mg / kg, 45 - 50 mg / kg, 0.1 - 45 mg / kg, 1 - 45 mg / kg, 2 - 45 mg / kg, 3 - 45 mg / kg, 4 - 45 mg / kg, 5 - 45 mg / kg, 6 - 45 mg / kg, 7 - 45 mg / kg, 8 - 45 mg / kg, 9 - 45 mg / kg, 10 - 45 mg / kg, 15 - 45 mg / kg, 20 - 45 mg / kg, 25 - 45 mg / kg, 30 - 45 mg / kg, 35 - 45 mg / kg, 40 - 45 mg / kg, 0.1 - 40 mg / kg, 1 - 40 mg / kg, 2 - 40 mg / kg, 3 - 40 mg / kg, 4 - 40 mg / kg, 5 - 40 mg / kg, 6 - 40 mg / kg, 7 - 40 mg / kg, 8 - 40 mg / kg, 9 - 40 mg / kg, 10 - 40 mg / kg, 15 - 40 mg / kg, 20 - 40 mg / kg, 25 - 40 mg / kg, 30 - 40 mg / kg, 35 - 40 mg / kg, 0.1 - 35 mg / kg, 1 - 35 mg / kg, 2 - 35 mg / kg, 3 - 35 mg / kg, 4 - 35 mg / kg, 5 - 35 mg / kg, 6 - 35 mg / kg, 7 - 35 mg / kg, 8 - 35 mg / kg, 9 - 35 mg / kg, 10 - 35 mg / kg, 15 - 35 mg / kg, 20 - 35 mg / kg, 25 - 35 mg / kg, 30 - 35 mg / kg, 0.1 - 30 mg / kg, 1 - 30 mg / kg, 2 - 30 mg / kg, 3 - 30 mg / kg, 4 - 30 mg / kg, 5 - 30 mg / kg, 6 - 30 mg / kg, 7 - 30 mg / kg, 8 - 30 mg / kg, 9 - 30 mg / kg, 10 - 30 mg / kg, 15 - 30 mg / kg, 20 - 30 mg / kg, 25 - 30 mg / kg, 0.1 - 25 mg / kg, 1 - 25 mg / kg, 2 - 25 mg / kg, 3 - 25 mg / kg, 4 - 25 mg / kg, 5 - 25 mg / kg, 6 - 25 mg / kg, 7 - 25 mg / kg, 8 - 25 mg / kg, 9 - 25 mg / kg, 10 - 25 mg / kg, 15 - 25 mg / kg, 20 - 25 mg / kg, 0.It may be 1 to 20 mg / kg, 1 to 20 mg / kg, 2 to 20 mg / kg, 3 to 20 mg / kg, 4 to 20 mg / kg, 5 to 20 mg / kg, 6 to 20 mg / kg, 7 to 20 mg / kg, 8 to 20 mg / kg, 9 to 20 mg / kg, 10 to 20 mg / kg, 15 to 20 mg / kg, 0.1 to 15 mg / kg, 1 to 15 mg / kg, 2 to 15 mg / kg, 3 to 15 mg / kg, 4 to 15 mg / kg, 5 to 15 mg / kg, 6 to 15 mg / kg, 7 to 15 mg / kg, 8 to 15 mg / kg, 9 to 15 mg / kg, 10 to 15 mg / kg, 0.1 to 10 mg / kg, 1 to 10 mg / kg, 2 to 10 mg / kg, 3 to 10 mg / kg, 4 to 10 mg / kg, 5 to 10 mg / kg, 6 to 10 mg / kg, 7 to 10 mg / kg, 8 to 10 mg / kg, 9 to 10 mg / kg, 0.1 to 9 mg / kg, 1 to 9 mg / kg, 2 to 9 mg / kg, 3 to 9 mg / kg, 4 to 9 mg / kg, 5 to 9 mg / kg, 6 to 9 mg / kg, 7 to 9 mg / kg, 8 to 9 mg / kg, 0.1 to 8 mg / kg, 1 to 8 mg / kg, 2 to 8 mg / kg, 3 to 8 mg / kg, 4 to 8 mg / kg, 5 to 8 mg / kg, 6 to 8 mg / kg, 7 to 8 mg / kg, 0.1 to 7 mg / kg, 1 to 7 mg / kg, 2 to 7 mg / kg, 3 to 7 mg / kg, 4 to 7 mg / kg, 5 to 7 mg / kg, 6 to 7 mg / kg, 0.1 to 6 mg / kg, 1 to 6 mg / kg, 2 to 6 mg / kg, 3 to 6 mg / kg, 4 to 6 mg / kg, 5 to 6 mg / kg, 0.1 to 5 mg / kg, 1 to 5 mg / kg, 2 to 5 mg / kg, 3 to 5 mg / kg, 4 to 5 mg / kg, 0.1 to 4 mg / kg, 1 to 4 mg / kg, 2 to 4 mg / kg, 3 to 4 mg / kg, 0.1 to 3 mg / kg, 1 to 3 mg / kg, 2 to 3 mg / kg, 0.1 to 2 mg / kg, 1 to 2 mg / kg, or 0.1 to 1 mg / kg.

[0396] Dosing Schedule The gene insertion system of the present invention and / or a pharmaceutical composition comprising the gene insertion system may be administered at a frequency (i.e., dosing schedule) that provides a desired effect (e.g., a desired therapeutic effect or research result, etc.) in a subject. In some embodiments, the dosing schedule may be determined by any of the methods used to determine the dosages described herein. In some embodiments, the gene insertion system may be administered only once.

[0397] In some embodiments, the gene insertion system of the present invention may be administered two or more times. For example, the gene insertion system of the present invention may be administered 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some embodiments, the gene insertion system of the present invention may be administered intermittently and / or continuously during the treatment course of a therapeutic indication in a subject. In some embodiments, the gene insertion system of the present invention may be repeatedly administered to a subject over a lifetime.

[0398] V. Methods of Use Target Regions, Target Tissues, or Target Cells in Delivering Gene Insertion System Preparations Provided herein is a method of delivering a pharmaceutical composition and / or pharmaceutical formulation described herein to at least one target location in a subject, the method being carried out by contacting at least one target (including one or more target cells), such as a physiological system, anatomical site, organ, tissue, a particular type of cell, cell population, etc., with at least one pharmaceutical composition and / or pharmaceutical formulation described herein.

[0399] The pharmaceutical composition and / or pharmaceutical formulation described herein contains an effective amount of an active ingredient (e.g., the gene insertion system of the present invention) sufficient to achieve a desired effect (e.g., the effect of inserting at least one transgene into the genome of a subject) in at least one cell located at the target.

[0400] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein typically include one or more cell permeabilizers, but "naked" formulations (e.g., formulations that do not contain cell permeabilizers or other agents) that may include a pharmaceutically acceptable carrier are also contemplated.

[0401] Physiological Systems In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target physiological systems.

[0402] In some embodiments, physiological systems include the auditory system, cardiovascular system, central nervous system, chemoreceptor system, circulatory system, digestive system, endocrine system, excretory system, exocrine system, genital system, integumentary system, lymphatic system, muscular system, musculoskeletal system, nervous system, peripheral nervous system, renal system, reproductive system, respiratory system, urinary system, and visual system.

[0403] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target the amine precursor uptake and decarboxylation system (APUD) (a series of cells that have endocrine functions and secrete various low molecular weight amines or polypeptide hormones), and APUD system tissues include, but are not limited to, pituitary tissue, parathyroid tissue, thyroid tissue, bronchial tissue, adrenal medulla tissue, pancreatic tissue, gastrointestinal tract, carotid body, chemoreceptor system tissue, and the like.

[0404] Organs In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target organs. Organs include the anal canal, artery, ascending colon, bladder, bone marrow, brain, bronchus, bronchiole, bulbourethral gland, capillary, cecum, cerebellum, cerebral hemisphere, cerebrum, cervix, choroid plexus, clitoris, cranial nerve, descending colon, diencephalon, duodenum, ear, enteric nervous system, epididymis, esophagus, external genitalia, fallopian tube, gallbladder, ganglion, gustatory organ, gut-associated lymphoid tissue, heart, ileum, internal genitalia, interstitial, jejunum, joint, kidney, large intestine, larynx, ligament, liver, lung, lymph node, lymphatic vessel, mammary gland, medulla oblongata, mesentery, midbrain, mouth, respiratory muscle, nasal cavity, nerve, olfactory organ, ovary, pancreas, parotid gland, penis, pharynx, placenta, pons, prostate, rectum, salivary gland, scrotum, seminal vesicle, sigmoid colon, skeleton, skin, small intestine, spinal nerve, spleen, stomach, subcutaneous tissue, sublingual gland, submandibular gland, tooth, tendon, testis, brainstem, spinal cord, ventricular system, thymus, tongue, tonsil, trachea, transverse colon, ureter, urethra, uterus, vagina, vas deferens, vein, and vulva.

[0405] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target one or both eyes.

[0406] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target the liver.

[0407] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target the brain.

[0408] Cells In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target specific cells and / or specific types of cells.

[0409] Examples of cells include adipocytes, adrenergic neurons, alpha cells, acinar cells, ameloblasts, anterior chamber lens epithelial cells, anterior / posterior pituitary cells, apocrine sweat gland cells, astrocytes, auditory inner hair cells of the organ of Corti, auditory outer hair cells of the organ of Corti, B cells, Bartholin gland cells, cornea, tongue, oral cavity, nasal cavity, distal anal canal, basal cells (stem cells) of the distal urethra and distal vagina, olfactory epithelial basal cells, basket cells, basophil granulocytes and their progenitor cells, beta cells, Betz cells, bone marrow reticular fibroblasts, border cells of the organ of Corti, border cells, Bowman gland cells, brown adipocytes, Brunner gland cells, bulbourethral gland cells, atrial cells, C cells, Cajal-Retzius cells, cardiomyocytes, cardiac muscle cells, cartwheel cells, fasciculata cells that produce glucocorticoids, zona glomerulosa cells that produce mineralocorticoids, zona reticularis cells that produce androgens, adrenal cortex cells, cementoblasts, atrial cells, ceruminous gland cells of the ear canal, Chandelier cells, chemoreceptor glomus cells of the carotid body, chief cells, cholinergic neurons, chromaffin cells, club cells, cold-sensitive primary sensory neurons, connective tissue macrophages (all types), corneal fibroblasts (corneal stromal cells of the cornea), luteal cells of the ruptured follicle that secrete progesterone, dermal hair stem cells, adrenocorticotropic hormone-secreting cells, lens fiber cells containing crystallins, epidermal hair stem cells, cytotoxic T cells, D cells, delta cells, dendritic cells, double bouquet cells, duct cells, eccrine sweat gland clear cells, eccrine sweat gland dark cells, testicular efferent duct cells, chondrocytes of elastic cartilage, endothelial cells, enteric glial cells, enterochromaffin cells, enterochromaffin-like cells, enteroendocrine cells, eosinophil granulocytes and their progenitor cells, epithelial cells, epidermal basal cells, epidermal Langerhans cells, suprarenal gland basal cells, suprarenal gland chief cells, epithelial reticular cells, epsilon cells, erythrocytes, chondrocytes of fibrous cartilage, fork neurons, crypt cells, G cells, gallbladder epithelial cells, germ cells, litter gland cells, Moll gland cells of the eyelid, glial cells, Golgi cells, gonadal interstitial cells, gonadotropin-secreting cells, granulosa cells, granulosa lutein cells, granule cells and mitral cells.

[0410] In some embodiments, the cells may be cancer cells. In some embodiments, the cells may be non-cancer cells.

[0411] In some embodiments, the eukaryotic cell may be a stem cell. Various types of stem cells are known in the art, and all of them may be used to practice the present disclosure. Examples of stem cells include, but are not limited to, embryonic stem cells, hematopoietic stem cells, neural stem cells, epidermal neural crest stem cells, induced pluripotent stem cells, mammary stem cells, intestinal stem cells, mesenchymal stem cells, olfactory adult stem cells, testicular cells, and progenitor cells (e.g., neural progenitor cells, angioblasts, osteoblasts, chondroblasts, pancreatic progenitor cells, epidermal progenitor cells, etc.). In some embodiments, the stem cell may be an embryonic stem cell line derived from cells collected from a subject.

[0412] In some embodiments, the eukaryotic cell is a cell found in the circulatory system of a human, non-human primate, and / or other mammal (including mice and / or rats). Exemplary circulatory system cells include, but are not limited to, platelets, plasma cells, red blood cells, B cells, T cells, natural killer cells, macrophages, neutrophils, and their progenitor cells. In some embodiments, at least one eukaryotic cell may be derived from any of these circulatory system eukaryotic cells.

[0413] In some embodiments, at least one eukaryotic cell is a natural killer cell or a progenitor cell of a natural killer cell.

[0414] In some embodiments, at least one eukaryotic cell is a B cell or a progenitor B cell.

[0415] In some embodiments, the eukaryotic cell may be a plant cell. In some embodiments, the plant cell is a cell of a monocotyledonous or dicotyledonous plant, specifically, zucchini, woody plants such as coniferous plants and deciduous trees, wheat, turnip, tomato, tobacco, sunflower, sugarcane, sugar beet, strawberry, spinach, soybean, sorghum, rye, rice, raspberry, rapeseed, radish, pumpkin, potato (including sweet potato), plum, pineapple, peanut, pea, papaya, oat, melon, mango, corn, lettuce, lentil, herb, hemp, forage grass, flower, eucalyptus, cucumber, cotton, coffee, citrus, chicory, blueberry, celery, cauliflower, carrot, canola, cabbage, broccoli, rape, blackberry, legume, barley, banana, avocado, asparagus, Arabidopsis thaliana, other fruit plants, ornamental plants, almond, alfalfa, perennial forage grass, feed crops, other vegetables, other stone fruits (such as peach, nectarine, apricot, pear, plum, etc.), other pome fruits (such as apple and pear, etc.), other fruits, other bulbous plants (such as garlic, onion, leek, etc.), other agricultural crops, parts of perennial plants (such as bulbs; tubers; roots; crowns; stems; stolons; tillers; shoots; cuttings without roots, cuttings, and cuttings with callus or young plants forming callus, etc.); apical meristems, etc.), and any combination thereof, or hybrids thereof, but not limited thereto. As used herein, "plant" means the physical parts of a plant, such as seeds, seedlings, saplings, roots, tubers, stems, petioles, leaves and fruits.

[0416] Tumors In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target tumors. The tumor may be a benign tumor, a pre-cancerous tumor or a malignant tumor.

[0417] Insertion of Transgenes The present invention provides a method for introducing a transgene into a subject, for example, a human subject. In some embodiments, the method includes introducing into the subject, in an effective amount, at least one gene insertion system described herein. In some embodiments, the method includes introducing into the subject, in an effective amount, at least one gene insertion system comprising a transgene.

[0418] In some embodiments, the method may include inserting the transgene into one or more target insertion sites. Referring to FIG. 8, a genomic region 500 of a subject having an inserted transgene is shown. In this example, the genomic DNA of the subject includes a target insertion site 120 and genomic DNA 110 therearound. For the sake of clarity, it should be noted that the target insertion site is a part of the DNA of the subject. The 5' junction 510 indicates the transition point between the DNA of the subject and the inserted transgene 520, which corresponds to the 5' end of the transgene. This junction 510 may have a partial or complete overlap of the upstream sequence of the target site that is present in both the genomic DNA of the subject and the 5' end of the template RNA. In contrast, the 3' junction 530 indicates the transition point between the 3' end of the transgene and the DNA of the subject. This junction 530 may have a partial or complete overlap of the downstream sequence of the target site that is present in both the genomic DNA of the subject and the 3'-side module of the template RNA. The junction 510 and / or the junction 530 may further include another nucleotide, which may be, for example, a nucleotide added to a primer before elongation or a cDNA 3' end primer by a reverse transcriptase with a nucleotide other than the template, or a nucleotide generated by dissociation from the double strand of the template product by an enzyme.

[0419] Target Insertion Sites In some embodiments, one or more target insertion sites include safe harbor sites. As used herein, a "safe harbor site" means a specific location in the genome of a subject where insertion of a transgene does not cause disruption of unintended cellular functions. Generally, a site in the genome may be identified as a safe harbor site if (a) insertion of genetic material at that site does not modify the expression of a gene of interest, or (b) insertion of genetic material at that site modifies the expression of a gene, but (e.g., because there are many repetitive sequences in the disrupted gene of the subject's genome) the change does not modify the normal function of the subject's cells. An example of case (b) includes, but is not limited to, the case where in the genome, the coding genes of ribosomal RNA (rRNA) are repeated in an amount such that normal cellular functions are not perturbed even if some rRNA genes are disrupted.

[0420] In some embodiments, at least one safe harbor site and / or target insertion site includes at least one ribosomal DNA (rDNA) sequence. As used herein, the term "ribosomal DNA" means a gene encoding rRNA. In some embodiments, at least one safe harbor site and / or target insertion site includes at least one 28S rDNA sequence.

[0421] Transgenes The methods and compositions of the invention can be used for insertion of any payload sequence (i.e., transgene), and the length and origin of the payload sequence are not limited.

[0422] In some embodiments, the transgene includes a therapeutically active gene. As used herein, the term "therapeutically active gene" means any gene capable of expressing an expression product useful for the treatment, alleviation or prevention of at least one therapeutic indication.

[0423] In some embodiments, at least one transgene may comprise at least one telomerase reverse transcriptase (TERT) gene. In some embodiments, at least one transgene may comprise at least one short-chain factor VIII gene. In some embodiments, at least one transgene may comprise at least one phenylalanine hydroxylase (PAH) gene.

[0424] In some embodiments, at least one transgene is a reporter gene. As used herein, "reporter gene" means any gene capable of expressing an expression product detectable by an assay.

[0425] In some embodiments, at least one reporter gene may include, but is not limited to, at least one green fluorescent protein (GFP), at least one red fluorescent protein (RFP), luciferase enzyme (LUC), β-galactosidase (LacZ), chloramphenicol acetyltransferase (cat), etc., and may encode these.

[0426] Non-Wild-Type Transgenes One of ordinary skill in the art will understand that many of the transgenes exemplified above are native or wild-type sequences, but the gene insertion system disclosed herein is not limited to the insertion of wild-type genes or native genes or a portion of their gene sequences. The gene insertion system of the present invention may be used, for example, to insert genes derived from wild-type genes, genes containing only a portion of a wild-type gene, genes assembled from portions of various types of wild-type genes, and / or genes whose sequences are not known to exist in nature. Further, the gene insertion system of the present invention may be used to insert transgenes whose expression products are not normally found in the target cells and / or transgenes whose gene expression is not normally observed.

[0427] Regulatory Factors of Transgenes In some embodiments, the gene insertion system of the present invention may be used to insert at least one regulatory factor or at least one transgene encoding the same. For example, the transgene may be designed and / or recombined such that any number of miRNA binding regions and / or siRNA binding regions are included in the expression product of the transgene. Usually, by incorporating miRNA and / or siRNA, cells containing complementary miRNA or siRNA in the transcriptome can be removed from the target of transgene expression.

[0428] In some embodiments, the transgene may include at least one miRNA and / or siRNA, or a first expression product encoding the same, and at least one miRNA binding site and / or siRNA binding site complementary to the first expression product, or a second expression product (or more than one number of expression products) encoding the same, and may encode the same. Without wishing to be bound by any theory, long-term expression of the second expression product may be blocked in such a manner.

[0429] Antibodies As used herein, the term "antibody" is used in the broadest sense and specifically includes various embodiments, including monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies formed from at least two full-length antibodies), and antibody fragments (e.g., diabodies), as long as they exhibit the desired biological activity (e.g., as long as they are "functional"), but are not limited thereto. An antibody is a molecule mainly composed of amino acids and is a monomeric polypeptide or multimeric polypeptide that includes at least one amino acid region derived from a known antibody sequence or a parental antibody sequence. An antibody may include amino acid motifs that recruit one or more endogenous modifications or non-natural modifications (including, but not limited to, the addition of carbohydrate moieties, fluorescent moieties, chemical tags, etc.). For the purpose of achieving the object of the present invention, an "antibody" may include a heavy chain variable domain, a light chain variable domain, and an Fc region.

[0430] The gene insertion system of the present invention may contain at least one or more functional antibodies or may be used to insert a transgene encoding the same.

[0431] Treatment of Therapeutic Indications The present invention provides a method for treating or preventing at least one therapeutic indication in a subject in need of treatment or prevention of at least one therapeutic indication. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system described herein into the subject. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system containing at least one therapeutically active transgene into the subject.

[0432] In some embodiments, at least one therapeutic indication includes at least one genetic disorder associated with loss of function. In some embodiments, at least one method of treating at least one therapeutic indication includes administering to a subject at least one transgene that restores the subject from a genetic disorder associated with loss of function. As used herein, the term "restore" means providing to the subject at least one composition that enables the subject to perform the original function that the subject has lost.

[0433] In some embodiments, the at least one method includes restoring insufficient telomerase activity in a subject by administering to the subject, in an effective amount, a gene insertion system comprising at least one TERT transgene.

[0434] In some embodiments, the methods and compositions of the invention may be used to treat or prevent conditions caused by insufficient telomerase function in a subject. In some embodiments, the at least one method includes administering to a subject presenting insufficient telomerase activity, in a therapeutically effective amount, at least one gene insertion system comprising at least one TERT gene. In some embodiments, the at least one method includes administering to a subject suspected of developing a disease due to insufficient telomerase activity, in a therapeutically effective amount, at least one gene insertion system comprising at least one TERT gene.

[0435] Regulation of Heterologous Genes The gene insertion system of the invention, including the pharmaceutical formulations and pharmaceutical compositions described herein, may be used in methods of regulating the expression of a heterologous gene. For the sake of clarity, the term "heterologous gene" as used herein with respect to the regulation of gene expression in this specification means a gene other than the gene inserted by the gene insertion system of the invention in the genome of the subject.

[0436] Typically, a method for regulating the expression of a heterologous gene may include inserting, using the gene insertion system of the present invention, a sequence whose expression product acts on the expression pathway of another gene. For example, the expression product of the inserted gene may affect transcription from the heterologous gene to mRNA, translation from the mRNA of the heterologous gene to polypeptide, the degradation rate or inactivation rate of the mRNA of the heterologous gene in the cytoplasm, or any combination thereof.

[0437] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one microRNA (miRNA), or a transgene encoding at least one microRNA (miRNA). In some embodiments, miRNAs suitable for the practice of the present disclosure may include miRNAs known in the art or miRNAs to be discovered in the future. In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one artificial miRNA, or a transgene encoding at least one artificial miRNA, which artificial miRNA is designed to bind to at least one gene expression product present within the subject. As used herein, the term "artificial miRNA" means an miRNA whose sequence has been modified or designed to bind to a desired target sequence. Artificial miRNAs may be designed by various methods known in the art.

[0438] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one small interfering RNA (siRNA), or a transgene encoding at least one small interfering RNA (siRNA). As used herein, the term "small interfering RNA" means a double-stranded ribonucleic acid (dsRNA) having a nucleotide sequence that is substantially identical to at least a portion of a target gene. Generally, siRNAs are usually 21 to 25 nucleotides in length, but may be shorter or longer, and interfere with (suppress) the expression of the target gene by promoting the degradation of the mRNA of the target gene. Known siRNAs or siRNAs discovered in the future may be suitable for use in the present invention.

[0439] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one artificial siRNA, or a transgene encoding at least one artificial siRNA. As used herein, the term "artificial siRNA" means an siRNA whose sequence is designed to have complementarity to at least one target gene.

[0440] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one transcription factor (TF), or a transgene encoding at least one transcription factor (TF). As used herein, the term "transcription factor" means a polypeptide that binds to DNA and modifies or affects the transcription of at least one gene. Known transcription factors or transcription factors discovered in the future may be suitable for use in the present invention.

[0441] The gene insertion system of the present invention may include any combination of miRNA, siRNA, and / or transcription factors, or may be used to insert a transgene encoding these. For example, at least one gene insertion system may include at least one miRNA and at least one siRNA; at least one miRNA and at least one transcription factor; at least one siRNA and at least one transcription factor; or at least one miRNA, at least one siRNA, and at least one transcription factor, or may be used to insert a transgene encoding these.

[0442] Preventive Uses In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for the prevention of diseases or the stabilization of the progression of therapeutic indications.

[0443] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as a preventive measure for preventing the onset of future therapeutic indications.

[0444] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used to prevent further progression of therapeutic indications.

[0445] Vaccines In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as a vaccine and / or in a manner similar to a vaccine. As used herein, "vaccine" means a biological preparation that improves immunity against a specific therapeutic indication or infectious agent.

[0446] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as a vaccine for treatment areas including, but not limited to, dermatology, neurology, cardiology, oncology, endocrinology, immunology, respiratory medicine, anti-infective medicine, etc., and / or may be used in a manner similar to a vaccine for such uses.

[0447] Antigens The gene insertion system of the present invention may be used to insert a transgene containing at least one antigen, or a transgene encoding at least one antigen, which antigen may be activated by at least one cell of a subject or presented on its surface. As used herein, the term "antigen" means a composition that elicits an immune response in an organism. For example, it means a composition that can cause an organism to produce antibodies against itself, and in particular, a composition that can elicit an acquired immune response in the organism by its antibodies. The antigen may be an immunogenic substance such as, for example, a polypeptide, a protein, a polysaccharide, a nucleic acid, a lipid, etc. In some embodiments, the antigen may be derived from an infectious agent such as a bacterium, a virus, a protozoan, a fungus, a prion, etc., but is not limited thereto.

[0448] In some embodiments, the antigen may contain a part or subunit of an infectious agent, for example, it may contain a coat, a coat component, a coat protein, a coat polypeptide, a surface component, a surface protein, a surface polypeptide, a capsular component, a cell wall component, a flagellum, a pilus, a toxin, or a toxoid.

[0449] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene containing at least one antigen for vaccination against at least one therapeutic indication, or a transgene encoding such an antigen.

[0450] Research In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as diagnostic purposes or research tools for the therapeutic indications disclosed herein.

[0451] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in any research experiment, for example, in in vivo or in vitro experiments.

[0452] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for the detection of research biomarkers.

[0453] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used on cultured cells. The cultured cells may be derived from any source known to those skilled in the art, and may be, but are not limited to, stable cell lines, animal models or cells derived from human patients or control subjects.

[0454] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in in vivo experiments in animal models (i.e., mice, rats, rabbits, cats, dogs, non-human primates, guinea pigs, Drosophila, ferrets, nematodes, zebrafish, or other animals used for research purposes known in the art).

[0455] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for stem cells and / or cell differentiation.

[0456] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in human research experiments or human clinical trials.

[0457] The present invention provides a method for conducting scientific research and / or medical research on a subject. In some embodiments, the method includes introducing an effective amount of at least one gene insertion system described herein into the subject. In some embodiments, the method includes introducing an effective amount of at least one gene insertion system comprising at least one reporter transgene into the subject.

[0458] Monotherapy and Combination Therapies In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as a monotherapy or combination therapy for the treatment of diseases.

[0459] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as monotherapy. In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in combination therapy. The combination therapy may combine one or more neuroprotective agents, and the neuroprotective agents may be, for example, small molecule compounds, growth factors, hormones, etc. whose neuroprotective effects against neuronal degeneration have been evaluated.

[0460] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in combination with one or more other therapeutic agents. "Combining" should not be construed as suggesting the co-administration of multiple therapeutic agents and / or formulation for co-delivery, but such delivery methods are also included within the scope of the present invention. The pharmaceutical compositions and / or pharmaceutical formulations described herein, as well as other therapeutic agents, can be administered simultaneously, before, or after one or more other desired therapeutic agents or medical procedures. Usually, each therapeutic agent is administered at the dosage and / or time schedule specified for that therapeutic agent.

[0461] The therapeutic agents that may be used in combination with the pharmaceutical compositions and / or pharmaceutical formulations described herein may be small molecule compounds, and the small molecule compounds may be antioxidants, anti-inflammatory agents, anti-apoptotic agents, calcium regulators, anti-glutamate agonists, structural protein inhibitors, compounds involved in muscle function, or compounds involved in metal ion regulation.

[0462] Synthesis of Gene Insertion Constructs (GICs) In Vivo The present invention provides a method for synthesizing a gene insertion system (GIS) biopolymer, for example, a gene insertion construct (GIC) biopolymer. In some embodiments, the method comprises administering at least one construct for GIC synthesis to a population of target cells, maintaining the cell population for a time sufficient for the at least one construct for GIS synthesis to be expressed by the target cells, and recovering and purifying the expression product of the construct for GIS synthesis by methods known in the art.

[0463] In some embodiments, the at least one construct for GIC synthesis comprises or encodes the GIC of the present invention. In some embodiments, the at least one construct for GIC synthesis comprises or encodes the GIC of the present invention and means for synthesizing at least one recombinant RNA in vivo. Such means may include providing or encoding an RNA polymerase promoter, sequences for selection and purification of the recombinant RNA, complementary GIC sequences, and post-recombinant RNA production processing signals. In some embodiments, the at least one construct for GIC synthesis is administered in the form of a DNA plasmid that enables the production of the encoded RNA by endogenous mechanisms within the cell.

[0464] An exemplary GIC synthesis construct 600 is shown in FIG. 9. At the 5' end of this GIC synthesis construct, the RNAP module 610 may include a suitable RNA polymerase promoter (e.g., the T7 RNAP promoter). If an optional 5' leader module 620 is present, this 5' leader module 620 is located on the 3' side of the RNAP module. Components included in this 5' leader module 620 include components that improve the folding and self-cleavage of the 5' side module of the template, and / or destabilize the transcript, and / or components that can rapidly remove GIC transcripts having an immunogenic 5' end (e.g., resulting from a failure of self-cleavage by RZ). Before use as a GIC, the expressed 5' leader module RNA is cleaved at the RZ self-cleavage site 630. The complementary strand 640 of the 5' side module, the complementary strand 650 of the template module, and the complementary strand 660 of the 3' side module encode the 5' side module, the template module, and the 3' side module of the GIC, respectively. Further, the 3' end may be a linearized restriction enzyme site 670, which is a cleavage point by a restriction enzyme. By being cleaved by the restriction enzyme at this cleavage point, the GIC RNA is linearized and all extra vector components remain on the vector.

[0465] VII. List of Embodiments Embodiment 1. A system for genome editing, comprising: (i) at least one reverse transcriptase construct (RTC) comprising a polynucleotide encoding a polypeptide having an enzyme activity for reverse transcribing a polynucleotide template; (ii) at least one gene insertion construct (GIC) comprising at least one polynucleotide template suitable for reverse transcription by the polypeptide encoded by the at least one RTC A system comprising.

[0466] Embodiment 2. The system according to Embodiment 1, wherein the at least one reverse transcriptase construct includes at least one biopolymer, and the biopolymer includes at least one nucleic acid, at least one amino acid, and any combination thereof.

[0467] Embodiment 3. The system according to Embodiment 1 or 2, wherein the at least one reverse transcriptase construct may include at least one reverse transcriptase module (RTC: RT module), at least one 5'-side module (RTC: 5'-side module), at least one 3'-side module (RTC: 3'-side module), and any combination thereof.

[0468] Embodiment 4. The system according to Embodiment 3, wherein the at least one reverse transcriptase module includes or encodes at least one reverse transcriptase.

[0469] Embodiment 5. The system according to Embodiment 3 or 4, wherein the at least one reverse transcriptase module includes or encodes at least one reverse transcriptase derived from a non-long terminal repeat (non-LTR) retroelement.

[0470] Embodiment 6. The system according to Embodiment 4 or 5, wherein the at least one reverse transcriptase module includes or encodes a non-natural translation initiation codon.

[0471] Embodiment 7. The system according to any one of Embodiments 4 to 6, wherein the at least one reverse transcriptase includes at least one DNA binding domain, at least one RNA binding domain, at least one cDNA synthesis domain, at least one endonuclease domain, and any combination thereof.

[0472] Embodiment 8. The system according to Embodiment 7, wherein at least one reverse transcriptase domain, at least one target DNA binding domain, at least one template RNA binding domain, and at least one endonuclease domain, and at least one of any combination thereof, are derived from a species different from the species from which the reverse transcriptase from which at least one of the remaining domains is derived is derived.

[0473] Embodiment 9. The system according to Embodiment 3, wherein at least one 5'-side module, which is any component of the reverse transcriptase construct, includes at least one RNA polymerase promoter, at least one 5' untranslated region (5'-UTR), at least one Kozak sequence, at least one 5' cap, and any combination thereof, or encodes these.

[0474] Embodiment 10. The system according to Embodiment 3, wherein at least one 3'-side module, which is any component of the reverse transcriptase construct, includes at least one reverse transcriptase translation stop codon, at least one 3' untranslated region (3'UTR), at least one polyA tail, and any combination thereof, or encodes these.

[0475] Embodiment 11. The system according to any one of Embodiments 1 to 10, wherein the at least one reverse transcriptase module includes at least one structure described in FIGS. 2 to 5 or any combination thereof, or encodes these.

[0476] Embodiment 12. The system according to any one of Embodiments 1 to 11, wherein the at least one reverse transcriptase construct includes at least one sequence of SEQ ID NOs: 1 to 57 and any combination thereof, encodes this sequence, or is encoded by this sequence.

[0477] Embodiment 13. The system according to Embodiment 1, wherein the at least one gene insertion construct comprises or encodes at least one nucleic acid biopolymer.

[0478] Embodiment 14. The system according to Embodiment 1 or 13, wherein the at least one gene insertion construct comprises or encodes at least one GIC:5'-side module which is an optional component, at least one GIC:payload module, at least one GIC:3'-side module which is an optional component, and any combination thereof.

[0479] Embodiment 15. The system according to Embodiment 14, wherein the at least one GIC:5'-side module comprises or encodes at least one sequence derived from the 5'-region of a natural retroelement, and may include or encode at least one rRNA sequence, at least one ribozyme sequence, at least one folding motif sequence, or any combination thereof.

[0480] Embodiment 16. The system according to Embodiment 15, wherein the at least one rRNA sequence which is an optional component of the GIC:5'-side module comprises or encodes a 1-30 nt rRNA of interest.

[0481] Embodiment 17. The system according to Embodiment 15, wherein the at least one ribozyme sequence which is an optional component of the GIC:5'-side module comprises or encodes at least one self-cleaving ribozyme, and the self-cleaving ribozyme may include or encode a ribozyme of hepatitis delta virus.

[0482] Embodiment 18. The system according to Embodiment 17, wherein the at least one ribozyme sequence which is an optional component of the GIC:5'-side module comprises or encodes a ribozyme derived from the 5'-region of at least one non-long terminal repeat retroelement.

[0483] Embodiment 19. At least one folding motif sequence, which is any component of the GIC:5'-side module, includes or encodes at least one self-folding RNA sequence motif, and the self-folding RNA sequence motif may include or encode at least one hairpin motif, at least one stem-loop motif, at least one paired stem 4 motif, or any combination thereof. The system according to Embodiment 15.

[0484] Embodiment 20. The GIC:5'-side module includes at least one of SEQ ID NOs: 60 to 154, 250 to 276, and 278 to 279, or any combination thereof, or encodes any one of them. The system according to any one of Embodiments 14 to 19.

[0485] Embodiment 21. The at least one GIC:3'-side module includes or encodes at least one reverse transcriptase recognition sequence, and may include or encode at least one rRNA sequence, at least one A-tract sequence, or any combination thereof. The system according to Embodiment 14.

[0486] Embodiment 22. At least one reverse transcriptase recognition sequence of the GIC:3'-side module includes or encodes at least one sequence that interacts with at least one reverse transcriptase. The system according to Embodiment 21.

[0487] Embodiment 23. At least one reverse transcriptase recognition sequence of the GIC:3'-side module is derived from the 3'-region of a natural retroelement. The system according to Embodiment 21 or 22.

[0488] Embodiment 24. At least one rRNA sequence, which is any component of the GIC:3'-side module, includes or encodes an rRNA of 1 to 30 nt. The system according to Embodiment 21.

[0489] Embodiment 25. The system according to embodiment 21, wherein at least one A tract array, which is an arbitrary component of the GIC: 3'-side module, includes or encodes a sequence consisting of 1 to 50 adenine bases.

[0490] Embodiment 26. The system according to any one of embodiments 14 and 21 to 25, wherein the at least one GIC: 3'-side module includes or encodes at least one of SEQ ID NOs: 300 to 329 or any combination thereof.

[0491] Embodiment 27. The system according to embodiment 14, wherein the at least one GIC: payload module includes or encodes at least one transgene sequence, and may include or encode at least one promoter sequence of the transgene, at least one 5'-untranslated sequence of the transgene, at least one 3'-untranslated sequence of the transgene, at least one polyadenylation signal sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, or any combination thereof.

[0492] Embodiment 28. The system according to embodiment 27, wherein the at least one transgene sequence includes or encodes at least one target sequence for insertion into the genome of a subject.

[0493] Embodiment 29. The system according to embodiment 27, wherein the at least one promoter sequence of the transgene includes or encodes at least one sequence that promotes the expression of the transgene in the genome of a subject.

[0494] Embodiment 30. The system according to embodiment 27, including at least one 5'-untranslated sequence of the transgene, and the at least one 5'-untranslated sequence includes or encodes at least one 5'-untranslated region of the mRNA of the transgene.

[0495] Embodiment 31. The system according to embodiment 27, wherein at least one 3'untranslated sequence of the transgene contains or encodes at least one 3'untranslated region of the mRNA of the transgene.

[0496] Embodiment 32. The system according to embodiment 27, wherein at least one polyadenylation signal sequence of the transgene contains or encodes at least one polyadenylation signal of the transgene.

[0497] Embodiment 33. The system according to embodiment 27, wherein at least one non-coding RNA (ncRNA) processing sequence of the transgene contains or encodes at least one termination signal, at least one 3'processing signal, and any combination thereof of at least one ncRNA expressed from the transgene.

[0498] Embodiment 34. The system according to any one of embodiments 14 and 27 to 33, wherein the at least one GIC:payload module contains or encodes at least one of SEQ ID NOs: 499 to 525 or any combination thereof.

[0499] Embodiment 35. The system according to any one of embodiments 13 to 34, wherein at least one of the at least one GIC:5'side module and the at least one GIC:3'side module contains or encodes at least one sequence derived from a species different from the species from which the non-long terminal repeat retroelement from which the other is derived.

[0500] Embodiment 36. The system according to any one of embodiments 1 and 13 to 35, wherein the at least one gene insertion construct contains or encodes at least one structure shown in FIGS. 6 to 9 and any combination thereof.

[0501] Embodiment 37. (i) At least one reverse transcriptase construct contained in at least one of SEQ ID NOs: 1 to 57 or encoded by at least one of these; (ii) At least one gene insertion construct contained in at least one of the sequences of SEQ ID NOs: 60 to 154, 250 to 276, 278 to 279, 280 to 289, 300 to 329, 400 to 404, 405 to 407, 411 to 422, and 499 to 536 or encoded by at least one of these sequences The system according to any one of Embodiments 1 and 13 to 36, comprising

[0502] Embodiment 38. A system according to any one of Embodiments 1 and 13 to 37, comprising at least one of the gene insertion constructs according to Embodiments 13 to 37 or a construct for synthesizing a gene insertion construct (GIC: construct for synthesis) encoding the same.

[0503] Embodiment 39. The system according to any one of Embodiments 1 to 38, wherein at least one of the at least one reverse transcriptase construct and the at least one gene insertion construct comprises at least one sequence derived from a species different from the species from which the retroelement from which the other is derived is derived or encodes this sequence.

[0504] Embodiment 40. A system according to any one of Embodiments 1 to 39, comprising at least one of combinations of (i) at least one reverse transcriptase construct according to Embodiments 2 to 12 and (ii) at least one gene insertion construct according to Embodiments 13 to 37.

[0505] Embodiment 41. A method for inserting at least one transgene into the genome of a subject, the method comprising the step of administering an effective amount of at least one of the gene insertion systems (GIS) according to Embodiments 1 to 40.

[0506] Embodiment 42. The method according to embodiment 41, wherein the introduced gene is inserted into one or more target sites of the genome of the subject, and the one or more target sites may include at least one safe harbor site.

[0507] Embodiment 43. The method according to embodiment 42, wherein the at least one safe harbor site, which is an arbitrary component, includes at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may include at least one 28S rDNA sequence.

[0508] Embodiment 44. The method according to any one of embodiments 40 to 43, comprising the step of administering at least one of the gene insertion systems formulated with at least one delivery agent.

[0509] Embodiment 45. The method according to embodiment 44, wherein the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may include at least one lipid nanoparticle.

[0510] Embodiment 46. A pharmaceutical composition comprising at least one of the gene insertion systems according to embodiments 1 to 40, and optionally including at least one additive, at least one delivery agent, at least one adjuvant, and any combination thereof.

[0511] Embodiment 47. A method for treating a therapeutic indication in a subject in need of treatment of the therapeutic indication, comprising the step of administering, in an effective amount, at least one of the gene insertion systems according to embodiments 1 to 40 or at least one of the pharmaceutical compositions according to embodiment 46, and optionally including at least one of the methods according to embodiments 41 to 45.

[0512] Embodiment 48. The method according to embodiment 47, wherein the therapeutic indication is caused by a deletion of telomerase activity.

[0513] Embodiment 49. The method according to Embodiment 46 or 47, wherein the at least one gene insertion system comprises at least one TERT transgene.

[0514] Embodiment 50. A kit for preparing a gene insertion system, comprising the method of the gene insertion system according to Embodiments 1 to 40, optionally comprising the pharmaceutical composition according to Embodiment 46, and further optionally comprising a buffer, a DNA plasmid, or a protocol for preparing the gene insertion system or the pharmaceutical composition.

[0515] VIII. Definitions 28S rDNA: As used herein, the term "28S rDNA" means a part of the genome that encodes the large ribosomal RNA (rRNA) that constitutes the large subunit (LSU) of the cytoplasmic ribosome of eukaryotes.

[0516] 3'-junction: As used herein, the term "3'-junction" means the position where the 3'-end of the inserted sequence is ligated to the 5'-end of the target genome.

[0517] 3'-region: As used herein, the term "3'-region" means the part located on the 3'-side of the open reading frame in a retroelement gene.

[0518] 5'-junction: As used herein, the term "5'-junction" means the position where the 3'-end of the target genome is ligated to the 3'-end of the inserted sequence.

[0519] 5'-region: As used herein, the term "5'-region" means the part located on the 5'-side of the open reading frame in a retroelement gene.

[0520] Activity: As used herein, the term "activity" means a state in which some event is occurring or a state in which some event is being performed. The proteins and nucleic acids of the present disclosure may have activity, and this activity may be involved in one or more biological events.

[0521] Altered: As used herein, the term "altered" means a change in a protein sequence or amino acid sequence to change, add, or remove its properties and / or activity.

[0522] Assay: As used herein, when the term "assay" is used as a verb, this term is used in the broadest sense and means the act of performing a test using an appropriate method known in the art. As used herein, when the term "assay" is used as a noun, this term means a test used to measure the properties, states, and / or activities of the assay target.

[0523] Biological property: As used herein, the terms "biological property" and "property" mean the characteristics or activities of an organism, physiological system, organ, tissue, cell, or molecule that are measurable or observable.

[0524] Cargo: In the context related to delivery vehicles, the terms "cargo" and "payload" usually mean a compound or structure (e.g., the gene insertion system of the present invention) intended for delivery to a target cell, tissue, organ, or physiological system, delivery to these, or delivery in the vicinity of these.

[0525] Cell: As used herein, the term "cell" has the broadest possible meaning and means a living membrane-bound structure.

[0526] Cell process: As used herein, the term "cell process" and grammatically equivalent terms thereto mean a process occurring at the cell level, where the cell level may be limited to one cell or may not be limited to one cell.

[0527] Feature: In this specification, the term "feature" usually refers to a characteristic or quality belonging to a human, place, or thing, and means something that serves to identify these. The terms "feature" and "characteristic" have the same meaning and may be used interchangeably.

[0528] Impart: In this specification, the term "impart" and terms grammatically equivalent thereto mean the process of adding a feature to an object.

[0529] Construct: In this specification, the noun "construct" means an artificially designed biopolymer. Examples of biopolymers include DNA, RNA, and polypeptides. Usually, the constructs described in this specification are designed for use in gene insertion systems.

[0530] Degradation: In this specification, "degradation" means the loss of function of a composition over time.

[0531] Delivery: In this specification, the term "delivery" means an act or method of delivering a compound, substance, object, part, cargo, or payload to a living cell or living organism. Unless otherwise specified, the term "delivery" and the term "biological delivery" may be used with the same meaning.

[0532] Delivery system: In this specification, the term "delivery system" means a composition, method, or combination thereof that can deliver each component of the gene insertion system into the cytoplasm of a target cell when prepared as a formulation in combination with the gene insertion system of the present invention. Examples of delivery systems include, but are not limited to, systems composed of a delivery medium and systems for direct transfection.

[0533] Derived from: As used herein, the term "derived from" means a nucleic acid sequence or a protein sequence, such as a non-long terminal repeat (non-LTR) retrotransposon, that has been isolated or obtained from a particular source. This term includes natural sequences that have been isolated or obtained from a particular source. Further, this term includes artificial variant sequences derived from sources having the same or similar functional characteristics, for example, the variant may include a nucleic acid sequence or an amino acid sequence that has been modified so that its functional characteristics are improved as compared to the molecule derived from the original source.

[0534] Designed: As used herein, the term "designed" means a composition that has been altered from its natural or current state to have novel and desired properties and / or activities.

[0535] DNA and RNA: As used herein, the terms "RNA", "RNA molecule" or "ribonucleic acid molecule" mean a polymer of ribonucleotides. As used herein, the terms "DNA", "DNA molecule" or "deoxyribonucleic acid molecule" mean a polymer of deoxyribonucleotides. DNA and RNA can be synthesized naturally. For example, DNA can be synthesized naturally by replication of DNA, and RNA can be synthesized naturally by transcription of DNA. Also, DNA and RNA can be synthesized chemically. DNA and RNA can be single-stranded (i.e., ssRNA or ssDNA) or multi-stranded (e.g., double-stranded, i.e., dsRNA or dsDNA). As used herein, the term "mRNA" or "messenger RNA" means a single-stranded RNA that encodes the amino acid sequence of one or more polypeptide chains. When an RNA sequence is described using deoxyribonucleotides, the DNA sequence can be converted to an RNA sequence by substituting thymidine ("T") with uridine ("U") or a uridine analog.

[0536] DNA repair: As used herein, the term "DNA repair" refers to an endogenous process carried out within a cell to correct damage that has occurred to the genome within the cell.

[0537] Efficient: As used herein, the term "efficient" and grammatically equivalent terms in connection with the insertion of a transgene mean the effectiveness of any combination of an RT protein and GIC:5'-side module and GIC:3'-side module when inserting the full length of a payload module into a desired target site.

[0538] Element: As used herein, the term "element" refers to an individual component of a molecule, system, or step of a method.

[0539] Expression product: As used herein, the term "expression product" refers to RNA transcribed from a sequence of interest (e.g., mRNA), or a polypeptide translated from mRNA transcribed from a sequence of interest.

[0540] Enclose: As used herein, the term "enclose" means to contain, surround, or store.

[0541] Encode: As used herein, the term "encode" broadly refers to a process of inducing the production of a second molecule different from a first molecule using information written in a polymeric macromolecule (the first molecule). The second molecule may have a chemical structure with chemical properties different from those of the first molecule.

[0542] Endonuclease: As used herein, the term "endonuclease" refers to a protein or a part thereof that cleaves a polynucleotide chain by degrading nucleotides other than both ends.

[0543] Exosome: As used herein, an "exosome" is a vesicle secreted by mammalian cells or a complex involved in the degradation of RNA.

[0544] Ex vivo: The term "ex vivo" means removing cells from a donor subject, modifying the cells using the methods described herein, and transplanting the cells into a recipient subject. This term includes autologous cells obtained from one individual subject (i.e., this subject is both the donor of the unmodified cells and the recipient of the ex vivo modified cells), and allogeneic cells obtained from a donor subject that is a different individual from the recipient subject. The allogeneic donor and recipient may have a HLA match.

[0545] Facilitate: As used herein, the term "facilitate" is used in its broadest sense and means making some action or process more likely to occur by adding a particular element.

[0546] Fidelity: As used herein, the term "fidelity" means the accuracy with which a gene of interest is inserted into the genome of a subject. The term "high fidelity" corresponds to the insertion of the gene of interest with relatively few errors with respect to nucleotide identity, sequence length, and the position of the target site. For example, if a template RNA containing approximately 5,000 nucleotides can be copied by an RT protein to produce cDNA without generating mismatched base pairs, this gene insertion has high fidelity. Although it depends on the purpose of the transgene insertion, even if a small number of mismatches occur, a sufficiently high fidelity can produce a functional transgene.

[0547] Adjacent: As used herein, the term "adjacent" means that one element is located on the 5' side (5' adjacent) or 3' side (3' adjacent) of another element. The adjacent elements may be directly linked to each other, or there may be another element between the adjacent elements.

[0548] Formulation: As used herein, "formulation" includes at least one component of the gene insertion system described herein and at least one delivery agent or pharmaceutically acceptable additive or both thereof.

[0549] Functional / active form: As used herein, the term "functional" when used in connection with a biomolecule means the biomolecule in a form that exhibits the properties and / or activities that characterize it.

[0550] Gene: As used herein, the term "gene" is used in the broadest sense and means a distinguishable nucleotide sequence that forms part of a chromosome or a distinguishable nucleotide sequence that may form part of a chromosome, and by its order, determines the order of monomers in a polypeptide or in a nucleic acid molecule.

[0551] Gene insertion construct: As used herein, the term "gene insertion construct" or "GIC" means an RNA construct that contains an RNA template for a reverse transcriptase protein.

[0552] Gene insertion system: As used herein, the term "gene insertion system" or "GIS" is a system consisting of each component (module) that may be used for the insertion of a gene sequence (transgene) into a specific position of a target genome via reverse transcription such as TPRT.

[0553] GIC: 3'-side module: As used herein, the term "3'-side module" means a part of a gene insertion construct (GIC) that contains at least one element derived from the 3'-region of a retroelement gene or at least one element that replaces the function of the 3'-region of a retroelement gene.

[0554] GIC: 5'-side module: As used herein, the term "5'-side module" refers to a part of a gene insertion construct (GIC) that facilitates the insertion of the full length of the transgene, and may or may not be derived from the 5' region of a retroelement gene.

[0555] Generate: As used herein, the verb "generate" and its conjugated forms are used in the broadest sense and refer to the process for obtaining a specific product.

[0556] Genome: As used herein, the term "genome" is used in the broadest sense and refers to all genetic material present within a cell.

[0557] Delta hepatitis virus (HDV) ribozyme (RZ) folding structure: As used herein, the term "HDV RZ folding structure" refers to an RNA sequence that can adopt the folding structure of the ribozyme of the delta hepatitis virus (HDV) and retains the function of the ribozyme.

[0558] Heterologous: As used herein, the term "heterologous" refers to the sequence or structure of a gene or protein that is not normally made in a certain cell when introduced into that cell. Further, this term includes individual elements, modules or parts of the reverse transcriptase construct or gene insertion construct of the present disclosure that contain nucleic acid sequences (DNA or RNA) or amino acid sequences derived from different biological species. For example, the 5'-side module of a reverse transcriptase construct or gene insertion construct may contain a sequence derived from one species (or the first species) of birds, and the 3'-side module of this reverse transcriptase construct or gene insertion construct may contain a sequence derived from another species (or the second species) of birds.

[0559] Homologous recombination: As used herein, the term "homologous recombination" refers to the process of inserting a transgene that depends on the sequence homology between the transgene and the target genome.

[0560] In vitro: As used herein, the term "in vitro" means a reaction or process that occurs outside of living cells or living organisms.

[0561] In vivo: As used herein, the term "in vivo" means a reaction or process that occurs inside or on the surface of living cells or living organisms.

[0562] Inactive form: As used herein, the term "inactive form" in relation to a biomolecule means a form of the biomolecule that does not exhibit the characteristics and / or activities that characterize it.

[0563] Inactive ingredient: As used herein, the term "inactive ingredient" means one or more agents that do not contribute to the activity of the active ingredient of a pharmaceutical composition contained in a formulation. In some embodiments, all of the inactive ingredients that may be used in the formulations of the present invention may be those approved by the US Food and Drug Administration (FDA), some of them may be those approved by the FDA, and all of the inactive ingredients that may be used in the formulations of the present invention may be those not approved by the FDA.

[0564] Induce: As used herein, the term "induce" and terms grammatically equivalent thereto mean a process that brings about a particular result without imposing a particular limitation on the process.

[0565] Introduce: As used herein, the term "introduce" means to add genetic material (often DNA) to a cell.

[0566] Insert: As used herein, the term "insert" means to add nucleotides to a DNA sequence.

[0567] Linking region: In this specification, the term "linking region" means the position where the cDNA of the transgene inserted into the target genome is linked to the DNA of the target genome at the insertion site.

[0568] At least one: In this specification, the term "at least one" means one, two, three, four, or five or more modified objects, for example, the constructs, modules, or sequences of the present disclosure.

[0569] Lipid nanoparticle: In this specification, the term "lipid nanoparticle" or "LNP" means a delivery medium containing one or more lipids (for example, cationic lipids, non-cationic lipids, PEG-modified lipids).

[0570] Liposome: In this specification, the term "liposome" usually means a vesicle composed of one or more bilayers or spherical bilayers made of lipids (for example, amphiphilic lipids).

[0571] Loss of function: In this specification, the term "loss of function" means a change in the target gene that results in a modified gene product in which the function of the wild-type gene is lost.

[0572] Modification: In this specification, "modification" means that a change has been made to the state or structure of a molecule. The molecule may be chemically, structurally, or functionally modified in various ways.

[0573] Module system: In this specification, the term "module system" means a system that can be divided into multiple sets composed of multiple parts that function relatively autonomously with respect to each other and interact strongly.

[0574] Motif: In this specification, the term "motif" means a sequence of a biopolymer having a recognizable structure, and this recognizable structure may or may not be defined by a specific chemical function or biological function.

[0575] Natural: As used herein, the term "natural" means a wild-type or native compound, biomolecule (e.g., protein or nucleic acid), or composition.

[0576] Non-LTR retroelement reverse transcriptase: As used herein, the term "non-LTR retroelement reverse transcriptase (RT)" means a protein having reverse transcriptase activity derived from a non-LTR retroelement.

[0577] Non-LTR retroelement: As used herein, the term "non-LTR retroelement" means a group of retroelement genes (also known as retrotransposons) that do not contain long terminal repeats.

[0578] Outside: As used herein, the term "outside" in relation to an insertion site means any portion of the genome that is more than about 60 bp away from the 5' or 3' end of the insertion site.

[0579] Pairing reverse transcriptase (RT): As used herein, "pairing RT" means a reverse transcriptase (RT) used in combination with at least one module containing an insertion payload module. The module may be homologous to the pairing RT, which means that all elements contained in this module and the RT are derived from the same retroelement gene. The module may be heterologous to the pairing RT, which means that at least one element contained in this module is not derived from the same retroelement gene as the RT.

[0580] Payload: The term "payload" may mean the sequence of a nucleic acid (e.g., a gene of interest) contained in a gene insertion system (GIS) intended to be inserted into the genome of a subject, except when used in the context of a delivery vehicle.

[0581] Percent identity: The terms "percent identity" or "identity (%)" mean the amount of identical or the same sequences between two nucleic acid sequences or amino acid sequences. As defined herein, the term "percent identity" can be used in the same sense as the terms "proportion of identity" or "proportion of sequence identity".

[0582] As used herein, "proportion of identity", "proportion of sequence identity" or "percent identity" is determined by comparing two sequences in an optimized alignment within a comparison window, and a portion of the sequences within the comparison window may have additions or deletions (i.e., gaps) as compared to a reference sequence (excluding additions and deletions) to optimize the alignment between the two sequences. The proportion of sequence identity can be calculated by measuring the number of positions where identical nucleic acid bases or amino acid residues are present in both sequences, calculating the number of positions where the nucleic acid bases or amino acid residues match, dividing the number of matching positions by the total number of positions within the comparison window, and multiplying the resulting value by 100.

[0583] The terms "identical," "identity," or "homology" in the context of two or more nucleic acid sequences or polypeptide sequences mean two or more identical sequences or subsequences. When measured using one of the sequence comparison algorithms described below, or when compared or aligned to maximize correspondence within a comparison window or a specified region by manual alignment and visual inspection, if the same nucleotide or amino acid residues are present in a given percentage in multiple sequences, these sequences are "substantially identical" to each other (e.g., having at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identity in a specified region, etc.). This definition also applies to the complementary strands of the test sequences. Thus, unless otherwise stated, all nucleic acid sequences and amino acid sequences provided herein include sequences that are substantially identical to the reference sequences.

[0584] In sequence comparison, usually one sequence is used as the reference sequence against which the test sequence is compared. When using a sequence comparison algorithm, the test sequence and the reference sequence are input into a computer, coordinates of subsequences are specified if necessary, and parameters of the sequence algorithm program are specified. Generally, default program parameters are used, and other parameters can also be specified. Next, the sequence comparison algorithm calculates the percentage of sequence identity or sequence similarity of the test sequence to the reference sequence based on the program parameters.

[0585] Suitable algorithms for determining sequence identity and sequence similarity include the BLAST algorithm described by Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977) and the BLAST 2.0 algorithm described by Altschul et al. (J. Mol. Biol. 215:403-10, 1990). Software for performing BLAST analysis is publicly available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). In this algorithm, first, high scoring sequence pairs (HSPs) are identified. A short word of length W in the query sequence is identified by aligning with words of the same length on the database sequence, and if this word matches a threshold T score that shows a certain positive value, or satisfies this T score, it is reported as an HSP. "T" is called the neighborhood word score threshold (Altschul et al., supra). The first identified neighborhood word hits are used as seeds for starting the search, and long HSPs containing these words are identified. The word hits are extended in both directions along each sequence, and the extension continues as long as the cumulative alignment score increases. The cumulative score for nucleotide sequences is calculated using M (reward score for a matched pair of residues; always >0) and N (penalty score for mismatched residues; always <0). For amino acid sequences, the cumulative score is calculated using a score matrix. The extension of the word hits in both directions is stopped when the cumulative alignment score begins to decrease by the value of X from the maximum value, when the alignment of one or more residues with a negative score accumulates and the cumulative score becomes 0 or less, or when the end of one of the sequences is reached. The sensitivity and speed of the alignment are determined by W, T, and X, which are parameters of the BLAST algorithm. In the BLASTN program (for nucleotide sequences), as default parameters, word length (W)=11, expectation value (E)=10, M=5, N=-4, and comparison of both strands are used.When comparing amino acid sequences, in the BLASTP program, as default parameters, an alignment using a word length = 3, an expectation value (E) = 10, the BLOSUM62 score matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915, 1989), (B) = 50, an expectation value (E) = 10, M = 5, and N = -4 is used.

[0586] Furthermore, in the BLAST algorithm, a statistical analysis of the similarity between two sequences is performed (see, for example, Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One of the similarity indicators provided by the BLAST algorithm is the sum of the minimum probabilities (P(N)), which provides an indicator of the probability that a match between two nucleoti...

Claims

**Claim 1** A genome editing system comprising: (i) at least one reverse transcriptase construct (RTC); and (ii) at least one gene insertion construct (GIC). The RTC includes at least one reverse transcriptase module (RTC: RT module), and the RTC: RT module includes messenger RNA (mRNA) encoding a reverse transcriptase (RT), at least one 5'-side module (RTC: 5'-side module), and / or at least one 3'-side module (RTC: 3'-side module). The GIC includes at least one RNA template suitable for reverse transcription by a polypeptide encoded by the at least one RTC, and may include at least one GIC: 5'-side module, at least one GIC: payload module, and at least one GIC: 3'-side module. A system. **Claim 2** i) The 5'-side module of the RTC includes a 5'-untranslated region (5'UTR), a Kozak sequence or an internal ribosome entry site, an unnatural translation start codon, and / or a 5'-cap; ii) The RT module includes mRNA encoding a reverse transcriptase derived from an organism selected from the group consisting of Zonotrichia albicollis (ZoAl), Taeniopygia guttata (TaGu), Tinamus guttatus (TiGu), Oryzias latipes (OrLa), and Tribolium castaneum (strain B) (TriCasB); iii) The 3'-side module of the RTC includes a translation stop codon of the reverse transcriptase, a 3'-untranslated region (3'UTR), and a polyA tail; iv) The GIC: 5'-side module includes a sequence derived from the 5'-region of a natural retroelement, an rRNA sequence, a ribozyme sequence, a folding motif sequence, and / or an RNA polymerase terminator sequence. ​ (v) The GIC: payload module includes at least one transgene ORF or non-coding RNA (ncRNA) sequence, a promoter sequence of the transgene, a 5' untranslated sequence of the transgene, a 3' untranslated sequence of the transgene, a polyadenylation signal sequence of the transgene, and / or an ncRNA processing sequence of the transgene; (iv) The GIC: 3'-side module includes a reverse transcriptase recognition sequence, an rRNA sequence, and / or an A-tract sequence. The system according to claim 1.

3. The system according to claim 1 or 2, wherein the at least one reverse transcriptase is derived from a non-long terminal repeat (non-LTR) type retroelement or a modified variant thereof.

4. The system according to any one of claims 1 to 3, wherein the at least one reverse transcriptase includes at least one DNA binding domain, at least one RNA binding domain, at least one cDNA synthesis domain, at least one endonuclease domain, and any combination thereof.

5. The system according to any one of claims 1 to 4, wherein the reverse transcriptase is derived from birds.

6. The system according to claim 5, wherein the reverse transcriptase is derived from ZoA1, TaGu or TiGU.

7. The system according to claim 6, wherein the reverse transcriptase includes an amino acid sequence having at least 90% identity with SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 29 or SEQ ID NO:

25.

8. The system according to any one of claims 1 to 7, wherein the at least one reverse transcriptase module includes or encodes at least one structure described in FIGS. 2 to 5 or any combination thereof.

9. The system according to any one of claims 1 to 8, wherein the at least one reverse transcriptase construct includes, encodes, or is encoded by at least one sequence selected from the group consisting of SEQ ID NOs: 1 to 57 and any combination thereof.

10. The system according to any one of claims 2 to 9, wherein at least one rRNA sequence, which is an arbitrary component of the GIC: 5'-side module, contains or encodes a 1-30 nt rRNA of interest.

11. The system according to claim 10, wherein the rRNA sequence contains a sequence selected from the group consisting of SEQ ID NOs: 250 to 276, or contains a sequence having one, two, or three nucleotide changes as compared with a sequence selected from the group consisting of SEQ ID NOs: 250 to 276.

12. The system according to claim 11, wherein the GIC: 5'-side module does not contain an rRNA sequence.

13. The system according to any one of claims 2 to 12, wherein the ribozyme sequence of the GIC: 5'-side module contains at least one self-cleaving ribozyme, and the self-cleaving ribozyme may contain a ribozyme folding structure of hepatitis delta virus (HDV).

14. The system according to claim 13, wherein the HDV-derived ribozyme contains a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 102 to 127 and 129 to 154.

15. The system according to any one of claims 2 to 12, wherein the ribozyme sequence of the GIC: 5'-side module contains a ribozyme derived from the 5'-region of at least one non-long terminal repeat retroelement.

16. The system according to claim 15, wherein the ribozyme contains a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 64 to 65, 67, 75 to 76, 86, 89 to 101, and 128.

17. The system according to any one of claims 2 to 16, wherein the folding motif sequence of the GIC: 5'-side module contains at least one autonomous folding RNA sequence motif, and the autonomous folding RNA sequence motif may contain at least one hairpin motif, at least one stem-loop motif, at least one paired stem 4 motif, or any combination thereof.

18. The system according to claim 17, wherein the folding motif sequence contains SEQ ID NO: 278 or 279, or contains a sequence having at least 90% identity with SEQ ID NO: 278 or 279.

19. The system according to any one of claims 2 to 18, wherein the GIC: 5'-side module comprises a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 60 to 154.

20. The system according to any one of claims 2 to 19, wherein the reverse transcriptase recognition sequence of the GIC: 3'-side module comprises at least one sequence that interacts with at least one reverse transcriptase.

21. The system according to claim 20, wherein the reverse transcriptase recognition sequence of the GIC: 3'-side module is derived from the 3'-region of a natural retroelement.

22. The system according to claim 20 or 21, wherein at least one reverse transcriptase recognition sequence of the GIC: 3'-side module comprises a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 200 to 224.

23. The system according to any one of claims 2 to 22, wherein the rRNA sequence of the GIC: 3'-side module comprises 1 to 30 nt of rRNA.

24. The system according to claim 23, wherein the rRNA sequence is selected from the group consisting of SEQ ID NOs: 280 to 289 and sequences comprising one or two nucleotide substitutions in SEQ ID NOs: 280 to 289.

25. The system according to any one of claims 2 to 24, wherein the A-tract sequence of the GIC: 3'-side module comprises 1 to 50 adenine bases.

26. The system according to any one of claims 2 to 25, wherein the GIC: 3'-side module comprises a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 300 to 329 and any combination thereof.

27. The system according to any one of claims 2 to 25, wherein the GIC: 3'-side module comprises a 3'-UTR sequence derived from ZoA1, TaGu, GeFo or TiGu.

28. The system according to claim 27, wherein the 3'-UTR sequence comprises a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 202 to 205 and SEQ ID NOs: 222 to 224.

29. The system according to any one of claims 2 to 28, wherein the at least one transgene sequence comprises or encodes at least one target sequence for insertion into the genome of a subject.

30. The system according to claim 29, wherein the introduced gene sequence comprises at least one of mRNA, microRNA, siRNA, rRNA, tRNA, long non-coding RNA, small cytoplasmic RNA, small nuclear RNA, small nucleolar RNA, Cajal body small RNA, circular RNA, regulatory RNA, peptide, polypeptide, protein, inhibitory protein, and / or a sequence that controls the expression of at least one introduced gene, or encodes these.

31. The system according to claim 30, wherein the introduced gene encodes a protein selected from hTERT, hPAH, human factor VIII, mutant human factor VIII with various lengths of B domain, and factor IX.

32. The system according to any one of claims 2 to 31, wherein the promoter sequence of the introduced gene comprises at least one sequence that promotes the expression of the introduced gene in the genome of the subject.

33. The system according to any one of claims 2 to 32, wherein the 5' untranslated sequence of the introduced gene comprises at least one 5' untranslated region of the mRNA of the introduced gene.

34. The system according to any one of claims 2 to 33, wherein the 3' untranslated sequence of the introduced gene comprises at least one 3' untranslated region of the mRNA of the introduced gene.

35. The system according to any one of claims 2 to 34, wherein the polyadenylation signal sequence of the introduced gene comprises at least one introduced gene polyadenylation signal.

36. The system according to any one of claims 2 to 35, wherein the non-coding RNA (ncRNA) processing sequence of the introduced gene comprises at least one termination signal, at least one 3' processing signal, and any combination thereof of at least one ncRNA expressed from the introduced gene.

37. The system according to any one of claims 2 to 36, wherein the at least one GIC:payload module comprises at least one sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 411-422, SEQ ID NOs: 499-536, and any combination thereof, or encodes this sequence.

38. At least one of the at least one GIC: 5'-side module and the at least one GIC: 3'-side module contains at least one sequence derived from a species different from the species from which the non-long terminal repeat retroelement from which the other is derived is derived, or encodes this sequence. The system according to any one of claims 1 to 37.

39. The system according to any one of claims 1 to 38, wherein the at least one gene insertion construct contains at least one structure shown in FIGS. 6 to 9 and any combination thereof, or encodes these structures.

40. The system according to any one of claims 1 to 39, comprising two different gene insertion constructs, wherein the ORFs of the transgenes of the GIC: payload modules contained in these gene insertion constructs are different from each other.

41. The system according to claim 40, wherein the two different gene insertion constructs are present on the same RNA template.

42. The system according to claim 40, wherein the two different gene insertion constructs are present on different RNA templates.

43. (i) at least one reverse transcriptase construct; (ii) at least one gene insertion construct comprising the reverse transcriptase construct contains at least one sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 1 to 57, or is encoded by this sequence; the at least one gene insertion construct a GIC: 5'-side module containing a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 60 to 154; an rRNA sequence which is any component containing a sequence selected from the group consisting of SEQ ID NOs: 250 to 276, or a sequence having one, two or three nucleotide changes compared to a sequence selected from the group consisting of SEQ ID NOs: 250 to 276; a GIC: payload module containing at least one transgene sequence; a GIC: 3'-side module containing a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 300 to 329; a reverse transcriptase recognition sequence of the GIC: 3'-side module containing a sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 200 to 224; A GIC: rRNA sequence of the 3'-side module selected from the group consisting of SEQ ID NOs: 280 to 289 and sequences containing one or two nucleotide substitutions in SEQ ID NOs: 280 to 289, and a GIC: A-tract sequence of the 3'-side module containing 1 to 100 adenine bases comprising The system according to any one of claims 1 to 42.

44. The system according to claim 43, wherein the GIC: payload module comprises at least one sequence having at least 90% identity with a sequence selected from the group consisting of SEQ ID NOs: 411 to 422 and 499 to 536.

45. The system according to any one of claims 1 to 44, wherein at least one of the at least one reverse transcriptase construct and the at least one gene insertion construct comprises or encodes at least one sequence derived from a species different from the species from which the retroelement from which the other is derived.

46. The system according to any one of claims 1 to 45, wherein the RNA of the reverse transcriptase construct and / or the RNA of the gene insertion construct comprises at least one modified uracil, or 100% of the uracils constituting the RNA of the reverse transcriptase construct and / or the RNA of the gene insertion construct are modified uracils.

47. wherein the modified uracil is selected from the group consisting of 5-methyluridine, 5-methoxyuridine, pseudouridine, N 1 -methylpseudouridine and / or 2-thiouridine, the system according to claim 46.

48. A method of inserting at least one transgene into the genome of a subject, the method comprising administering to the subject an effective amount of at least one of the gene insertion systems (GIS) according to any one of claims 1 to 47.

49. The method according to claim 48, wherein the transgene is inserted into one or more target sites of the genome of the subject, and the one or more target sites may comprise at least one safe harbor site.

50. The method according to claim 49, wherein the at least one safe harbor site, which is an arbitrary component, comprises at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may comprise at least one 28S rDNA sequence.

51. The method according to any one of claims 48 to 50, comprising administering at least one of the gene insertion systems formulated with at least one delivery agent.

52. The method according to claim 51, wherein the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may comprise at least one lipid nanoparticle.

53. The method according to any one of claims 48 to 52, wherein the transgene is inserted with a target site specificity exceeding 90%.

54. The method according to claim 53, wherein the RNA of the reverse transcriptase construct encodes a reverse transcriptase derived from ZoA1, TaGu or TiGU, or comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 29 or SEQ ID NO:

25.

55. The method according to any one of claims 48 to 54, wherein the transgene is expressed at the target site for 3 months or more.

56. A pharmaceutical composition comprising at least one of the gene insertion systems according to claims 1 to 47, and optionally comprising at least one additive, at least one delivery agent, at least one adjuvant, and any combination thereof.

57. A method for treating a therapeutic indication in a subject in need of treatment of the therapeutic indication, the method comprising administering an effective amount of at least one of the gene insertion systems according to claims 1 to 47 or the pharmaceutical composition according to claim 56, and optionally comprising at least one of the methods according to claims 48 to 52.

58. The method according to claim 57, wherein the therapeutic indication is caused by a deletion of telomerase activity.

59. The method according to claim 57 or 58, wherein the at least one gene insertion system comprises at least one TERT transgene.

60. A kit for producing a gene insertion system, comprising the gene insertion system according to claims 1 to 47, optionally comprising the pharmaceutical composition according to claim 56, and further optionally comprising a buffer, a DNA plasmid, or a protocol for producing the gene insertion system or the pharmaceutical composition.

61. A method comprising de novo design of a 5'-side module that mobilizes a host mechanism for introducing a nick into the second strand to synthesize the second strand.

62. The method according to claim 61, wherein the insertion efficiency is increased by de novo design of the 5'-side module configured to (a) contain rRNA of a predetermined length at a predetermined position, (b) promote folding of ribozyme (RZ), and / or (c) mobilize host cell machinery.

63. A method for inserting at least one transgene into the genome of a cell, the method comprising the step of contacting at least one of the gene insertion systems (GIS) according to any one of claims 1 to 47 with the cell.

64. The method according to claim 63, wherein the transgene is inserted into one or more target sites of the genome of interest, and the one or more target sites may contain at least one safe harbor site.

65. The method according to claim 64, wherein the at least one safe harbor site, which is an arbitrary component, contains at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may contain at least one 28S rDNA sequence.

66. The method according to any one of claims 63 to 65, comprising the step of administering at least one of the gene insertion systems formulated with at least one delivery agent.

67. The method according to claim 66, wherein the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may contain at least one lipid nanoparticle.

68. The method according to any one of claims 63 to 67, wherein the transgene is inserted with a target site specificity of more than 90%.

69. The RNA of the reverse transcriptase construct encodes a reverse transcriptase derived from Zoanthus sansibaricus (ZoA1), Tagetes patula (TaGu) or Tagetes erecta (TiGU), or contains an amino acid sequence having at least 90% identity with SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 29 or SEQ ID NO:

25. The method according to claim 68.

70. The method according to any one of claims 63 to 69, wherein the transgene is expressed at the target site for 3 months or more.

71. The method according to any one of claims 63 to 70, wherein the molar ratio of the reverse transcriptase construct to the gene insertion construct is about 10:1 to 1:

20.

72. The method according to any one of claims 63 to 71, which is an in vitro method, an ex vivo method or an in vivo method. **Claim 73** The method according to any one of claims 63 to 72, wherein the cell is selected from the group consisting of a primary cell, a transformed cell, an epithelial cell, a fibroblast, a human cell, a monkey cell and a mouse cell. **Claim 74** The method according to any one of claims 63 to 73, wherein the cell is an allogeneic cell or an autologous cell. **Claim 75** The method according to claim 74, wherein the autologous cell is a cell with a matched HLA.