Multi-element system for site-specific genome modification
A genome editing system using non-LTR retrotransposons enables site-specific transgene insertion in eukaryotic cells via TPRT, addressing integration challenges and achieving high efficiency and prolonged expression.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RGT UNIV OF CALIFORNIA
- Filing Date
- 2023-05-02
- Publication Date
- 2026-04-28
AI Technical Summary
Current methods for introducing genetic material into cells face challenges such as immune responses and non-site-directed integration, particularly in higher eukaryotes, leading to potential genome disruption and inefficiencies in transgene insertion.
A genome editing system using non-long-chain terminal repeat (non-LTR) retrotransposons, comprising reverse transcriptase constructs and gene insertion constructs, facilitates site-specific transgene insertion through target-primed reverse transcription (TPRT) into eukaryotic cells, utilizing components derived from various organisms to enhance biostability and specificity.
The system achieves high specificity and efficiency in inserting transgenes into target genomes, including safe harbor sites, with insertion efficiencies exceeding 90% and prolonged expression up to three months.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] true Introducing transgenes into the genomes of nuclei allows for the improvement, modification, and / or alteration of gene expression, while also being useful in treating and alleviating disease symptoms. If transgenes can be successfully inserted, it is possible to achieve recovery of loss-of-function mutations, inhibition of gain-of-function mutations, exogenous regulation of RNA and / or protein expression, introduction of isoform specificity, expression of recombinant genes and recombinant proteins, and other useful results.
[0002] However, current methods for introducing genetic material into cells and inserting it into the genome still face significant hurdles to overcome. For example, methods for delivering DNA to target cells require the DNA to pass through the cytoplasm, which often induces destructive or harmful immune responses. Furthermore, site-directed integration methods, such as homologous recombination (HR), can introduce mutagenic double-strand DNA breaks, potentially disrupting the target genome or epigenome at the integration site. In higher eukaryotes, particularly in post-mitotic cells, DNA integration is often non-site-directed. This is because homologous recombination is suppressed for most of the cell cycle, and non-homologous end joining is preferred.
[0003] A method that can flexibly accommodate the length of the DNA to be introduced and effectively and site-specifically insert a transgene into the genome of a living cell without introducing DNA into the cytoplasm would make a significant contribution to human biology, animal biology, microbial biology, and plant biology, and is expected to lead to strong progress in research and clinical applications.
[0004] One possible method involves introducing the transgene sequence as RNA and using it as a template for complementary DNA (cDNA) synthesis by reverse transcriptase (RT). However, currently, no molecular signal has been identified that can induce RNA introduced into mammalian cells to replicate as a template for inserting the transgene into the genome.
[0005] In addressing these challenges, a group of genes known as non-long-chain terminal repeats (non-LTR) retroelements (REs), or synonymous non-LTR retrotransposons, may offer a promising solution. These genes can autoamplify within the host genome and exert their function by expressing non-LTR retrotransposon reverse transcriptase (RT) proteins. This RT protein binds to its retroelement transcript RNA, using it as a template, and synthesizes cDNA by using the nick introduced into the genomic DNA (catalyzed by the endonuclease (EN) domain of this RT protein) as a primer for initiating cDNA synthesis (RT primer elongation). This process is known as target-primed reverse transcription (TPRT), and it results in the addition of a copy of the double-stranded DNA retroelement within the genome.
[0006] WO2022 / 155055 describes a two-component system for site-specific transgene insertion into safe harbor regions of the human genome. These two components are a non-LTR retroelement reverse transcriptase (RT) and template RNA. This RT has been recombined to allow full-length transgene insertion, replacing the tendency of natural retroelements to insert and cleave at the 5' end. The template RNA is adapted to this RT. The synthesis mechanism of the first DNA strand is target-primed reverse transcription (TPRT), induced by the 3' module of the template RNA, and this reverse transcription is enhanced by a non-natural 3' tail, which is part of the 3' module. The 5' module of the template RNA confers biostability to the template RNA, increases its bioavailability when bound to the RT protein, and induces the synthesis of the second strand.
[0007] This disclosure provides compositions and methods for inserting and expressing transgenes into the genomes of eukaryotic cells, particularly human cells, by constructing biopolymer constructs from a portion of retroelement sequences. [Overview of the Initiative] [Means for solving the problem]
[0008] The present invention provides compositions, methods, and / or uses of proteins and nucleotides, as well as modified proteins and modified polynucleotides, for inserting a transgene into a target genome by target primed reverse transcription (TPRT) using components derived from non-long-chain terminal repeat (non-LTR) retrotransposons.
[0009] The present invention is a genome editing system, (i) at least one reverse transcriptase construct (RTC) comprising a polynucleotide encoding a polypeptide having enzymatic activity for reverse transcription of a polynucleotide template, (ii) at least one gene insertion construct (GIC) comprising at least one polynucleotide template suitable for reverse transcription by the polypeptide encoded by the at least one RTC. We provide a system that includes this.
[0010] In some embodiments, the genome editing system is (i) at least one reverse transcriptase construct (RTC), (ii) at least one gene insertion construct (GIC) Includes, The RTC comprises at least one reverse transcriptase module (RTC:RT module), the RTC:RT module comprises mRNA encoding reverse transcriptase (RT), at least one 5' module (RTC:5' module), and / or at least one 3' module (RTC:3' module), The GIC comprises at least one RNA template suitable for reverse transcription by a polypeptide encoded by the at least one RTC, the GIC comprising at least one GIC:5' side module, at least one GIC:payload module, and / or at least one GIC:3' side module.
[0011] In some embodiments, the RT module includes mRNA encoding a reverse transcriptase derived from an organism selected from birds, arthropods, fish, urochordates, and other animals (including mammals and humans).
[0012] In some embodiments, the genome editing system is i) The RTC 5' side module containing the 5' untranslated region (5'UTR), Kozak sequence, non-natural translation start codon, and / or 5' cap; ii) An RT module containing mRNA encoding reverse transcriptase derived from an organism selected from the group consisting of the White-throated Bunting (Zonotrichia albicollis) (ZoAl), the Zebra Finch (Taeniopygia guttata) (TaGu), the White-throated Swann (Tinamus guttatus) (TiGu), the Japanese Killifish (Oryzias latipes) (OrLa), and the Confused Flour Beetle (Tribolium castaneum) (Planet B) (TriCasB); iii) The RTC 3' side module containing the reverse transcriptase translation termination codon, the 3' untranslated region (3'UTR), and the poly-A tail; iv) GIC:5' side module containing sequences derived from the 5' region of a natural retroelement, rRNA sequences, ribozyme sequences, folding motif sequences, and / or RNA polymerase terminator sequences; (v) A GIC:payload module comprising at least one transgene ORF or non-coding RNA (ncRNA) sequence, a promoter sequence of the transgene, an internal ribosome entry site (IRES), a 5' untranslated sequence of the transgene, a 3' untranslated sequence of the transgene, a polyadenylation signal sequence of the transgene, and / or an ncRNA processing sequence of the transgene; and (iv) GIC:3' side module containing reverse transcriptase recognition sequence, rRNA sequence, and / or A tract sequence Includes.
[0013] In some embodiments, the at least one reverse transcriptase construct comprises at least one biopolymer, the biopolymer comprising at least one nucleic acid, at least one amino acid, and any combination thereof. In some embodiments, the polynucleotide of the reverse transcriptase construct described in (i) comprises mRNA encoding reverse transcriptase. In some embodiments, the polynucleotide template of the gene insertion construct described in (ii) comprises RNA. In some embodiments, the polynucleotide of the reverse transcriptase construct described in (i) comprises mRNA encoding reverse transcriptase, and the polynucleotide template of the gene insertion construct described in (ii) comprises another (different) RNA. In some embodiments, the gene insertion construct comprises an RNA template different from the mRNA encoding reverse transcriptase described in (i).
[0014] In some embodiments, the at least one reverse transcriptase construct may include at least one reverse transcriptase open reading frame (ORF) module (RTC:RT module), at least one 5' untranslated region (UTR) module (RTC:5' side module), at least one 3' UTR module (RTC:3' side module), and any combination thereof.
[0015] In some embodiments, the at least one reverse transcriptase module comprises at least one reverse transcriptase or encodes at least one reverse transcriptase.
[0016] In some embodiments, the at least one reverse transcriptase module comprises or encodes at least one reverse transcriptase derived from a non-long terminal repeat (non-LTR) type retroelement.
[0017] In some embodiments, the at least one reverse transcriptase module comprises or encodes a non-native translation initiation codon.
[0018] In some embodiments, the at least one reverse transcriptase comprises at least one DNA binding domain, at least one RNA binding domain, at least one cDNA synthesis domain, at least one endonuclease domain, and any combination thereof.
[0019] In some embodiments, at least one of the at least one reverse transcriptase domain, the at least one target DNA binding domain, the at least one template RNA binding domain, and the at least one endonuclease domain, and any combination thereof, is derived from a species different from the species from which at least one of the remaining domains is derived.
[0020] In some embodiments, at least one 5'-side module of the reverse transcriptase construct comprises or encodes at least one RNA polymerase promoter, at least one 5' untranslated region (5'-UTR), at least one Kozak sequence, at least one 5' cap, and any combination thereof.
[0021] In some embodiments, at least one 3' end module of the reverse transcriptase construct includes or encodes at least one reverse transcriptase translation termination codon, at least one 3' untranslated region (3'UTR), at least one polyA tract and / or polyA tail, and any combination thereof.
[0022] In some embodiments, the at least one reverse transcriptase module includes or codes for at least one structure shown in Figures 2-5 or any combination thereof.
[0023] In some embodiments, the at least one reverse transcriptase construct comprises at least one of SEQ ID NOs: 1-57, codes for at least one of SEQ ID NOs: 1-57, or is coded by at least one of SEQ ID NOs: 1-57. In some embodiments, the at least one reverse transcriptase construct comprises mRNA encoding a reverse transcriptase protein derived from a species selected from the group consisting of TriCasB, NaViB, OrLa, ZoAl, TiGu, TaGu, GeFo, DroSi, BoMo, DrMerc, DrMe, GaAc, PuPu, AdVa, HyMaA, CiIn, LiPo, TriCan, LeCo, and any combination thereof.
[0024] In some embodiments, the at least one gene insertion construct includes or encodes at least one nucleic acid biomolecule. In some embodiments, the gene insertion construct includes template RNA.
[0025] In some embodiments, the at least one gene insertion construct includes or encodes at least one GIC:5' side module, which is an optional component, at least one GIC:payload module, at least one GIC:3' side module, which is an optional component, and any combination thereof.
[0026] In some embodiments, the at least one GIC:5' side module includes or encodes at least one sequence derived from the 5' region of a native retroelement, and may include, or encode, at least one rRNA sequence, at least one ribozyme (RZ) sequence, at least one folding motif sequence, or any combination thereof.
[0027] In some embodiments, at least one rRNA sequence, which is an optional component of the GIC:5' side module, contains or encodes a 1-30 nt rRNA of the subject.
[0028] In some embodiments, at least one ribozyme sequence, which is an optional component of the GIC:5' side module, comprises or encodes at least one autocleaved ribozyme, which may comprise a ribozyme of hepatitis delta virus (HDV).
[0029] In some embodiments, at least one ribozyme sequence, which is an optional component of the GIC:5' side module, comprises or encodes a ribozyme derived from the 5' region of at least one non-long-chain terminal repeat retroelement. In some embodiments, at least one folding motif sequence, which is an optional component of the GIC:5' side module, comprises or encodes at least one autonomous folding RNA sequence motif, which may comprise at least one hairpin motif, at least one stem-loop motif, at least one paired stem motif, or any combination thereof within the RZ.
[0030] In some embodiments, the GIC:5' side module is sequence number 60~ 153 At least one of the above, or sequence number 60~ 153The GIC:5' side module contains or encodes a sequence having at least 90% identity with at least one of the above (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity). In some embodiments, the GIC:5' side module contains a sequence derived from a species selected from the group consisting of OrLa, TriCasB, TriCasA, ZoAl, TiGu, DroSi, LeCo, CiIn, FoRa, TriCan, HDV-28, HDV-24, HDV-21, HDV-13, HDV-36, and any combination thereof.
[0031] In some embodiments, the at least one GIC:3' side module may include or encode at least one reverse transcriptase recognition sequence, and may also include, or encode, at least one rRNA sequence, at least one A tract sequence, or any combination thereof.
[0032] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module includes or encodes at least one sequence that interacts with at least one reverse transcriptase. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is sequence number 154 ~ 178 Includes sequences selected from the group consisting of [the specified elements].
[0033] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is derived from the 3' region of a native retroelement.
[0034] In some embodiments, at least one rRNA sequence, which is an optional component of the GIC:3' side module, contains or encodes 1 to 30 nt of rRNA.
[0035] In some embodiments, at least one A tract sequence, which is an optional component of the GIC:3' side module, contains or encodes a sequence consisting of approximately 1 to 50 adenine bases.
[0036] In some embodiments, the at least one GIC:3' side module is, 154 ~ 178 at least one or array number 225 ~ 253 It includes at least one of the following, or codes for any one of them. In some embodiments, the GIC:3' side module includes a sequence derived from a species selected from the group consisting of OrLa, TriCasB, TaGu, GeFo, ZoAl, NaViB, DroSi, PuPu, LiPo, BoMo, GaAc, LeCo, CiIn, DrMe, DrNa, DrMer, TriCan, AdVa, HyMaA, and any combination thereof.
[0037] In some embodiments, the at least one GIC:payload module may include or encode at least one transgene ORF sequence, and may also include, and encode, at least one promoter sequence of the transgene, at least one 5' untranslated sequence of the transgene, at least one 3' untranslated sequence of the transgene, at least one polyadenylation signal sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, at least one ncRNA processing sequence, and / or other 3' end processing sequences or stabilization signals, or any combination thereof.
[0038] In some embodiments, the at least one transgene sequence includes or encodes at least one target sequence for insertion into the genome of interest.
[0039] In some embodiments, at least one promoter sequence of the transgene includes or encodes at least one sequence that promotes the expression of the transgene in the target genome.
[0040] In some embodiments, the at least one GIC:payload module includes at least one 5' untranslated sequence of the transgene, which includes or encodes at least one 5' untranslated region of the mRNA of the transgene.
[0041] In some embodiments, at least one 3' untranslated sequence of the transgene includes or encodes at least one 3' untranslated region of the mRNA of the transgene.
[0042] In some embodiments, at least one polyadenylation signal sequence of the transgene includes or encodes at least one polyadenylation signal of the transgene.
[0043] In some embodiments, at least one non-coding RNA (ncRNA) processing sequence and / or other 3' end processing sequence or stabilization signal of the transgene includes or encodes at least one stop signal, at least one 3' processing signal, and any combination thereof of at least one ncRNA expressed from the transgene.
[0044] In some embodiments, the at least one GIC:payload module is, 284 ~ 295 and Scalar 296 ~ 332 Furthermore, it includes or encodes a sequence that has at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with at least one of any combination thereof.
[0045] In some embodiments, at least one of the at least one GIC:5' side module and the at least one GIC:3' side module contains or encodes at least one sequence derived from a different species than the species from which the non-long-chain terminal repeat retroelement from the other originates.
[0046] In some embodiments, the at least one gene insertion construct comprises or codes for at least one structure shown in the drawings attached herein, for example, at least one structure shown in Figures 6-9 and any combination thereof.
[0047] In some embodiments, the genome editing system is (i) at least one reverse transcriptase construct, (ii) at least one gene insertion construct Includes, The at least one reverse transcriptase construct includes, codes for, or is coded by, a sequence selected from the group consisting of SEQ ID NOs: 1 to 57, having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity). The aforementioned at least one gene insertion construct corresponds to Sequence ID No. 60~ 153 , 179 ~ 205 , 206 ~ 207 , 208 ~ 217 , 225 ~ 253 , 275 ~ 278 , 279 ~ 281 , 284 ~ 295 and 296 ~ 332It includes at least one sequence that has at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with a sequence selected from the group consisting of the above. In some embodiments, the mRNA sequence transfected to generate a reverse transcriptase protein is split from a plasmid and expressed, encoding the amino acid sequences of multiple proteins.
[0048] In some embodiments, the genome editing system is (i) at least one reverse transcriptase construct, (ii) at least one gene insertion construct Includes, The at least one reverse transcriptase construct includes, or is encoded by, at least one sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) with a sequence selected from the group consisting of SEQ ID NOs: 1 to 57. The aforementioned at least one gene insertion construct is Sequence ID 60~ 153 A GIC:5' side module containing a sequence selected from the group consisting of and having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), Scalar 179 ~ 205 A sequence or sequence number selected from the group consisting of the following: 179 ~ 205 An rRNA sequence is any component that includes a sequence having one, two, or three nucleotide changes compared to a sequence selected from the group consisting of the following: GIC: Payload module containing at least one transgene sequence, Scalar 225 ~ 253A GIC:3' side module containing a sequence selected from the group consisting of and having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), Scalar 154 ~ 178 A reverse transcriptase recognition sequence of the GIC:3' side module, comprising a sequence selected from the group consisting of the following, and a sequence having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), Scalar 208 ~ 217 , and sequence number 208 ~ 217 The rRNA sequence of the GIC:3' side module, selected from a group consisting of sequences containing one, two, or three nucleotide substitutions, A tract sequence of the GIC:3' side module containing 1 to 100 adenine bases Includes.
[0049] In some embodiments, the 5'UTR of the RTC 5' side module includes a sequence having at least 90% identity with sequence number 58 (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity).
[0050] In some embodiments, the 3'UTR of the RTC 3' side module includes a sequence having at least 90% identity with sequence number 59 (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity).
[0051] In some embodiments, the genome editing system includes at least one of the gene insertion constructs described herein, or a synthesis construct (GIC) for a gene insertion construct that encodes it.
[0052] In some embodiments, at least one of the at least one reverse transcriptase construct and the at least one gene insertion construct includes or encodes at least one sequence derived from a different species than the species from which the retroelement from which the other originates.
[0053] In some embodiments, the genome editing system comprises (i) at least one reverse transcriptase construct described herein and (ii) at least one combination of gene insertion constructs described herein.
[0054] Furthermore, the present invention provides a method for inserting at least one transgene into a target genome, comprising the step of administering at least one of the gene insertion systems (GIS) of the present disclosure to the target in an effective amount.
[0055] In some embodiments, the transgene is inserted into one or more target sites in the target genome, and these one or more target sites may include at least one safe harbor site.
[0056] In some embodiments, the at least one safe harbor site, which is an optional component, comprises at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may comprise at least one 28S rDNA sequence.
[0057] In some embodiments, the at least one method includes the step of administering at least one of the gene insertion systems formulated with at least one delivery agent.
[0058] In some embodiments, the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may include at least one lipid nanoparticle.
[0059] Furthermore, the present invention provides a pharmaceutical composition comprising at least one of the gene insertion systems described in the claims, and which may also comprise at least one of the following: at least one additive, at least one delivery agent, at least one auxiliary agent, and any combination thereof.
[0060] Furthermore, the present invention provides a method for treating a therapeutic indication in a subject requiring treatment of the therapeutic indication, comprising the step of administering at least one of the gene insertion systems of the present disclosure or a pharmaceutical composition of the present disclosure in an effective amount to the subject.
[0061] In some embodiments, the therapeutic indications are caused by a deficiency in telomerase activity.
[0062] In some embodiments, the at least one gene insertion system includes at least one TERT-transformed gene.
[0063] Furthermore, a kit for constructing the gene insertion system of the Disclosure is provided. In some embodiments, the kit comprises the pharmaceutical composition of the Disclosure. In some embodiments, the kit may further comprise a buffer, a DNA plasmid, or a protocol for constructing the gene insertion system or pharmaceutical composition.
[0064] Furthermore, the present invention provides a method including a de novo design of a 5' end module that mobilizes a host mechanism to synthesize the second strand by introducing a nick into the second strand. In some embodiments, the insertion efficiency is increased by the de novo design of the 5' end module, which is configured to (a) contain rRNA of a predetermined length (as described herein) at a predetermined position, (b) facilitate ribozyme (RZ) folding, and / or (c) mobilize a host cellular mechanism.
[0065] In another embodiment, the present disclosure provides a method for inserting at least one transgene into the genome of a cell, comprising the step of bringing at least one gene insertion system (GIS) of the present disclosure into contact with the cell.
[0066] In some embodiments, the transgene is inserted into one or more target sites in the target genome, and these one or more target sites may include at least one safe harbor site. In some embodiments, the at least one safe harbor site, which is an optional component, includes at least one ribosomal DNA (rDNA) sequence, and this at least one ribosomal DNA sequence may include at least one 28S rDNA sequence.
[0067] In some embodiments, the method includes administering at least one of the gene insertion systems formulated with at least one delivery agent. In some embodiments, the at least one delivery agent is at least one nanoparticle, which may include at least one lipid nanoparticle.
[0068] In some embodiments, the transgene is inserted with on-target site specificity exceeding 90% (e.g., site specificity exceeding 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%).
[0069] In some embodiments, the reverse transcriptase construct includes RNA encoding a reverse transcriptase derived from a white-throated sparrow (ZoA1), a zebra finch (TaGu), or a white-throated swan (TiGU), or includes an amino acid sequence having at least 90% identity with SEQ ID NO: 18, SEQ ID NO: 20, SEQ ID NO: 27, SEQ ID NO: 29, or SEQ ID NO: 25.
[0070] In some embodiments, the transgene is expressed at the target site for a period of three months or longer.
[0071] In some embodiments, the cells and the gene insertion system are brought into contact, and the molar ratio of the reverse transcriptase construct to the gene insertion construct is approximately 10:1 to 1:20.
[0072] In some embodiments, the method is an in vitro method, an ex vivo method, or an in vivo method.
[0073] In some embodiments, the cells are selected from the group consisting of primary cells, transformed cells, epithelial cells, fibroblasts, human cells, monkey cells, and mouse cells.
[0074] In some embodiments, the cells are allogeneic cells or autologous cells. In some embodiments, the autologous cells are HLA-matched cells.
[0075] The present invention encompasses any combination of the specific embodiments described herein, as described in detail. [Brief explanation of the drawing]
[0076] [Figure 1] This diagram shows an example of a target genome containing the target insertion site and the natural retroelement. The enlarged view (below) shows the structure of an exemplary component of the natural R2 retroelement.
[0077] [Figure 2] This is a diagram showing the structure of an example of a reverse transcriptase construct (RTC).
[0078] [Figure 3] This diagram shows exemplary domains of the reverse transcriptase (RT) protein of the present invention.
[0079] [Figure 4]This figure illustrates exemplary organisms as sources of reverse transcriptase (RT) protein domains, including DNA-binding domains (DB), RNA-binding domains (RB), reverse transcriptase (RT) domains, and endonuclease (EN) domains. Furthermore, a diagram showing a small set of example combinations of each RT protein domain is provided. The identification of each domain is defined for each organism, and wild-type RT is indicated by the following abbreviations. A1 is a white-throated sparrow (Zonotrichia albicollis), A2 is a zebra finch (Taeniopygia guttata), A3 is a white-throated sandpiper (Tinamus guttatus), A4 is a Galapagos finch (Geospiza fortis), B1 is a northern stickleback (Pungitis pungitis), B2 is a medaka (Oryzias latipes), B3 is a three-spined stickleback (Gasterosteus aculeatus), C1 is a parasitic wasp (Nasonia vitripennis), C2 is a yellow fruit fly (Drosophila melanogaster), C3 is a confused flour beetle (Tribolium castaneum) (B strain), C4 is a silkworm (Bombyx mori), and C5 is a common fruit fly (Drosophila C6 is Drosophila mercatorum, D1 is Lepidurus couseii, D2 is Triops cancriformis, E1 is Hydra magnipapillata, E2 is Limulus polyphemus, E3 is Adineta vaga, and E4 is Ciona intestinalis.
[0080] [Figure 5]This is a series of diagrams illustrating an exemplary set of reverse transcriptase constructs (RTCs) of the present invention. Each reverse transcriptase construct (RTC) includes a sequence containing a reverse transcriptase protein (RT) containing the translation start codon (M) of the reverse transcriptase, or a sequence encoding it. The reverse transcriptase construct (RTC) may also include a 5' untranslated sequence (5'-UTR), a translation stop codon (SC), and / or a 3' untranslated sequence (3'-UTR).
[0081] [Figure 6] This diagram shows the structure of an example of a gene insertion construct (center). The enlarged view shows the structures of an example of the 5' module (bottom left), 3' module (bottom right), and payload module (top).
[0082] [Figure 7]This figure illustrates exemplary organisms that serve as sources for the components of the 5' module (5'M) and 3' module (3'M) of a gene insertion construct (GIC), and the components of the RT module (RT) of a reverse transcriptase construct (RTC). Furthermore, the figure shows several sets of hypothetical gene insertion constructs with possible combinations of the 5' and 3' modules positioned to flank the payload module, along with their corresponding reverse transcriptase constructs (paired RTs). Module identification is defined by organismal name, with wild-type retroelements and / or reverse transcriptases listed below. A1 is a white-throated sparrow (Zonotrichia albicollis), A2 is a zebra finch (Taeniopygia guttata), A3 is a white-throated sandpiper (Tinamus guttatus), A4 is a Galapagos finch (Geospiza fortis), B1 is a northern stickleback (Pungitis pungitis), B2 is a medaka (Oryzias latipes), B3 is a threadfin stickleback (Gasterosteus aculeatus), C1 is a parasitic wasp (Nasonia vitripennis), C2 is a yellow fruit fly (Drosophila melanogaster), C3 is a confused flour beetle (Tribolium castaneum), C4 is a silkworm (Bombyx mori), and C5 is a common fruit fly (Drosophila C6 is Drosophila mercatorum, D1 is Lepidurus couseii, D2 is Triops cancriformis, E1 is Hydra magnipapillata, E2 is Limulus polyphemus, E3 is Adineta vaga, and E4 is Ciona intestinalis.
[0083] [Figure 8]This figure shows the structure of an example of a target genome after the introduction gene has been inserted using the gene insertion system (GIS) of the present invention.
[0084] [Figure 9] This figure shows the structure of an example of a synthetic gene insertion construct.
[0085] [Figure 10] This image shows radioactive DNA synthesis products separated on a denatured PAGE gel. The solid black border indicates the gel region corresponding to the expected product length. The lane numbers correspond to the various RT proteins tested, as detailed in Table 3 of Example 10. The reaction in lane 1 includes a negative control purified from cells that do not express the RT protein.
[0086] [Figure 11] Figure 11A shows an example of an experimental design to verify the specificity of RT proteins to template RNA derived from homogeneous R2 factor 3'UTR and non-homogeneous R2 factor 3'UTR. Figure 11B shows spot blot results examining the selectivity of RT from silkworms, Drosophila melanogaster, and medaka to homogeneous and non-homogeneous 3'UTR.
[0087] [Figures 12A-12B]Figure 12A shows the results of denatured PAGE gels of TPRT reaction products. Arrows indicate the expected size for the correct TPRT product. Lane B contains the reaction product of silkworm RT, lane D contains the reaction product of Drosophila melanogaster RT, lane O contains the reaction product of medaka RT, and lane N contains the reaction product in the absence of the enzyme. Figure 12A shows the results of the TPRT reaction, showing the reaction products of various RT proteins with a template containing the 3'UTR of Drosophila melanogaster (lanes labeled "single"), and the reaction products of various RT proteins with a template containing the 3'UTR and 4nt rRNA of Drosophila melanogaster (lanes labeled "with R4"). Figure 12B shows the results of the TPRT reaction, showing the reaction products of various RT proteins with a template containing the 3'UTR of medaka (lanes labeled "single"), and the reaction products of various RT proteins with a template containing the 3'UTR and 4nt rRNA of medaka (lanes labeled "with R4").
[0088] [Figure 13] The results of modified PAGE gels for TPRT reaction products derived from silkworm moths and various templates are shown. Arrows indicate the expected size of the correct TPRT product. Circles indicate the length of the product if synthesis is initiated in the internal sequence.
[0089] [Figures 14A-14B] The results of denatured PAGE gels of reticular reaction products derived from medaka fish and TPRT using various templates are shown.
[0090] [Figure 15] The results of denatured PAGE gels of RT derived from confused flour beetle and TPRT reaction products of various templates are shown. The length of the target TPRT product is indicated by an arrow.
[0091] [Figure 16]The results of denatured PAGE gels of TPRT reaction products from white-throated mitchinid RT proteins are shown. Table 8, shown in Example 17, shows the identification names of the gene insertion constructs used in each lane shown in the photograph of the denatured PAGE gel. The solid box (top) shows the TPRT product of the expected length, the dashed box shows the precipitate recovery operation control of the expected length (middle), and the dashed box shows the radiolabeled target site oligonucleotide of the expected length (bottom).
[0092] [Figure 17] The results of a denatured PAGE gel of TPRT reaction products from zebra finch-derived RT protein are shown. Lane 1 contained a ladder for length reference. Lane 2 contained only the RT protein (without template RNA). Table 11, shown in Example 19, shows the identification names of the gene insertion constructs used in the other lanes shown in the photograph of the denatured PAGE gel. The solid box (top) shows the TPRT product of expected length, the dashed box shows the precipitate recovery operation control of expected length (middle), and the dotted-dotted box shows the radiolabeled target site oligonucleotide of expected length (bottom).
[0093] [Figures 18A-18B] Figure 18A shows the PCR amplification products of genomic DNA after insertion of transgenes templated using RT protein derived from the confused flour beetle and various templates. In Figure 18A, the expected product length is indicated by a box. All accurately inserted PCR products should be the same size. In Figure 18B, the expected product length is indicated by an arrow. The lengths of accurately inserted PCR products differ considerably between templates without a 5' module (3) and templates with a 5' module (5_3).
[0094] [Figure 19]The results of PCR amplification of genomic DNA are shown. The upper panel shows the expected amplification of the 3' junction, and the lower panel shows the expected amplification of the 5' junction. Lanes marked "L" contained ladders for length reference. Lanes 1 and 9 contained PCR products that were not transfected with plasmids expressing TriCasB-derived RT or gene insertion constructs. Lanes 2-8 contained PCR products that were transfected with the gene insertion constructs shown in Table 13 of Example 21, without transfecting with RT expression plasmids. Lanes 10-16 contained PCR products that were transfected with both the gene insertion constructs shown in Table 13 of Example 21 and the RT expression plasmids. Some of the expected lengths of the PCR products are marked with asterisks. For all images containing asterisks, please refer to the figures in the supplementary materials.
[0095] [Figure 20] The results of PCR amplification of genomic DNA are shown. Lanes A-J show the expected size of PCR products when the target 5' ligament was detected after co-transfecting the mRNA of the reverse transcriptase construct and the RNA of the gene insertion system shown in Table 16 of Example 24.
[0096] [Figure 21] The images show exemplary FACS analysis results for a population of GFP-negative cloned cells derived from the transgene (the top two panels) and a population of GFP-positive cloned cells derived from the transgene (the bottom two panels). [Modes for carrying out the invention]
[0097] I. Introduction Throughout the following detailed description and throughout this specification, unless otherwise indicated or otherwise appropriate, the terms "a" and "an" mean one or more, and the term "or" means "and / or". The examples and embodiments described herein are for illustrative purposes only. It is suggested to those skilled in the art that various improvements and modifications can be made, and these improvements and modifications are included in the essence and scope of this application and the scope of the appended claims. All publications, patents and patent applications cited herein (including the references within these documents) are incorporated herein by reference in their entirety for all purposes.
[0098] The present invention provides systems and methods for genome editing and / or gene modification, comprising inserting a transgene into a target genome. The systems of the present invention, referred to herein as “Gene Insertion System (GIS),” may comprise at least two components: (a) at least one reverse transcriptase (RT) construct (RTC) comprising or encoding at least one reverse transcriptase; and (b) at least one gene insertion construct (GIC) comprising or encoding an RNA construct used as a template for reverse transcription, and expressed separately from the RTC (i.e., a GIS comprising two components). The term “construct” herein means any biomolecule artificially designed or synthesized. This biomolecule may consist, for example, nucleic acids (e.g., DNA or RNA), amino acids, or any combination thereof. In some embodiments, both (a) and (b) are RNA constructs. In some embodiments, (a) is an amino acid construct (i.e., a protein) and (b) is an RNA construct.
[0099] Furthermore, a recombinant reverse transcriptase construct (RTC) capable of performing target-primed reverse transcription (TPRT) is provided. In this specification, the term "target-primed reverse transcription (TPRT)" means the process by which a reverse transcriptase initiates cDNA synthesis by utilizing the 3' end of available DNA at a target site as a primer.
[0100] Furthermore, the systems and methods provided herein may enable the insertion of a transgene into a sequence-specific location of the target DNA (hereinafter referred to as the “target site”), such as a safe harbor site. In this specification, the terms “safe harbor” and “safe harbor site” refer to a location in the target genome where, for example, the insertion of a heterologous sequence disrupts the target DNA sequence without adversely affecting the function of the target cell. Exemplary safe harbor sites used in the present invention are located in the portion of the target genome that encodes ribosomal RNA (rRNA). Examples of safe harbor sites include rRNA precursors transcribed by RNA polymerase I at loci referred to herein as “ribosomal DNA (rDNA) loci,” which include sequences encoding 5,8S rRNA, 18S rRNA, or 28S rRNA.
[0101] This disclosure demonstrates that it is possible to program a DNA transgene to be inserted into a safe harbor region of the genome of a cell (e.g., a human cell) by delivering RNA alone. In some embodiments, an RNA template encoding the transgene to be inserted and a messenger RNA encoding the reverse transcriptase necessary to convert the RNA template to genomic DNA are delivered to the cell. The RNA-only delivery method of the present invention is expected to further facilitate the transition to human gene therapy by utilizing RNA delivery mechanisms that are still evolving, are non-toxic, and highly efficient in targeting specific types of cells.
[0102] In some embodiments, the expression of reverse transcriptase (RT) from a plasmid is combined with the transfection of an RNA template. In some embodiments, the use of a heterologous combination of the reverse transcriptase and the 5' end module of the transgene template, which contains a native R2 retroelement sequence or a portion thereof, has the advantage of site-specific insertion of the full sequence rather than inserting a cleaved retroelement sequence. In some embodiments, the template RNA includes a 3' end module having a 3' UTR sequence of a retroelement derived from the same species as the reverse transcriptase. In some embodiments, this 3' UTR further includes a 3' polyA tract that increases the efficiency of specific insertion into the target site.
[0103] This disclosure offers the following improvements and advantages compared to systems and methods of the prior art.
[0104] The inventors were able to demonstrate the following: (i) Avian reverse transcriptase (RT) protein showed remarkable activity for the insertion of transgenes, and the transgenes were functionally expressed in more than 20% of transfected cells. Avian reverse transcriptase exhibited extremely high selectivity for the replication of template RNA containing avian 3'UTR and 3' poly(A) tract.
[0105] (ii) By combining the 3'UTR of the avian R2 retroelement with a reverse transcriptase protein from a different species, it may be more effective than the natural combination.
[0106] (iii) De novo-fabricated and optimized non-natural 5' module is even more effective, resulting in a site-specific insertion efficiency that increases by more than an order of magnitude.
[0107] (iv) Natural 5' modules from the confused flour beetle (TriCasA) (TCA, TCA5, TCARZ, etc.) are derived from the R2 retroelement of a clade completely different from that of avian reverse transcriptase proteins, and such natural 5' modules from confused flour beetles may be even more effective.
[0108] (v) After expressing reverse transcriptase in a plasmid, instead of transfecting with template RNA, the introduced gene is inserted by co-transfecting and delivering two types of RNA systems.
[0109] (vi) By transfecting with two types of RNA systems, multiple transgenes can be inserted into a single cell, enabling multiplexing of gene delivery with a single RNA administration. This method makes it possible to insert multiple therapeutic transgenes into the genome of a single cell. These multiple therapeutic transgenes include multiple transgenes that each encode a therapeutic protein, multiple transgenes that each encode different subunits of multiple therapeutic proteins, combinations of therapeutic proteins, and combinations of therapeutic RNAs.
[0110] (vii) By delivering two RNA systems, transgenes can be expressed in a wide range of cell types, including primary cell lines, non-dividing cells, and slow-dividing cells, such as mouse cells, monkey cells, and human cells.
[0111] (viii) Genome sequencing has shown that the insertion of transgenes is site-specific.
[0112] (ix) The inserted transgene expression cassette will be stably expressed for several months.
[0113] Components derived from retro elements
[0114] The reverse transcriptase construct (RTC) and / or gene insertion construct (GIC) of the present invention may originate from a portion of at least one non-long-terminant repeat (non-LTR) retroelement and / or may include components (also called "modules") that do not exist in nature. While we do not wish to be bound by any theory, Figure 1 (top) shows a target genome containing a natural retroelement 100, in this case a non-long-terminant repeat (non-LTR) retroelement. As can be seen from this figure, the target DNA 110 may contain at least one target insertion site 120, to which the natural retroelement 130 may be inserted. In the enlarged view (bottom), the structure of an example natural retroelement is examined in more detail. In this figure, the 5' region 131 of the natural retroelement is located before the translation initiation site 132. The 5' region of the natural retroelement is not typically translated into amino acid biomolecules. The 5' region of the natural retroelement may contain nucleic acid sequences that, in subsequent insertions, are recognized by the reverse transcriptase (RT) of the natural retroelement itself and / or affect the synthesis of the second strand of the natural retroelement. The translation start site 132 is the first nucleotide translated into an amino acid. The open reading frame 133 of the reverse transcriptase of the natural retroelement encodes the reverse transcriptase, which recognizes and binds to the RNA transcript of the natural retroelement itself and can perform reverse transcription using it as a template. The open reading frame of the reverse transcriptase of the natural retroelement extends to the translation termination site 134, although this open reading frame and the translation termination site are distinct. The 3' region 135 of the natural retroelement may contain nucleic acid sequences that are not typically translated into amino acid biomolecules and are recognized by the reverse transcriptase of the natural retroelement itself. The 5' region 131 and the 3' region 135 may or may not be present, and if present, they may contain sequences that replicate the surrounding target site sequences and / or are not encoded by the RNA template of the natural retroelement.
[0115] Suitable retroelements from which each component of the gene insertion system (GIS) may originate include, but are not limited to, RLE, APE, or Penelope type non-LTR retroelements. RLE type non-LTR retrotransposons may originate from any one of a number of clades, including, but are not limited to, R2, R4, CRE, Genie, HERO, and NeSL. APE type non-LTR retrotransposons may originate from any one of a number of clades, including, but are not limited to, I, R1, L1, Tx1, CR1, Rex1, Jockey, L2, Tad, RTE, RTEX, Ingi, Vingi, TRAS, SART, or any combination thereof. In some embodiments, the components of the gene insertion system may originate from retroelements inserted into rDNA, i.e., so-called R factors, for example, retroelements of the R1 or R2 clade. In some embodiments, the retroelements of the R2 clade may be specific to the insertion site of a standard R2 retroelement, or they may be derived from R8 and / or R9 retroelements derived from a higher-level R2 clade by a modification of the target sequence of a standard R2 retroelement, or they may be derived from R2NS retroelements derived by losing specificity to the target site.
[0116] Each component of a gene insertion system (GIS) may be derived from a part of a retroelement or domain from some biological species, including species that are distantly related to the target species. For example, suitable retroelements from which each component of the gene insertion system may originate include birds (e.g., white-throated sparrow (Zonotrichia albicollis), zebra finch (Taeniopygia guttata), white-throated sandpiper (Tinamus guttatus), and Galapagos finch (Geospiza fortis)), fish (e.g., northern stickleback (Pungitis pungitis), medaka (Oryzias latipes), zebrafish (Danio rerio), Oryzias melastigma, sea lamprey (Petromyzon marinus), brown trout (Salmo trutta), Atlantic salmon (Salmo salar), or three-spined stickleback (Gasterosteus aculeatus)), insects (e.g., Drosophila mercatorum, yellow fruit fly (Drosophila melanogaster), and parasitic wasp (Nasonia spp.) Examples of retroelements found in mammals such as *Vitripennis*, *Tribolium castaneum*, *Drosophila simulans*, *Apis cerana*, and *Bombyx mori*, crustaceans (e.g., *Lepidurus couesii* and *Triops cancriformis*), other invertebrates (e.g., *Limulus polyphemus*, *Hydra magnipapillata*, or *Adineta vaga*), chordates (e.g., *Ciona intestinalis*), mammals, and any combination thereof.
[0117] In some embodiments, each component of the gene insertion system may be derived from any sequence or domain disclosed herein.
[0118] II. Composition of Gene Insertion Systems (GIS) Throughout this disclosure, the system of the present invention for inserting genetic material (e.g., a transgene) into a target genome is referred to as a “Gene Insertion System (GIS)”. The Gene Insertion System of the present disclosure may consist of a plurality of biomolecular constructs, which are co-administered to insert at least one transgene via target primed reverse transcription (TPRT). These biomolecular constructs may be biomolecules consisting of amino acids, biomolecules consisting of nucleic acids, hybrid biomolecules containing both amino acids and nucleic acids, or any combination thereof. In some examples, the Gene Insertion System of the present disclosure consists of at least two biomolecules, namely at least one reverse transcriptase construct (RTC) and at least one gene insertion construct (GIC). In such examples, the reverse transcriptase construct (RTC) includes means for performing reverse transcription, for example, by including or encoding a reverse transcriptase, and the gene insertion construct (GIC) includes or encodes at least one RNA sequence which may be used as a template for synthesizing cDNA by the reverse transcriptase construct (RTC).
[0119] Since the biopolymer constructs of the present invention are themselves composed of multiple modules, the gene insertion system of this disclosure may be modified to perform desired functions by combining these modules as needed. In this specification, the term "module" means a part of a construct defined by its function (e.g., a functional domain of a protein) or a part of a construct defined by its sequence (e.g., an amino acid sequence or a nucleic acid sequence).
[0120] Reverse transcriptase construct (RTC) The gene insertion system of the present invention comprises an active reverse transcriptase protein, such as a reverse transcriptase derived from a non-LTR retroelement, or at least one reverse transcriptase construct (RTC) encoding such a protein. In this specification, the term "RTC" means a biopolymer construct comprising at least one reverse transcriptase (RT) or encoding at least one reverse transcriptase (RT). In some embodiments, the at least one reverse transcriptase construct (RTC) used in the gene insertion system of the present invention may comprise, but are not limited to, amino acid biopolymers such as polypeptides, proteins, proproteins, or any combination thereof. In some embodiments, the at least one reverse transcriptase construct (RTC) used in the gene insertion system of the present invention may comprise, but are not limited to, nucleic acid biopolymers such as RNA, DNA, or any combination thereof. In some embodiments, the at least one reverse transcriptase construct (RTC) may comprise at least one mRNA construct.
[0121] Structure of reverse transcriptase construct (RTC) The reverse transcriptase construct (RTC) of the present invention may comprise at least one RTC:reverse transcriptase module (RTC:RT module), at least one optional 5' module (RTC:5' module), at least one optional 3' module (RTC:3' module), and any combination thereof. In some examples of the RTC, the RTC:5' module and the RTC:3' module may be optional components, and one or both may be absent. In some embodiments, at least one RTC may contain a linear RNA biopolymer, or may be delivered to the target as a linear RNA biopolymer. In some embodiments, at least one RTC may contain an mRNA biopolymer, or may be delivered to the target as an mRNA biopolymer.
[0122] Referring to Figure 2, an exemplary structure of a reverse transcriptase construct (RTC) 200, which is a linear RNA biomolecule (e.g., mRNA), is shown. As shown in this figure, in an mRNA biomolecule RTC, the RTC:5' end module 210 is an optional component of the RTC and, if present within the RTC, may contain sequences that modify the immunogenicity of the RTC and / or sequences that control the expression of the RTC:RT module 220. For example, the RTC:5' end module may contain, and encode, at least one 5' cap (e.g., TriLink Clean Cap AG or m7(3'OMeG)(5')ppp(5')(2'OMeA)pG), at least one 5' untranslated region (5'-UTR), at least one Kosack sequence, at least one promoter, and any combination thereof. The start codon is a 3-nucleotide nucleic acid sequence known to initiate translation and defines the 5' end of the RTC:RT module. The RTC:RT module (described in detail below) includes a region extended from the start codon to the stop codon, but does not include the stop codon. The optional component, the RTC:3' side module 230, if present within the RTC, includes a region extended from the stop codon to the 3' end of the RTC. If present within the RTC, the RTC:3' side module may include sequences that modify the immunogenicity of the RTC and / or sequences that control the expression of the RTC:RT module. For example, the RTC:3' side module may include, and encode, a translation stop codon, a 3'UTR, a polyadenosine sequence, a polyadenylation signal, or any combination thereof.
[0123] In some embodiments, at least one RTC may contain a plasmid or be delivered to the subject as a plasmid. In some embodiments, at least one RTC may contain mRNA or pro-mRNA or be delivered to the subject as mRNA or pro-mRNA. In some embodiments, at least one RTC may contain a protein or be delivered to the subject as a protein. In some embodiments, at least one RTC may contain a proprotein or be delivered to the subject as a proprotein.
[0124] RTC: RT module The RT module of RTC comprises or encodes at least one compound or composition having reverse transcription activity, and specific examples of such compounds or compositions include, but are not limited to, a group of enzyme proteins known as reverse transcriptases (RTs). In some embodiments, the RT module may also comprise and encode a biomolecule derived from at least one reverse transcriptase (i.e., retroelement reverse transcriptase) found in retroelement genes. In some embodiments, the RTC:RT module comprises or encodes at least one reverse transcriptase derived from a non-long-chain terminal repeat (non-LTR) type retroelement.
[0125] reverse transcriptase In this specification, the term “reverse transcriptase (RT)” is used in its broadest sense to mean any biomolecule that possesses reverse transcription activity. In some embodiments, the RT used in the present invention includes the white-throated sparrow (Zonotrichia albicollis), zebra finch (Taeniopygia guttata), white-throated sandpiper (Tinamus guttatus), Galapagos finch (Geospiza fortis), northern stickleback (Pungitis pungitis), medaka (Oryzias latipes), zebrafish (Danio rerio), Oryzias melastigma, sea lamprey (Petromyzon marinus), brown trout (Salmo trutta), Atlantic salmon (Salmo salar), three-spined stickleback (Gasterosteus aculeatus), Drosophila mercatorum, yellow fruit fly (Drosophila melanogaster), parasitic wasp (Nasonia vitripennis), and confused flour beetle (Tribolium). The non-LTR type RT or its genome may originate from, or be derived from, these non-LTR type RTs or their genomes, of, the following animals: Drosophila simulans, Apis cerana, Bombyx mori, Lepidurus couesii, Triops cancriformis, Limulus polyphemus, Hydra magnipapillata, Adineta vaga, Ciona intestinalis, other birds, other arthropods, other fish, other tunicates, and other animals (including mammals and humans).
[0126] In some embodiments, at least one RTC:RT module used in the gene insertion system (GIS) of the present disclosure may include, encode, or be encoded by at least one of sequence numbers 1 to 57. In some embodiments, at least one RTC:RT module may include, encode, or be encoded by at least one of sequence numbers 1 to 57, having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of sequence numbers 1 to 57. In some embodiments, the RTC:RT module includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with sequence numbers 1 to 57.
[0127] In some embodiments, at least one RTC:RT module may include, or be coded by, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of sequence numbers 17-21 (ZoA1 RT sequences), and may encode this sequence or be coded by this sequence.
[0128] In some embodiments, at least one RTC:RT module may include, or be coded by, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of sequence numbers 26-29 (TaGu RT sequences), and may encode this sequence or be coded by this sequence.
[0129] In some embodiments, at least one RTC:RT module may include, or be coded by, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of sequence numbers 1-5 (TriCasB RT sequences), and may encode this sequence or be coded by this sequence.
[0130] In some embodiments, the RTC:RT module may contain, and may encode, a protein that has been shown to have TPRT activity by a suitable TPRT assay. A suitable TPRT assay, as an example, includes, but is not limited to, (i) transfecting a cell population with an expression plasmid encoding an RT protein having a suitable tag for affinity purification (e.g., a FLAG tag); (ii) lysing the cell population and recovering and purifying the expressed protein product by a suitable method known in the art; (iii) preparing recombinant template RNA by any method known in the art (e.g., T7 RNA polymerase); (iv) combining the purified RT protein, the recombinant template, and a nucleotide solution containing a target site oligonucleotide double-stranded DNA having a lower strand with radiolabeled ends, in a medium that promotes reverse transcription by RT; and (v) recovering and analyzing the product by any suitable method known in the art (e.g., denatured PAGE).
[0131] A reverse transcriptase (RT) suitable for use in the present invention may consist of multiple functional domains. In some embodiments, for example, as shown in Figure 3, at least one reverse transcriptase 300 includes at least one DNA-binding domain 310, at least one RNA-binding domain 320, at least one cDNA synthesis domain 330, at least one endonuclease domain 340, and any combination thereof. Note that this figure shows only one possible configuration that can be placed in each domain. In some embodiments, each illustrated domain may be included in this reverse transcriptase (RT) with varying frequencies, and / or these domains may be included in any order. In some embodiments, the DNA-binding domain or RNA-binding domain may be derived from a different type of polypeptide than the reverse transcriptase (RT), and may be a sequence unknown in the eukaryotic genome (e.g., a de novo-constructed DNA-binding domain or RNA-binding domain).
[0132] Start codon At least one non-natural translation start codon may be added to a nucleic acid sequence encoding a reverse transcriptase using various methods known in the art. This non-natural translation start codon may be added to any position in the sequence derived from a non-LTR retroelement that can induce the production of a functional reverse transcriptase. For example, in a wild-type non-LTR retroelement, at least one non-natural start codon may be added at a position approximately 1, 10, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, or 1000 bases or more away from a known reference point (e.g., from the amino acid sequence motif of the ORF of a natural retroelement reverse transcriptase). The location of the translation start codon may be selected based on various considerations that are obvious to those skilled in the art who are engaged in optimizing or regulating protein expression in target cells using recombinant technology, depending on the results of optimizing polypeptide length, sequence composition, activity, biological stability, avoidance of aggregation, or localization, and / or in order to improve the biological stability of the protein-coding mRNA.
[0133] The translation start codon may be any three nucleotides known to initiate translation by the ribosome, either in response to or independently of another sequence or structure in the mRNA. In some embodiments, the non-natural translation start codon is AUG.
[0134] RTC: 5' side module The reverse transcriptase construct (RTC) of the present invention may include at least one RTC:5' side module. Typically, the RTC:5' side module includes an untranslated biomolecular component, which may, but is not limited to, modifying the immunogenicity of the insertion construct, assisting in the localization of the insertion construct to a target intracellular region, controlling or modifying the expression of the RTC:RT module contained in the insertion construct, labeling the insertion construct for identification, assisting in the purification of the insertion construct, controlling the degradation of the insertion construct, regulating the exogenous or endogenous activity and / or function of the insertion construct, and any combination thereof.
[0135] In some embodiments, at least one RTC:5' side module may include and encode at least one 5'UTR. In some embodiments, at least one RTC:5' side module may include and encode at least one 5' cap. In some embodiments, at least one RTC:5' side module may include and encode at least one microRNA binding sequence. In some embodiments, at least one RTC:5' side module may include and encode at least one RNA polymerase promoter.
[0136] In some embodiments, at least one RTC:5' side module used in the gene insertion system of the present disclosure includes the 5'UTR of SEQ ID NO: 58.
[0137] In several embodiments, the inventors used one 5'UTR and one 3'UTR for transfected mRNA, obtained from a BioNTech vaccine sequence reported by the WHO. Furthermore, they used the polyA region encoded by the template (instead of using polyA polymerase after transcription). This polyA region consisted of 30 adenosines, a 10nt linker, and 70 adenosines. In addition, a Type IIS restriction site was introduced to cleave the mRNA transcription template without adding extra 3' terminal bases. All mRNAs were capped with TriLink AG clean cap, i.e., m7(3'OMeG)(5')ppp(5')(2'OMeA)pG). The UTRs were selected to allow tissue-specific reverse transcriptase expression, for example, to perform cell-type-specific translational regulation.
[0138] In some embodiments, the RTC:5' side module may include, and may encode, or be coded by, this sequence, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with sequence number 58.
[0139] RTC: 3' side module The reverse transcriptase construct (RTC) of the present invention may include at least one RTC:3' side module. Typically, the RTC:3' side module includes an untranslated biomolecular component, which may, but is not limited to, modifying the immunogenicity of the insertion construct, assisting in the localization of the insertion construct to a target intracellular region, controlling or modifying the expression of the RTC:RT module contained in the insertion construct, labeling the insertion construct for identification, assisting in the purification of the insertion construct, controlling the degradation of the insertion construct, regulating the exogenous or endogenous activity and / or function of the insertion construct, and any combination thereof.
[0140] In some embodiments, at least one RTC:3' side module may include at least one 3'UTR. In some embodiments, at least one RTC:3' side module may include at least one polyA tract, i.e., polyA tail, which may encode. In some embodiments, at least one RTC:3' side module may include at least one microRNA binding sequence, which may encode.
[0141] In some embodiments, at least one RTC:3' side module used in the gene insertion system of the present disclosure includes a 3'UTR and a poly-A tail as shown in SEQ ID NO: 59.
[0142] In some embodiments, the RTC:3' side module includes a 3'UTR having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with sequence number 59.
[0143] Modularity of reverse transcriptase constructs (RTCs) The reverse transcriptase construct (RTC) of the present invention may be designed by combining at least one RTC:RT module, at least one RTC:5' side module as an optional component, and / or at least one RTC:3' side module as an optional component, to obtain a desired function or activity. In some embodiments, the reverse transcriptase construct includes at least one RTC:5' side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:3' side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:RT module. In some embodiments, the reverse transcriptase construct includes at least one RTC:5' side module, at least one RTC:RT module, and at least one RTC:3' side module. In some embodiments, the reverse transcriptase construct includes at least one RTC:5' side module and at least one RTC:RT module. In some embodiments, the reverse transcriptase construct includes at least one RTC:RT module and at least one RTC:3' side module.
[0144] In some embodiments, the reverse transcriptase construct of the present invention may not include both at least one RTC:5' side module and at least one RTC:3' side module. In some embodiments, the reverse transcriptase construct of the present invention may not include either at least one RTC:5' side module or at least one RTC:3' side module. In some embodiments, the reverse transcriptase construct of the present invention may not include at least one RTC:5' side module. In some embodiments, the reverse transcriptase construct of the present invention may not include at least one RTC:3' side module.
[0145] In some embodiments, at least one reverse transcriptase construct may include any combination of (a) at least one RTC:5' side module selected from or encoding SEQ ID NO: 58, or encoded by SEQ ID NO: 58; (b) at least one RTC:RT module selected from or encoding any of SEQ ID NOs: 1 to 57, or encoded by any of SEQ ID NOs: 1 to 57; and / or (c) at least one RTC:3' side module selected from or encoding SEQ ID NO: 59, or encoded by SEQ ID NO: 59.
[0146] Exemplary reverse transcriptase construct (RTC) The reverse transcriptase construct (RTC) used in the present invention may contain at least one of sequence numbers 1 to 57, may encode at least one of sequence numbers 1 to 57, or may be encoded by at least one of sequence numbers 1 to 57. In some embodiments, the reverse transcriptase construct (RTC) may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of sequence numbers 1 to 57, may encode this sequence, or may be encoded by this sequence.
[0147] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, with at least one of sequence numbers 17-21 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0148] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, with at least one of sequence numbers 26-29 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0149] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with sequence number 24 or 25.
[0150] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, with at least one of sequence numbers 1-5 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0151] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, with at least one of sequence numbers 35-37 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0152] In some embodiments, at least one reverse transcriptase construct (RTC) may contain, and may encode, or be encoded by, this sequence, with at least one of sequence numbers 32-34 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0153] In some embodiments, at least one reverse transcriptase construct (RTC) includes the structure shown in Figure 5.
[0154] Regulators of reverse transcriptase constructs (RTCs) The reverse transcriptase construct (RTC) of the present invention may further contain any number of regulatory factors, which may be contained in any of the RTC modules. In this specification, the term “regulatory factor” means a sequence, region or domain that enables control of the expression or activity of a biomolecule that is part of the construct.
[0155] For example, an RNA-based RTC may contain any number of microRNA (miRNA) binding sites or small interfering RNA (siRNA) binding sites. While we do not wish to be bound by any theory, the presence of RNA interference (RNAi) binding sites may inhibit the expression of RT proteins in certain types of cells based on the presence of an RNAi transcriptome. In this way, the gene insertion system of the present invention can exclude cells of a target type. In this specification, the term “miRNA binding site or siRNA binding site” means an RNA sequence complementary to at least one miRNA or siRNA.
[0156] In some embodiments, the RTC may include at least one miRNA and / or siRNA contained in the transgene inserted by the gene insertion system, or at least one miRNA-binding site and / or siRNA-binding site complementary to the at least one miRNA and / or siRNA encoded by the transgene. By including such miRNA-binding sites or siRNA-binding sites, the gene insertion system of the present invention may generally be able to regulate the number of transgene insertions achieved by a single administration of the gene insertion system and / or prevent repeated insertions of the transgene after the initial administration. Thus, the gene insertion system may have an improved ability to be re-administered or co-administered to a subject.
[0157] Gene insertion construct (GIC) The gene insertion system of the present invention comprises at least one gene insertion construct (GIC), which typically comprises or encodes at least one target sequence (i.e., “payload sequence”) intended for insertion into a genome of interest. In this specification, the term “gene insertion construct (GIC)” means a biopolymer construct comprising or encoding at least one RNA sequence, characterized in that the RNA sequence is included in or recognized by at least one reverse transcriptase encoded by at least one RTC:RT module, thereby enabling the RNA sequence to function as a template for reverse transcription. In some embodiments, the at least one gene insertion construct used in the gene insertion system of the present invention may comprise, but is not limited to, a nucleic acid biopolymer such as RNA, DNA, or any combination thereof.
[0158] Structure of a gene insertion construct (GIC) The gene insertion construct (GIC) of the present invention may include, and may encode, at least one GIC:5' module, at least one GIC:payload module, at least one GIC:3' module, and any combination thereof. In some embodiments, at least one GIC may include a plasmid or be delivered to the target as a plasmid. In some embodiments, at least one GIC may include a linear RNA or be delivered to the target as a linear RNA.
[0159] In some embodiments, the at least one GIC:5' side module is an optional component. In some embodiments, the at least one GIC:3' side module is an optional component. In some embodiments, the gene insertion construct (GIC) of the present invention includes or codes for at least one GIC:payload module but does not include, and does not code for, at least one GIC:5' side module and / or at least one GIC:3' side module.
[0160] Referring to Figure 6, an exemplary linear RNA gene insertion construct (GIC) 400 is shown, where an optional component, the GIC:5' module 410, extends from the 5' end of the GIC sequence to the end 420 of the GIC:5' module. The GIC:payload module 430 extends toward the 3' end of the GIC:5' module (if present) to the end 440 of the GIC:payload module. Furthermore, the GIC:3' module 450 extends to the 3' end of the GIC. Details of each of these features are described below.
[0161] GIC: 5' side module The GIC:5' side module used in the gene insertion construct (GIC) of this disclosure may include, and may encode, at least one sequence derived from the 5' region of a natural retroelement. While we do not wish to be bound by any theory, this 5' side module may include, and may encode, an RNA sequence that interacts with at least one RNA-binding domain of a reverse transcriptase, an RNA sequence that can synthesize a second strand upon transgene insertion, an RNA sequence that reduces the immunogenicity of the gene insertion construct, an RNA sequence that provides features useful for the stability and / or purification of the gene insertion construct, and any combination thereof.
[0162] GIC: 5' side module structure In several embodiments, the 5' end module includes a 5' rRNA sequence and a ribozyme (RZ) sequence. In some embodiments, the 5' rRNA sequence and the RZ sequence do not necessarily have to be completely separated. In some embodiments, the 5' end module includes a “folding sequence” which may be separated from the RZ sequence. In some embodiments, the GIC:5' end module may include, and encode, at least one rRNA sequence (or other target site sequence), at least one ribozyme (RZ) sequence, at least one folding sequence, and any combination thereof.
[0163] Referring again to Figure 6, the enlarged view (bottom left) of the GIC:5' side module 410 shows the structure of an exemplary GIC:5' side module. If the GIC:5' rRNA sequence 411 is present at the 5' end of the 5' side module, this GIC:5' rRNA sequence 411 may contain, and may encode, an RNA sequence complementary to the target DNA sequence located 5' to or near the target insertion site. If a ribozyme (RZ) sequence 412 is present in the GIC:5' side module, this RZ sequence may contain at least one RNA sequence having a self-cleaving ribozyme folding structure, which may release a functional GIC from the transcribed 5' leader sequence by self-cleaving, or such self-cleaving may not occur. The RZ sequence of the GIC:5' side module folds and, if active, autocleaves to incorporate the GIC:5' rRNA sequence as part of the RZ at or near the 5' end of the GIC. The folding motif sequence 413, which is an optional component of the GIC:5' side module, may contain at least one RNA sequence that is predicted or demonstrated to fold autonomously, and this RNA sequence is considered useful for physically and / or kinetically separating the folding of the RZ of the GIC:5' side module from the folding of the payload sequence. Furthermore, within region 414, or at position 420 between the GIC 5' side module 410 and the payload module 430, the GIC sequence may be appended to halt transcription initiated from an endogenous promoter sequence adjacent to the target site in the cell, or to perform other control. In some embodiments, an endogenous promoter sequence adjacent to the target site within the cell may be used for payload expression, one example of which involves regulating payload expression by adding a GIC sequence to positions 420 and / or 440 (e.g., initiating or terminating translation of an RNA transcript containing the payload sequence by the host promoter).Furthermore, region 414 may contain an RNA polymerase (RNAP) termination sequence to prevent RNA polymerase from reading through (skipping) the gene at the target insertion site. In some embodiments, RNAP is RNAP I (Pol I), and when the GIC payload module is incorporated into the gene target site of ribosomal DNA, this termination sequence prevents Pol I read-through transcription. In some embodiments, the RNAP transcription termination sequence is 5'-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3'(Sequence ID: 5'-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3'). 333 Includes the array indicated by ).
[0164] GIC: rRNA sequence of the 5' side module At least one rRNA sequence in the GIC:5' side module is an optional component of the GIC:5' side module. If an rRNA sequence is present in the GIC:5' side module, this rRNA sequence may contain, and may encode, a human ribosomal RNA (rRNA) sequence or other sequence that is homogeneous and / or complementary to at least one target DNA sequence located 5' to the target insertion site. While we do not wish to be bound by any theory, this rRNA sequence may induce the synthesis of the second strand of the inserted cDNA transgene by mobilizing at least one endogenous DNA repair mechanism. In some embodiments, the rRNA sequence in the GIC:5' side module is located 5' to the RZ sequence of the GIC:5' side module. In some embodiments, the GIC:5' side module does not contain a sequence containing an rRNA genome sequence.
[0165] In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1 to 36 nt of rRNA and may encode approximately 1 to 36 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1 to 30 nt of rRNA and may encode approximately 1 to 30 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1 to 28 nt of rRNA and may encode approximately 1 to 28 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1 to 26 nt of rRNA and may encode approximately 1 to 26 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1 to 13 nt of rRNA and may encode approximately 1 to 13 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain approximately 1–11 nt of rRNA and may encode approximately 1–11 nt of rRNA.
[0166] In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain rRNA of approximately 1nt, 2nt, 3nt, 4nt, 5nt, 6nt, 7nt, 8nt, 9nt, 10nt, 11nt, 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, 30nt, 31nt, 32nt, 33nt, 34nt, 35nt, or 36nt, which may encode rRNA of such length.
[0167] In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 30 nt of rRNA and may encode about 30 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 36 nt of rRNA and may encode about 36 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 28 nt of rRNA and may encode about 28 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 26 nt of rRNA and may encode about 26 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 13 nt of rRNA and may encode about 13 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:5' side module may contain about 11 nt of rRNA and may encode about 11 nt of rRNA. In some embodiments, the rRNA sequence of the GIC:5' side module includes a 5' G nucleotide.
[0168] In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number. 179 ~ 205 It may contain at least one of the sequences, may encode at least one of these sequences, or may be encoded by at least one of these sequences. In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number 179 ~ 205 It may include, and may encode, or be encoded by, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% homology with at least one of the above. In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number 179 ~ 205The sequence includes one, two, or three nucleotide changes or substitutions compared to a sequence selected from the group consisting of the above.
[0169] In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number. 181 It may also contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% homology, which may encode or be encoded by this sequence. In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number 181 It includes sequences that have one, two, or three nucleotide changes compared to the above.
[0170] In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number. 183 It may also contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% homology, which may encode or be encoded by this sequence. In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number 183 It includes sequences that have one, two, or three nucleotide changes compared to the above.
[0171] In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number. 184 It may also contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% homology, which may encode or be encoded by this sequence. In some embodiments, at least one rRNA sequence of the GIC:5' side module is the sequence number 184 It includes sequences that have one, two, or three nucleotide changes compared to the above.
[0172] GIC:5' side module RZ array The RZ sequence of the GIC:5' module is an optional component of the GIC:5' module and, if present in the GIC:5' module, contains or encodes at least one sequence having a self-cleaving ribozyme or a folded structure of a self-cleaving ribozyme (collectively referred to as "RZ"). While we do not wish to be bound by any theory, such an RZ motif generates a 5'OH end in the GIC, which may be, for example, a 5' end produced by self-cleavage, and in a stable tertiary structure, the generated 5'OH end may reduce the innate immune response to exogenous RNA, reduce the degradation of the GIC by 5'-3' exonucleases that initiate cleavage in a 5' monophosphate-dependent manner, and reduce the likelihood that the GIC may be recognized by the target cell as mRNA or other undesirable types of RNA rather than as template RNA.
[0173] In some embodiments, at least one RZ sequence of the GIC:5' side module contains or encodes a ribozyme derived from the 5' region of at least one non-LTR type retroelement. In some embodiments, at least one RZ sequence of the GIC:5' side module contains or encodes a ribozyme derived from the 5' region of a non-LTR type retroelement obtained from the genomes of three-spined stickleback, American horseshoe crab, northern stickleback, parasitic wasp, Galapagos finch, medaka, white-throated bunting, zebra finch, confused flour beetle (e.g., R2 derived from strain A or B), white-throated swan, other birds, other arthropods, other fish, other urochordates, other animals or similar genomes.
[0174] In some embodiments, the RZ sequence of the GIC:5' side module contains or encodes an RZ capable of forming the secondary and tertiary structures of the hepatitis delta virus (HDV) RZ, which may be modified from a naturally occurring sequence and / or de novo designed without using a known genome sequence. In some embodiments, the folded RZ sequence that bridges paired stem P1 and paired stem P2 in HDV is also called junction (J)1 / 2, and part or all of it is contained in a target site sequence of a desired length (e.g., 5' rRNA) or in a desired target site sequence further protected by the formation of a stem-loop. In some embodiments, by incorporating the paired stem 4 (P4) of the folded RZ sequence of HDV into the design, it may be possible to purify the gene insertion construct (GIC) without denaturation by, for example, binding it to a native or modified sequence of the coat protein of a PP7 phage or MS2 phage. In some embodiments, the RZ sequence is designed and optimized to minimize or eliminate unproductive folding. In some embodiments, the RZ sequence is designed and optimized to minimize the number of uridine nucleotides. In some embodiments, the RZ sequence is designed and optimized so that all or some of the standard ribonucleotides can be replaced by nucleotide analogs incorporated during the synthesis of the template RNA.
[0175] In some embodiments, at least one RZ sequence of the GIC:5' side module is sequence number 60~ 153 It may include at least one of the following, and sequence number 60~ 153 You may code at least one of the following, array index 60~ 153It may be encoded by at least one of the following. In some embodiments, the RZ sequence autonomously folds to form an active ribozyme. In some embodiments, the RZ sequence contains an internal rRNA sequence at its 5' end. In some embodiments, the RZ sequence is an extended sequence at the 5' or 3' end. In some embodiments, the RZ sequence is a catalytically inactive RZ sequence. In some embodiments, at least one RZ sequence of the GIC:5' side module is sequence number 60~ 153 It may include sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of the above, and may encode or be encoded by this sequence. In some embodiments, the RZ sequence of the GIC:5' side module is sequence number 60~ 153 It includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0176] In some embodiments, at least one RZ sequence of the GIC:5' side module may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to sequence 60, which may encode or be encoded by this sequence.
[0177] In some embodiments, at least one RZ sequence of the GIC:5' side module may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to sequence 64, which may encode or be encoded by sequence.
[0178] In some embodiments, at least one RZ sequence of the GIC:5' side module may contain a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to sequence 67, which may encode or be encoded by this sequence.
[0179] In some embodiments, at least one RZ sequence of the GIC:5' side module is an array sequence. 100 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0180] In some embodiments, at least one RZ sequence of the GIC:5' side module is an array sequence. 120 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0181] In some embodiments, at least one RZ sequence of the GIC:5' side module is an array sequence. 121 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0182] In some embodiments, at least one RZ sequence of the GIC:5' side module is an array sequence. 136 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0183] GIC:5' side module folding arrangement The folding sequence of the GIC:5' side module is an optional component of the 5' side module, and if present, it includes at least one RNA sequence motif having a specially designed structure. In some embodiments, the autonomous folding RNA sequence motif includes at least one hairpin motif, which may be present after, for example, the RZ has blocked misfolding of the RZ sequence by forming a base pair with a payload region to be transcribed later. In some embodiments, the 5' side module region designed to improve the folding of a productive template RNA directly or indirectly forms a base pair or interacts with another template RNA region of the payload module or the 3' side module. In some embodiments, the at least one RNA sequence motif that induces the folding of the template RNA may include at least one stem-loop motif that causes a protein crosslink to another stem-loop motif. In some embodiments, the folding sequence of the 5' module may facilitate the pairing of the reverse transcriptase-encoding mRNA with the template RNA, for example, by promoting the packaging of the reverse transcriptase-encoding mRNA and the template RNA together in a 1:1 stoichiometric ratio on individual delivery media. In some embodiments, the folding sequence of the 5' module may facilitate the pairing of the endogenous RNA of a target cell with the template RNA, for example, with the aim of achieving template RNA stabilization, localization, and / or other useful outcomes.
[0184] In some embodiments, at least one folding sequence of the GIC:5' side module is the sequence number. 206 ~ 207 It may include at least one of these sequences, may encode at least one of these sequences, or may be encoded by at least one of these sequences. In some embodiments, at least one folding sequence of the GIC:5' side module is the sequence number 206 ~ 207It may include, and may encode, or be encoded by, a sequence having at least one of the following homologies: 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5%. In some embodiments, the folding sequence of the GIC:5' side module is the sequence number. 206 ~ 207 It includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0185] In some embodiments, at least one folding sequence of the GIC:5' side module is the sequence number. 206 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0186] In some embodiments, at least one folding sequence of the GIC:5' side module is the sequence number. 207 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0187] Modularity of GIC:5' side modules Each component of the 5' module disclosed herein may be used combinatorially and interchangeably to design a 5' module having the functionality required or desired for a particular gene insertion system.
[0188] In some embodiments, at least one GIC:5' side module includes at least one rRNA sequence of the GIC:5' side module. In some embodiments, at least one GIC:5' side module includes at least one RZ sequence of the GIC:5' side module. In some embodiments, at least one GIC:5' side module includes at least one folding sequence of the GIC:5' side module. In some embodiments, at least one GIC:5' side module includes at least one rRNA sequence of the GIC:5' side module and at least one RZ sequence of the GIC:5' side module. In some embodiments, at least one GIC:5' side module includes at least one rRNA sequence of the GIC:5' side module, at least one RZ sequence of the GIC:5' side module and at least one folding sequence of the GIC:5' side module.
[0189] In some embodiments, at least one GIC:5' side module is, (a) Sequence ID 179 ~ 205 At least one rRNA sequence selected from any one of the following, or encoding any one of these, or encoded by any one of these; (c) Sequence ID 60~ 153 At least one RZ array selected from any one of the following, or encoding any one of these, or encoded by any one of these; and / or (d) Sequence ID 206 ~ 207 A folding array that is selected from any one of the following, codes for any one of these, or is coded by any one of these. It may include any combination of these elements.
[0190] Example GIC:5' side module In some embodiments, at least one GIC:5' side module is sequence number 60~153 It may include at least one of the following, and sequence number 60~ 153 You may code at least one of the following, array index 60~ 153 It may be coded by at least one of the following. In some embodiments, at least one GIC:5' side module is sequence number 60~ 153 It may include sequences having at least one of the above and at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by this sequence. In some embodiments, the GIC:5' side module includes sequence numbers 60~ 153 It includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0191] In some embodiments, at least one GIC:5' side module may contain, and may encode, or be coded by, this sequence, with at least one of sequence numbers 60, 61, 77 and 79-83 having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0192] In some embodiments, at least one GIC:5' side module may contain, and may encode, or be coded by, this sequence, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with sequence number 62 or 63.
[0193] In some embodiments, at least one GIC:5' side module is array number 120 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0194] In some embodiments, at least one GIC:5' side module is array number 116 ~ 118 It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0195] GIC:3' side module
[0196] The 3' module used in the gene insertion construct (GIC) of this disclosure may include, and may encode, at least one sequence derived from the 3'UTR of a natural retroelement. Typically, the 3' module includes components that facilitate recognition and binding of the gene insertion construct by reverse transcriptase, components for positioning the payload module for reverse transcription, and components for stabilizing the RNA of the gene insertion construct.
[0197] GIC:3' side module structure In some embodiments, the GIC:3' side module may include, and may encode, at least one reverse transcriptase (RT) recognition sequence, at least one rRNA sequence which is an optional component, at least one A tract sequence which is an optional component, and any combination thereof.
[0198] Referring again to Figure 6, the enlarged view (bottom right) shows an example of the structure of the GIC:3' side module 450. The 5' end of the GIC:3' side module is the reverse transcriptase recognition sequence 451 of the GIC:3' side module, which may contain, and may encode, a sequence that is recognized or bound by at least one reverse transcriptase. If an rRNA sequence 452 of the GIC:3' side module is present, this rRNA sequence 452 may be located on the 3' side of the reverse transcriptase recognition sequence of the GIC:3' side module, and may contain, and may encode, a sequence of the same type as the target site region, for example, a 28S rRNA nucleotide that can base pair with the 3' end of the TPRT primer. Furthermore, if an A tract sequence 453 exists in the GIC:3' side module, this A tract sequence 453 may contain an adenosine-rich sequence or a sequence of adenosines arranged in series, which may have a limited length, and the length of this sequence may be, for example, 10 to 60 nt, and the A tract sequence 453 may be located at the 3' end of the GIC:3' side module.
[0199] Reverse transcriptase recognition sequence of the GIC:3' module The reverse transcriptase recognition sequence of the GIC:3' side module may include, and may encode, at least one sequence that interacts with at least one reverse transcriptase, or at least one sequence that is recognized by at least one reverse transcriptase. While not wishing to be constrained by any theory, at least one RNA sequence included in the reverse transcriptase recognition sequence of the GIC:3' side module may at least transiently bind to at least one template RNA-binding domain of a reverse transcriptase, such as a retroelement reverse transcriptase. The length and sequence identity of the reverse transcriptase recognition sequence of the GIC:3' side module may be configured to position a reverse transcriptase in the gene insertion construct (GIC), thereby allowing the first nucleotide reverse-transcribed by the reverse transcriptase to be positioned at the intended 3' end of the transgene being inserted. In this specification, the reverse transcriptase recognition sequence of the GIC:3' side module may also be referred to as the "GIC:3' side module 3'UTR".
[0200] In some embodiments, the reverse transcriptase recognition sequence of at least one GIC:3' side module is derived from or includes the 3' region of a natural retroelement. In some embodiments, the reverse transcriptase recognition sequence of at least one GIC:3' side module is derived from the 3' region of a non-LTR type retroelement obtained from the genomes of three sticklebacks, Drosophila melanogaster, American horseshoe crab, three stickleback, parasitic wasp, Galapagos finch, medaka, white-throated sablefish, zebra finch, confused flour beetle, white-throated swan, Drosophila melanogaster, silkworm moth, A. vaga, other birds, other arthropods, other fish, other urochordates, other animals, or similar organisms. In some embodiments, the reverse transcriptase recognition sequence of the GIC:3' side module is modified from the 3' region of a natural retroelement by increasing folding stability or uniformity. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is designed and / or selected to have a desired affinity and / or specificity in interaction with reverse transcriptase, or to have another mechanism that confers a desired function as a template for reverse transcription. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is designed and / or selected not to interact with or affect endogenous components of the target cell and / or not to have a harmful effect on the host cell.
[0201] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module (i.e., the 3'UTR sequence of the GIC:3' side module) may include at least one of sequence numbers 200-224, may encode at least one of sequence numbers 200-224, or may be encoded by at least one of sequence numbers 200-224. In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module may be sequence number 154 ~ 175 It may include, and may encode, or be encoded by, a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology with at least one of the above. In some embodiments, the reverse transcriptase recognition sequence of the GIC:3' side module is the sequence number 154 ~ 178 It is a non-natural sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with the original sequence.
[0202] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is sequence number 156 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0203] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is sequence number 158 , 176 , 177 or 178It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0204] In some embodiments, at least one reverse transcriptase recognition sequence of the GIC:3' side module is sequence number 157 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0205] In some embodiments, the GIC:3' side module includes a reverse transcriptase recognition sequence derived from a different species than the species from which the reverse transcriptase encoded by the reverse transcriptase construct (RTC) originates. For example, in some embodiments, the reverse transcriptase recognition sequence may originate from one species of bird, and the reverse transcriptase may originate from another species of bird. In some embodiments, the reverse transcriptase recognition sequence originates from a bird selected from one of the following species: white-throated sparrow, zebra finch, white-throated swan, and Galapagos finch, and the reverse transcriptase is selected from a different species of bird than the selected bird (e.g., white-throated sparrow, zebra finch, white-throated swan, or Galapagos finch). In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from one of the following species: white-throated sparrow, zebra finch, white-throated swan, and Galapagos finch, and the reverse transcriptase recognition sequence is selected from a different bird species than the selected bird (e.g., white-throated sparrow, zebra finch, white-throated swan, or Galapagos finch). In some embodiments, the reverse transcriptase encoded by the RTC construct is selected from an amino acid sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 18 or 20, and the reverse transcriptase recognition sequence is SEQ ID NO: 157 , 158 , 159 or 176 ~ 178 The amino acid sequence is selected from sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 27 or 29, and the reverse transcriptase recognition 156 , 158 , 159 or 176 ~ 178The amino acid sequence is selected from sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 25, and the reverse transcriptase recognition 156 , 157 , 158 or 176 ~ 178 The amino acid sequence is selected from sequences having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with SEQ ID NO: 31, and the reverse transcriptase recognition. 156 , 157 or 159 The amino acid sequences are selected from those having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity.
[0206] GIC: rRNA sequence of the 3' side module The rRNA sequence of the GIC:3' module, i.e., the sequence that forms a base pair with the TPRT primer immediately downstream of the nick introduced at a non-rDNA target site, is an optional component of the 3' module, and if present in the 3' module, this rRNA sequence may include a human ribosomal RNA (rRNA) sequence. While we do not wish to be bound by any theory, the length and sequence identity of the rRNA sequence of the GIC:3' module affect how accurately and efficiently the gene insertion systems disclosed herein can insert the transgene into the target genome. For example, depending on the length of the selected GIC:3' module rRNA sequence, reverse transcription may be initiated from the internal sequence, efficiently shortening the inserted transgene or allowing insertion into an off-target site, in either case reducing the insertion efficiency and specificity of the transgene at the target site. The RTC and GIC are configured such that the base pair formed between the primer sequence immediately downstream of the nick introduced at the target site and the rRNA sequence of the GIC:3' module must be of a specific length. This further improves fidelity in utilizing target sites, allowing for more efficient acquisition of precise junctions where the transgene is inserted. The optimal length of GIC:3' rRNA is less than 20 nt, and 4 nt in particular allows for strong stimulation from the entire 4 bp base pair formed at the target site nick. Therefore, if the RTC randomly introduces nicks, using 4 nt GIC:3' rRNA, it is thought that only one out of 256 nicks will achieve optimal transgene insertion.
[0207] In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain and encode approximately 1 to 30 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain and encode approximately 1 to 20 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain and encode approximately 1 to 10 nt of rRNA. In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain and encode approximately 1 to 5 nt of rRNA.
[0208] In some embodiments, at least one rRNA sequence of the GIC:3' side module may contain a portion of rRNA of approximately 1 nt, 2 nt, 3 nt, 4 nt, 5 nt, 6 nt, 7 nt, 8 nt, 9 nt, 10 nt, 11 nt, 12 nt, 13 nt, 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt, 22 nt, 23 nt, 24 nt, 25 nt, 26 nt, 27 nt, 28 nt, 29 nt, or 30 nt, which may encode a portion of rRNA of such length.
[0209] In some embodiments, at least one rRNA sequence in the GIC:3' side module may contain about 20 nt of rRNA and encode about 20 nt of rRNA. In some embodiments, at least one rRNA sequence in the GIC:3' side module may contain about 4 nt of rRNA and encode about 4 nt of rRNA. In some embodiments, at least one rRNA sequence in the GIC:3' side module may contain about 10 nt of rRNA and encode about 10 nt of rRNA.
[0210] In some embodiments, at least one rRNA sequence of the GIC:3' side module is the sequence number. 208 ~213 may include at least one of. In some embodiments, at least one rRNA sequence of the GIC:3'-side module is SEQ ID NO. 208 ~ 217 , and SEQ ID NO. 208 ~ 217 and is selected from the group consisting of sequences containing one, two or three nucleotide substitutions.
[0211] GIC:3' side module A tract sequence The A-tract sequence of the GIC:3'-side module is an optional component of the 3'-side module. When present in the 3'-side module, this A-tract sequence includes a terminal poly sequence containing a plurality of adenosines (A) arranged in series. Without wishing to be bound by any theory, the A-tract sequence of the GIC:3'-side module may stabilize or protect the GIC from further processing on the 3'-side. Furthermore, in cells, the GIC is recognized as mRNA, making it less likely to undergo degradation of the GIC related to ribonucleoprotein assembly, transport and translation. Additionally, at least one A-tract sequence of the GIC:3'-side module may protect the GIC from binding by common single-stranded RNA-binding proteins and may assist in arranging the GIC:3' rRNA sequence to form base pairs with the target site primer. For the sake of clarity, this A-tract sequence is different from the natural mRNA polyA tail sequence in which usually more than about 100 - 200 nt of adenosines are arranged in series.
[0212] In some embodiments, at least one A tract sequence, which is an optional component of the GIC:3' side module, contains or encodes a sequence consisting of approximately 1 to 50 adenosines. For example, an A tract sequence, which is an optional component of the GIC:3' side module, may contain approximately 1 to 50 adenosines, approximately 5 to 50 adenosines, approximately 10 to 50 adenosines, approximately 15 to 50 adenosines, approximately 20 to 50 adenosines, approximately 25 to 50 adenosines, approximately 30 to 50 adenosines, approximately 35 to 50 adenosines, approximately 40 to 50 adenosines, approximately 45 to 50 adenosines, approximately 1 to 45 adenosines, approximately 5 to 45 adenosines, and approximately 10 to 45 adenosine molecules, approximately 15-45 adenosine molecules, approximately 20-45 adenosine molecules, approximately 25-45 adenosine molecules, approximately 30-45 adenosine molecules, approximately 35-45 adenosine molecules, approximately 40-45 adenosine molecules, approximately 1-40 adenosine molecules, approximately 5-40 adenosine molecules, approximately 10-40 adenosine molecules, approximately 15-40 adenosine molecules, approximately 20-40 adenosine molecules, approximately 25-40 adenosine molecules, approximately 30-40 adenosine molecules, approximately 35-40 adenosine molecules, approximately 1- 35 adenosine molecules, approximately 5-35 adenosine molecules, approximately 10-35 adenosine molecules, approximately 15-35 adenosine molecules, approximately 20-35 adenosine molecules, approximately 25-35 adenosine molecules, approximately 30-35 adenosine molecules, approximately 1-30 adenosine molecules, approximately 5-30 adenosine molecules, approximately 10-30 adenosine molecules, approximately 15-30 adenosine molecules, approximately 20-30 adenosine molecules, approximately 25-30 adenosine molecules, approximately 1-25 adenosine molecules, approximately 5-25 adenosine molecules, approximately 10- The sequence may contain, and may encode, a sequence consisting of 25 adenosines, approximately 15-25 adenosines, approximately 20-25 adenosines, approximately 1-20 adenosines, approximately 5-20 adenosines, approximately 10-20 adenosines, approximately 15-20 adenosines, approximately 1-15 adenosines, approximately 5-15 adenosines, approximately 10-15 adenosines, approximately 1-10 adenosines, approximately 5-10 adenosines, or approximately 1-5 adenosines.In some embodiments, the A tract sequence of the GIC:3' side module contains approximately 1 to 100, 1 to 90, 1 to 80, 1 to 70, or 1 to 60 adenosines.
[0213] In some embodiments, at least one A tract sequence, which is an optional component of the GIC:3' side module, contains a sequence consisting of approximately 20 to 25 adenosines or encodes a sequence consisting of approximately 20 to 25 adenosines.
[0214] In some embodiments, at least one A tract sequence, which is an arbitrary component of the GIC:3' side module, is approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25. It contains or codes for a sequence consisting of approximately 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 adenosines. In some embodiments, the A tract sequence of the GIC:3' side module contains 22 adenosines.
[0215] Modularity of GIC:3' side modules Each component of the 3' module disclosed herein may be used combinatorially and interchangeably to design a 3' module having the functionality required or desired for a particular gene insertion system.
[0216] In some embodiments, at least one GIC:3' side module includes at least one reverse transcriptase recognition sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module includes at least one rRNA sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module includes at least one A tract sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module includes at least one reverse transcriptase recognition sequence of the GIC:3' side module and at least one rRNA sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module includes at least one reverse transcriptase recognition sequence of the GIC:3' side module and at least one A tract sequence of the GIC:3' side module. In some embodiments, at least one GIC:3' side module includes at least one reverse transcriptase recognition sequence of the GIC:3' side module, at least one rRNA sequence of the GIC:3' side module, and at least one A tract sequence of the GIC:3' side module.
[0217] In some embodiments, at least one GIC:3' side module is, (a) Sequence ID 154 ~ 175 At least one reverse transcriptase recognition sequence selected from, encoding, or encoded by any one of the following; (b) Sequence ID 208 ~ 217 At least one rRNA sequence selected from any one of the following, or encoding any one of these, or encoded by any one of these; and / or (c) at least one A tract sequence It may include any combination of these elements.
[0218] An example GIC:3' side module In some embodiments, at least one GIC:3' side module is array number 225 ~ 253 It may contain at least one of the array sequences. 225 ~ 253 You may code at least one of the arrays 225 ~ 253 It may be coded by at least one of the following. In some embodiments, at least one GIC:3' side module is array number 225 ~ 253 The GIC:3' side module may include at least one sequence selected from the group consisting of the following and a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode this sequence or be encoded by this sequence. In some embodiments, at least one GIC:3' side module is sequence number 225 ~ 253 The sequence includes sequences selected from the group consisting of or any combination thereof, and sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity). In some embodiments, the GIC:3' side module includes sequence numbers. 225 ~ 253 It includes a non-natural sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.
[0219] In some embodiments, at least one GIC:3' side module is array number 238 ~ 244 It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0220] In some embodiments, at least one GIC:3' side module is "GACGGTAGC TAGGTTCGCA AGGCAGCCAC AAGCCAAAGA TAGGTAGGGT GCTCATAGTG AGTAGGGACA GTGCCTTTTG ATTCACAACG CGTCAATACC ATCTGACACG GATACCCTTA CCGGACTTGT CATGATCTCC CAGACTTGTC CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO 176 ), "CCGGACTTGT CATGATCTCC CAGACTTGTC CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO 177 ), and "CAAGGTGGAC GGGCCACCTT TACTTAACCC GGAAAAGGAA CATATATTAA TTATATGTGT TCGGAAAA" (SEQ ID NO 178 ), and may include a sequence having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10% or 5% homology with a sequence selected from the group consisting of. In some embodiments, such a sequence includes the 3' side sequence represented by TAGCaaaaaaaaaaaaaaaaaaaaaa (SEQ ID NO 334 ).
[0221] In some embodiments, at least one GIC:3' side module is SEQ ID NO 239It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0222] In some embodiments, at least one GIC:3' side module is array number 232 It may also contain sequences having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, which may encode or be encoded by such sequences.
[0223] In some embodiments, at least one GIC:3' side module is array number 240 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0224] GIC: Payload Module The GIC:payload module used in the gene insertion construct (GIC) of the present invention functions as part of a template for reverse transcription and comprises or encodes at least one payload sequence to be inserted into the target genome by the gene insertion system disclosed herein. In this specification, the terms “payload sequence” or simply “payload” mean a biomolecular sequence intended for insertion into the target genome by at least one gene insertion system of the present invention. The payload sequence of the present invention may comprise at least one transgene.
[0225] In this specification, the term “transgene” is used in its broadest sense and means any gene sequence inserted into the target genome by the gene insertion system of the present invention. For example, a transgene may be a sequence not normally found in the target genome, or a sequence normally found in the target genome but not normally found at the target insertion site. A transgene may include, but is not limited to, sequences comprising or encoding a desired expression product (e.g., at least one mRNA, microRNA, siRNA, rRNA, tRNA, long non-coding RNA, cytoplasmic small RNA, nuclear small RNA, nucleolar small RNA, Cajal small RNA, circular RNA, peptide, polypeptide, and / or protein), and / or sequences that control the expression of at least one transgene. In some embodiments, the transgene encodes a protein selected from telomerase reverse transcriptase (TERT, e.g., human TERT), phenylalanine hydroxylase (PAH, e.g., human PAH), factor VIII (e.g., human factor VIII), mutant factor VIII with varying B-domain lengths (e.g., hFactor VIII N6 or hFactor VIII N6 mutants), and factor IX (e.g., human factor IX). In some embodiments, the transgene encodes a regulatory RNA. In some embodiments, the transgene encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody. In some embodiments, the transgene encodes a protein that can be used to treat diseases selected from the genes listed in Table X below.
[0226] [Table 1] JPEG2023215727000002.jpg210163
[0227] Payload module structure
[0228] GIC: The payload module may include at least one (e.g., one, two, three, or more) transgene sequences, and may further include at least one promoter sequence of the transgene, at least one 5' untranslated sequence of the transgene, at least one 3' untranslated sequence of the transgene, at least one polyadenylation signal sequence or poly-A tail sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, and any combination thereof.
[0229] Referring again to Figure 6, the upper enlarged view shows the structure of an exemplary payload module 430. If a promoter sequence 431, which is an optional component of the transgene, is present, this promoter sequence 431 may contain, and may encode, at least one promoter that may control the expression of the inserted transgene in the target cell. The 5'UTR sequence 432, which is an optional component of the transgene, may contain, and may encode, a sequence that codes for the 5'UTR of the transgene mRNA when the inserted transgene is expressed. The transgene sequence 433 of the payload module may contain, and may encode, at least one transgene sequence that is reverse transcribed and inserted by the gene insertion system of this disclosure. For example, this sequence may contain, and may encode, the ORF of the gene of interest. The 3'UTR sequence 434, which is an optional component of the transgene, may contain, and may encode, at least one 3'UTR of the expressed transgene mRNA. Similarly, polyadenylation signal sequence 435, which is an optional component of the transgene, may contain a polyadenylation signal of the expressed transgene mRNA and may encode this polyadenylation signal. Furthermore, non-coding RNA (ncRNA) processing sequence 436, which is an optional component of the transgene, may contain a termination signal and / or 3' processing signal of the nrRNA expressed by the transgene and may encode such signals.
[0230] Promoter sequence of the transgene and RNAP II 5' UTR sequence If a promoter sequence exists for the transgene, this promoter sequence may include, and may encode, at least one promoter sequence that includes means for promoting the expression of the transgene in the genome of interest. Many such means for promoting the expression of a gene and / or transgene are known in the art, including the insertion of a known promoter sequence into the 5' end of the gene of interest. Those skilled in the art will understand that the type of promoter sequence may be selected based on the transgene and other specific factors used, and therefore any suitable promoter may be used in carrying out this disclosure.
[0231] The exemplary promoters used in this disclosure may be constitutive or inductive promoters. In some embodiments, the promoter sequence of the transgene may include, and may encode, at least one promoter of RNA polymerase I-III (RNAP I, RNAP II, or III). In some embodiments, instead of, or in addition to, the same region as the promoter in at least one transgene may include, and may encode, at least one ribozyme or other motif that enables the release of the RNA transcript of the transgene transtranscribed by rDNA RNAP I of the host cell.
[0232] In some embodiments, at least one promoter sequence of the transgene includes or encodes at least one human U1 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene includes or encodes at least one human U3 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene includes or encodes at least one human U6 snRNA promoter. In some embodiments, at least one promoter sequence of the transgene includes or encodes at least one tRNA promoter.
[0233] If a 5' UTR sequence exists for the transgene, this 5' UTR sequence contains or encodes at least one mRNA 5' UTR of the inserted transgene. Typically, this 5' UTR sequence contains or encodes a sequence that is not translated into amino acid biomolecules by intracellular ribosomes when the inserted transgene is expressed by a cell. Examples of such sequences include, for example, the 5' UTR associated with the transgene in its native state, the transgene's non-native 5' UTR (including sequences derived from the 5' side of the retroelement), a "synthetic" 5' UTR that may not be associated with a known wild-type gene, and any combination thereof.
[0234] Those skilled in the art will understand that the selection of the 5' UTR sequence of an introduced gene depends on the identity of the introduced gene and other specific factors used, and that any known or newly discovered 5' UTR sequence may be suitable for use in the 5' sequence of the introduced gene's payload module.
[0235] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 275 ~ 278 and 282 ~ 283It may contain at least one of the array sequences. 275 ~ 278 and 282 ~ 283 You may code at least one of the arrays 275 ~ 278 and 282 ~ 283 It may be encoded by at least one of the following. In some embodiments, at least one promoter sequence of the transgene is the sequence number. 275 ~ 278 and 282 ~ 283 It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0236] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 275 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0237] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 276 It may also contain sequences having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, which may encode or be encoded by such sequences.
[0238] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 277It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0239] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 278 and include sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0240] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 282 and include sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0241] In some embodiments, at least one promoter sequence of the transgene is the sequence number. 283 and include sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0242] In some embodiments, the GIC payload module includes an RNA polymerase (RNAP) terminator sequence located 5' to the promoter sequence of the transgene. In some embodiments, the RNAP is RNAP I (Pol I), and this stop sequence prevents read-through transcription of Pol I when the GIC payload module is incorporated into the target site of the ribosomal DNA gene. In some embodiments, the RNAP terminator sequence is 5'-AGGTCGACCAGATGTCCGAGGTCGACCAGTTGTCCG-3'(Sequence ID 333 Includes the array indicated by ).
[0243] Transgene sequence The transgene sequence of the payload module contains or encodes at least one target sequence for insertion into the target genome. In this specification, “target sequence” means a biomolecular sequence containing or encoding at least one desired expression product. In some embodiments, the transgene encodes a protein selected from hTERT, hPAH, hFactor VIII, mutant human factor VIII with varying B-domain lengths (e.g., hFactor VIII N6 and hFactor VIII N6 variants), and factor IX (e.g., human factor IX). In some embodiments, the transgene encodes a regulatory RNA. In some embodiments, the transgene encodes an inhibitor of another protein. In some embodiments, the inhibitor is a single-chain antibody. In some embodiments, the transgene encodes a protein that can be used to treat a disease selected from the genes listed in Table X above.
[0244] Any sequence for any purpose may be suitable for the implementation of this disclosure, and is not limited to the source of the sequence (i.e., whether it is a species of organism from which it originates, or whether it is a natural or artificial sequence), nor is it limited to the length of the sequence.
[0245] In some embodiments, at least one transgene sequence is the sequence number. 284 ~ 295 It may contain at least one of the array sequences. 284 ~ 295 You may code at least one of the arrays 284 ~ 295 It may be encoded by at least one of the following. In some embodiments, at least one transgene sequence is the sequence number. 284 ~ 295 It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0246] In some embodiments, at least one transgene sequence is the sequence number. 292 or 293 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0247] In some embodiments, at least one transgene sequence is the sequence number. 294 ~ 295 It may include a sequence having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to at least one of the sequences, which may encode or be encoded by this sequence.
[0248] In some embodiments, at least one transgene sequence is the sequence number. 314 ~ 332It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0249] 3'UTR sequence of the transgene and polyadenylation signal If a 3'UTR sequence exists for a transgene, this 3'UTR sequence contains or encodes at least one mRNA 3'UTR of the inserted transgene. Typically, this 3'UTR sequence contains or encodes a sequence that is not translated into amino acid biomolecules by intracellular ribosomes when the inserted transgene is expressed by a cell. Examples of such sequences include the 3'UTR associated with the transgene in its native state, the non-native 3'UTR of the transgene (including sequences derived from the 3' side of the retroelement), a “synthetic” 3'UTR not associated with a known wild-type gene, and any combination thereof.
[0250] Those skilled in the art will understand that the selection of the 3'UTR sequence of an introduced gene depends on the identity of that introduced gene and other specific factors used, and that any known or newly discovered 3'UTR sequence may be suitable for use in the 3' sequence of the introduced gene's payload module.
[0251] If a polyadenylation signal sequence is present in the transgene, this polyadenylation signal sequence contains or encodes at least one polyadenylation signal of the transgene's mRNA. Any known or novel suitable polyadenylation signal may be used in the template module of this disclosure. For clarification, at least one polyadenylation signal present in or encoded within the inserted transgene provides RNAP II, which adds a poly(A) tail to the transgene's mRNA expression product or ncRNA expression product.
[0252] In some embodiments, at least one 3' UTR sequence of the transgene is sequence number 279 ~ 281 It may include a sequence selected from at least one of the following. In some embodiments, at least one 3' UTR sequence of the transgene is the sequence number. 279 ~ 281 It may also include sequences having at least one of the above and at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology.
[0253] In some embodiments, at least one 3' UTR sequence of the transgene is sequence number 279 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0254] In some embodiments, at least one 3' UTR sequence of the transgene is sequence number 280It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0255] In some embodiments, at least one 3' UTR sequence of the transgene is sequence number 281 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0256] Non-coding RNA (ncRNA) processing sequence of the transgene If an ncRNA processing sequence is present in the transgene, this ncRNA processing sequence includes, or encodes, a sequence that controls the expression or processing of ncRNAs (e.g., transfer RNA (tRNA), rRNA, microRNA, siRNA, snRNA, etc.) expressed from the transgene. In some embodiments, at least one non-coding RNA (ncRNA) processing sequence includes, or encodes, at least one stop signal, at least one 3' processing signal, and any combination thereof of at least one ncRNA expressed from the transgene.
[0257] In some embodiments, at least one ncRNA processing sequence of the transgene includes or encodes at least one MALAT1 3' end processing signal and / or MALAT1 3' end protection signal. In some embodiments, at least one ncRNA processing sequence of the transgene includes or encodes at least one RNA triple-forming end protection structure. In some embodiments, at least one ncRNA processing sequence of the transgene includes or encodes at least one endonuclease recruitment structure, endonuclease recruitment site, or endonuclease recruitment motif. In some embodiments, at least one ncRNA processing sequence of the transgene includes or encodes at least one polythymidine tract. In some embodiments, at least one RNA 3' end termination sequence and / or RNA 3' end processing sequence of the transgene includes the SalI termination box of RNAP I.
[0258] Modularity of payload modules Each component of the GIC:payload module disclosed herein may be used combinatorially and interchangeably to design a 3' module having the functionality required or desired for a particular gene insertion system.
[0259] In some embodiments, at least one GIC:payload module may include and encode at least one transgene sequence. In some embodiments, at least one GIC:payload module may include and encode at least one promoter sequence of the transgene. In some embodiments, at least one GIC:payload module may include and encode at least one 5' UTR sequence of the transgene. In some embodiments, at least one GIC:payload module may include and encode at least one 3' UTR sequence of the transgene. In some embodiments, at least one GIC:payload module may include and encode at least one polyadenylation signal sequence of the transgene. In some embodiments, at least one GIC:payload module may include and encode at least one ncRNA processing sequence of the transgene.
[0260] In some embodiments, the at least one GIC:payload module may include, and encode, at least one transgene sequence, at least one promoter sequence of the transgene, at least one 5' UTR sequence of the transgene, at least one 3' UTR sequence of the transgene, at least one polyadenylation signal sequence of the transgene, and / or at least one ncRNA processing sequence.
[0261] In some embodiments, at least one GIC:payload module is (a) Sequence ID 275 ~ 278 One of the following, and the sequence number. 275 ~ 278At least one promoter sequence and 5'UTR sequence of the transgene, selected from sequences having at least 90% identity with any one of the following (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity); (b) Sequence ID 284 ~ 295 and Scalar 296 ~ 332 Any one of the following, and the sequence number. 284 ~ 295 or 296 ~ 332 At least one transgene sequence selected from sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), or encoding a sequence selected from these sequences, or being encoded by a sequence selected from these sequences; and (c) Sequence ID 279 ~ 281 , and sequence number 279 ~ 281 At least one 3' UTR sequence and polyadenylation signal of the transgene, selected from sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity). It may include any combination of these elements.
[0262] Exemplary GIC: Payload Module In some embodiments, at least one GIC:payload module is array number 296 ~ 332 It may contain at least one array selected from the array array, 296 ~ 332 You may code at least one array selected from the array array 296 ~ 332It may be coded by at least one array selected from. In some embodiments, at least one GIC:payload module is array number 296 ~ 332 It may include, and may encode, or be encoded by, at least one sequence selected from and having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology to, and may encode, or be encoded by, this sequence.
[0263] In some embodiments, at least one GIC:payload module is array number 292 , 293 , 314 or 315 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0264] In some embodiments, at least one GIC:payload module is array number 294 , 295 , 316 or 317 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0265] In some embodiments, at least one GIC:payload module is array number 318 , 319 , 320 or 321It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0266] Modularity of gene insertion constructs (GICs) Each component of the gene insertion construct (GIC) disclosed herein (i.e., GIC:5' module, GIC:3' module, and GIC:payload module) may be used combinatorially and interchangeably to design a GIC having the functionality required or desired for a particular gene insertion system.
[0267] In some embodiments, at least one gene insertion construct (GIC) includes at least one GIC:5' side module. In some embodiments, at least one gene insertion construct includes at least one GIC:payload module. In some embodiments, at least one gene insertion construct includes at least one GIC:3' side module. In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module and at least one GIC:payload module. In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module and at least one GIC:3' side module. In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module, at least one GIC:payload module and at least one GIC:3' side module.
[0268] In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module containing a GIC:5' side module retroelement sequence derived from a retroelement of the same species as the source of the reverse transcriptase recognition sequence of the GIC:3' side module. In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module containing a GIC:5' side module retroelement sequence derived from a retroelement of a different species than the source of the reverse transcriptase recognition sequence of the GIC:3' side module. In some embodiments, at least one gene insertion construct includes at least one GIC:5' side module containing a GIC:5' side module sequence that is generally useful for at least one gene insertion construct containing the reverse transcriptase recognition sequence of the GIC:3' side module and is not naturally found in eukaryotic ecosystems.
[0269] In some embodiments, the gene insertion construct includes a combination of a GIC:5' side module sequence and a GIC:3' side module sequence derived from the source shown in Figure 7. In Figure 7, A1 is a white-throated sparrow (Zonotrichia albicollis), A2 is a zebra finch (Taeniopygia guttata), A3 is a white-throated swan (Tinamus guttatus), A4 is a Galapagos finch (Geospiza fortis), B1 is a northern stickleback (Pungitis pungitis), B2 is a medaka (Oryzias latipes), B3 is a three-spined stickleback (Gasterosteus aculeatus), C1 is a parasitic wasp (Nasonia vitripennis), C2 is a yellow fruit fly (Drosophila melanogaster), C3 is a confused flour beetle (Tribolium castaneum), C4 is a silkworm (Bombyx mori), and C5 is a common fruit fly (Drosophila C6 is Drosophila mercatorum, D1 is Lepidurus couseii, D2 is Triops cancriformis, E1 is Hydra magnipapillata, E2 is Limulus polyphemus, E3 is Adineta vaga, and E4 is Ciona intestinalis.
[0270] In some embodiments, at least one gene insertion construct is (a) at least one GIC:5' side module, with array number 179 ~ 205 , array 179 ~ 205 In comparison, sequences having one, two, or three nucleotide changes or substitutions, sequence number 60~ 153 , Sequence ID 60~ 153Sequences and array keys that have at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) 206 ~ 207 , and sequence number 206 ~ 207 At least one GIC:5' side module selected from sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), or encoding a sequence selected from these sequences, or being encoded by a sequence selected from these sequences; (b) at least one GIC:payload module, with array number 284 ~ 295 and one of 499-525, and the array number. 284 ~ 295 or 296 ~ 318 At least one GIC:payload module selected from sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), or encoding an sequence selected from these sequences, or being encoded by an sequence selected from these sequences; and / or (c) at least one GIC:3' side module, with array number 225 ~ 253 One of the following, and the sequence number. 225 ~ 253 At least one GIC:3' side module that is selected from sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), codes for a sequence selected from these sequences, or is coded by a sequence selected from these sequences. It may include any combination of these elements, and such combinations may be coded, or may be coded by such combinations.
[0271] Exemplary gene insertion constructs (GICs) In some embodiments, at least one gene insertion construct (GIC) is sequence number 284 ~ 295 and may include at least one of 499-525, and the sequence number 284 ~ 295 And may code at least one of 499-525, array index 284 ~ 295 and may be encoded by at least one of 499-525. In some embodiments, at least one gene insertion construct is sequence number 284 ~ 295 and 296 ~ 332 It may include, and may encode, or be encoded by, this sequence, having at least one of the above homology with at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% of the above, and may encode this sequence or be encoded by this sequence.
[0272] In some embodiments, at least one gene insertion construct is sequence number 292 , 293 , 314 or 315 It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0273] In some embodiments, at least one gene insertion construct is sequence number 294 , 295 , 316 or 317It may also contain sequences having at least 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, and may encode or be encoded by such sequences.
[0274] In some embodiments, at least one gene insertion construct is sequence number 318 , 319 , 320 or 321 It may also contain sequences having at least 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 10%, or 5% homology, which may encode or be encoded by such sequences.
[0275] Design and modularity of gene insertion systems The components of the gene insertion system (GIS) disclosed herein (i.e., the reverse transcriptase construct (RTC) and the gene insertion construct (GIC)) may be used combinatorially and interchangeably to design a gene insertion system (GIS) having the required or desired functionality.
[0276] In some embodiments, the gene insertion system may include at least one reverse transcriptase construct (RTC). In some embodiments, the gene insertion system may include at least one gene insertion construct (GIC). In some embodiments, the gene insertion system may include at least one reverse transcriptase construct (RTC) and at least one gene insertion construct (GIC).
[0277] Composition of biomolecules including gene insertion systems The composition of biomolecules including each component of the gene insertion system (GIS) disclosed herein may be selected combinatorially from those disclosed herein to design a gene insertion system (GIS) having the required or desired functionality.
[0278] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as an RNA biomolecule.
[0279] In some embodiments, at least one gene insertion construct (GIC) may be introduced into at least one target as an RNA biomolecule. In some embodiments, at least one gene insertion construct (GIC) may be introduced into at least one target as a linear RNA biomolecule.
[0280] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as an RNA biomolecule, and at least one gene insertion construct (GIC) may be introduced into at least one target as an RNA biomolecule.
[0281] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as an mRNA biomolecule, and at least one gene insertion construct (GIC) may be introduced into at least one target as an RNA biomolecule.
[0282] In some embodiments, at least one reverse transcriptase construct (RTC) and / or at least one gene insertion construct (GIC) may be introduced into at least one subject as a DNA biomolecule. In some embodiments, at least one reverse transcriptase construct (RTC) and / or at least one gene insertion construct (GIC) may be introduced into at least one subject as a plasmid.
[0283] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as an amino acid biomolecule. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as a protein.
[0284] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as an amino acid biopolymer, and at least one gene insertion construct (GIC) may be introduced into at least one target as an RNA biopolymer. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as a plasmid, and at least one gene insertion construct (GIC) may be introduced into at least one target as an RNA biopolymer.
[0285] In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as a plasmid, and at least one gene insertion construct (GIC) may be introduced into at least one target as a plasmid. In some embodiments, at least one reverse transcriptase construct (RTC) may be introduced into at least one target as RNA (e.g., mRNA), and at least one gene insertion construct (GIC) may be introduced into at least one target as a plasmid.
[0286] Paired reverse transcriptase The gene insertion system of the present invention may be optimized for a desired function by controlling the interaction between the GIC and the RTC by designing or selecting the composition of at least one of the gene insertion construct (GIC) and reverse transcriptase construct (RTC) or both included in the gene insertion system. For example, by modifying the composition of the GIC and / or RTC, the efficiency, rate and / or fidelity of full-length payload insertion can be altered, as observed by detecting the insertion using PCR, sequencing and / or the expression of the transgene that is the payload; the sequence specificity and / or chromosomal position in the selection of the target site to which the payload is inserted can be altered, as observed by sequencing, hybridization or other methods that visualize the position of the inserted DNA on the gene; and the selectivity of the RTC to utilize only the GIC administered as a template for reverse transcription can be altered. In this specification, the term “paired reverse transcriptase” means a specific RTC:RT module sequence administered in combination with a specific GIC sequence.
[0287] While we do not wish to be bound by any particular theory, modifications to the interaction between RTC and GIC may be achieved by selecting an RTC:RT module and a GIC:5' side module and / or a GIC:3' side module. For example, the specificity of RTC to GIC may be modified by selecting components derived from retroelements of the same or different species. In this specification, two components of a gene insertion system are considered homogeneous if they originate from retroelements of the same species. Conversely, two components of a gene insertion system are considered heterogeneous if they originate from retroelements of different species.
[0288] In some embodiments, at least one RTC:RT module includes, or encodes, at least one sequence derived from a retroelement of a different species from the at least one retroelement from which the GIC:5' side module sequence and / or the GIC:3' side module sequence originate (hereinafter referred to as a “heterogeneous paired reverse transcriptase”).
[0289] In some embodiments, the retroelement-derived sequences contained in the RTC and GIC are both derived from retroelements of the same species (referred to herein as "homogeneous pair reverse transcriptases").
[0290] In some embodiments, heterogeneous paired reverse transcriptases may have increased specificity compared to homogeneous paired reverse transcriptases.
[0291] In this specification, the term "specificity" means the ability of a pair of reverse transcriptases to efficiently and / or selectively utilize the template RNA intended for the insertion of the transgene.
[0292] In some embodiments, the gene insertion system comprises at least one combination of a GIC and at least one paired reverse transcriptase, such combinations are shown in Figure 7.
[0293] Exemplary gene insertion system In some embodiments, at least one gene insertion system is (a) at least one reverse transcriptase construct (RTC) which is selected from any one of sequence numbers 1 to 59 and sequences having at least 90% identity with any one of sequence numbers 1 to 59 (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), or which encodes a sequence selected from these sequences, or which is encoded by a sequence selected from these sequences; and (b) at least one gene insertion construct (GIC) and sequence number 179 ~ 205 , array 179 ~ 205 In comparison, sequences having one, two, or three nucleotide changes or substitutions, sequence number 60~ 153 , Sequence ID 60~ 153 Sequences and array keys that have at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) 206 ~ 207 , array 206 ~ 207 Sequences and array keys that have at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) 284 ~ 295 or 296 ~ 332 , array 284 ~ 295 or 296 ~ 332 Sequences and array keys that have at least 90% identity (for example, at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity) 225 ~ 253 , and / or array 225 ~ 253 A sequence selected from, or encoding, or being encoded by, one of the sequences having at least 90% identity (e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity), or at least one gene insertion construct (GIC) It may include any combination of these, may code them in any combination, and may be coded by any combination of these.
[0294] III. Formulation and Delivery Mechanism nucleic acid In some embodiments, the reverse transcriptase construct (RTC) or gene insertion construct (GIC) may contain one or more modified nucleotides, including, but not limited to, nucleic acid base modifications, sugar-modified nucleotides, and / or main chain modifications. In some embodiments, the RTC construct or GIC construct may contain a combination of modifications, for example, a combination of nucleic acid base modifications and main chain modifications.
[0295] In some embodiments, the modified nucleotide may be a nucleotide in which the nucleic acid base has been modified. "Modified base" means, but is not limited to, adenine, cytosine, thymine, guanine, uracil, xanthine, inosine, cuosine, and other nucleotide bases that have been modified by the substitution or addition of one or more groups or atoms. In some embodiments, the modified nucleotide may be a nucleotide in which the backbone has been modified.
[0296] The RTC construct and / or GIC construct may include one or more substitutions, insertions and / or additions, deletions, and covalent modifications compared to the reference sequence, in particular the sequence of interest, and such RTC constructs and / or GIC constructs are also included within the scope of the present invention.
[0297] In some embodiments, the RTC construct and / or GIC construct includes one or more post-transcriptional modifications (e.g., capping, cleavage, polyadenylation, splicing, polyA sequence, methylation, acylation, phosphorylation, methylation of lysine and arginine residues, acetylation, nitrosylation of thiol and tyrosine residues, etc.).
[0298] The RTC construct and / or GIC construct may include useful modifications to sugars, nucleic acid bases, or nucleoside bonds (e.g., phosphate bonds, phosphate diester bonds, or phosphodiester backbones).
[0299] In some embodiments, the modifications may include chemically inducible or cell-inducible modifications. For example, some examples of intracellular RNA modifications are reported by Lewis and Pan in “RNA modifications and structures cooperate to guide RNA-protein interactions” in Nat Reviews Mol Cell Biol, 2017, 18:202-210, but are not limited to these.
[0300] In some embodiments, immune evasion may be enhanced by chemical modification of RNA. RNA may be synthesized and / or modified by methods established in the art.
[0301] In some embodiments, at least one RNA construct may contain at least one modified uracil. Examples of uracil modifications include 5-methyluridine, 5-methoxyuridine, pseudouridine, and N 1 Examples include methylpseudridine and / or 2-thiouridine. In some embodiments, at least one RNA construct may contain at least one modified adenosine. An example of an adenosine modification is 2,6-diaminopurine deoxyribonucleotide.
[0302] In some embodiments, sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions or main chain modifications in one or more RNAs include modifications or substitutions of phosphate diester bonds.
[0303] delivery mechanism The gene insertion system (GIS) of the present invention may be introduced into a subject via any delivery mechanism known in the art. In this specification, “delivery mechanism” means a method or composition used to introduce the gene insertion system, its components, or the products of the gene insertion system into a subject. Examples of delivery mechanisms include, but are not limited to, delivery media, direct transfection (using transfection reagents, etc.), transplantation of cells pre-transfected with the gene insertion system, and any combination thereof.
[0304] delivery vehicle In some embodiments, the gene insertion system of the present invention may be formulated with a delivery medium. The delivery medium may generally facilitate transfection of the target cell in vivo or in vitro by protecting each component of the gene insertion system from degradation in the extracellular environment, promoting uptake by the target cell, promoting endosomal escape, and any combination thereof. Examples of delivery mediums include, but are not limited to, nanoparticles such as lipid nanoparticles (e.g., lipid nanoparticles (LNPs), liposomes, and micelles) and non-lipid nanoparticles (e.g., virus-like particles (VLPs) and polymer delivery particles).
[0305] nanoparticles In some embodiments, the delivery medium may contain at least one type of nanoparticle. In this specification, the term "nanoparticle" may generally mean particles with a size of 10 to 1000 nm, for example, 10 nm, 15 nm, 20 nm, 25 nm, 30 nm, 35 nm, 40 nm, 45 nm, 50 nm, 55 nm, 60 nm, 65 nm, 70 nm, 75 nm, 80 nm, 85 nm, 90 nm, 95 nm, 100 nm, 105 nm, 110 nm, 115 nm, 120 nm, 125 nm, 130 nm, 135 nm, 140 nm, 145 nm, 150 nm, 155 nm, 160 nm, 165 nm, 170 nm, 175 nm. m, 180nm, 185nm, 190nm, 195nm, 200nm, 205nm, 210nm, 215nm, 220nm, 225 nm, 230nm, 235nm, 240nm, 245nm, 250nm, 255nm, 260nm, 265nm, 270nm, 275 nm, 280nm, 285nm, 290nm, 295nm, 300nm, 305nm, 310nm, 315nm, 320nm, 325 nm, 330nm, 335nm, 340nm, 345nm, 350nm, 355nm, 360nm, 365nm, 370nm, 375 nm, 380nm, 385nm, 390nm, 395nm, 400nm, 405nm, 410nm, 415nm, 420nm, 42 5nm, 430nm, 435nm, 440nm, 445nm, 450nm, 455nm, 460nm, 465nm, 470nm, 47 5nm, 480nm, 485nm, 490nm, 495nm, 500nm, 505nm, 510nm, 515nm, 520nm, 5 25nm, 530nm, 535nm, 540nm, 545nm, 550nm, 555nm, 560nm, 565nm, 570nm, 5 75nm, 580nm, 585nm, 590nm, 595nm, 600nm, 605nm, 610nm, 615nm, 620nm, 625nm, 630nm, 635nm, 640nm, 645nm, 650nm, 655nm, 660nm, 665nm, 670nm, 675nm, 680nm, 685nm, 690nm, 695nm, 700nm, 705nm, 710nm, 715nm, 720nm, 725nm, 730nm, 735nm, 740nm, 745nm, 750nm, 755nm, 760nm, 765nm, 770nm,Particles of the following sizes may be used: 775nm, 780nm, 785nm, 790nm, 795nm, 800nm, 805nm, 810nm, 815nm, 820nm, 825nm, 830nm, 835nm, 840nm, 845nm, 850nm, 855nm, 860nm, 865nm, 870nm, 875nm, 880nm, 885nm, 890nm, 895nm, 900nm, 905nm, 910nm, 915nm, 920nm, 925nm, 930nm, 935nm, 940nm, 945nm, 950nm, 955nm, 960nm, 965nm, 970nm, 975nm, 980nm, 985nm, 990nm, 995nm, or 1000nm.
[0306] lipid-based particles In some embodiments, the delivery medium may include, but is not limited to, lipid nanoparticles (LNPs), liposomes, micelles, and any combination thereof.
[0307] Lipid nanoparticles In some embodiments, the delivery medium may be lipid nanoparticles (LNPs). Generally, an LNP has an outer lipid layer including a hydrophilic outer surface in contact with a non-LNP environment, a non-aqueous or aqueous internal space (i.e., micelle-like LNPs and vesicle-like LNPs, respectively), and at least one hydrophobic intermembrane space. The LNP membrane may be non-lamellar or lamellar, and may consist of one, two, three, four, five, or more layers. The LNP may be solid or semi-solid. In some embodiments, at least one cargo or payload (such as a gene insertion system (GIS)) may be contained in the internal space, intermembrane space, outer surface, or any combination thereof of the LNP.
[0308] LNPs useful in the present invention are known in the art and generally include ionized (cationic) lipids, phospholipids, cholesterol, and polymer-modified lipids. While we do not wish to be bound by any theory, cholesterol may promote membrane fusion and aid in the stability of LNPs; phospholipids may promote endosomal escape and provide structure to the LNP bilayer; polymer-modified lipids may suppress LNP aggregation and "protect" LNPs from nonspecific endocytosis by immune cells; and ionized (cationic) lipids may promote endosomal escape and form complexes with negatively charged cargo (such as polynucleotides in gene insertion systems).
[0309] In some embodiments, the gene insertion system of the present invention may be incorporated into lipid nanoparticles (LNPs). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid), at least one non-cationic lipid (e.g., a phospholipid), at least one sterol (e.g., cholesterol), at least one polymer-modified lipid (e.g., a PEG lipid), or any combination thereof. In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid), at least one non-cationic lipid (e.g., a phospholipid), at least one sterol (e.g., cholesterol), and at least one polymer-modified lipid (e.g., a PEG lipid). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid), at least one non-cationic lipid, and at least one sterol (e.g., cholesterol). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid), at least one non-cationic lipid (e.g., a phospholipid), and at least one polymer-modified lipid (e.g., a PEG lipid). In some embodiments, the LNP may consist of at least one noncationic lipid (e.g., a phospholipid), at least one sterol (e.g., cholesterol), and at least one polymer-modified lipid (e.g., a PEG lipid). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid) and at least one noncationic lipid (e.g., a phospholipid). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid) and at least one sterol. In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid) and at least one polymer-modified lipid (e.g., a PEG lipid).In some embodiments, the LNP may consist of at least one noncationic lipid (e.g., a phospholipid) and at least one sterol (e.g., cholesterol). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid) and at least one polymer-modified lipid (e.g., a PEG lipid). In some embodiments, the LNP may consist of at least one sterol (e.g., cholesterol) and at least one polymer-modified lipid (e.g., a PEG lipid). In some embodiments, the LNP may consist of at least one cationic lipid (e.g., an ionized cationic lipid). In some embodiments, the LNP may consist of at least one noncationic lipid (e.g., a phospholipid). In some embodiments, the LNP may consist of a sterol (e.g., cholesterol). In some embodiments, the LNP may consist of a polymer-modified lipid (e.g., a PEG lipid).
[0310] The LNPs described herein may be formed using techniques known in the art. For example, a delivery medium on which the gene insertion system is supported can be formed by mixing an acidic aqueous solution containing the gene insertion system with an organic solution containing lipids in a microfluidic channel, but the method is not limited to this.
[0311] Micelle In some embodiments, the delivery medium comprises at least one micelle. In some embodiments, the micelle may be composed of the same components as lipid nanoparticles, but fundamentally different in its manufacturing method. In this specification, “micelle” means a small particle that does not have an aqueous intraparticle space. While we do not wish to be bound by any theory, the intraparticle space of a micelle does not contain lipid head groups, but rather contains hydrophobic tails of lipids constituting the micelle membrane and gene insertion systems (GIS) that may bind to them.
[0312] Liposomes In some embodiments, the delivery medium comprises at least one type of liposome. In some embodiments, the liposome may consist of the same components and amounts as the lipid nanoparticles, but the manufacturing method is fundamentally different. In this specification, “liposome” means a vesicle composed of at least one lipid bilayer surrounding an aqueous internal nanoparticle space. Furthermore, unlike extracellular vesicles, liposomes are generally not derived from progenitor cells / host cells. Examples of liposomes include those comprising multiple concentric bilayers separated by a narrow aqueous space, with diameters of several hundred nanometers (i.e., (large) multilayer vesicles (MLVs)), those with diameters smaller than 50 nm (small monolayer vesicles (SUVs)), and those with diameters of 50 to 500 nm (large monolayer vesicles (LUVs)).
[0313] Exosomes In some embodiments, the delivery medium includes at least one type of exosome. Generally, “exosome” refers to an extracellular vesicle, which is a small membrane-bound organism derived from endocytosis. The exosome membrane generally has a lamellar structure composed of a lipid bilayer and contains an aqueous internal nanoparticle space. Exosomes tend to contain components of the host cell membrane / progenitor cell membrane in addition to pre-defined components. While we do not wish to be bound by any theory, exosomes are generally released from host cells / progenitor cells into the extracellular environment after the multivesicular fusion of the polyvesicle to the cell plasma membrane.
[0314] Virus-like particles In some embodiments, the delivery medium comprises at least one virus-like particle (VLP). Generally, a virus-like particle is a non-infectious vesicle composed primarily of a virus-derived protein capsid, coat, shell, or sheath (all of which are used herein as the same and in the same sense) capable of carrying a gene insertion system (GIS). In some embodiments, the VLP may be synthesized by the self-assembly of a viral capsid protein sequence expressed by an intracellular mechanism to incorporate the gene insertion system (GIS). In some embodiments, the VLP may be formed by preparing the components of the capsid and gene insertion system (GIS) and allowing them to self-assemble without utilizing intracellular mechanisms related to expression.
[0315] Examples of virological families and virus species from which VLPs originate include, but are not limited to, parvoviridae, retroviridae, flaviviridae, paramyxoviridae, adeno-associated viruses, HIV, hepatitis C virus, HPV, bacteriophages, or combinations thereof.
[0316] Polymeric delivery particles In some embodiments, the delivery medium may include at least one polymeric delivery particle. In this specification, “polymeric delivery particle” means a non-aggregating delivery particle composed of a soluble polymer bound to a portion constituting the gene insertion system via various linking groups. In some embodiments, the polymeric delivery particle may include any of the polymers described herein.
[0317] In some embodiments, the delivery medium may include nucleic acid nanoparticles (NANPs). Generally, "nucleic acid nanoparticles" are small particles formed from non-coding nucleic acid sequences that, through interactions, can form three-dimensional structures capable of carrying cargo (e.g., each component of a gene insertion system).
[0318] Encapsulation In some embodiments, the delivery medium may completely encapsulate the gene insertion system disclosed herein. In some embodiments, the delivery medium may partially encapsulate the gene insertion system disclosed herein. In some embodiments, in the final formulation, substantially 0% of the gene insertion system contained in the delivery medium is exposed to the external environment of the delivery medium (i.e., the gene insertion system is completely encapsulated). In some embodiments, the gene insertion system is bound to the delivery medium, but at least a portion of it is exposed to the external environment of the delivery medium.
[0319] In some embodiments, the delivery medium may be characterized by its encapsulation efficiency, that is, by the percentage of the gene insertion system that is not exposed to the external environment of the delivery medium. For clarification, an encapsulation efficiency of about 100% means a delivery medium formulation in which substantially the entire gene insertion system is completely encapsulated by the delivery medium, and an encapsulation efficiency of about 0% means a delivery medium in which substantially no gene insertion system is encapsulated, such as a delivery medium in which the gene insertion system is bound to the external surface of the delivery medium. In some embodiments, the encapsulation efficiency of the delivery medium may be less than about 100%, less than about 95%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than 50%, less than 45%, less than 40%, less than about 35%, less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, or less than 5%. In some embodiments, the encapsulation efficiency of the delivery medium is approximately 90-100%, 80-100%, 70-100%, 60-100%, 50-100%, 40-100%, 30-100%, 20-100%, 10-100%, 80-90%, 70-90%, 60-90%, 50-90%, 40-90%, 30-90%, 20-90%, 10-90%, 70-80%, and 60-90%. It may also be 0-80%, approximately 50-80%, approximately 40-80%, approximately 30-80%, approximately 20-80%, approximately 10-80%, approximately 60-70%, approximately 50-70%, approximately 40-70%, approximately 30-70%, approximately 20-70%, approximately 10-70%, approximately 40-50%, approximately 30-50%, approximately 20-50%, approximately 10-50%, approximately 30-40%, approximately 20-40%, approximately 10-40%, approximately 20-30%, approximately 10-30%, or approximately 10-20%.
[0320] Physical properties of nanoparticles used as a delivery medium In some embodiments, the delivery medium is characterized by its shape. In some embodiments, the delivery medium may be substantially spherical, substantially cylindrical (i.e., cylindrical), or substantially disc-shaped, but is not limited to these.
[0321] In some embodiments, the delivery medium is characterized by its size. In some embodiments, the size of the delivery medium can be defined as its diameter. As used herein in relation to the size of the delivery medium, “diameter” means the diameter of the largest circular cross-section of the delivery medium. In some embodiments, the diameter of the delivery medium may be between 30 nm and about 150 nm. For example, the diameter of the delivery medium is approximately 40-150 nm, approximately 50-150 nm, approximately 60-150 nm, approximately 70-150 nm, approximately 80-150 nm, approximately 90-150 nm, approximately 100 nm or more, approximately 110-150 nm, approximately 120-150 nm, approximately 130-150 nm, approximately 140-150 nm, approximately 30-30-140 nm, approximately 40-140 nm, approximately 50-140 nm, approximately 60-140 nm, approximately 70-140 nm, approximately 80-140 nm, and approximately 90-140 nm. , approximately 100~140nm, approximately 110~140nm, approximately 120~140nm, approximately 130~140nm, approximately 140~140nm, approximately 30~140nm, approximately 40~130nm, approximately 50~130nm, approximately 60~130nm, approximately 70~ 130nm, about 80-130nm, about 90-130nm, about 100-130nm, about 110-130nm, about 120-130nm, about 30-120nm, about 40-120nm, about 50-120nm, about 60-120nm, about 70~120nm, approx. 80~120nm, approx. 90~120nm, approx. 100~120nm, approx. 110~120nm, approx. 30~110nm, approx. 40~110nm, approx. 50~110nm, approx. 60~110nm, approx. 70~110nm , about 80-110nm, about 90-110nm, about 100-110nm, about 30-100nm, about 40-100nm, about 50-100nm, about 60-100nm, about 70-100nm, about 80-100nm, about 90-100nm m may be approximately 30-90nm, approximately 40-90nm, approximately 50-90nm, approximately 60-90nm, approximately 70-90nm, approximately 80-90nm, approximately 30-80nm, approximately 40-80nm, approximately 50-80nm, approximately 60-80nm, approximately 70-80nm, approximately 30-70nm, approximately 40-70nm, approximately 50-70nm, approximately 60-70nm, approximately 30-60nm, approximately 40-60nm, approximately 50-60nm, approximately 30-50nm, approximately 40-50nm, or approximately 30-40nm.
[0322] In some embodiments, a group of delivery media, for example, all delivery media derived from a single formulation, may be characterized by measuring the uniformity of the physical properties (e.g., size, shape, or mass) of the particles contained in the group. In some embodiments, this uniformity may be expressed as the polydispersity index (PI) of the group. In some embodiments, this uniformity may be expressed as the imbalance (D) of the group. In this specification, the terms "polydispersity index" and "imbalance" may be used interchangeably and have the same meaning.
[0323] In some embodiments, the PI of a group of delivery media derived from a certain formulation is about 0.1 to 1. In some embodiments, the PI of a group of delivery media derived from a certain formulation is about 0.1 to 1, about 0.1 to 0.8, about 0.1 to 0.6, about 0.1 to 0.4, about 0.1 to 0.2, about 0.2 to 1, about 0.2 to 0.8, about 0.2 to 0.6, about 0.2 to 0.4, about 0.4 to 1, about 0.4 to 0.8, about 0.4 to 0.6, about 0.6 to 1, about 0.6 to 0.8, or about 0.8 to 1. In some embodiments, the PI of a group of delivery media derived from a certain formulation is less than about 1, less than about 0.5, less than about 0.4, less than about 0.3, less than about 0.2, or less than about 0.1.
[0324] Targeting of delivery In some embodiments, a delivery medium formulated to include the gene insertion system of the present invention may facilitate the localization of the gene insertion system to any of the target regions, tissues, cells, or physiological systems described herein (i.e., the delivery medium “targets” a particular site). In some embodiments, targeting may be achieved by components of the delivery medium contained in the formulation. In some embodiments, the delivery medium may contain a targeting agent.
[0325] Targeting agent In some embodiments, the delivery medium may contain at least one targeting agent. In this specification, the term “targeting agent” may, in some embodiments, mean a moiety, compound, antibody (e.g., a moiety that targets a particular cell or a particular type of cell) that specifically binds to a particular type or classification of cells and / or other particular types of compounds. In some embodiments, the targeting agent may have affinity for (i.e., be specific to) the surface of a particular target cell, a targeted cell surface antigen, a targeted cell receptor, or a combination thereof.
[0326] In some embodiments, “targeting agent” may mean an agent that can target a delivery medium to a particular type or classification of cells by exerting a specific action (e.g., a cleavage action) when exposed to a particular type or classification of substance and / or cells.
[0327] In some embodiments, the term “targeting agent” may mean a drug that is part of the delivery medium and does not necessarily have specificity for a particular type or classification of cells, but plays a certain role in the specificity of the delivery medium to its target.
[0328] In some embodiments, the inclusion of at least one targeting agent in the delivery medium may increase the efficiency (e.g., total amount or rate) of the gene insertion system delivered by the delivery medium being taken up by cells. In some embodiments, the inclusion of at least one targeting agent in the delivery medium may increase the specificity (e.g., total amount or rate) of the gene insertion system delivered by the delivery medium being taken up by cells. In this specification, “specificity” means that the efficiency of uptake by target cells is higher than that by non-target cells.
[0329] In some embodiments, suitable targeting agents include, but are not limited to, one or more of the following: small molecular weight targeting agents (e.g., glycan moieties), antibodies, antibody-like molecules, peptides, vitamins (e.g., folic acid), sugars (e.g., lactose or galactose), artificial affinity molecules (e.g., peptide mimes or aptamers), antibody fragments, single-chain variable region fragments (scFv), cell surface receptors (e.g., T cell receptors (TCRs), B cell receptors (BCRs), or chimeric antigen receptors (CARs)), and any combination thereof.
[0330] In some embodiments, cell surface molecules of the target cell may be used as cell surface antigens that can be targeted by the targeting agent. Suitable cell surface molecules include, but are not limited to, proteins, sugars, lipids, or other antigens on the cell surface. In some embodiments, the cell surface antigens migrate into the cell.
[0331] In some specific embodiments, the delivery medium may contain two or more targeting agents.
[0332] In some embodiments, at least one targeting agent may be incorporated into the lipid membrane of the nanoparticle. In some embodiments, at least one targeting agent may be presented on the outer surface of the nanoparticle. In some embodiments, at least one targeting agent may be bound to the lipid component of the nanoparticle. In some embodiments, at least one targeting agent may be bound to the polymer component of the nanoparticle. In some embodiments, polymer-modified lipids can be formed to form a delivery medium by copolymerizing a monomer containing a portion of the targeting agent (e.g., a polymerizable derivative of the targeting agent, e.g., an (alkyl)acrylic acid derivative of a peptide). In some embodiments, at least one targeting agent may be linked to the nanoparticle by being tethered to the nanoparticle via hydrophobic or hydrophilic interactions, thereby linking the at least one targeting agent to the membrane structure of the nanoparticle and the aqueous environment inside or outside the nanoparticle. In some embodiments, at least one targeting agent is bound to the peptide / protein component of the membrane structure of the nanoparticle. In some embodiments, at least one targeting agent is bound to a suitable linker portion bound to a component of the membrane structure of the nanoparticle. In some embodiments, any combination of forces or bonds can bind the targeting agent to the nanoparticle.
[0333] In some embodiments, one or more targeting agents may be linked to at least one polymer of the delivery medium via a linking portion. In some embodiments, this linking portion may be a cleavable linking portion (e.g., a linking portion containing a cleavable bond). In some embodiments, the linking portion may contain a bond that may be cleaved by a specific enzyme (e.g., a phosphatase or protease). In some embodiments, the linking portion may contain a bond that may be cleaved by a change in intracellular pH, redox potential or other intracellular parameters. In some embodiments, the linking portion may contain a bond that may be cleaved by exposure to a matrix metalloproteinase (MMP).
[0334] Direct transfection In some embodiments, the gene insertion system (GIS) disclosed herein may be directly transfected into target cells without using the delivery medium. In some embodiments, the gene insertion system (GIS) disclosed herein may be transfected into target cells using any technique known in the art. Such techniques include, but are not limited to, chemical transfection methods (e.g., calcium phosphate contact) and physical transfection methods (e.g., electroporation, microinjection, gene gun delivery). In some embodiments, direct transfection may be carried out using lipid-based transfection reagents such as Lipofectamine, Lipofectamine 2000, and any combination thereof, but are not limited to those listed below.
[0335] Transplantation of transfected cells In some embodiments, the gene insertion system of the present invention may be introduced into a cell population in vitro (for example, via direct transfection as described herein) and subsequently transplanted into a subject. In some embodiments, the cell population for transplantation may be stem cells. In some embodiments, the cell population for transplantation may be cells derived from the subject. In some embodiments, transplantation may be carried out by methods known in the art.
[0336] IV. Pharmaceutical Compositions and Routes of Administration The present invention provides pharmaceutical compositions for administering a gene insertion system (GIS) to a target. In some embodiments, the present invention provides pharmaceutical compositions for use as pharmaceuticals in the treatment of therapeutic indications. In some embodiments, the pharmaceutical composition comprises at least one active ingredient (e.g., the gene insertion system (GIS) of the present invention) and at least one pharmaceutically acceptable additive, auxiliary agent, carrier, diluent, or any combination thereof. In some embodiments, the pharmaceutical composition is prepared as a formulation suitable for at least one route of administration. In some embodiments, the pharmaceutical composition is prepared as a formulation such that a predetermined dose of at least one active ingredient (e.g., the gene insertion system (GIS)) is delivered. It may also be prepared as a formulation for delivery on a predetermined schedule.
[0337] In this specification, the term “pharmaceutical composition” means a composition comprising at least one active ingredient and which may further comprise one or more pharmaceutically acceptable excipients. In this specification, the term “active ingredient” typically means a gene insertion system (GIS) as described herein, a gene payload delivered by a gene insertion system (GIS) for insertion into a genome of interest, or an expression product of a gene payload delivered by a gene insertion system (GIS) as described herein.
[0338] Pharmaceutical preparations and pharmaceutical compositions The gene insertion system of the present invention may achieve the following by being formulated using one or more additives. (1) Increase the stability of the gene insertion system or the delivery mechanism including the gene insertion system. (2) Increased transfection or transduction into cells, (3) Continuous or delayed introduction of a gene insertion system into target cells, (4) Modification of in vivo distribution (e.g., targeting of gene insertion systems to specific tissues or specific types of cells), (5) Increased expression of the encoded gene, (6) Modification of the release properties of the encoded protein, and / or (7) A gene insertion system and / or controllable expression of its payload.
[0339] Pharmaceutical formulations may include, but are not limited to, saline solution, liposomes, lipid nanoparticles, polymers, peptides, proteins, cells transfected with gene insertion systems (e.g., for transfer to or transplantation into a target), and any combination thereof.
[0340] In some embodiments, formulations of the pharmaceutical compositions described herein may be prepared by any method known or hereafter developed in the field of pharmacology. Generally, methods for preparing such formulations include the step of combining the active ingredient with an additive and / or one or more other auxiliary components.
[0341] The gene insertion systems and pharmaceutical compositions described herein may be prepared by any method known or to be developed in the field of pharmacology. Generally, the preparation of such preparations includes a step of mixing the active ingredient with an additive and / or one or more other auxiliary ingredients, and, depending on necessity and / or desire, a step of dividing the resulting mixture, a step of molding, and / or a step of packaging into a desired single-dose or multi-dose dosage form.
[0342] The pharmaceutical compositions described herein may be prepared, packaged and / or sold as a single unit dose and / or multiple unit doses. In this specification, the term “unit dose” means a specific amount of a pharmaceutical composition containing a predetermined amount of the active ingredient. Generally, the amount of the active ingredient is equal to and / or a convenient fraction of such a dose, for example, half or one-third of such a dose.
[0343] In some embodiments, the additives are approved for human and veterinary use. In some embodiments, the additives may be approved by the U.S. Food and Drug Administration. In some embodiments, the additives may conform to the standards of the United States Pharmacopeia (USP), European Pharmacopeia (EP), British Pharmacopeia, and / or International Pharmacopoeia. In some embodiments, the pharmaceutically acceptable purity of the additive may be at least 100%, at least 99%, at least 98%, at least 97%, at least 96%, or 95%. In some embodiments, the additives may be pharmaceutical grade.
[0344] In some embodiments, the relative amounts of pharmaceutically acceptable additives, active ingredients, and / or additional ingredients may vary considerably from one pharmaceutical composition to another. In some embodiments, the relative amounts may vary considerably depending on the body length, condition, and / or identity of the person being treated. In some embodiments, the relative amounts may vary considerably depending on the route of administration of the pharmaceutical composition of the present invention. For example, the pharmaceutical composition of the present invention may contain 0.1% to 100% (e.g., 0.1% to 99%, 0.5% to 50%, 1% to 30%, 5% to 80%, or at least 80% (w / w)) of the active ingredient.
[0345] Additives, diluents, and inert components In some embodiments, the pharmaceutical compositions of the present invention may contain additives known or novel in the art. Examples of suitable additives include, but are not limited to, all kinds of preservatives, isotonic agents, thickeners, emulsifiers, solvents, dispersion media, diluents or other liquid solvents, dispersion or suspension aids, surfactants, and combinations thereof. In some embodiments, the additives may be selected to suit a desired specific dosage form.
[0346] In some embodiments, the formulations described herein may contain at least one inactive component. In this specification, the term "inactive component" means one or more substances contained in the formulation that do not contribute to the activity of the active ingredient of the pharmaceutical composition of the present invention. In some embodiments, all or some of the inactive components contained in the pharmaceutical composition of the present invention may be approved by the U.S. Food and Drug Administration (FDA), or all of the inactive components contained in the pharmaceutical composition of the present invention may not be approved by the U.S. Food and Drug Administration (FDA).
[0347] In some embodiments, the pharmaceutical formulations disclosed herein may contain a cation or anion. In some embodiments, the pharmaceutical formulations disclosed herein contain Ca 2+ Zn 2+ Mn 2+ Cu 2+ Mg + This includes, but is not limited to, metal cations such as any combination thereof. In some embodiments, the pharmaceutical formulations of this disclosure may include polymers complexed with metal cations.
[0348] In some embodiments, the pharmaceutical compositions of the present disclosure may contain one or more pharmaceutically acceptable salts. In this specification, the term “pharmaceutically acceptable salt” means a derivative of the compound of the present disclosure modified by converting an acidic or basic moiety present in the parent compound into a salt form (for example, by reacting a free base with a suitable organic acid). Examples of pharmaceutically acceptable salts of the present invention include conventional non-toxic salts of the parent compound formed using a non-toxic inorganic or organic acid. Furthermore, pharmaceutically acceptable salts include, but are not limited to, alkali or organic salts of acidic residues such as carboxylic acids, and inorganic or organic acid salts of basic residues such as amines.
[0349] In some embodiments, the pharmaceutical compositions of the present disclosure may contain at least one solvent. In some embodiments, when the solvent is water, the solvate is generally referred to as a “hydrate.”
[0350] Route of administration The gene insertion system (GIS) of the present invention, for example, a pharmaceutical composition comprising the gene insertion system (GIS) described herein, may be administered by any delivery route that allows the gene insertion system (GIS) to be incorporated into the target cells. Acceptable routes of administration include transaural administration (administration into or via the ear), biliary irrigation, buccal administration (administration into the inner cheek), cardiac irrigation, sacral block, conjunctival administration, transdermal administration, transdental administration (administration into one or more teeth), intracoronal administration, diagnostic administration, ear drops, electroosmosis, endometrial administration, intrasinal administration, intratracheal administration, enema, intraintestinal administration (administration into the intestinal tract), epidermal administration (puncture from the skin), epidural administration (administration into the epidural space), extraamniotic administration, administration via an extracorporeal circulation device, and ophthalmic administration (on the conjunctiva). Administration of (intravenous), gastrointestinal administration, hemodialysis, infiltration administration, nasal aspiration, interstitial space administration, intraperitoneal administration, amniotic membrane administration, intra-arterial administration, intra-arterial administration, intra-bile duct administration, intra-bronchial administration, intracapsular administration, intracardiac administration, intrachondral administration, cauda equina administration, intracauda equina injection, intracavitary injection, intracavernosal sinus administration (administration to the base of the penis), intracerebral administration, intraventricular administration, intracisional administration, cisterna magna administration (administration into the cisterna magna / cerebellar-medulla oblongata cistern), Intracorneal administration (administration into the cornea), intracoronary artery administration (administration into the coronary artery), intracavernosal administration (administration into the expandable space of the corpus cavernosum of the penis), intradermal administration (administration into the skin itself), intradiscal administration (administration into the intervertebral disc), intraductal administration (administration into the glandular duct), intraduodenal administration (administration into the duodenum), intradural administration (administration into or subdural), intraepidermal administration (administration into the epidermis), intraesophageal administration (administration into the esophagus), intragastric administration (administration into the stomach), intragingival administration (administration into the gums), intraileal administration (administration into the distal part of the small intestine), Intra-lesional administration (administration into or direct introduction into a local lesion), intracavitary administration (administration into the lumen of a tube), intralymphatic administration (administration into a lymphatic vessel), intramedullary administration (administration into the medullary cavity of bone), intrameningeal administration (administration into the meninges), intramuscular administration (administration into the muscle), intramyocardial administration (administration into the myocardium), intraocular administration (administration into the inside of the eye), intraosseous injection (administration into the bone marrow), intraovarian administration (administration into the ovary), intraparenchymal administration (administration into brain tissue), intrapericardial administration (administration into the pericardium), intraperitoneal administration (injection or injection into the peritoneal cavity),Intrapleural administration (administration into the pleura), intraprostatic administration (administration into the prostate), intrapulmonary administration (administration into the lung or its bronchi), intracavitary administration (administration into the paranasal sinuses or periorbital space), intraspinal administration (administration into the spinal column), intrabursal administration (administration into the synovial fluid space of a joint), intratendinous administration (administration into the tendon), intratesticular administration (administration into the testis), intrathecal administration (administration into the spinal canal), intrathecal administration (administration into the cerebrospinal fluid at any point on the cerebrospinal axis), intrathoracic administration (administration into the thoracic cavity), intraluminal administration (administration into the lumen of an organ), intratumoral administration (administration into a tumor) , intra-ear administration (administration into the middle ear), intrauterine administration, intravaginal administration, intravascular administration (administration into one or more blood vessels), intravenous administration (administration into a vein), intravenous bolus, intravenous drip infusion, intracardiac administration (administration into the ventricle), intravesical instillation, intravitreal administration (administration into the inside of the eyeball), iontophoresis (a method of introducing ions of a soluble salt into living tissue using electric current), lavage (immersion lavage or flushing of an open wound or body cavity), laryngeal administration (direct administration into the larynx), transnasal administration (administration through the nasal cavity), transnasogastric administration (administration through the nasal cavity into the stomach) Methods of administration include: nerve block, occlusive therapy (administering via the external route and then covering the site with a covering material), transocular administration (administering to the external eye), oral administration (administering via the oral cavity), oropharyngeal administration (direct administration to the oral cavity and pharynx), parenteral administration, transcutaneous administration, periarticular administration, peridural administration, perineurial administration, periodontal administration, photopheresis, rectal administration, respiratory administration (administering into the respiratory tract by oral or nasal inhalation to obtain local or systemic effects), retrobulbar administration (administering behind the pons or behind the eyeball), and soft tissue administration. Administration methods include, but are not limited to, intraoral, subarachnoid, subconjunctival, subcutaneous (administration under the skin), suboral, sublingual, submucosal, topical, transdermal, transdermal (diffusion through intact skin for systemic distribution), transmucosal (diffusion through mucosa), transplacental (administration through the placenta or via the placenta), transtracheal (administration through the tracheal wall), transtympanic (administration through the tympanic cavity or through the tympanic cavity), transvaginal, ureteral (administration into the ureter), urethral (administration into the urethra), vaginal, and spinal administration.
[0351] In some embodiments, the pharmaceutical compositions of the present disclosure may be administered in a manner that allows them to cross vascular barriers, the blood-brain barrier, or other epithelial barriers. The gene insertion system may be administered in any suitable form, including, but not limited to, solutions, suspensions, solids, solids suitable for dissolution in solution, solids suspendable in solution, and any combination thereof.
[0352] In some embodiments, the gene insertion system may be delivered to the target via multiple administration routes. It may be administered to two, three, four, five, or six or more sites in the target.
[0353] In some embodiments, the gene insertion system may be delivered to the target via a single administration route.
[0354] In some embodiments, a bolus injection may be used to target and administer the gene insertion system.
[0355] In some embodiments, the gene insertion system may be administered to the target over several minutes, hours, or days using a continuous delivery method (i.e., intravenous infusion). The infusion rate may be modified depending on delivery parameters such as the characteristics of the target, the desired distribution, and the formulation used, but is not limited to the following.
[0356] In some embodiments, the gene insertion system may be delivered by an intramuscular delivery route, such as subcutaneous injection or intravenous injection, but is not limited to the following.
[0357] In some embodiments, the gene insertion system may be delivered by oral administration, such as by gastrointestinal administration or buccal administration, but is not limited to the following.
[0358] In some embodiments, the gene insertion system may be delivered by an intraocular delivery route, such as intravitreal injection or instillation of eye drops, but is not limited to the following.
[0359] In some embodiments, the gene insertion system may be delivered by an intranasal delivery route, such as a nasal spray or nasal drops, but is not limited to the following.
[0360] In some embodiments, the gene insertion system may be administered to the subject by peripheral injection, such as intramuscular injection, intraperitoneal injection, intravenous injection, conjunctival injection, or joint injection, but is not limited to the following.
[0361] In some embodiments, the gene insertion system may be delivered by injection into the cerebrospinal fluid pathway, such as intrathecal or intraventricular administration, but is not limited to the following.
[0362] In some embodiments, the gene insertion system may be delivered by a systemic delivery route, such as intravascular administration, but is not limited to the following.
[0363] In some embodiments, the gene insertion system may be administered to the subject by intracellular administration.
[0364] In some embodiments, the gene insertion system may be administered to the subject by topical administration.
[0365] In some embodiments, the gene insertion system may be administered to the subject by intracranial delivery.
[0366] In some embodiments, the gene insertion system may be administered to the subject by intramuscular injection.
[0367] In some embodiments, the gene insertion system may be administered to the subject intravenously.
[0368] In some embodiments, the gene insertion system may be administered to the subject by subcutaneous injection.
[0369] In some embodiments, the gene insertion system may be delivered by two or more administration routes.
[0370] Injectable and parenteral administration In some embodiments, the pharmaceutical compositions described herein may be administered parenterally. Liquid dosage forms for parenteral and oral administration include, but are not limited to, pharmaceutically acceptable liquids, emulsions, microemulsions, elixirs, suspensions, and / or syrups. In addition to the active ingredient, the liquid dosage forms may contain inert diluents commonly used in the art, such as solubilizers, water or other solvents, and emulsifiers (e.g., polyethylene glycol, propylene glycol, 1,3-butylene glycol, tetrahydrofurfuryl alcohol, isopropyl alcohol, ethyl alcohol, ethyl carbonate, ethyl acetate, benzyl alcohol, benzyl benzoate, dimethylformamide, oils, glycerol, and sorbitan fatty acid esters), and any combination thereof. Examples of oils include cottonseed oil, peanut oil, corn oil, germ oil, olive oil, castor oil, and sesame oil, as well as mixtures thereof. In some embodiments, the pharmaceutical compositions of the present disclosure include solubilizers such as alcohols, oils, glycols, CREMOPHOR®, modified oils, polysorbates, polymers, cyclodextrins, and / or combinations thereof. In some embodiments, surfactants such as hydroxypropylcellulose are included.
[0371] In some embodiments, the injectable preparation may comprise a sterile aqueous suspension or a sterile oily suspension for injection. The sterile solution for injection may be formulated by known techniques using a suitable wetting agent, dispersant, and / or suspending agent. The sterile injectable preparation may be a sterile suspension, sterile solution, and / or emulsion for injection in a parenterally acceptable, non-toxic diluent and / or solvent. In some embodiments, the sterile injectable preparation may be a 1,3-butanediol solution. In some embodiments, acceptable media and solvents include, but are not limited to, Ringer's solution (United States Pharmacopeia), water, isotonic saline, and sterile non-volatile oils. In some embodiments, non-volatile oils may be non-irritating non-volatile oils (e.g., synthetic monoglycerides or synthetic diglycerides). In some embodiments, fatty acids such as oleic acid may also be used in the preparation of the injectable preparation.
[0372] In some embodiments, the injectable formulation may be sterilized by filtration through a bacterial capture filter and / or by the addition of a sterilizing agent. In some embodiments, the sterilizing agent may be in the form of a sterile solid composition that can be dissolved or dispersed in a sterile injectable medium such as sterile water before use.
[0373] To prolong the effects of the active ingredient, it is often desirable to delay the absorption of the active ingredient from subcutaneous or intramuscular injection. In some embodiments, the absorption of a parenterally administered pharmaceutical composition can be delayed by dissolving or suspending the pharmaceutical composition of this disclosure in an oily solvent. In some embodiments, the absorption of the active ingredient may be delayed by using a liquid suspension of a poorly water-soluble amorphous or crystalline component. The absorption rate of the active ingredient depends on its solubility, which depends on the size and morphology of the crystals.
[0374] Oral administration In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be administered orally. Examples of solid dosage forms for oral administration include tablets, capsules, powders, pills, and granules. Generally, in solid dosage forms, the active ingredient is mixed with at least one pharmaceutically acceptable inert additive, such additives include, but are not limited to, dicalcium phosphate or sodium citrate; binders (e.g., carboxymethylcellulose, alginates, gelatin, polyvinylpyrrolidone, sucrose, and gum arabic); fillers or bulking agents (e.g., starch, lactose, sucrose, glucose, mannitol, and silicic acid); disintegrants (e.g., agar, calcium carbonate, potato starch or tapioca starch, alginic acid, certain silicates, and sodium carbonate); absorption enhancers (e.g., quaternary ammonium compounds); humectants (e.g., glycerol); liquid retarders (e.g., paraffin); absorbents (e.g., kaolin and bentonite clay); wetting agents (e.g., cetyl alcohol and glyceryl monostearate); lubricants (e.g., talc, calcium stearate, magnesium stearate, solid polyethylene glycol, sodium lauryl sulfate); and any combination thereof. In the case of tablets, capsules, and pills, these dosage forms may contain a buffering agent.
[0375] Liquid dosage forms for oral administration may contain the parenteral additives described above. In addition to inert diluents, oral compositions may also contain auxiliary agents such as emulsifiers, wetting agents, suspending agents, flavoring agents, sweeteners, and / or fragrances.
[0376] Topical or transdermal administration In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be formulated for topical administration. The skin can be an ideal target site for delivery due to its easy accessibility. In some embodiments, routes of delivery to or via the skin of the pharmaceutical compositions described herein include, but are not limited to, topical application (e.g., for cosmetic and / or topical / regional treatment), intradermal injection (e.g., for cosmetic and / or topical / regional treatment), and systemic delivery (e.g., for the treatment of skin diseases affecting both the cutaneous and extradermal regions).
[0377] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be delivered using various covering bandages (e.g., adhesive bandages or wound dressings) to effectively and / or conveniently carry out the methods described herein. In some embodiments, the covering or bandage may contain a sufficient amount of the pharmaceutical composition described herein so that the user can perform multiple treatments themselves.
[0378] Dosage forms for topical and / or transdermal administration include lotions, creams, ointments, gels, sprays, pastes, powders, solutions, inhalants, and / or patches. Generally, dosage forms for topical and / or transdermal administration may be formulated under sterile conditions by mixing the active ingredient with pharmaceutically acceptable additives, buffers, and / or preservatives as needed.
[0379] In some embodiments, transdermal patches may be used. Transdermal patches may have the further advantage of being able to deliver the pharmaceutical compositions described herein to the body in a controlled manner. Generally, transdermal patches may be prepared by dissolving and / or dispersing the pharmaceutical compositions described herein in a suitable medium. In some embodiments, the delivery rate may be controlled by dispersing the pharmaceutical compositions of the present invention in a polymer matrix and / or gel, by providing a membrane that controls the release rate, or by a combination thereof.
[0380] In some embodiments, formulations suitable for topical administration include liquid preparations and / or semi-liquid preparations (e.g., liniments and lotions), oil-in-water emulsions and / or water-in-oil emulsions (e.g., ointments, creams and / or pastes), liquid preparations and / or suspensions, and any combination thereof.
[0381] Intraocular or intraaural administration In some embodiments, the pharmaceutical compositions described herein may be in the form of formulations suitable for transocular, transaural, or both. Generally, such formulations may be in the form of eye drops and / or ear drops, specifically including, but not limited to, liquid and / or suspension formulations in which the active ingredient is added to an aqueous liquid additive and / or an oily liquid additive. In some embodiments, such eye and / or ear drops may contain salts, buffers, one or more other additional components described herein, and combinations thereof. In some embodiments, formulations suitable for transocular administration include liposome preparations and / or active ingredients in microcrystalline form. In some embodiments, the pharmaceutical compositions described herein may be administered subretinally.
[0382] Transpulmonary administration In some embodiments, the pharmaceutical compositions described herein may be in the form of formulations suitable for intrapulmonary administration. In some embodiments, intrapulmonary administration is performed via the buccal cavity. In some embodiments, the pharmaceutical compositions described herein may include dry particles containing the active ingredient. In some embodiments, the diameter of the dry particles for intrapulmonary administration may be in the range of about 0.5 to 7 nm or about 1 to 6 nm.
[0383] In some embodiments, the pharmaceutical compositions described herein may be administered using a self-propelled solvent / powder supply container. Generally, the active ingredient may be dissolved and / or suspended in a low-boiling point propellant in a sealed container. In some embodiments, the pharmaceutical compositions described herein may be in the form of a dry powder administered using an apparatus equipped with a dry powder reservoir in which such powder is dispersed by a flow of propellant. In some embodiments in which a dry powder is used, the powder may contain particles in which at least 98% by weight have a diameter greater than 0.5 nm and at least 95% by number have a diameter less than 7 nm. In some embodiments, at least 95% by weight have a diameter greater than 1 nm and at least 90% by number have a diameter less than 6 nm. In some embodiments, the dry pharmaceutical composition containing the powder may contain a diluent in the form of a solid fine powder (e.g., sugars) and may be conveniently provided in unit dose form.
[0384] In some embodiments, low-boiling-point propellants include liquid propellants having a boiling point below 65°F at atmospheric pressure. In some embodiments, the propellant content in the pharmaceutical composition of the present invention may be 50% to 99.9% (w / w), and the active ingredient content in the pharmaceutical composition of the present invention may be 0.1% to 20% (w / w). In some embodiments, the propellant may contain additional components, specifically, but are not limited to, liquid nonionic surfactants, solid anionic surfactants, solid diluents (for example, solid diluents having the same order of magnitude particle size as the particles containing the active ingredient), and any combination thereof.
[0385] In some embodiments, the pharmaceutical composition formulated for transpulmonary delivery may be in the form of droplets of a liquid, a suspension, or a combination thereof. When such formulations are prepared, packaged, and / or sold as a liquid, a suspension, or a combination thereof, they may be administered using a sprayer and / or nebulizer. In some embodiments, these liquids and / or suspensions may be sterile. Exemplary liquids and / or suspensions include aqueous alcohol compositions and / or diluted alcohol compositions. In some embodiments, the pharmaceutical composition formulated for transpulmonary delivery may contain fragrances (e.g., sodium saccharin), volatile oils, surfactants, buffers, preservatives (e.g., methyl hydroxybenzoate), and any combination thereof. In some embodiments, the average diameter of the droplets delivered by the transpulmonary administration route may be in the range of about 0.1 nm to about 200 nm.
[0386] Intranasal administration, intranasal administration, or buccal administration In some embodiments, the pharmaceutical compositions described herein may be administered intranasally, intranasally, or both. In some embodiments, the pharmaceutical compositions for intranasal delivery may contain the additives described herein for intrapulmonary delivery. In some embodiments, the pharmaceutical compositions for intranasal delivery may contain a coarse powder having an average particle size of about 0.2 μm to 500 μm. In some embodiments, the pharmaceutical compositions for intranasal delivery may be administered by rapid inhalation through the nasal cavity from a powder container held near the nose, i.e., by snorting. Exemplary pharmaceutical formulations may contain about 0.1% (w / w) to 100% (w / w) of the active ingredient and may further contain one or more additional ingredients described herein.
[0387] In some embodiments, the pharmaceutical compositions described herein may be formulations suitable for buccal administration, such formulations include, but are not limited to, tablets, lozenges, and combinations thereof. Generally, such tablets or lozenges may be prepared using conventional methods, may contain (as a non-limiting example) 0.1% to 20% (w / w) of the active ingredient, may contain any combination of orally soluble compositions and orally disintegrable compositions, and may contain one or more additional ingredients described herein as optional components. In some embodiments, the pharmaceutical compositions suitable for buccal administration may contain any combination of powders, aerosolized liquids and / or suspensions, or sprayed liquids and / or suspensions, each containing the active ingredient, and these dosage forms may be in the form of dispersed particles with an average particle size of about 0.1 nm to 200 nm and / or dispersed droplets with a droplet diameter of about 0.1 nm to 200 nm. In some embodiments, the pharmaceutical compositions for buccal administration may further contain one or more additional ingredients described herein.
[0388] Depot administration In some embodiments, the pharmaceutical compositions described herein may be formulated in the form of a sustained-release depot preparation. In some embodiments, the pharmaceutical compositions described herein are spatially retained within or near a target tissue.
[0389] Depot injection formulations are generally prepared by forming a microencapsulation matrix of a pharmaceutical composition encapsulated in a biodegradable polymer (e.g., polylactic acid-polyglycolide). Generally, the release rate of the pharmaceutical composition can be controlled by changing the ratio of the pharmaceutical composition to the polymer and by changing the properties of the specific polymer used. Suitable biodegradable polymers include, but are not limited to, poly(orthoesters) and poly(anhydride). Depot injection formulations are prepared by encapsulating the pharmaceutical composition within liposomes or microemulsions that are compatible with biological tissues.
[0390] Rectal administration and vaginal administration In some embodiments, the pharmaceutical compositions described herein may be administered rectally, vaginally, or in combination thereof. Generally, compositions for rectal or vaginal administration are suppositories, which can be prepared by mixing an active ingredient with a suitable non-irritating additive (e.g., polyethylene glycol, cocoa butter, or suppository wax) that is solid at ambient temperature but becomes liquid at body temperature. The active ingredient is released when the suppository dissolves in the rectal or vaginal cavity.
[0391] dose The gene insertion system (GIS) and / or pharmaceutical compositions comprising the gene insertion system (GIS) of the present invention may be administered in an amount (i.e., dose) that produces a desired effect (e.g., a desired therapeutic effect or research outcome) in a subject. In some embodiments, the desired dose may be determined based on parameters of the subject (e.g., body length, condition, or nature of the subject), parameters of the effect (e.g., the degree of the required response, the threshold of the therapeutic effect, the duration of the effect, or the side effects presented), or any combination thereof. In some embodiments, the appropriate dose may be determined before the first administration and based on at least one assay testing at least one parameter of the subject. In some embodiments, the appropriate dose may be determined after the first administration and based on at least one assay testing at least one parameter of the effect. In some embodiments, the dose may be maintained without change throughout the course of administration. In some embodiments, the dose may be changed once, twice, or multiple times throughout the course of administration.
[0392] In some embodiments, the dose may be expressed as the mass ratio of the active ingredient to the mass of the substance (for example, the mass in units of mg / kg). For example, the dose may be 0.1-100 mg / kg, 1-100 mg / kg, 2-100 mg / kg, 3-100 mg / kg, 4-100 mg / kg, 5-100 mg / kg, 6-100 mg / kg, 7-100 mg / kg, 8-100 mg / kg, 9-100 mg / kg, 10-100 mg / kg, 15-100 mg / kg, 20-100 mg / kg, 25-100 mg / kg, 30-100 mg / kg, 35-100 mg / kg, 40-100 mg / kg, 45-100 mg / kg, 50-100 mg / kg, 55-100mg / kg, 60-100mg / kg, 65-100mg / kg, 70-100mg / kg, 75-100mg / kg, 80-100mg / kg, 85-100mg / kg, 90-100mg / kg, 95-100mg / kg, 0.1-95mg / kg, 1-95mg / kg, 2-95mg / kg, 3-95mg / kg, 4-95mg / kg, 5-95mg / kg, 6-95mg / kg, 7-95mg / kg, 8-95mg / kg, 9-95mg / kg, 10-95mg / kg, 15 ~95mg / kg, 20~95mg / kg, 25~95mg / kg, 30~95mg / kg, 35~95mg / kg, 40~95mg / kg, 45~95mg / kg, 50~95mg / kg, 55~95mg / kg, 60~95mg / kg, 65~95 mg / kg, 70~95mg / kg, 75~95mg / kg, 80~95mg / kg, 85~95mg / kg, 90~95mg / kg, 0.1~90mg / kg, 1~90mg / kg, 2~90mg / kg, 3~90mg / kg, 4~90mg / kg, 5 ~90mg / kg, 6~90mg / kg, 7~90mg / kg, 8~90mg / kg, 9~90mg / kg, 10~90mg / kg, 15~90mg / kg, 20~90mg / kg, 25~90mg / kg, 30~90mg / kg, 35~90mg / k g, 40~90mg / kg, 45~90mg / kg, 50~90mg / kg, 55~90mg / kg, 60~90mg / kg, 65~90mg / kg, 70~90mg / kg, 75~90mg / kg, 80~90mg / kg, 85~90mg / kg, 0.1~85mg / kg、1~85mg / kg、2~85mg / kg、3~85mg / kg、4~85mg / kg、5~85mg / kg、6~85mg / kg、7~85mg / kg、8~85mg / kg、9~85mg / kg、10~85mg / kg、15~85mg / kg、20~85mg / kg、25~85mg / kg、30~85mg / kg、35~85mg / kg、40~85mg / kg、45~85mg / kg、50~85mg / kg、55~85mg / kg、60~85mg / kg、65~85mg / kg、70~85mg / kg、75~85mg / kg、80~85mg / kg、0.1~80mg / kg、1~80mg / kg、2~80mg / kg、3~80mg / kg、4~80mg / kg、5~80mg / kg、6~80mg / kg、7~80mg / kg、8~80mg / kg、9~80mg / kg、10~80mg / kg、15~80mg / kg、20~80mg / kg、25~80mg / kg、30~80mg / kg、35~80mg / kg、40~80mg / kg、45~80mg / kg、50~80mg / kg、55~80mg / kg、60~80mg / kg、65~80mg / kg、70~80mg / kg、75~80mg / kg、0.1~75mg / kg、1~75mg / kg、2~75mg / kg、3~75mg / kg、4~75mg / kg、5~75mg / kg、6~75mg / kg、7~75mg / kg、8~75mg / kg、9~75mg / kg、10~75mg / kg、15~75mg / kg、20~75mg / kg、25~75mg / kg、30~75mg / kg、35~75mg / kg、40~75mg / kg、45~75mg / kg、50~75mg / kg、55~75mg / kg、60~75mg / kg、65~75mg / kg、70~75mg / kg、0.1~70mg / kg、1~70mg / kg、2~70mg / kg、3~70mg / kg、4~70mg / kg、5~70mg / kg、6~70mg / kg、7~70mg / kg、8~70mg / kg、9~70mg / kg、10~70mg / kg、15~70mg / kg、20~70mg / kg、25~70mg / kg、30~70mg / kg、35~70mg / kg、40~70mg / kg、45~70mg / kg、50~70mg / kg、55~70mg / kg、60~70mg / kg、65~70mg / kg、0.1~65mg / kg、1~65mg / kg、2~65mg / kg、3~65mg / kg、4~65mg / kg、5~65mg / kg、6~65mg / kg、7~65mg / kg、8~65mg / kg、9~65mg / kg、10~65mg / kg、15~65mg / kg、20~65mg / kg、25~65mg / kg、30~65mg / kg、35~65mg / kg、40~65mg / kg、45~65mg / kg、50~65mg / kg、55~65mg / kg、60~65mg / kg、0.1~60mg / kg、1~60mg / kg、2~60mg / kg、3~60mg / kg、4~60mg / kg、5~60mg / kg、6~60mg / kg、7~60mg / kg、8~60mg / kg、9~60mg / kg、10~60mg / kg、15~60mg / kg、20~60mg / kg、25~60mg / kg、30~60mg / kg、35~60mg / kg、40~60mg / kg、45~60mg / kg、50~60mg / kg、55~60mg / kg、0.1~55mg / kg、1~55mg / kg、2~55mg / kg、3~55mg / kg、4~55mg / kg、5~55mg / kg、6~55mg / kg、7~55mg / kg、8~55mg / kg、9~55mg / kg、10~55mg / kg、15~55mg / kg、20~55mg / kg、25~55mg / kg、30~55mg / kg、35~55mg / kg、40~55mg / kg、45~55mg / kg、50~55mg / kg、0.1~50mg / kg、1~50mg / kg、2~50mg / kg、3~50mg / kg、4~50mg / kg、5~50mg / kg、6~50mg / kg、7~50mg / kg、8~50mg / kg、9~50mg / kg、10~50mg / kg、15~50mg / kg、20~50mg / kg、25~50mg / kg、30~50mg / kg、35~50mg / kg、40~50mg / kg、45~50mg / kg、0.1~45mg / kg、1~45mg / kg、2~45mg / kg、3~45mg / kg、4~45mg / kg、5~45mg / kg、6~45mg / kg、7~45mg / kg、8~45mg / kg、9~45mg / kg、10~45mg / kg、15~45mg / kg、20~45mg / kg、25~45mg / kg、30~45mg / kg、35~45mg / kg、40~45mg / kg、0.1~40mg / kg、1~40mg / kg、2~40mg / kg、3~40mg / kg、4~40mg / kg、5~40mg / kg、6~40mg / kg、7~40mg / kg、8~40mg / kg、9~40mg / kg、10~40mg / kg、15~40mg / kg、20~40mg / kg、25~40mg / kg、30~40mg / kg、35~40mg / kg、0.1~35mg / kg、1~35mg / kg、2~35mg / kg、3~35mg / kg、4~35mg / kg、5~35mg / kg、6~35mg / kg、7~35mg / kg、8~35mg / kg、9~35mg / kg、10~35mg / kg、15~35mg / kg、20~35mg / kg、25~35mg / kg、30~35mg / kg、0.1~30mg / kg、1~30mg / kg、2~30mg / kg、3~30mg / kg、4~30mg / kg、5~30mg / kg、6~30mg / kg、7~30mg / kg、8~30mg / kg、9~30mg / kg、10~30mg / kg、15~30mg / kg、20~30mg / kg、25~30mg / kg、0.1~25mg / kg、1~25mg / kg、2~25mg / kg、3~25mg / kg、4~25mg / kg、5~25mg / kg、6~25mg / kg、7~25mg / kg、8~25mg / kg、9~25mg / kg、10~25mg / kg、15~25mg / kg、20~25mg / kg、0.1~20mg / kg、1~20mg / kg、2~20mg / kg、3~20mg / kg、4~20mg / kg、5~20mg / kg、6~20mg / kg、7~20mg / kg、8~20mg / kg、9~20mg / kg、10~20mg / kg、15~20mg / kg、0.1~15mg / kg、1~15mg / kg、2~15mg / kg、3~15mg / kg、4~15mg / kg、5~15mg / kg、6~15mg / k g、7~15mg / kg、8~15mg / kg、9~15mg / kg、10~15mg / kg、0.1~10mg / kg、1~10mg / kg、2~10mg / kg、3~10mg / kg、4~10mg / kg、5 ~10mg / kg、6~10mg / kg、7~10mg / kg、8~10mg / kg、9~10mg / kg、0.1~9mg / kg、1~9mg / kg、2~9mg / kg、3~9mg / kg、4~9mg / kg、5 ~9mg / kg、6~9mg / kg、7~9mg / kg、8~9mg / kg、0.1~8mg / kg、1~8mg / kg、2~8mg / kg、3~8mg / kg、4~8mg / kg、5~8mg / kg、6~8mg / kg、7~8mg / kg、0.1~7mg / kg、1~7mg / kg、2~7mg / kg、3~7mg / kg、4~7mg / kg、5~7mg / kg、6~7mg / kg、0.1~6mg / kg、1~6mg / kg 、2~6mg / kg、3~6mg / kg、4~6mg / kg、5~6mg / kg、0.1~5mg / kg、1~5mg / kg、2~5mg / kg、3~5mg / kg、4~5mg / kg、0.1~4mg / kg、1 ~4mg / kg、2~4mg / kg、3~4mg / kg、0.1~3mg / kg、1~3mg / kg、2~ 3mg / kg, 0.1~2mg / kg, 1~2mg / kg.
[0393] Dosage schedule The gene insertion system and / or pharmaceutical composition comprising the gene insertion system of the present invention may be administered at a frequency (i.e., a dosing schedule) that produces a desired effect (e.g., a desired therapeutic effect or research result) in a subject. In some embodiments, the dosing schedule may be determined by any of the methods used to determine the dose described herein. In some embodiments, the gene insertion system may be administered only once.
[0394] In some embodiments, the gene insertion system of the present invention may be administered two or more times. For example, the gene insertion system of the present invention may be administered two, three, four, five, six, seven, eight, nine, ten, or more times. In some embodiments, the gene insertion system of the present invention may be administered intermittently and / or continuously in the course of treatment for the therapeutic indication in the subject. In some embodiments, the gene insertion system of the present invention may be repeatedly administered to the subject over a lifetime.
[0395] V. How to use Target region, target tissue, or target cell when delivering gene insertion system formulations This specification provides a method for delivering a pharmaceutical composition and / or pharmaceutical formulation described herein to at least one target location of a subject, the method being carried out by contacting at least one target (including one or more target cells) such as a physiological system, anatomical site, organ, tissue, specific type of cell, or cell population with the at least one pharmaceutical composition and / or pharmaceutical formulation described herein.
[0396] The pharmaceutical compositions and / or pharmaceutical formulations described herein contain an amount of an active ingredient (e.g., the gene insertion system of the present invention) sufficient to achieve the desired effect (e.g., the effect of inserting at least one transgene into the target genome) in at least one cell located at the target.
[0397] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein typically comprise one or more cell membrane permeabilizing agents, but “naked” formulations (e.g., formulations without cell membrane permeabilizing agents or other agents) may also be envisioned, which may include a pharmaceutically acceptable carrier.
[0398] Physiological system In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target the physiological system.
[0399] In some embodiments, physiological systems include the auditory system, cardiovascular system, central nervous system, chemoreceptor system, circulatory system, digestive system, endocrine system, excretory system, exocrine system, genital system, cutaneous system, lymphatic system, muscular system, musculoskeletal system, nervous system, peripheral nervous system, renal system, reproductive system, respiratory system, urinary system, and visual system.
[0400] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target the amine precursor uptake decarboxylation system (APUD) (a series of cells with endocrine function that secrete various low molecular weight amines or polypeptide hormones), and APUD system tissues include, but are not limited to, the pituitary tissue, parathyroid tissue, thyroid tissue, bronchial tissue, adrenal medulla tissue, pancreatic tissue, gastrointestinal tract, carotid body, chemoreceptor system tissue, etc.
[0401] organs In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target organs. These organs include the anal canal, arteries, ascending colon, bladder, bone marrow, brain, bronchi, bronchioles, bulbourethral glands, capillaries, cecum, cerebellum, cerebral hemispheres, cerebrum, cervix, choroid plexus, clitoris, cranial nerves, descending colon, diencephalon, duodenum, ear, enteric nervous system, epididymis, esophagus, external genitalia, fallopian tubes, gallbladder, ganglia, taste organs, intestinal lymphoid tissue, heart, ileum, internal genitalia, interstitium, jejunum, joints, kidneys, large intestine, larynx, ligaments, liver, lungs, lymph nodes, and ligaments. These include the pulmonary vessels, mammary glands, medulla oblongata, mesentery, midbrain, oral cavity, respiratory muscles, nasal cavity, nerves, olfactory organs, ovaries, pancreas, parotid gland, penis, pharynx, placenta, pons, prostate, rectum, salivary glands, scrotum, seminal vesicles, sigmoid colon, skeleton, skin, small intestine, spinal nerves, spleen, stomach, subcutaneous tissue, sublingual gland, submandibular gland, teeth, tendons, testes, brainstem, spinal cord, ventricular system, thymus, tongue, tonsils, trachea, transverse colon, ureter, urethra, uterus, vagina, vas deferens, veins, and vulva.
[0402] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target one or both eyes.
[0403] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target the liver.
[0404] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target the brain.
[0405] cell In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may target specific cells and / or specific types of cells.
[0406] Cells include adipocytes, adrenergic neurons, α cells, amacrine cells, ameloblasts, anterior chamber lenticular epithelial cells, anterior / middle pituitary cells, apocrine sweat gland cells, astrocytes, auditory hair cells of the organ of Corti, auditory hair cells of the organ of Corti, B cells, Bartholin's gland cells, basal cells (stem cells) of the cornea, tongue, oral cavity, nasal cavity, distal anal canal, distal urethra and distal vagina, olfactory epithelial basal cells, basket cells, basophilic granulocytes and their precursor cells, β cells, Betz cells, bone marrow reticular tissue fibroblasts, and organ of Corti Border cells, border cells, Bowman's gland cells, brown adipose tissue, Brunner's gland cells, bulbourethral gland cells, tuft cells, C cells, Cajal-Retius cells, cardiomyocytes, cardiomyocytes, Cartwheel cells, glucocorticoid-producing zona fasciculata cells, mineralocorticoid-producing zona glomerulosa cells, androgen-producing zona reticularis cells, adrenal cortical cells, cementoblasts, atrial cells, ear canal gland cells, chandelier cells, chemoreceptor glomus cells of carotid bodies, chief cells, cholinergic neurons, chromophilic cells, club cells Follicles, cold-sensitive primary sensory neurons, connective tissue macrophages (all types), corneal fibroblasts (corneal stromal cells of the cornea), progesterone-secreting luteal cells of ruptured follicles, cortical hair stem cells, adrenocorticotropic hormone-secreting cells, lens fiber cells containing crystallin, epidermal hair stem cells, cytotoxic T cells, D cells, δ cells, dendritic cells, double bouquet cells, ductal cells, eccrine sweat gland clear cells, eccrine sweat gland dark cells, efferent duct cells of the testis, chondrocytes of the cartilage elastocartilage, endothelial cells, intestinal glial cells, enterochromaffin cells, intestinal Examples include chromophilic cell-like cells, enteroendocrine cells, eosinophilic granulocytes and their precursor cells, ependymal cells, epidermal basal cells, epidermal Langerhans cells, epididymal basal cells, epididymal chief cells, epithelial reticular cells, ε cells, erythrocytes, chondrocytes of fibrocartilage, fork neurons, foveal cells, G cells, gallbladder epithelial cells, germ cells, litter gland cells, eyelid Moll gland cells, glial cells, Golgi cells, gonadal stromal cells, gonadotropin-secreting cells, granule cells, granulosa cells, granulosa luteal cells, lattice cells, and cephalic azimuthal cells.
[0407] In some embodiments, the cells may be cancer cells. In some embodiments, the cells may be non-cancerous cells.
[0408] In some embodiments, eukaryotic cells may be stem cells. Various types of stem cells are known in the art and all may be used to carry out the disclosure. Examples of stem cells include, but are not limited to, embryonic stem cells, hematopoietic stem cells, neural stem cells, epidermal neural crest stem cells, induced pluripotent stem cells, mammary gland stem cells, intestinal stem cells, mesenchymal stem cells, olfactory organ adult stem cells, testicular cells, and progenitor cells (e.g., neural progenitor cells, angioblasts, osteoblasts, chondrocytes, pancreatic progenitor cells, epidermal progenitor cells, etc.). In some embodiments, the stem cells may be embryonic stem cell lines derived from cells isolated from a subject.
[0409] In some embodiments, eukaryotic cells are cells found in the circulatory system of humans, non-human primates, and / or other mammals (including mice and / or rats). Exemplary circulatory system cells include, but are not limited to, platelets, plasma cells, erythrocytes, B cells, T cells, natural killer cells, macrophages, neutrophils, and their progenitor cells. In some embodiments, at least one eukaryotic cell may originate from any of these circulatory system eukaryotic cells.
[0410] In some embodiments, at least one eukaryotic cell is a natural killer cell or a natural killer cell precursor.
[0411] In some embodiments, at least one eukaryotic cell is a B cell or a precursor B cell.
[0412] In some embodiments, the eukaryotic cells may be plant cells. In some embodiments, the plant cells are cells of monocots or dicots, specifically including zucchini, woody plants such as conifers and deciduous trees, wheat, turnips, tomatoes, tobacco, sunflowers, sugarcane, sugar beets, strawberries, spinach, soybeans, sorghum, rye, rice, raspberries, rapeseed, radishes, pumpkins, potatoes (including sweet potatoes), plums, pineapples, peanuts, peas, papayas, oats, melons, mangoes, corn, lettuce, lentils, herbs, hemp, pasture grass, flowers, eucalyptus, cucumbers, cotton, coffee, citrus fruits, chicory, cherries, celery, cauliflower, carrots, canola, cabbage, broccoli, rapeseed, blackberries, legumes, barley, and bananas. This includes, but is not limited to, avocados, asparagus, Arabidopsis thaliana, other fruit plants, ornamental plants, almonds, alfalfa, perennial pastures, fodder crops, other vegetables, other drupes (e.g., peaches, nectarines, apricots, pears, plums, etc.), other pome fruits (e.g., apples, pears, etc.), other fruits, other bulbs (e.g., garlic, onions, chives, etc.), other crops, parts of perennial plants (e.g., corms; tubers; roots; crowns; stems; stolons; suckers; shoots; rootless cuttings, cuttings, and cuttings with callus or callus-forming young plants; apical meristems, etc.), and any combination of these or hybrids thereof. In this specification, “plant” means a physical part of a plant, such as seeds, seedlings, saplings, roots, tubers, stems, stalks, leaves and fruits.
[0413] tumor In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein target a tumor, which may be a benign tumor, a precancerous tumor, or a malignant tumor.
[0414] Insertion of introduced genes The present invention provides a method for introducing a transgene into a subject, for example, a human subject. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system described herein into the subject. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system containing a transgene into the subject.
[0415] In some embodiments, the method may include the step of inserting a transgene into one or more target insertion sites. Referring to Figure 8, a target genomic region 500 having the inserted transgene is shown. In this example, the target genomic DNA includes the target insertion site 120 and the surrounding genomic DNA 110. For clarification, the target insertion site is a portion of the target DNA. The 5' junction 510 indicates a transition point between the target DNA and the inserted transgene 520, which corresponds to the 5' end of the transgene. This junction 510 may have an overlap of part or all of the upstream sequence of the target site, present in both the target genome and the 5' end of the template RNA. In contrast, the 3' junction 530 indicates a transition point between the 3' end of the transgene and the target DNA. This junction 530 may have an overlap of part or all of the downstream sequence of the target site, present in both the target genome and the 3' module of the template RNA. The ligation portion 510 and / or ligation portion 530 may further contain another nucleotide, which may be, for example, a nucleotide produced when a nucleotide other than the template is added to the pre-extension primer or cDNA 3' terminal primer by reverse transcriptase and then enzymatically dissociated from the double helix of the template product.
[0416] Target insertion site In some embodiments, one or more target insertion sites include safe harbor sites. In this specification, “safe harbor site” means a specific location in the target genome where insertion of a transgene will not cause unintended disruption of cellular function. Generally, a site in the target genome may be identified as a safe harbor site if (a) insertion of genetic material at that site does not alter the expression of the target gene, or (b) insertion of genetic material at that site alters gene expression, but the change does not alter the normal function of the target cell (for example, due to the presence of many repetitive sequences in the interfered gene in the target genome). An example of case (b) is, but is not limited to, cases where ribosomal RNA (rRNA) coding genes are repeated in the genome in amounts such that interference with some rRNA genes does not perturb normal cellular function.
[0417] In some embodiments, at least one safe harbor site and / or target insertion site comprises at least one ribosomal DNA (rDNA) sequence. In this specification, the term "ribosomal DNA" means a gene encoding rRNA. In some embodiments, at least one safe harbor site and / or target insertion site comprises at least one 28S rDNA sequence.
[0418] Transgene The methods and compositions of the present invention can be used for the insertion of any payload sequence (i.e., transgene), and the length and origin of the payload sequence are not limited.
[0419] In some embodiments, the transgene includes a therapeutically active gene. In this specification, the term “therapeutically active gene” means any gene capable of expressing an expression product useful for treating, alleviating, or preventing at least one therapeutic indication.
[0420] In some embodiments, at least one transgene may contain at least one telomerase reverse transcriptase (TERT) gene. In some embodiments, at least one transgene may contain at least one short-chain factor VIII gene. In some embodiments, at least one transgene may contain at least one phenylalanine hydroxylase (PAH) gene.
[0421] In some embodiments, at least one transgene is a reporter gene. In this specification, “reporter gene” means any gene capable of expressing an expression product detectable by an assay.
[0422] In some embodiments, at least one reporter gene may include, but is not limited to, at least one green fluorescent protein (GFP), at least one red fluorescent protein (RFP), luciferase enzyme (LUC), β-galactosidase (LacZ), chloramphenicol acetyltransferase (cat), etc., and may encode these.
[0423] Non-wild-type introduced genes Those skilled in the art will understand that while many of the transgenes exemplified above are natural or wild-type sequences, the gene insertion systems disclosed herein are not limited to the insertion of wild-type genes or natural genes or parts of their gene sequences. The gene insertion systems of the present invention may be used, for example, to insert genes derived from wild-type genes, genes containing only a portion of wild-type genes, genes assembled from parts of various types of wild-type genes, and / or genes whose sequences are not known to exist in nature. Furthermore, the gene insertion systems of the present invention may be used to insert transgenes whose expression products are not normally observed in target cells, and / or transgenes whose gene expression is not normally observed.
[0424] Regulators of transgenes In some embodiments, the gene insertion system of the present invention may be used to insert at least one transgene containing or encoding at least one regulator. For example, the transgene may be designed and / or recombined such that the expression product of the transgene contains any number of miRNA-binding regions and / or siRNA-binding regions. Typically, by incorporating miRNA and / or siRNA, cells containing complementary miRNA or siRNA in their transcriptome can be excluded from the target of transgene expression.
[0425] In some embodiments, the transgene may include a first expression product comprising or encoding at least one miRNA and / or siRNA, and a second expression product (or more expression products) comprising or encoding at least one miRNA-binding site and / or siRNA-binding site complementary to the first expression product. While we do not wish to be bound by any theory, such embodiments may prevent the long-term expression of the second expression product.
[0426] antibody In this specification, the term “antibody” is used in its broadest sense and specifically includes, but is not limited to, monoclonal antibodies, polyclonal antibodies, multispecific antibodies (e.g., bispecific antibodies formed from at least two full-length antibodies), and antibody fragments (e.g., diabodies), encompassing various embodiments and exhibiting desired biological activity (e.g., “functional”). An antibody is a molecule primarily composed of amino acids, and is a monomeric polypeptide or a polymeric polypeptide containing at least one amino acid region derived from a known antibody sequence or a parental antibody sequence. An antibody may contain an amino acid motif that recruits one or more endogenous or non-natural modifications (including, but not limited to, the addition of sugar moieties, fluorescent moieties, chemical tags, etc.). To achieve the objectives of the present invention, an “antibody” may include a heavy chain variable domain and a light chain variable domain, as well as an Fc region.
[0427] The gene insertion system of the present invention may be used to insert a transgene that includes or encodes at least one functional antibody.
[0428] Treatment of indications for treatment The present invention provides a method for treating or preventing at least one therapeutic indication in a subject requiring treatment or prevention of at least one therapeutic indication. In some embodiments, the method includes introducing an effective amount of at least one gene insertion system described herein into the subject. In some embodiments, the method includes introducing an effective amount of at least one gene insertion system comprising at least one therapeutically active transgene into the subject.
[0429] In some embodiments, at least one therapeutic indication includes at least one genetic disorder involving loss of function. In some embodiments, at least one method of treating at least one therapeutic indication includes administering at least one transgene that restores a subject from a genetic disorder involving loss of function. In this specification, the term “restores” means providing the subject with at least one composition that enables the subject to perform the original function that it had lost.
[0430] In some embodiments, the method includes the step of restoring insufficient telomerase activity in a subject by administering an effective amount of a gene insertion system containing at least one TERT transgene to the subject.
[0431] In some embodiments, the methods and compositions of the present invention may be used to treat or prevent conditions caused by insufficient telomerase function in a subject. In some embodiments, the at least one method includes administering a therapeutically effective dose of at least one gene insertion system containing at least one TERT gene to a subject exhibiting insufficient telomerase activity. In some embodiments, the at least one method includes administering a therapeutically effective dose of at least one gene insertion system containing at least one TERT gene to a subject suspected of developing a disease due to insufficient telomerase activity.
[0432] Regulation of heterogeneous genes The gene insertion systems of the present invention, encompassing the pharmaceutical formulations and pharmaceutical compositions described herein, may be used in methods for regulating the expression of heterogeneous genes. For clarification, the term “heterogeneous gene,” as used herein in relation to the regulation of gene expression, means a gene in the genome of interest other than the gene inserted by the gene insertion system of the present invention.
[0433] Typically, methods for regulating the expression of heterologous genes may involve using the gene insertion system of the present invention to insert a sequence whose expression product acts on the expression pathway of another gene. For example, the expression product of the inserted gene may affect the transcription from the heterologous gene to mRNA, the translation of the heterologous gene's mRNA to polypeptides, the rate of degradation or inactivation of the heterologous gene's mRNA in the cytoplasm, or any combination thereof.
[0434] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene containing at least one microRNA (miRNA) or a transgene encoding at least one microRNA (miRNA). In some embodiments, miRNAs suitable for the implementation of this disclosure may include miRNAs known in the art or miRNAs to be discovered in the future. In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene containing at least one artificial miRNA or a transgene encoding at least one artificial miRNA, the artificial miRNA being designed to bind to at least one gene expression product present in the subject. In this specification, the term “artificial miRNA” means a miRNA whose sequence has been modified or designed to bind to a desired target sequence. Artificial miRNAs may be designed in various ways known in the art.
[0435] In some embodiments, at least one gene insertion system of the present invention may be used to insert a transgene containing at least one small interfering RNA (siRNA), or a transgene encoding at least one small interfering RNA (siRNA). In this specification, the term "small interfering RNA" means double-stranded ribonucleic acid (dsRNA) having a nucleotide sequence substantially identical to at least a portion of the target gene. Generally, siRNAs are typically 21–25 nt in length, but may be shorter or longer, and interfere with (suppress) the expression of the target gene by promoting the degradation of the target gene's mRNA. Known or future siRNAs may be suitable for use in the present invention.
[0436] In some embodiments, the at least one gene insertion system of the present invention may be used to insert a transgene comprising at least one artificial siRNA, or a transgene encoding at least one artificial siRNA. In this specification, the term "artificial siRNA" means an siRNA whose sequence is designed to be complementary to at least one gene of interest.
[0437] In some embodiments, the at least one gene insertion system of the present invention may be used to insert a transgene containing at least one transcription factor (TF), or a transgene encoding at least one transcription factor (TF). Herein, the term “transcription factor” means a polypeptide that binds to DNA and modifies or affects the transcription of at least one gene. Known or future transcription factors may be suitable for use in the present invention.
[0438] The gene insertion system of the present invention may include any combination of miRNA, siRNA and / or transcription factors, or may be used to insert a transgene encoding these. For example, at least one gene insertion system may include at least one miRNA and at least one siRNA; at least one miRNA and at least one transcription factor; at least one siRNA and at least one transcription factor; or at least one miRNA, at least one siRNA, and at least one transcription factor, or may be used to insert a transgene encoding these.
[0439] Preventive use In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used to prevent disease or stabilize the progression of a therapeutic indication.
[0440] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as preventive measures to prevent the development of future therapeutic indications.
[0441] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used to prevent further progression of a therapeutic indication.
[0442] vaccine In some embodiments, the pharmaceutical compositions and / or pharmaceutical preparations described herein may be used as and / or in the same manner as vaccines. In this specification, “vaccine” means a biological preparation that enhances immunity against a specific therapeutic indication or infectious agent.
[0443] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as vaccines for therapeutic areas such as dermatology, central nervous system, cardiovascular, oncology, endocrinology, immunology, respiratory medicine, and anti-infective medicine, and / or may be used in the same manner as vaccines for such uses.
[0444] antigen The gene insertion system of the present invention may be used to insert a transgene containing at least one antigen, or a transgene encoding at least one antigen, which may be activated by or presented on the surface of at least one target cell. In this specification, the term “antigen” means a composition that elicits an immune response in an organism. For example, a composition that can cause an organism to produce antibodies against itself, and more particularly, a composition that can induce an adaptive immune response within the target organism by such antibodies. Antigens may be immunogenic substances such as polypeptides, proteins, polysaccharides, nucleic acids, or lipids. In some embodiments, antigens may be derived from infectious agents such as bacteria, viruses, protozoa, fungi, or prions, but are not limited to these.
[0445] In some embodiments, the antigen may include a part or subunit of the infectious agent, for example, a coat, coat components, coat proteins, coat polypeptides, surface components, surface proteins, surface polypeptides, capsule components, cell wall components, flagella, cilia, toxins, or toxoids.
[0446] In some embodiments, the at least one gene insertion system of the present invention may be used to insert a transgene containing at least one antigen for use in vaccination for at least one therapeutic indication, or a transgene encoding such an antigen.
[0447] the study In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for diagnostic purposes or as research tools for the therapeutic indications disclosed herein.
[0448] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in any research experiment, for example, in vivo or in vitro experiments.
[0449] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for the detection of research biomarkers.
[0450] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in cultured cells. The cultured cells may be derived from any source known to those skilled in the art, and may be, but are not limited to, stable cell lines, animal models, or cells derived from human patients or controls.
[0451] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in vivo experiments in animal models (i.e., mice, rats, rabbits, cats, dogs, non-human primates, guinea pigs, fruit flies, ferrets, nematodes, zebrafish, or other animals known in the art for research purposes).
[0452] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used for stem cell and / or cell differentiation.
[0453] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in human research experiments or human clinical trials.
[0454] The present invention provides a method for conducting scientific and / or medical research on a subject. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system described herein into the subject. In some embodiments, the method includes the step of introducing an effective amount of at least one gene insertion system containing at least one reporter transgene into the subject.
[0455] Monotherapy and combination therapy In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as monotherapy or combination therapy for the treatment of a disease.
[0456] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used as monotherapy. In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in combination therapy. Combination therapy may involve a combination of one or more neuroprotective agents, which may be, for example, small molecule compounds, growth factors, or hormones that have been evaluated for their neuroprotective effects against neuronal degeneration.
[0457] In some embodiments, the pharmaceutical compositions and / or pharmaceutical formulations described herein may be used in combination with one or more other therapeutic agents. “Combined” should not be interpreted as implying formulation for simultaneous administration and / or co-delivery of multiple therapeutic agents, but such delivery methods are also within the scope of the present invention. The pharmaceutical compositions and / or pharmaceutical formulations described herein, as well as other therapeutic agents, may be administered simultaneously with, before, or after one or more other desired therapeutic agents or medical procedures. Typically, each therapeutic agent is administered in a dose and / or time schedule specified for that therapeutic agent.
[0458] The therapeutic agents that may be used in combination with the pharmaceutical compositions and / or pharmaceutical preparations described herein may be low molecular weight compounds, which are antioxidants, anti-inflammatory agents, anti-apoptotic agents, calcium regulators, anti-glutamate agonists, structural protein inhibitors, compounds involved in muscle function, or compounds involved in metal ion regulation.
[0459] In vivo synthesis of gene insertion constructs (GICs) The present invention provides a method for synthesizing gene insertion system (GIS) biopolymers, such as gene insertion construct (GIC) biopolymers. In some embodiments, the method includes the steps of: administering at least one GIC synthesis construct to a target cell population; maintaining the cell population for a time sufficient to allow at least one GIS synthesis construct to be expressed by the target cells; and recovering and purifying the expression product of the GIS synthesis construct by a method known in the art.
[0460] In some embodiments, at least one GIC synthesis construct comprises or encodes the GIC of the present invention. In some embodiments, at least one GIC synthesis construct comprises or encodes the GIC of the present invention and means for synthesizing at least one recombinant RNA in vivo. Such means may include providing or encoding an RNA polymerase promoter, a sequence for selection and purification of the recombinant RNA, a complementary GIC sequence, and a post-recombinant RNA production processing signal. In some embodiments, at least one GIC synthesis construct is administered in the form of a DNA plasmid that enables the encoded RNA to be produced by an endogenous intracellular mechanism.
[0461] An exemplary GIC synthesis construct 600 is shown in Figure 9. At the 5' end of this GIC synthesis construct, the RNAP module 610 may contain a suitable RNA polymerase promoter (e.g., the T7 RNAP promoter). If an optional component, a 5' leader module 620, is present, this 5' leader module 620 is located on the 3' side of the RNAP module. Components of this 5' leader module 620 may include components that improve the folding and self-cleavage of the template 5' module and / or destabilize the transcript and / or rapidly remove GIC transcripts having an immunogenic 5' end (e.g., which may result from a failed RZ self-cleavage). Before use as a GIC, the expressed 5' leader module RNA is cleaved at the RZ self-cleavage site 630. The complementary strand 640 of the 5' module, the complementary strand 650 of the template module, and the complementary strand 660 of the 3' module encode the 5' module, template module, and 3' module of the GIC, respectively. Furthermore, the 3' end may be a linearization restriction enzyme site 670, which is a restriction enzyme cleavage site. Cleavage at this site by the restriction enzyme linearizes the GIC RNA, leaving all excess vector components on the vector.
[0462] VII. List of Embodiments Embodiment 1. A genome editing system, (i) at least one reverse transcriptase construct (RTC) comprising a polynucleotide encoding a polypeptide having enzymatic activity for reverse transcription of a polynucleotide template, (ii) at least one gene insertion construct (GIC) comprising at least one polynucleotide template suitable for reverse transcription by the polypeptide encoded by the at least one RTC. A system that includes this.
[0463] Embodiment 2. The system according to Embodiment 1, wherein the at least one reverse transcriptase construct comprises at least one biopolymer, the biopolymer comprising at least one nucleic acid, at least one amino acid, and any combination thereof.
[0464] Embodiment 3. The system according to Embodiment 1 or 2, wherein the at least one reverse transcriptase construct comprises at least one reverse transcriptase module (RTC:RT module), which may include at least one 5' module (RTC:5' module), at least one 3' module (RTC:3' module), and any combination thereof.
[0465] Embodiment 4. The system according to Embodiment 3, wherein the at least one reverse transcriptase module includes or encodes at least one reverse transcriptase.
[0466] Embodiment 5. The system according to Embodiment 3 or 4, wherein the at least one reverse transcriptase module comprises or encodes at least one reverse transcriptase derived from a non-long-chain terminal repeat (non-LTR) retroelement.
[0467] Embodiment 6. The system according to Embodiment 4 or 5, wherein the at least one reverse transcriptase module includes or encodes a non-natural translation start codon.
[0468] Embodiment 7. The system according to any one of Embodiments 4 to 6, wherein the at least one reverse transcriptase comprises at least one DNA-binding domain, at least one RNA-binding domain, at least one cDNA synthesis domain, at least one endonuclease domain, and any combination thereof.
[0469] Embodiment 8. The system according to Embodiment 7, wherein at least one reverse transcriptase domain, at least one target DNA binding domain, at least one template RNA binding domain, and at least one endonuclease domain, and at least one of any combination thereof, are derived from a different species than the species from which the reverse transcriptase from which at least one of the remaining domains originates.
[0470] Embodiment 9. The system according to Embodiment 3, wherein at least one 5' end module, which is an optional component of the reverse transcriptase construct, comprises or encodes at least one RNA polymerase promoter, at least one 5' untranslated region (5'-UTR), at least one Kozak sequence, at least one 5' cap, and any combination thereof.
[0471] Embodiment 10. The system according to Embodiment 3, wherein at least one 3' end module, which is an optional component of the reverse transcriptase construct, includes or encodes at least one reverse transcriptase translation termination codon, at least one 3' untranslated region (3'UTR), at least one polyA tail, and any combination thereof.
[0472] Embodiment 11. The system according to any one of Embodiments 1 to 10, wherein the at least one reverse transcriptase module includes or codes for at least one structure shown in Figures 2 to 5 or any combination thereof.
[0473] Embodiment 12. The system according to any one of Embodiments 1 to 11, wherein the at least one reverse transcriptase construct comprises, codes for, or is coded by, at least one sequence from Sequence ID Nos. 1 to 57 and any combination thereof.
[0474] Embodiment 13. The system according to Embodiment 1, wherein the at least one gene insertion construct comprises or encodes at least one nucleic acid biomolecule.
[0475] Embodiment 14. The system according to Embodiment 1 or 13, wherein the at least one gene insertion construct comprises or codes for at least one GIC:5' side module which is an optional component, at least one GIC:payload module, at least one GIC:3' side module which is an optional component, and any combination thereof.
[0476] Embodiment 15. The system according to Embodiment 14, wherein the at least one GIC:5' side module includes or encodes at least one sequence derived from the 5' region of a natural retroelement, and may include, or encode, at least one rRNA sequence, at least one ribozyme sequence, at least one folding motif sequence, or any combination thereof.
[0477] Embodiment 16. The system according to Embodiment 15, wherein at least one rRNA sequence, which is an optional component of the GIC:5' side module, contains or encodes a 1-30 nt rRNA of the subject.
[0478] Embodiment 17. The system according to Embodiment 15, wherein at least one ribozyme sequence, which is an optional component of the GIC:5' side module, contains or encodes at least one self-cleaving ribozyme, the self-cleaving ribozyme may contain or encode a ribozyme of hepatitis delta virus.
[0479] Embodiment 18. The system according to Embodiment 17, wherein at least one ribozyme sequence, which is an optional component of the GIC:5' side module, comprises or encodes a ribozyme derived from the 5' region of at least one non-long-chain terminal repeat retroelement.
[0480] Embodiment 19. The system according to Embodiment 15, wherein at least one folding motif sequence, which is an optional component of the GIC:5' side module, includes or encodes at least one autonomous folding RNA sequence motif, the autonomous folding RNA sequence motif may include, and encode at least one hairpin motif, at least one stem-loop motif, at least one paired stem-4 motif, or any combination thereof.
[0481] Embodiment 20. The GIC:5' side module is sequence number 60~ 153 , 179 ~ 205 and 206 ~ 207 A system according to any one of embodiments 14 to 19, comprising or coding at least one of the following, or any combination thereof.
[0482] Embodiment 21. The system according to Embodiment 14, wherein the at least one GIC:3' side module includes or encodes at least one reverse transcriptase recognition sequence, and may also include, or encode, at least one rRNA sequence, at least one A tract sequence, or any combination thereof.
[0483] Embodiment 22. The system according to Embodiment 21, wherein at least one reverse transcriptase recognition sequence of the GIC:3' side module includes or encodes at least one sequence that interacts with at least one reverse transcriptase.
[0484] Embodiment 23. The system according to Embodiment 21 or 22, wherein at least one reverse transcriptase recognition sequence of the GIC:3' side module is derived from the 3' region of a native retroelement.
[0485] Embodiment 24. The system according to Embodiment 21, wherein at least one rRNA sequence, which is an optional component of the GIC:3' side module, contains or encodes 1 to 30 nt of rRNA.
[0486] Embodiment 25. The system according to Embodiment 21, wherein at least one A tract sequence, which is an optional component of the GIC:3' side module, contains or encodes a sequence consisting of 1 to 50 adenine bases.
[0487] Embodiment 26. The at least one GIC:3' side module is array number 225 ~ 253 A system according to any one of embodiments 14 and 21-25, comprising or coding at least one of these or any combination thereof.
[0488] Embodiment 27. The system according to Embodiment 14, wherein the at least one GIC:payload module includes or codes for at least one transgene sequence, and may include, and code for, at least one promoter sequence of the transgene, at least one 5' untranslated sequence of the transgene, at least one 3' untranslated sequence of the transgene, at least one polyadenylation signal sequence of the transgene, at least one non-coding RNA (ncRNA) processing sequence of the transgene, or any combination thereof.
[0489] Embodiment 28. The system according to Embodiment 27, wherein the at least one transgene sequence includes or encodes at least one target sequence for insertion into the target genome.
[0490] Embodiment 29. The system according to Embodiment 27, wherein at least one promoter sequence of the transgene includes or encodes at least one sequence that promotes the expression of the transgene in the target genome.
[0491] Embodiment 30. The system according to Embodiment 27, comprising at least one 5' untranslated sequence of the transgene, wherein the at least one 5' untranslated sequence comprises or encodes at least one 5' untranslated region of the mRNA of the transgene.
[0492] Embodiment 31. The system according to Embodiment 27, wherein at least one 3' untranslated sequence of the transgene includes or encodes at least one 3' untranslated region of the mRNA of the transgene.
[0493] Embodiment 32. The system according to Embodiment 27, wherein at least one polyadenylation signal sequence of the transgene contains or encodes at least one polyadenylation signal of the transgene.
[0494] Embodiment 33. The system according to Embodiment 27, wherein at least one non-coding RNA (ncRNA) processing sequence of the transgene includes or encodes at least one stop signal, at least one 3' processing signal, and any combination thereof of at least one ncRNA expressed from the transgene.
[0495] Embodiment 34. The at least one GIC:payload module is, 296 ~ 321 A system according to any one of embodiments 14 and 27-33, comprising or coding at least one of these or any combination thereof.
[0496] Embodiment 35. The system according to any one of Embodiments 13 to 34, wherein at least one of the at least one GIC:5' side module and the at least one GIC:3' side module contains or encodes at least one sequence derived from a different species than the species from which the non-long-chain terminal repeat retroelement from the other originates.
[0497] Embodiment 36. The system according to any one of Embodiments 1 and 13 to 35, wherein the at least one gene insertion construct includes or codes for at least one structure shown in Figures 6 to 9 and any combination thereof.
[0498] Embodiment 37. (i) at least one reverse transcriptase construct included in or encoded by at least one of sequence numbers 1 to 57, (ii) Sequence ID 60~ 153 , 179 ~ 205 , 206 ~ 207 , 208 ~ 217 , 225 ~ 253 , 275 ~ 278 , 279 ~ 281 , 284 ~ 295 and 296 ~ 332 At least one gene insertion construct that is contained in at least one sequence of or encoded by at least one of these sequences A system according to any one of Embodiments 1 and 13 to 36, including the above.
[0499] Embodiment 38. A system according to any one of Embodiments 1 and 13-37, comprising a synthetic construct (GIC) for gene insertion constructs (GIC) that includes or encodes at least one of the gene insertion constructs described in Embodiments 13-37.
[0500] Embodiment 39. The system according to any one of Embodiments 1 to 38, wherein at least one of the at least one reverse transcriptase construct and the at least one gene insertion construct includes or encodes at least one sequence derived from a different species than the species from which the retroelement from which the other originates.
[0501] Embodiment 40. A system according to any one of Embodiments 1 to 39, comprising (i) at least one reverse transcriptase construct described in Embodiments 2 to 12 and (ii) at least one combination of at least one gene insertion construct described in Embodiments 13 to 37.
[0502] Embodiment 41. A method for inserting at least one transgene into a target genome, comprising the step of administering at least one of the gene insertion systems (GIS) described in Embodiments 1 to 40 in an effective amount.
[0503] Embodiment 42. The method according to Embodiment 41, wherein the introduced gene is inserted into one or more target sites of the target genome, and the one or more target sites may include at least one safe harbor site.
[0504] Embodiment 43. The method according to Embodiment 42, wherein the at least one safe harbor site, which is an optional component, comprises at least one ribosomal DNA (rDNA) sequence, and the at least one ribosomal DNA sequence may comprise at least one 28S rDNA sequence.
[0505] Embodiment 44. The method according to any one of Embodiments 40 to 43, comprising the step of administering at least one of the gene insertion systems formulated with at least one delivery agent.
[0506] Embodiment 45. The method according to Embodiment 44, wherein the at least one delivery agent is at least one nanoparticle, and the at least one nanoparticle may contain at least one lipid nanoparticle.
[0507] Embodiment 46. A pharmaceutical composition comprising at least one of the gene insertion systems described in Embodiments 1 to 40, and which may further comprise at least one additive, at least one delivery agent, at least one auxiliary agent, and at least one of any combination thereof.
[0508] Embodiment 47. A method for treating a therapeutic indication in a subject requiring treatment of the therapeutic indication, comprising the step of administering at least one of the gene insertion systems described in Embodiments 1 to 40 or at least one of the pharmaceutical compositions described in Embodiment 46 in an effective amount, and which may also include at least one of the methods described in Embodiments 41 to 45.
[0509] Embodiment 48. The method according to Embodiment 47, wherein the therapeutic indication is caused by a deficiency in telomerase activity.
[0510] Embodiment 49. The method according to Embodiment 46 or 47, wherein the at least one gene insertion system includes at least one TERT transgene.
[0511] Embodiment 50. A kit for preparing a gene insertion system, comprising the method for preparing a gene insertion system described in Embodiments 1 to 40, and which may further comprise the pharmaceutical composition described in Embodiment 46, and which may further comprise a buffer, a DNA plasmid, or a protocol for preparing the gene insertion system or the pharmaceutical composition.
[0512] VIII. Definitions 28S rDNA: In this specification, the term "28S rDNA" refers to a portion of the genome of a target organism that encodes large ribosomal RNA (rRNA) that constitutes the large subunit (LSU) of the cytoplasmic ribosome of a eukaryote.
[0513] 3' junction: In this specification, the term "3' junction" refers to the location where the 3' end of the inserted sequence junctions with the 5' end of the target genome.
[0514] 3' region: In this specification, the term "3' region" refers to the portion of a retroelement gene located at the 3' end of the open reading frame.
[0515] 5' junction: In this specification, the term "5' junction" means the location where the 3' end of the genome in question junctions with the 3' end of the insertion sequence.
[0516] 5' region: In this specification, the term "5' region" means the portion of a retroelement gene located at the 5' end of the open reading frame.
[0517] Activity: In this specification, the term "activity" means a state in which an event is occurring or is taking place. The proteins and nucleic acids of this disclosure may be active, and this activity may be involved in one or more biological events.
[0518] Modifications made: In this specification, the term “modifications made” means changes to the protein sequence or amino acid sequence to alter, add, or remove its properties and / or activity.
[0519] Assay: In this specification, when the term “assay” is used as a verb, it is used in its broadest sense to mean the act of performing a test using appropriate methods known in the art. In this specification, when the term “assay” is used as a noun, it means a test used to measure the properties, state and / or activity of the subject of the assay.
[0520] Biological properties: In this specification, the terms “biological properties” and “properties” mean measurable or observable characteristics or activities of an organism, physiological system, organ, tissue, cell, or molecule.
[0521] Cargo: In the context of delivery media, the terms “cargo” and “payload” typically refer to compounds or structures (e.g., the gene insertion system of the present invention) intended for delivery to, to, or near, target cells, tissues, organs, or physiological systems.
[0522] Cell: In this specification, the term "cell" has the broadest possible meaning and refers to a living, membrane-bound structure.
[0523] Cellular process: In this specification, the term “cellular process” and its grammatical equivalents mean a process that takes place at the cellular level, which may or may not be limited to a single cell.
[0524] Features: In this specification, the term “feature” usually means a characteristic or quality belonging to a person, place, or thing that serves to identify them. The terms “feature” and “characteristic” are synonymous and may be used interchangeably.
[0525] To grant: In this specification, the term “to grant” and its grammatical equivalents mean the process of adding characteristics to an object.
[0526] Construct: In this specification, the noun “construct” means an artificially designed biomolecule. Examples of biomolecules include DNA, RNA, and polypeptides. Typically, the constructs described herein are designed for use in gene insertion systems.
[0527] Decomposition: In this specification, “decomposition” means the loss of function of a composition over time.
[0528] Delivery: In this specification, the term “delivery” means the act or method of delivering a compound, substance, object, part, cargo, or payload to a living cell or living organism. Unless otherwise specified, the terms “delivery” and “biological delivery” may be used interchangeably.
[0529] Delivery System: In this specification, the term "delivery system" means a composition, method, or combination thereof that, when prepared as a formulation in combination with the gene insertion system of the present invention, can deliver the components of the gene insertion system into the cytoplasm of target cells. Examples of delivery systems include, but are not limited to, systems comprising a delivery medium and systems for direct transfection.
[0530] Derived: In this specification, the term “derived” means nucleic acid sequences or protein sequences, such as non-long-chain terminal repeat (non-LTR) retrotransposons, isolated or obtained from a particular source. This term includes native sequences isolated or obtained from a particular source. Furthermore, this term includes artificial variant sequences derived from a source having the same or similar functional properties, for example, the variant may include nucleic acid sequences or amino acid sequences that have been modified to have improved functional properties compared to molecules derived from the original source.
[0531] Designed: In this specification, the term "designed" means a composition that has been modified from its natural or existing state to have novel and desired properties and / or activity.
[0532] DNA and RNA: In this specification, the terms “RNA,” “RNA molecule,” or “ribonucleic acid molecule” mean polymers of ribonucleotides. In this specification, the terms “DNA,” “DNA molecule,” or “deoxyribonucleic acid molecule” mean polymers of deoxyribonucleotides. DNA and RNA can be synthesized naturally. For example, DNA can be synthesized naturally by DNA replication, and RNA can be synthesized naturally by DNA transcription. DNA and RNA can also be synthesized chemically. DNA and RNA may be single-stranded (i.e., ssRNA or ssDNA) or multi-stranded (e.g., double-stranded, i.e., dsRNA or dsDNA). In this specification, the terms “mRNA” or “messenger RNA” mean single-stranded RNA encoding an amino acid sequence of one or more polypeptide chains. When an RNA sequence is described using deoxyribonucleotides, the DNA sequence can be converted to an RNA sequence by substituting thymidine ("T") with uridine ("U") or a uridine analog.
[0533] DNA Repair: In this specification, the term "DNA repair" refers to the endogenous processes that occur within a cell to correct damage to the genome within the cell.
[0534] Efficient: In this specification, the term “efficient” and its grammatical equivalents in relation to transgene insertion mean the effectiveness of any combination of RT protein and GIC:5' and GIC:3' modules in inserting the entire payload module into a desired target site.
[0535] Element: In this specification, the term "element" means an individual component of a molecule, system, or step in a method.
[0536] Expression product: In this specification, the term “expression product” means RNA transcribed from the sequence of interest (e.g., mRNA), or polypeptide translated from mRNA transcribed from the sequence of interest.
[0537] To enclose: In this specification, the term “enclose” means to contain, surround, or store.
[0538] Code: In this specification, the term “code” broadly refers to a process that uses information written into a polymer macromolecule (the first molecule) to induce the production of a second molecule distinct from the first molecule. The second molecule may have a chemical structure having different chemical properties from the first molecule.
[0539] Endonucleases: In this specification, the term "endonuclease" means a protein or part thereof that cleaves a polynucleotide chain by degrading nucleotides other than those at the ends.
[0540] Exosome: In this specification, “exosome” refers to a vesicle or complex secreted by mammalian cells that is involved in the degradation of RNA.
[0541] Ex vivo: The term "ex vivo" means taking cells from a donor subject, modifying the cells using the methods described herein, and transferring the cells to a recipient subject. This term includes autologous cells obtained from one individual subject (i.e., this subject is both a donor of unmodified cells and a recipient of cells modified ex vivo), and allogeneic cells obtained from a donor subject that is a different individual from the recipient subject. Allogeneic donors and recipients may be HLA-matched.
[0542] To facilitate: In this specification, the term “facilitate” is used in its broadest sense to mean making a certain action or process more likely to occur by adding a particular element.
[0543] Fidelity: In this specification, the term "fidelity" refers to the accuracy with which the target gene is inserted into the target genome. "High fidelity" corresponds to the insertion of the target gene with relatively few errors in nucleotide identity, sequence length, and target site location. For example, if a template RNA containing approximately 5,000 nucleotides can be copied by an RT protein to produce cDNA without producing mismatched base pairs, this gene insertion has high fidelity. Depending on the purpose of the transgene insertion, a small number of mismatches may occur, but a sufficiently high fidelity can still be achieved to produce a functional transgene.
[0544] Adjacent: In this specification, the term “adjacent” means that one element is located on the 5' side (5' adjacent) or 3' side (3' adjacent) of another element. Adjacent elements may be directly connected to each other, or another element may exist between adjacent elements.
[0545] Formulation: In this specification, “formulation” comprises at least one component of the gene insertion system described herein and at least one delivery agent or a pharmaceutically acceptable excipient or both.
[0546] Functional / Active Form: In this specification, the term “functional” as used in relation to biomolecules means a form of biomolecule that exhibits its characteristic properties and / or activity.
[0547] Gene: In this specification, the term “gene” is used in its broadest sense to mean a distinguishable nucleotide sequence that forms part of a chromosome, or a distinguishable nucleotide sequence that may form part of a chromosome, the order of which determines the order of monomers within a polypeptide or nucleic acid molecule.
[0548] Gene insertion construct: In this specification, the term “gene insertion construct” or “GIC” means an RNA construct containing an RNA template for a reverse transcriptase protein.
[0549] Gene Insertion System: In this specification, the term "gene insertion system," or "GIS," refers to a system consisting of components (modules) that may be used to insert a gene sequence (transgene) into a specific location in the target genome via reverse transcription such as TPRT.
[0550] GIC:3' side module: In this specification, the term “3' side module” means a portion of a gene insertion construct (GIC) that includes at least one element derived from the 3' region of a retroelement gene or at least one element that replaces the function of the 3' region of a retroelement gene.
[0551] GIC: 5' module: In this specification, the term “5' module” means a portion of a gene insertion construct (GIC) that facilitates the insertion of the entire length of the transgene, and may or may not originate from the 5' region of a retroelement gene.
[0552] To generate: In this specification, the verb “to generate” and its conjugations are used in their broadest sense to mean the process of obtaining a particular product.
[0553] Genome: In this specification, the term “genome” is used in its broadest sense to mean all genetic material present within a cell.
[0554] Hepatitis delta virus (HDV) ribozyme (RZ) folding structure: In this specification, the term "HDV RZ folding structure" means an RNA sequence that can take the folded structure of the ribozyme of hepatitis delta virus (HDV) and that retains the function of the ribozyme.
[0555] Heterogeneous: In this specification, the term “heterogeneous” means the gene or protein sequence or structure being introduced into a cell that does not normally produce such a sequence or structure. Furthermore, the term includes individual elements, modules, or parts of the reverse transcriptase construct or gene insertion construct of this disclosure that include nucleic acid sequences (DNA or RNA) or amino acid sequences derived from different species. For example, the 5' module of the reverse transcriptase construct or gene insertion construct may include a sequence derived from one species of bird (or a first species), and the 3' module of the reverse transcriptase construct or gene insertion construct may include a sequence derived from another species of bird (or a second species).
[0556] Homologous recombination: In this specification, the term “homologous recombination” means a transgene insertion process that depends on sequence homology between the transgene and the target genome.
[0557] In vitro: In this specification, the term "in vitro" means a reaction or process that takes place outside of a living cell or living organism.
[0558] In vivo: In this specification, the term "in vivo" means a reaction or process that takes place inside or on the surface of a living cell or living organism.
[0559] Inactive form: In this specification, the term “inactive form” as used in relation to biomolecules means a form of biomolecule that does not exhibit its characteristic properties and / or activity.
[0560] Inactive components: In this specification, the term “inactive component” means one or more agents that do not contribute to the activity of the active ingredient in the pharmaceutical composition contained in the formulation. In some embodiments, all of the inactive components that may be used in the formulation of the present invention may be approved by the U.S. Food and Drug Administration (FDA), some of them may be approved by the FDA, and all of the inactive components that may be used in the formulation of the present invention may not be approved by the FDA.
[0561] Induce: In this specification, the term "induce" and its grammatical equivalents mean a process that brings about a particular result without imposing any specific restrictions on the process.
[0562] To introduce: In this specification, the term “to introduce” means to add genetic material (often DNA) to a cell.
[0563] Insertion: In this specification, the term "insert" means adding a nucleotide to a DNA sequence.
[0564] Linkage: In this specification, the term "linkage" refers to the location where the cDNA of the transgene inserted into the target genome is linked to the DNA of the target genome at the insertion site.
[0565] At least one: In this specification, the term “at least one” means one, two, three, four, or five or more modified objects, e.g., a construct, module, or array of the present disclosure.
[0566] Lipid nanoparticles: In this specification, "lipid nanoparticles," or "LNPs," means a delivery medium containing one or more lipids (e.g., cationic lipids, non-cationic lipids, PEG-modified lipids).
[0567] Liposome: In this specification, "liposome" usually means a vesicle composed of one or more bilayers or spherical bilayers made of lipids (e.g., amphiphilic lipids).
[0568] Loss of function: In this specification, the term "loss of function" means a change in the target gene that results in a modified gene product in which the function of the wild-type gene is lost.
[0569] Modification: In this specification, “modification” means that a change has been made to the state or structure of the molecule. Molecules may be modified chemically, structurally, or functionally in a variety of ways.
[0570] Modular system: In this specification, “modular system” means a system that can be divided into multiple sets of parts that function relatively autonomously to each other but interact strongly with one another.
[0571] Motif: In this specification, the term "motif" means a sequence of biomolecules having a recognizable structure, which may or may not be defined by a specific chemical or biological function.
[0572] Natural: In this specification, the term “natural” means a wild-type or naturally occurring compound, biomolecule (e.g., protein or nucleic acid) or composition.
[0573] Non-LTR retroelement reverse transcriptase: In this specification, the term "non-LTR retroelement reverse transcriptase (RT)" means a protein having reverse transcription activity derived from a non-LTR retroelement.
[0574] Non-LTR retroelements: In this specification, the term "non-LTR retroelement" refers to a group of retroelement genes (also known as retrotransposons) that do not contain long-chain terminal repeats.
[0575] Outer: In this specification, the term “outer” in relation to an insertion site means any portion of the genome that is more than approximately 60 bp away from the 5' or 3' end of the insertion site.
[0576] Paired reverse transcriptase (RT): In this specification, “paired RT” means a reverse transcriptase (RT) used in combination with at least one module containing an insertion payload module. The module may be homogeneous with the paired RT, meaning that all elements contained in this module and RT originate from the same retroelement gene. The module may be heterogeneous with the paired RT, meaning that at least one element contained in this module does not originate from the same retroelement gene as RT.
[0577] Payload: The term "payload" may mean the sequence of nucleic acids (e.g., the target gene) contained within a gene insertion system (GIS) intended to be inserted into the target genome, unless used in a context related to the delivery medium.
[0578] Homologousity: The term “homologousity” or “homology (%)” refers to the amount of identical or same sequence between two nucleic acid sequences or amino acid sequences. As defined herein, the term “homologousity” may be used interchangeably with the terms “percentage of identity” or “percentage of sequence identity.”
[0579] In this specification, “degree of identity,” “degree of sequence identity,” or “homology” is determined by comparing two sequences in an optimized alignment within a comparison window, where some of the sequences within the comparison window may have additions or deletions (i.e., gaps) compared to a reference sequence (without additions or deletions) in order to optimize the alignment between the two sequences. The degree of sequence identity can be calculated by measuring the number of positions where identical nucleic acid bases or amino acid residues exist in both sequences, calculating the number of positions where nucleic acid bases or amino acid residues match, dividing the number of matched positions by the total number of positions in the comparison window, and multiplying the result by 100.
[0580] In the context of two or more nucleic acid sequences or polypeptide sequences, the terms “identical,” “identity,” or “homology” mean two or more identical sequences or subsequences. If, by measurement using one of the sequence comparison algorithms described below, or by measurement by manual alignment and visual inspection, multiple sequences have a predetermined proportion of the same nucleotide or amino acid residues within a comparison window or a given region, these sequences are "substantially identical" (for example, if they have at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.9% identity in a given region). This definition also applies to the complementary strand of the test sequence. Therefore, unless otherwise stated, all nucleic acid sequences and amino acid sequences provided herein include sequences that are substantially identical to the reference sequences.
[0581] In sequence comparison, typically one sequence is used as a reference sequence, and the test sequence is compared against it. When using a sequence comparison algorithm, the test sequence and reference sequence are entered into the computer, the coordinates of subsequences are specified as needed, and the parameters of the sequence algorithm program are specified. Generally, default program parameters are used, but different parameters can also be specified. The sequence comparison algorithm then calculates the ratio of sequence identity or sequence similarity of the test sequence to the reference sequence based on the program parameters.
[0582] Suitable algorithms for determining sequence identity and similarity include the BLAST algorithm described by Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977) and the BLAST 2.0 algorithm described by Altschul et al. (J. Mol. Biol. 215:403-10, 1990). Software for performing BLAST analysis is publicly available from the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). In this algorithm, high-scoring sequence pairs (HSPs) are first identified. Short words with length W in the query sequence are identified by alignment with words of the same length in the database sequence, and if these words match or satisfy a T-score, which is a threshold indicating a certain level of positive value, they are reported as HSPs. "T" is called the neighbor word score threshold (Altschul et al., cited above). The first identified neighbor word hits are used as a seed to start the search, and long HSPs containing these words are identified. Word hits are extended bidirectionally along each sequence, and the extension continues as long as the cumulative alignment score increases. For nucleotide sequences, the cumulative score is calculated using M (reward score for a matched pair of residues; always > 0) and N (penalty score for mismatched residues; always < 0). For amino acid sequences, the score matrix is used to calculate the cumulative score. The bidirectional extension of word hits stops when the cumulative alignment score begins to decrease by X from the maximum value, when the alignment of one or more residues with negative scores accumulates and the cumulative score becomes 0 or less, or when the end of one of the sequences is reached. The parameters W, T, and X of the BLAST algorithm determine the sensitivity and speed of alignment. The BLASTN program (for nucleotide sequences) uses the default parameters of word length (W) = 11, expected value (E) = 10, M = 5, N = -4, and comparison of both strands.When comparing amino acid sequences, the BLASTP program uses the following default parameters: word length = 3, expected value (E) = 10, alignment using the BLOSUM62 score matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915, 1989), (B) = 50, expected value (E) = 10, M = 5, and N = -4.
[0583] Furthermore, the BLAST algorithm performs a statistical analysis of the similarity between two sequences (see, for example, Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One of the similarity metrics provided by the BLAST algorithm is the sum of minimum probabilities (P(N)), which provides an indicator of the probability that a match between two nucleotide sequences or two amino acid sequences occurs by chance. For example, a nucleic acid is considered similar to a reference sequence if the sum of minimum probabilities when comparing the test nucleic acid to the reference nucleic acid is less than approximately 0.2, generally less than approximately 0.01, and more generally less than approximately 0.001.
[0584] Peptide: In this specification, "peptide" means an amino acid chain of 50 amino acids or less in length, for example, about 5 amino acids, about 10 amino acids, about 15 amino acids, about 20 amino acids, about 25 amino acids, about 30 amino acids, about 35 amino acids, about 40 amino acids, about 45 amino acids, or about 50 amino acids.
[0585] Pharmaceutical composition: In this specification, the term "pharmaceutical composition" means a composition comprising at least one active ingredient and which may further comprise one or more pharmaceutically acceptable excipients.
[0586] Polyadenosine: In this specification, the term "polyadenosine" means an adenosine nucleotide sequence of any length.
[0587] Polyadenosine tail: In this specification, the term "polyadenosine tail" or "poly A tail" means an adenosine nucleotide sequence of approximately 80 nucleotides or longer.
[0588] Polyadenosine tract: In this specification, the terms "polyadenosine tract," "poly-A tract," and "A tract" (all abbreviated as PA) refer to the same thing and are used interchangeably, meaning an adenosine nucleotide sequence approximately 1 to 50 nucleotides in length.
[0589] Promoter: In this specification, the term "promoter" means a DNA sequence that, upon protein binding, initiates transcription.
[0590] Proprotein: In this specification, the terms “protein precursor,” “proprotein,” and “propeptide” mean an inactive protein that can be converted into an active form through post-translational modification.
[0591] Protect: In this specification, the term “protect” and its grammatical equivalents mean a composition or process that prevents the degradation of all or part of a biomolecule.
[0592] Protein: In this specification, "protein" means a biomacromolecule consisting of more than 50 amino acids in length. Examples of proteins in this specification include, but are not limited to, enzymes, reverse transcriptases, and endonucleases.
[0593] Region: In this specification, the term “region” means a portion of a nucleotide sequence or a portion of an amino acid sequence. A region may be of unknown length or undefined, and in such cases, it may be defined by the function it performs or by its relative position to other elements in the sequence.
[0594] Retroelements / Retrotransposons: In this specification, the terms “retroelements” and “retrotransposons” are used interchangeably and refer to a group of eukaryotic cell genes that can replicate at new locations within their own genome via RNA intermediates.
[0595] Reverse transcriptase: In this specification, the term “reverse transcriptase” means a protein capable of synthesizing cDNA from an RNA template sequence.
[0596] Reverse transcriptase construct: In this specification, the term “reverse transcriptase construct (RTC)” means a biomolecular construct comprising or encoding at least one reverse transcriptase, as described above.
[0597] RTC:RT Module: In this specification, the terms “RTC:RT Module” or “reverse transcriptase module” mean a biomolecular construct comprising at least one reverse transcriptase or encoding at least one reverse transcriptase.
[0598] Ribosomal DNA: In this specification, the term "ribosomal DNA (rDNA)" means the portion of the target genome that encodes a ribosomal RNA precursor synthesized by RNAP I.
[0599] Ribosomal RNA: In this specif...
Claims
1. A genome editing system, (i) at least one reverse transcriptase construct (RTC), (ii) at least one gene insertion construct (GIC) Includes, The aforementioned RTC, (A) At least one reverse transcriptase module (RTC: RT module) containing mRNA encoding reverse transcriptase (RT); (B) at least one reverse transcriptase construct 5' module (RTC: 5' module); and (C) At least one reverse transcriptase construct 3' module (RTC: 3' module) Includes, The aforementioned GIC, It comprises at least one RNA template suitable for reverse transcription by the RT encoded by the at least one RTC, and (A) At least one GIC:5' side module comprising a sequence, rRNA sequence, ribozyme sequence, or folding motif sequence derived from the 5' region of a natural retroelement; (B) At least one GIC:payload module including an open reading frame (ORF) for the transgene; and (C) At least one GIC:3' side module containing a reverse transcriptase (RT) recognition sequence Includes, The RT recognition sequence included in the at least one GIC is derived from a different species of organism than the species of organism from which the RT encoded by the at least one RTC originates. The aforementioned RT is a system derived from birds.
2. The system according to claim 1, wherein the RT is derived from a zebra finch (Taeniopygia guttata), a white-throated sparrow (Zonotrichia albicollis), a white-throated swan (Tinamus guttatus), or a Galapagos finch (Geospiza fortis).
3. The system according to claim 2, wherein the RT includes an amino acid sequence having at least 90% identity with the sequence shown in SEQ ID NO: 29, SEQ ID NO: 27, SEQ ID NO: 20, SEQ ID NO: 18, or SEQ ID NO:
25.
4. The system according to claim 1, wherein the RT recognition sequence included in the GIC is derived from different birds.
5. The system according to claim 4, wherein the RT recognition sequence is derived from the Galapagos finch (Geospiza fortis), the zebra finch (Taeniopygia guttata), the white-throated bunting (Zonotrichia albicollis), or the white-throated swan (Tinamus guttatus).
6. The system according to claim 5, wherein the RT recognition sequence includes a sequence having at least 90% identity with the sequence shown in sequence number 177, sequence number 158, sequence number 176, or sequence number 178.
7. The system according to claim 1, wherein the at least one GIC:5' side module includes a ribozyme sequence.
8. The system according to claim 7, wherein the ribozyme sequence includes a hepatitis delta virus (HDV) ribozyme folding structure.
9. The system according to claim 7, wherein the ribozyme sequence encodes a ribozyme derived from the 5' region of a non-long-chain terminal repeat (non-LTR) retroelement.
10. The system according to claim 9, wherein the non-LTR type retro element is derived from the confused flour beetle, the three-spined stickleback, the American horseshoe crab, the northern stickleback, the parasitic wasp, the Galapagos finch, the medaka, the white-throated bunting, the zebra finch, or the white-throated swan.
11. The system according to claim 1, wherein the at least one GIC:5' side module includes a sequence having at least 90% identity with any one of the sequences shown in sequence number 141, sequence number 98, sequence numbers 60-97, sequence numbers 99-140, and sequence numbers 142-153.
12. The system according to claim 1, wherein the at least one RTC or the at least one GIC comprises at least one modified uracil.
13. The modified uracil is 5-methyluridine, 5-methoxyuridine, pseudouridine, N 1 -The system according to claim 12, wherein the system is methylpseudridine or 2-thiouridine.
14. A method for inserting at least one transgene into the genome of a cell in vitro or ex vivo, comprising the step of bringing at least one of the systems described in any one of claims 1 to 13 into contact with the cell in vitro or ex vivo, wherein the cell is not a human germline cell.
15. A system for therapeutic use, according to any one of claims 1 to 13.