Systems for targeted manipulation of polynucleotides and uses thereof
The TnpB system addresses the limitations of current gene-editing technologies by providing a smaller and more active polypeptide for targeted polynucleotide manipulation, enhancing therapeutic efficacy.
Patent Information
- Application Number
- PCT/CN2025/112793
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-06
- Filing Date
- 2025-08-05
- Publication Date
- 2026-02-12
AI Technical Summary
Current gene-editing systems, such as those derived from the CRISPR-Cas system, are limited by size constraints and low endonuclease activity in mammalian cells, and their PAM sequences restrict therapeutic applications.
A system utilizing TnpB, a smaller and more versatile polypeptide with nickase activity, combined with a guide polynucleotide and scaffold, enables targeted manipulation of polynucleotides, overcoming size and activity limitations.
The TnpB-based system facilitates efficient and targeted editing of polynucleotides, including insertion, deletion, and substitution, with potential therapeutic applications in various diseases and conditions.
Smart Images

Figure PCTCN2025112793-FTAPPB-I100001 
Figure PCTCN2025112793-FTAPPB-I100002 
Figure PCTCN2025112793-FTAPPB-I100003
Abstract
Description
SYSTEMS FOR TARGETED MANIPULATION OF POLYNUCLEOTIDES AND USES THEREOFCROSS REFERENCE TO RELATED APPLICATION
[0001] This application claims priority to International Application No. PCT / CN2024 / 110103, filed on August 6, 2024, the content of which is incorporated by reference in its entirety. FIELD OF DISCLOSURE
[0002] The present disclosure relates to systems for targeted manipulation of polynucleotides, molecules associated therewith, and uses thereof. SEQUENCE LISTING
[0003] The present application includes an XML-format Sequence Listing named Seq. xml, created on August 6, 2024 with a size of 486, 397 bytes, which is incorporated herein in its entirety.BACKGROUND
[0004] Mobile Genetic Elements (MGEs) are DNA sequences capable of moving within a genome or between different genomes. MGEs exist in a wide variety of organisms and play a key role in evolution. Insertion Sequence (IS) is one of the simplest kinds of MGEs and can be categorized based on the characteristics of the transposases involved. IS200 / 605 transposase belongs to the HuH super family (Frost et al., Nat Rev Microbiol, 2005, 3: 722-732) . Based on the presence of TnpA and TnpB in the IS, IS200 / 605 can be divided into three subgroups, namely IS200 (containing only TnpA) , IS1341 (containing only TnpB) , and IS605 (containing both TnpA and TnpB) (He et al., Microbiol Spectr, 2015, 3 (4) : MDNA3-0039-2014) . IS605, as illustrated in Fig. 1, has both TnpA-and TnpB-coding sequences flanked by two palindromic elements, left end (LE) and right end (RE) . There is a tetra-or penta-nucleotide cleavage site upstream of LE (CL) and downstream of RE (CR) , respectively, and a tetra-nucleotide guide sequence between the cleavage site and the LE (GL) and between TnpB-encoding sequence and RE (GR) respectively, in which CL and GL, as well as CR and GR, are reversely complementary to each other respectively (He et al., Nucleic Acids Res, 2011, 39 (19) : 8503-8512; He et al., Microbiol Spectr, 2015, 3 (4) ) .
[0005] Different from other ISs, which transpose via double-stranded DNA (dsDNA) intermediates, IS200 / 605 uses single-stranded DNA (ssDNA) as transposition intermediates (Guynet et al., Mol Cell, 2008, 29: 302-312; He et al., 2015) . IS200 / IS605 transposes by a TnpA dimer-mediated “peel and paste” process, in which the transposon is excised from the lagging strand and form a circular ssDNA intermediate, which is then inserted to the target site (Siguier et al., FEMS Microbiol Rev, 2014, 38 (5) : 865-891; Barabas et al., Cell, 2008, 132 (2) : 208-120) . Certain TnpB was considered as an ancestor of Cas12 based on phylogenetic analyses (Shmakov et al., Nat Rev Microbiol, 2017, 15 (3) : 169-182; Altae-Tran et al., Science, 2021, 374 (6563) : 57-65) and was later found to share 3D structural similarity with Cas12f (Shmakov et al., 2017; Makarova et al., Nat Rev Microbiol, 2020, 18 (2) : 67-83) . Certain TnpB is an RNA-guided endonuclease capable of cutting dsDNA in eukaryotes (Karvelis et al., Nature, 2021, 599 (7886) : 692-696) . Non-coding RNAs (ncRNAs) overlapping the 3’ region of the tnpB gene have been identified in bacteria and archaea. Such ncRNAs are referred to as “right-end element RNAs (reRNAs) ” or ωRNAs and considered to provide a guiding function analogous to the CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA) in CRISPR-Cas ( et al., BMC Genomics, 2014, 15 (1) : 684; Gomes-Filho et al., RNA Biol, 2015, 12 (5) : 490-500) . Therefore, the reRNA may also be considered as guide RNA. TnpB-mediated cleavage requires a 4-or 5-nt sequence, named Transposon-Associated Motif (TAM) . A TnpB-mediated “peel, paste and copy” process has been proposed, in which TnpB cleaves the donor joint generated after transposon excision to form a double-strand break (DSB) , which triggers homology-directed repair to reinstate the transposon into its original site (Karvelis et al., 2021) .
[0006] A wide variety of gene-editing systems derived from the CRISPR-Cas system have been developed, which enable deletion, insertion, and base-editing in the genome and have great therapeutic potential (Adli, Nat Commun, 2018, 9 (1) : 1911) . The gene-editing systems are usually delivered into a subject by adeno-associated virus (AAV) -based methods with a capacity of no more than approximately 4.7 kb, which limits the therapeutic applications of those systems. Although the later discovered Cas12f is smaller in size (400-700 amino acids) , wild-type Cas12f has low endonuclease activity in mammalian cells. Moreover, the PAMs of Cas12f are mostly TTN, which further limits its therapeutic application (Kim et al., Nat Biotechnol, 2022, 40 (1) : 94-102; Xu et al., Mol Cell, 2021, 81 (20) : 4333-4345) . TnpB, on the other hand, has an even smaller size (typically 300-400 amino acids) and a variety of TAM sequences, and has the potential of overcoming the limitations of the currently available gene-editing tools.SUMMARY
[0007] The present disclosure relates to a system for targeted polynucleotide manipulation and / or a molecule associated therewith. It also discloses methods and applications of the system.
[0008] In a first embodiment, a system for manipulating a polynucleotide of interest, comprising (i) a polypeptide comprising a TnpB or variant thereof or a first polynucleotide encoding the polypeptide; and (ii) a guide polynucleotide comprising a targeting sequence and a scaffold or a second polynucleotide encoding the guide polynucleotide; wherein the polypeptide is capable of specifically binding to the guide polynucleotide via the scaffold to form a complex; and wherein the targeting sequence is capable of specifically binding to the polynucleotide of interest.
[0009] In a second embodiment, in the system disclosed herein, (i) the first polypeptide is provided by the first polynucleotide; and (ii) the guide polynucleotide is provided by second polynucleotide.
[0010] In a third embodiment, the TnpB or variant thereof in the system disclosed herein has nickase activity.
[0011] In a fourth embodiment, the targeting sequence in the system disclosed herein is about 10 to about 50 nt in length.
[0012] In a fifth embodiment, the scaffold in the system disclosed herein is about 50 nt to about 250 nt in length.
[0013] In a sixth embodiment, the polypeptide in the system disclosed herein comprises an amino acid sequence sharing at least 80%sequence identity to one or more of SEQ ID NOs: 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90.
[0014] In a seventh embodiment, the polypeptide in the system disclosed herein comprises: (i) SEQ ID NO: 3 or a variant thereof, wherein the variant comprises one or more mutations selected from Q72K, T88R, T112R, L160K, I172R, Q210K, and L351R; (ii) SEQ ID NO: 8 or a variant thereof, wherein the variant comprises one or more mutations selected from S54K and A66K; (iii) SEQ ID NO: 15 or a variant thereof, wherein the variant comprises one or more mutations selected from G56K, E64K, T87K, Q90K, Q97R, N111K, E202R, Q258K, and Q261K; (iv) SEQ ID NO: 18 or a variant thereof, wherein the variant comprises one or more mutations selected from S88K, E97K, S105R, D108K, I110R, I232K, D250R, D316R, and E323R; (v) SEQ ID NO: 19 or a variant thereof, wherein the variant comprises one or more mutations selected from D44H, S88K, S105K, E217K, and I237K; (vi) SEQ ID NO: 22 or a variant thereof, wherein the variant comprises one or more mutations selected from L224R and A326K; (vii) SEQ ID NO: 28 or a variant thereof, wherein the variant comprises one or more mutations selected from D173K and E313K; (viii) SEQ ID NO: 39 or a variant thereof, wherein the variant comprises one or more mutations selected from S83K, S88H, S92K, D100K, G109K, S111R, F120K, Y144R, Q234R, Q248R, A277K, and Q310R; (ix) SEQ ID NO: 48 or a variant thereof, wherein the variant comprises one or more mutations selected from S83K, S88H, S92K, D100K, G109K, S111R, F120K, Y144R, Q234R, Q248R, A277K, and Q310R; (x) SEQ ID NO: 55 or a variant thereof, wherein the variant comprises one or more mutations selected from E40K, G90R, G95K, T121H, S251K, S281K, and I301K; (xi) SEQ ID NO: 56 or a variant thereof, wherein the variant comprises one or more mutations selected from Q90K, S108H, N111K, E202K, L225R, T282K, and E333K; (xii) SEQ ID NO: 57 or a variant thereof, wherein the variant comprises one or more mutations selected from N44K, A59K, Q80K, Q87K, E97K, Q131K, E201H, A209R, A216R, Q221K, Q238K, A274K, E295K, T313K, and V350H; (xiii) SEQ ID NO: 59 or a variant thereof, wherein the variant comprises one or more mutations selected from A135K, D184K, V240R, T254K, and A298K; (xiv) SEQ ID NO: 63 or a variant thereof, wherein the variant comprises a mutation of E291K; (xv) SEQ ID NO: 65 or a variant thereof, wherein the variant comprises one or more mutations selected from E42K, N45K, C62K, A96R, I99R, C214R, N224K, L235R, N245R, I254R, D292R, and T310R; (xvi) SEQ ID NO: 72 or a variant thereof, wherein the variant comprises one or more mutations selected from N93K, N223K, and E306K; (xvii) SEQ ID NO: 76 or a variant thereof, wherein the variant comprises one or more mutations selected from D124G, E144K, L148K, N155K, T197K, A212R, I217K, S272K, N279K, T309K, and Q326K; and / or (xviii) SEQ ID NO: 90 or a variant thereof, wherein the variant comprises one or more mutations selected from S188K, V232K, D233R, N243R, M272K, and T360R.
[0015] In an eighth embodiment, the polypeptide in the system disclosed herein comprises: (i) SEQ ID NO: 18 or a variant thereof, wherein the variant comprises a mutation combination of E97K and S88K; E97K and S105R; E97K and D108K; E97K and I110R; E97K and I232K; E97K and D250R; E97K and D316R; E97K and E323R; D108K and S88K; D108K and S105R; D108K and I110R; D108K and I232K; D108K and D250R; D108K and D316R; D108K and E323R; D108K, I232K, and E97K; E97K, I110R, and D108K; E97K, I110R, and D250R; or E97K, I110R, and D316R; (ii) (SEQ ID NO: 76 or a variant thereof, wherein the variant comprises a mutation combination of D40K and D97K; D40K and N155K; D40K and T309K; or N155K and T197K; and / or (iii) SEQ ID NO: 90 or a variant thereof, wherein the variant comprises a mutation combination of V232K and N243R; D233R and N243R; V232K, D233R, and N243R, M272K and N243R; or T360R and N243R.
[0016] In a ninth embodiment, the scaffold in the system disclosed herein has at least 80%sequence identity to a polynucleotide sequence selected from SEQ ID NOs: 189, 194, 201, 204, 205, 208, 214, 225, 234, 241, 242, 243, 245, 249, 251, 258, 261.262, 264, and 276 or a fragment thereof.
[0017] In a tenth embodiment, the scaffold in the system disclosed herein comprises: (i) about 160 nt from 3’ -end of SEQ ID NO: 189; (ii) about 120 nt to about 160 nt from 3’ -end of SEQ ID NO: 194; (iii) about 125 nt to about 200 nt from 3’ -end of SEQ ID NO: 201; (iv) about 125 nt to about 200 nt from 3’ -end of SEQ ID NO: 204; (v) about 125 nt to about 200 nt from 3’ -end of SEQ ID NO: 205; (vi) about 130 nt to about 200 nt from 3’ -end of SEQ ID NO: 208; (vii) about 120 nt to about 140 nt from 3’ -end of SEQ ID NO: 214; (viii) about 180 nt from 3’ -end of SEQ ID NO: 225; (ix) about 160 nt to about 180 nt from 3’ -end of SEQ ID NO: 234; (x) about 120 nt to about 180 nt from 3’ -end of SEQ ID NO: 241; (xi) about 140 nt to about 180 nt from 3’ -end of SEQ ID NO: 242; (xii) about 120 nt to about 180 nt from 3’ -end of SEQ ID NO: 243; (xiii) about 140 nt to about 180 nt from 3’ -end of SEQ ID NO: 245; (xiv) about 140 nt to about 180 nt from 3’ -end of SEQ ID NO: 249; (xv) about 140 nt to about 180 nt from 3’ -end of SEQ ID NO: 251; (xvi) about 140 nt to about 180 nt from 3’ -end of SEQ ID NO: 258; (xvii) about 140 nt to about 200 nt from 3’ -end of SEQ ID NO: 261; (xviii) about 100 nt to about 200 nt from 3’ -end of SEQ ID NO: 262; (xix) about 140 nt to about 200 nt from 3’ -end of SEQ ID NO: 264; (xx) about 140 nt to about 200 nt from 3’ -end of SEQ ID NO: 276; (xxi) SEQ ID NO: 285; (xxii) SEQ ID NO: 286; (xxiii) SEQ ID NO: 292; (xxiv) SEQ ID NO: 305; (xxv) SEQ ID NO: 306; (xxvi) SEQ ID NO: 310; (xxvii) SEQ ID NO: 311; (xxviii) SEQ ID NO: 312; (xxix) SEQ ID NO: 314; (xxx) SEQ ID NO: 326; and / or (xxxi) SEQ ID NO: 333.
[0018] In an eleventh embodiment, in the system disclosed herein, (i) the polypeptide comprises SEQ ID NO: 3 or a variant thereof and the scaffold comprises SEQ ID NO: 189 or a fragment thereof; (ii) the polypeptide comprises SEQ ID NO: 8 or a variant thereof and the scaffold comprises SEQ ID NO: 194 or a fragment thereof; (iii) the polypeptide comprises SEQ ID NO: 15 or a variant thereof and the scaffold comprises SEQ ID NO: 201 or a fragment thereof; (iv) the polypeptide comprises SEQ ID NO: 18 or a variant thereof and the scaffold comprises SEQ ID NO: 204 or a fragment thereof; (v) the polypeptide comprises SEQ ID NO: 19 or a variant thereof and the scaffold comprises SEQ ID NO: 205 or a fragment thereof; (vi) the polypeptide comprises SEQ ID NO: 22 or a variant thereof and the scaffold comprises SEQ ID NO: 208 or a fragment thereof; (vii) the polypeptide comprises SEQ ID NO: 28 or a variant thereof and the scaffold comprises SEQ ID NO: 214 or a fragment thereof; (viii) the polypeptide comprises SEQ ID NO: 39 or a variant thereof and the scaffold comprises SEQ ID NO: 225 or a fragment thereof; (ix) the polypeptide comprises SEQ ID NO: 48 or a variant thereof and the scaffold comprises SEQ ID NO: 234 or a fragment thereof; (x) the polypeptide comprises SEQ ID NO: 55 or a variant thereof and the scaffold comprises SEQ ID NO: 241 or a fragment thereof; (xi) the polypeptide comprises SEQ ID NO: 56 or a variant thereof and the scaffold comprises SEQ ID NO: 242 or a fragment thereof; (xii) the polypeptide comprises SEQ ID NO: 57 or a variant thereof and the scaffold comprises SEQ ID NO: 243 or a fragment thereof; (xiii) the polypeptide comprises SEQ ID NO: 59 or a variant thereof and the scaffold comprises SEQ ID NO: 245 or a fragment thereof; (xiv) the polypeptide comprises SEQ ID NO: 63 or a variant thereof and the scaffold comprises SEQ ID NO: 249 or a fragment thereof; (xv) the polypeptide comprises SEQ ID NO: 65 or a variant thereof and the scaffold comprises SEQ ID NO: 251 or a fragment thereof; (xvi) the polypeptide comprises SEQ ID NO: 72 or a variant thereof and the scaffold comprises SEQ ID NO: 258 or a fragment thereof; (xvii) the polypeptide comprises SEQ ID NO: 75 or a variant thereof and the scaffold comprises SEQ ID NO: 261 or a fragment thereof; (xviii) the polypeptide comprises SEQ ID NO: 76 or a variant thereof and the scaffold comprises SEQ ID NO: 262 or a fragment thereof; (xix) the polypeptide comprises SEQ ID NO: 78 or a variant thereof and the scaffold comprises SEQ ID NO: 264 or a fragment thereof; and / or (xx) the polypeptide comprises SEQ ID NO: 90 or a variant thereof and the scaffold comprises SEQ ID NO: 276 or a fragment thereof.
[0019] In a twelfth embodiment, the first polynucleotide and / or the second polynucleotide in the system disclosed herein is operably linked to a regulatory sequence.
[0020] In a thirteenth embodiment, the system disclosed herein further comprises a donor polynucleotide.
[0021] In a fourteenth embodiment, the system disclosed herein comprises a third polynucleotide comprising the first polynucleotide and the second polynucleotide.
[0022] In a fifteenth embodiment, the third polynucleotide in the system disclosed herein further comprises the donor polynucleotide.
[0023] In a sixteenth embodiment, the system disclosed herein comprises a third polynucleotide comprising the first polynucleotide and a fourth polynucleotide comprises the second polynucleotide.
[0024] In a seventeenth embodiment, at least one of the third polynucleotide and the fourth polynucleotide in the system disclosed herein further comprise (s) the donor polynucleotide.
[0025] In an eighteenth embodiment, the system disclosed herein comprises a fifth polynucleotide comprising the donor polynucleotide.
[0026] In a nineteenth embodiment, the first polynucleotide, the second polynucleotide, the donor polynucleotide, or a combination thereof in the system disclosed herein is on a vector.
[0027] In a twentieth embodiment, the vector in the system disclosed herein is a viral vector.
[0028] In a twenty-first embodiment, the viral vector in the system disclosed herein is a lentiviral vector, a retroviral vector, an adenoviral vector, an adeno-associated viral vector, a herpes simplex viral vector, or a combination thereof.
[0029] In a twenty-second embodiment, the vector in the system disclosed herein is a non-viral vector.
[0030] In a twenty-third embodiment, the non-viral vector in the system disclosed herein is a plasmid.
[0031] In a twenty-fourth embodiment, the system disclosed herein further comprises a reporter polynucleotide capable of being cleaved by the complex that specifically binds to the polynucleotide of interest.
[0032] In a twenty-fifth embodiment, the vector in the system disclosed herein is attached to or embedded in an auxiliary substance.
[0033] In a twenty-sixth embodiment, the auxiliary substance in the system disclosed herein comprises a viral capsid.
[0034] In a twenty-seventh embodiment, the auxiliary substance in the system disclosed herein comprises a lipid layer, a polymer, an inorganic molecule, a nanoparticle, or a combination thereof
[0035] In a twenty-eighth embodiment, the system disclosed herein is used for introducing an insertion, a deletion, and / or a substitution of at least one nucleotide to the polynucleotide of interest.
[0036] In a twenty-ninth embodiment, the system disclosed herein is used for inserting the donor polynucleotide or fragment thereof to the polynucleotide of interest.
[0037] In a thirtieth embodiment, the system disclosed herein is attached to or embedded in a support.
[0038] In a thirty-first embodiment, a cell and / or progeny thereof comprises the system disclosed herein.
[0039] In a thirty-second embodiment, the system disclosed herein is transiently present in the cell and / or progeny thereof of the thirty-first embodiment.
[0040] In a thirty-third embodiment, the system disclosed herein is constitutively present in the cell and / or progeny thereof of the thirty-first embodiment.
[0041] In a thirty-fourth embodiment, the cell that comprises the system disclosed herein is a prokaryotic cell.
[0042] In a thirty-fifth embodiment, the cell that comprises the system disclosed herein is a eukaryotic cell.
[0043] In a thirty-sixth embodiment, the cell that comprises the system disclosed herein is a plant cell.
[0044] In a thirty-seventh embodiment, the cell that comprises the system disclosed herein is a non-human animal cell.
[0045] In a thirty-eighth embodiment, the cell that comprises the system disclosed herein is a human cell.
[0046] In a thirty-ninth embodiment, an organoid comprises the cell or progeny thereof that comprises the system disclosed herein.
[0047] In a fortieth embodiment, an organism comprises the cell or progeny thereof that comprises the system disclosed herein.
[0048] In a forty-first embodiment, a method for editing a polynucleotide of interest comprises (i) providing the polynucleotide of interest; and (ii) contacting the system disclosed herein with the polynucleotide of interest.
[0049] In a forty-second embodiment, a method for editing a polynucleotide of interest comprises introducing the system disclosed herein to a cell comprising the polynucleotide of interest.
[0050] In a forty-third embodiment, a method for preventing or treating a disease or condition in a subject in need thereof comprises introducing the system disclosed herein to at least one cell in the subject.
[0051] In a forty-fourth embodiment, a method for preventing or treating a disease or condition in a subject in need thereof comprises introducing to the subject one or more of the cells or progenies thereof that comprises the system disclosed herein.
[0052] In a forty-fifth embodiment, a method for preventing or treating a disease or condition in a subject in need thereof comprises introducing the organoid that comprises the system disclosed herein or one or more cells derived therefrom to the subject.
[0053] In a forty-sixth embodiment, the subject in the method disclosed herein is a human.
[0054] In a forty-seventh embodiment, the disease or condition in the prevention or treatment method disclosed herein is one or more of Alzheimer’s disease, myotonic dystrophy type 1, spinal muscular atrophy, Huntington’s disease, sickle cell disease, β-thalassemia, Duchenne muscular dystrophy, and hereditary tyrosinemia.
[0055] In a forty-eighth embodiment, the polynucleotide of interest in the prevention or treatment method disclosed herein is one or more of APP, Bace1, GSAP, APOE, CD33, GMF, CysLT1R, DMPK, SMN1, SMN2, HTT, BCL11A, ESE, HBB, Dmd, and FAH.
[0056] In a forty-ninth embodiment, in the prevention or treatment method disclosed herein, (i) the disease or condition in is Alzheimer’s disease, and the polynucleotide of interest comprises APP, Bace1, GSAP, APOE, CD33, GMF, CysLT1R, EMX1, VEGFA, DNMT1, ALDH1A3, TET1, TET2, RNF2, RUNX1, PCSK9, CXCR4, APOB, DNMT3b, PGK1, MECP2, and / or AGBL1; (ii) the disease or condition is myotonic dystrophy type 1, and the polynucleotide of interest comprises DMPK; (iii) the disease or condition is spinal muscular atrophy, and the polynucleotide of interest comprises SMN1 and / or SMN2; (iv) the disease or condition is Huntington’s disease, and the polynucleotide of interest comprises HTT; (v) the disease or condition is sickle cell disease, and the polynucleotide of interest comprises BCL11A, ESE, and / or HBB; (vi) the disease or condition is β-thalassemia, and the polynucleotide of interest comprises BCL11A, ESE, and / or HBB; (vii) the disease or condition is Duchenne muscular dystrophy, and the polynucleotide of interest comprises Dmd; and / or (viii) the disease or condition is hereditary tyrosinemia, and the polynucleotide of interest comprises FAH.
[0057] In a fiftieth embodiment, a method for detecting a target polynucleotide comprises contacting the system disclosed herein with an analyte.
[0058] In a fifty-first embodiment, the analyte in the detection method disclosed herein comprises a sample derived from a biopsy, a body fluid, feces, or a combination thereof from a subject.
[0059] In a fifty-second embodiment, the detection method disclosed herein is for diagnosis in the subject.
[0060] In a fifty-third embodiment, the subject in the detection method disclosed herein is a human.
[0061] In a fifty-fourth embodiment, a system for manipulating a polynucleotide of interest comprises (i) a fusion protein comprising a TnpB or variant thereof and an effector or a first polynucleotide encoding the fusion protein; and (ii) a guide polynucleotide comprising a targeting sequence and a scaffold or a second polynucleotide encoding the guide polynucleotide; wherein the fusion protein is capable of specifically binding to the guide polynucleotide via the scaffold to form a complex; and wherein the targeting sequence is capable of specifically binding to the polynucleotide of interest.
[0062] In a fifty-fifth embodiment, the TnpB variant in the system disclosed herein has weakened nuclease activity compared to the TnpB.
[0063] In a fifty-sixth embodiment, the TnpB variant in the system disclosed herein comprises the amino acid of SEQ ID NO: 8 and one or more mutations of D185A, R258A, and E279A.
[0064] In a fifty-seventh embodiment, the effector in the system disclosed herein comprises a nuclear localization signal (NLS) .
[0065] In a fifty-eighth embodiment, the effector in the system disclosed herein comprises a cell penetrating peptide.
[0066] In a fifty-ninth embodiment, the system disclosed herein is attached to or embedded in a support.
[0067] In a sixtieth embodiment, the effector in the system disclosed herein comprises one or more of a nuclease that is not the TnpB or variant thereof, a base editor, a prime editor, an epigenetic modifier, a polymerase, a transposase, a recombinase, a reverse transcriptase, a label, a transcription modulator, a transcription factor, a topoisomerases, a helicase, a kinase, a phosphatase, an integrase, a ligase, an antibody, and an antigen.
[0068] In a sixty-first embodiment, the base editor in the system disclosed herein is a deaminase, optionally a cytidine deaminase and / or an adenine deaminase.
[0069] In a sixty-second embodiment, the label in the system disclosed herein is a fluorescent, luminescent, and / or a chromogenic protein.
[0070] In a sixty-third embodiment, a method for manipulating a polynucleotide of interest and / or one or more molecules associated therewith comprises contacting the system disclosed herein with the polynucleotide of interest.
[0071] In a sixty-fourth embodiment, a method for modifying the polynucleotide of interest and / or one or more molecules associated therewith comprises contacting the system disclosed herein with the polynucleotide of interest.
[0072] In a sixty-fifth embodiment, a method for modulating transcription, replication, repair, and / or recombination of a polynucleotide of interest comprises contacting the system disclosed herein with the polynucleotide of interest
[0073] In a sixty-sixth embodiment, a method for promoting or preventing direct or indirect binding of a molecule to the polynucleotide of interest comprises contacting the system disclosed herein with the polynucleotide of interest.
[0074] In a sixty-seventh embodiment, a method for labeling a polynucleotide of interest comprises contacting the system disclosed herein with the polynucleotide of interest.BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Fig. 1. Illustration of components of IS605. LE: Left End; RE: Right End; CL and CR: tetra-or penta-nucleotide cleavage sites near LE and inside RE, respectively; GL and GR: tetra-nucleotide guide sequences inside LE and RE, respectively. A subterminal palindromes (astem-loop structure) is predicted to be between GL and TnpA and between GR and CR, respectively.
[0076] Figs. 2A-2D. Two different sets of plasmids for confirming the RNA-guided endonuclease activity of TnpB. The first set includes a TnpB-reRNA plasmid (Fig. 2A) and an EGFxxFP reporter plasmid (Fig. 2B) ; the second set includes a TnpB-reRNA plasmid (Fig. 2C) and an RFP-EGFP reporter plasmid (Fig. 2D) . U6: U6 promoter; reRNA: right-end RNA; EF-1α: Elongation Factor-1α promoter; NLS: Nuclear Localization Sequence; CMV: cytomegalovirus promoter; T2A: T2A peptide; TAM: Transposon-Associated Motif; RFP: red fluorescent protein; EGFP: enhanced green fluorescent protein.
[0077] Figs. 3A-3H. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells or RFP+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding TnpB 1-93 (Fig. 3A: TnpB 1-10, Fig. 3B: TnpB 11-25, Fig. 3C: TnpB 26-35 Fig. 3D: TnpB 36-45 Fig. 3E: TnpB 46-56 Fig. 3F: TnpB 57-65, Fig. 3G: TnpB 66-74, and Fig. 3H: TnpB 75-93) together with the reporter plasmids. Target 1, 2, or 3 was used as the targeting sequence as noted in these figures. Data are shown as the mean ± standard deviation (s. d. ) of three biological replicates.
[0078] Figs. 4A-4D. Fluorescent microscopic images of cells co-transfected with TnpB-reRNA plasmids encoding TnpB 8 (Fig. 4A) , 18 (Fig. 4B) , 19 (Fig. 4C) , and 22 (Fig. 4D) respectively together with the EGFxxFP reporter plasmids (upper panels) or with only the TnpB-reRNA plasmids as controls (lower panels) .
[0079] Figs. 5A-5X. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells or RFP+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, or 90 (Figs. 5A-5T: TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90, respectively; Figs. 5U-5X: TnpB 15, 18, 19, and 22, respectively) and their corresponding reRNAs with inferred wild-type scaffolds or 5’-truncated scaffolds together with the EGFxxFP reporter plasmids. The lengths of the truncated reRNA scaffolds are indicated in the figures. Target 1, 2, or 3 was used as the targeting sequence as indicated in the figures. Data are shown as the mean ± s. d. of three biological replicates.
[0080] Figs. 6A-6S. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells or RFP+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 56, 57, 59, 63, 65, 72, 75, 76, 78, or 90 (Figs. 6A-6S respectively) and their corresponding with inferred wild-type reRNA scaffolds or reRNA scaffolds containing deletions together with the corresponding reporter plasmids. Target 1, 2, or 3 was used as the targeting sequence as indicated in the figures. Data are shown as the mean ± s. d. of three biological replicates.
[0081] Figs. 7A-7T. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells or RFP+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding variants of TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, or 90 (Figs. 7A-7T, respectively) together with the corresponding reporter plasmids. Each of the variants contains one mutation. Target 1, 2, or 3 was used as the targeting sequence as indicated in the figures. Data are shown as the mean ± s. d. of three biological replicates.
[0082] Figs. 8A-8D. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells or RFP+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding variants of TnpB 18 (Figs. 8A and 8B) , 76 (Fig. 8C) , or 90 (Fig. 8D) together with the corresponding reporter plasmids. Each of the variant contains more than one mutations. Data are shown as the mean ± s. d. of three biological replicates.
[0083] Figs. 9A-9C. Endogenous gene editing efficiencies of TnpBs. The percentages of indels resulted from TnpB 8 (Fig. 9A) , 18 (Fig. 9B) , and 78 (Fig. 9C) -mediated endogenous gene editing in different target genes are shown. Data are shown as the mean ± s. d. of three biological replicates.
[0084] Fig. 10. The percentages of GFP+ cells in total co-transfected cells (mCherry+ cells) in cultures co-transfected with TnpB-reRNA plasmids encoding TnpB 8 variants with a D185A, R258A, or E279A mutation together with the corresponding reporter plasmids. Data are shown as the mean ± s. d. of three biological replicates.
[0085] Figs. 11A-11T. Predicted structures of 180-nt reRNA scaffolds (180 nt from 3’ -end of the inferred wild-type scaffolds in Table 1) of TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90 (Figs. 11A-11T, respectively) .DETAILED DESCRIPTION
[0086] All publications, patents, and patent applications referred to herein are incorporated by reference in their entirety to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference in its entirety.
[0087] Any embodiment of the disclosure described herein, including those described only in one section of the specification describing a specific aspect of the disclosure, and those described only in the examples or drawings, can be combined with any other one or more embodiments, unless explicitly disclaimed or improper.
[0088] In the present disclosure, unless otherwise specified, the scientific and technical terms used herein have meanings generally understood by a person skilled in the art. Although any methods and materials similar or equivalent to those described herein find use in the practice of the present disclosure, the preferred methods and materials are described herein. Accordingly, the terms defined herein are more fully described by reference to the Specification as a whole. Definitions
[0089] Words using the singular include the plural, and vice versa, unless the context clearly dictates otherwise.
[0090] In this disclosure, many terms and abbreviations are used. The following definitions apply unless specifically stated otherwise.
[0091] As used herein, the singular forms “a, ” “an, ” and “the” include plural references unless the context clearly dictates otherwise. For example, the term “acompound” or “at least one compound” may include a plurality of compounds, including mixtures thereof. The terms “a, ” “an, ” “the, ” “one or more, ” and “at least one, ” for example, can be used interchangeably herein.
[0092] As used herein, the term “about” will be understood by persons of ordinary skill in the art and will vary to some extent depending on the context in which it is used. In some embodiments, the term “about” when referring to a value is meant to encompass art-accepted variations. In some embodiments, the term “about” when referring to such values, is meant to encompass variations of ±20%or ±10%, more preferably ±5%, even more preferably ±1%, and still more preferably ±0.1%from the specified value, as such variations are appropriate in the context in which the term “about” is used.
[0093] The terms “and / or” and “or” are used interchangeably herein and refer to a specific disclosure of each of the two specified features or components with or without the other. Thus, the term “and / or” as used in a phrase such as “A and / or B” herein is intended to include “A and B, ” “A or B, ” “A” (alone) , and “B” (alone) . Likewise, the term “and / or” as used in a phrase such as “A, B and / or C” is intended to encompass each of the following aspects: “A and B and C” ; “A or B or C” ; “A or C” ; “A or B” ; “B or C” ; “A and C” ; “A and B” ; “B and C” ; “A” (alone) ; “B” (alone) ; and “C” (alone) .
[0094] Throughout this application, various embodiments can be presented in a range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the embodiments described herein. Accordingly, the description of a range should be considered to have specifically disclosed all the possible subranges as well as individual numerical values within that range. For example, description of a range, such as from 1 to 6 should be considered to have subranges such as from 1 to 2, from 1 to 3, from 1 to 4 and from 1 to 5, from 2 to 3, from 2 to 4, from 2 to 5, from 2 to 6, from 3 to 4, from 3 to 5, from 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5 and 6. This applies regardless of the breadth of the range.
[0095] The terms “specifically bind to, ” “bind to, ” and “recognizes” refer to binding that is measurably different from a non-specific interaction. For example, in some embodiments, a binding molecule, such as a receptor, specifically binds to a target molecule, such as a ligand, when the binding molecule reacts or associates more frequently, more rapidly, with greater duration, and / or with greater affinity with the particular target molecule than it does with alternative molecules. A binding molecule “specifically binds” to a target molecule if it binds with greater affinity, avidity, more readily, and / or with greater duration than it binds to other molecules. It is understood that a binding molecule that specifically binds to a first target molecule may or may not specifically bind to a second target molecule. As such, “specific binding” does not necessarily require (although it can include) exclusive binding. In some embodiments, specific binding can be determined, for example, by comparing binding of a particular binding molecule with binding of another molecule that does not bind to a particular target molecule. Specific binding for a particular target molecule can be shown, for example, when a binding molecule has a KD for the target molecule of at least about 10-4 M, at least about 10-5 M, at least about 10-6 M, at least about 10-7 M, at least about 10-8 M, at least about 10-9 M, at least about 10 -10 M, at least about 10-11 M, at least about 10-12 M, or less, where KD refers to a dissociation rate of the binding. In some embodiments, a binding molecule that specifically binds a target molecule will have a KD that is 20, 50, 100, 500, 1000, 5,000, 10,000 or more times greater than the KD of a molecule that does not bind to the same target molecule. In some embodiments, the binding between a binding molecule and a target molecule can be shown by an EC50 value, determined using suitable methods known in the art.
[0096] The terms “polynucleotide” and “nucleic acids” are used interchangeably herein to refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term includes, but is not limited to, single, double, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. For example, a polynucleotide disclosed herein may be modified with a labeling group such as a fluorophore, a biotin, or a phosphorothioate.
[0097] The terms “peptide, ” “polypeptide, ” and “protein” are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones. The terms also include polypeptides that have co-translational (e.g., signal peptide cleavage) and post-translational modifications of the polypeptide, such as, for example, disulfide-bond formation, glycosylation, acetylation, phosphorylation, proteolytic cleavage, and the like. A peptide disclosed herein may be modified, e.g., with a labeling group such as a fluorophore, a biotin, a His tag, or phosphorothioate.
[0098] The terms “right end element RNA (reRNA) ” (also referred to as “guide polynucleotide, ” “guide sequence, ” or “guide RNA” ) refers to a polynucleotide that comprises a “scaffold” and a “targeting sequence, ” or a polynucleotide encoding thereof. In some embodiment, the TnpB system disclosed herein comprises a scaffold that contains several subterminal palindromes, stem-loop structures (see, e.g., Figs. 11A-11T) and is capable of binding with a TnpB protein, and a targeting sequence comprising a sequence substantially complementary to and thus is capable of specifically binding to a polynucleotide of interest. Without being bound by theory, sequences of the scaffolds may be predicted based on experimental / empirical evidence and / or analysis thereof, and such scaffold sequences are referred to as “inferred” scaffold sequences herein.
[0099] The terms “variant” refers to a polypeptide / protein or polynucleotide differing from another (i.e., parental) polypeptide / protein or polynucleotide or from one another due to changes in one or more nucleic acids or amino acid residues but retain at least a degree of one functional property of the parent molecule. For example, a variant may include one or more amino acid changes such as one or more amino acid deletions / truncations, insertions, or substitutions as compared to the parental protein from which it is derived. A variant may include one or more changes such as deletions / truncations, insertions, or substitutions as compared to the parental molecule from which it is derived. The parental molecule may be a wild-type polypeptide / protein or polynucleotide. A variant may have a specified degree (percentage) of sequence identity with a parental polypeptide / protein or nucleic acid using the BLAST percent identity algorithms. The degree (percentage) of sequence identity may be at least about 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or integer percentage therebetween.
[0100] The term “wild-type” refers to an amino acid sequence or nucleic acid sequence indicates that is a native or naturally-occurring sequence (e.g., SEQ ID NO: 1) . As used herein, the term “naturally-occurring” refers to anything (e.g., proteins or polynucleotides) that is found in nature. Conversely, the term “non-naturally occurring” refers to anything that is not found in nature (e.g., recombinant / engineered amino acid or nucleic acid sequences produced in the laboratory or modification of the wild-type sequences) .
[0101] The term “regulatory sequence” used herein refers to a polynucleotide sequence associated with a protein or RNA-encoding DNA sequence in such a way that the polynucleotide sequence can regulate the transcription of the DNA sequence. Such association between the regulatory sequence and the corresponding DNA sequence may be referred to as “operably linked, ” herein. The regulatory elements may be promoters, enhancers, internal ribosome entry sites, and other expression control elements. One regulatory sequence may be operably linked to one or more protein or RNA-encoding DNA sequences; multiple regulatory sequences may be operably linked to one protein or RNA-encoding DNA sequence. These can be selected depending on the cell type.
[0102] The term “donor polynucleotide” used herein refers to a polynucleotide comprising a specific sequence to be inserted into the polynucleotide of interest. In the donor polynucleotide, the specific sequence is usually flanked by two homology arms (i.e., two flanking sequences that are respectively homologous to the upstream and downstream sequences near the site of the insertion) , which facilitate an insertion.
[0103] The term “analyte” refers to a composition or substance that is being identified and analyzed.
[0104] The term “endonuclease” refers to an enzyme that cleaves a phosphodiester bond within a polynucleotide chain.
[0105] The term “polynucleotide manipulation” encompasses binding, nicking one strand, or cleaving (i.e., cutting) both strands of the polynucleotide of interest. It may further encompass altering the modification of the polynucleotide, editing the sequence of the polynucleotide, and changing the secondary and / or higher structure of the polynucleotide. Manipulation of a DNA molecule can silence, activate, or modulate (increase or decrease) the expression of an RNA or polypeptide encoded by the DNA. It can also prevent or enhance the direct or indirect binding of a molecule to or bring the molecule to the proximity of the polynucleotide of interest. Manipulation of a DNA molecule can also modulate replication, repair, or recombination of the molecule. The term “targeted polynucleotide manipulation” refers to when such polynucleotide manipulation takes place at a specific (sometimes predetermined) locus of the polynucleotide.
[0106] By “complementary” or “substantially complementary, " it is meant that a nucleic acid comprises a sequence of nucleotides that enables it to non-covalently bind, e.g., form Watson-Crick base pairs and / or G / U base pairs, "anneal" , or "hybridize, " to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions, such as appropriate temperature and solution ionic strength. As is known in the art, standard Watson-Crick base-pairing includes: adenine (A) pairing with thymidine (T) , adenine (A) pairing with uracil (U) , and guanine (G) pairing with cytosine (C) [DNA, RNA] . In addition, it is also known in the art that for hybridization between two RNA molecules (e.g., dsRNA) , guanine (G) base pairs with uracil (U) . For example, G / U base-pairing is partially responsible for the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anti-codon base-pairing with codons in mRNA. In the context of this disclosure, a guanine (G) of a protein-binding segment (dsRNA duplex) of a targeting RNA molecule is considered complementary to a uracil (U) , and vice versa. As such, when a G / U base-pair can be made at a given nucleotide position a protein-binding segment (dsRNA duplex) of a guide RNA molecule, the position is not considered to be non-complementary, but is instead considered to be complementary.
[0107] The term “support” used herein refers to a substance, to / in which the system disclosed herein is attached / embedded. Such attachment / embedment may be covalent or non-covalent and does not impair the function of the system in any significant manner. Such attachment / embedment may be fixed or reversible. Non-limiting examples of the support include chips, strips, beads, flow cells, microwells, films, gels, and foams. The support may be made of naturally-occurring and man-made polymers (e.g., agar, chitosan, nitrocellulose, polyacrylamide, etc. ) , glass, silicon, metal, paper, etc. System
[0108] The present disclosure provides systems for targeted manipulation of a polynucleotide of interest and / or a molecule associated therewith. In some embodiments, the system comprises (i) a polypeptide comprising a TnpB or variant thereof or a first polynucleotide encoding the polypeptide; and (ii) a guide polynucleotide comprising a targeting sequence and a scaffold or a second polynucleotide encoding the guide polynucleotide; wherein the polypeptide is capable of specifically binding to the guide polynucleotide via the scaffold to form a complex; and wherein the targeting sequence is capable of specifically binding to the polynucleotide of interest. A non-limiting list of TnpB and their corresponding reRNA scaffolds are shown in Table 1. In some embodiments, the system further comprises a donor polynucleotide. In some embodiments, the system further comprises a reporter polynucleotide.
[0109] In some embodiments, the system comprises a TnpB or variant thereof and a corresponding reRNA scaffold, e.g., a TnpB selected from TnpB 1-93 and its corresponding reRNA scaffold listed in Table 1 or a fragment of the listed reRNA scaffold. In some embodiments, the system comprises one of TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90 and its corresponding reRNA listed in Table 1 or a fragment of the listed reRNA scaffold.
[0110] In some embodiments, the polypeptide is a fusion protein comprising the TnpB or variant thereof and one or more effectors. Non-limiting examples of such effector include a nuclear localization sequence (NLS) , a cell penetrating peptide, a nuclease that is not the TnpB or variant thereof, a polymerase, a base editor, a prime editor, an epigenetic modifier, a polymerase, a transposase, a recombinase, a reverse transcriptase, a label, a transcription modulator, a transcription factor, a topoisomerases, a helicase, a kinase, a phosphatase, a ligase, an antibody, an antigen, or a combination thereof. The nuclease may be an endonuclease, an exonuclease, or a nickase. The base editor may be a deaminase. The polymerase may be a DNA or RNA polymerase. A transcription modulator may be a transcription factor or co-factor, a methyltransferases, or a histone acetyltransferase. The label may be a fluorescent, luminescent, and / or a chromogenic protein.
[0111] In some embodiments, the system comprises (i) the first polynucleotide encoding the polypeptide comprising the TnpB or variant thereof; and (ii) the second polynucleotide encoding the guide polynucleotide. In some embodiments, the first polynucleotide and the second polynucleotide are in the same polynucleotide. In some embodiments, the first polynucleotide and the second polynucleotide are in separate polynucleotides. In some embodiments, the system further comprises a donor polynucleotide, wherein the donor polynucleotide may be in a separate polynucleotide or in the same polynucleotide as the first and / or second polynucleotide.
[0112] In some embodiments, the polynucleotide (s) harboring the first polynucleotide, the second polynucleotide, the donor polynucleotide, or a combination thereof are vectors, such as viral and / or non-viral vectors. In some embodiments, the vector is a viral vector, such as a lentiviral vector, a retroviral vector, an adenoviral vector, an adeno-associated viral vector, a herpes simplex viral vector, or a combination thereof. In some embodiments, the vector is a non-viral vector such as a plasmid. In some embodiments, the vector is attached to or embedded in an auxiliary substance, such as a viral capsid, a lipid layer, a polymer, an inorganic molecule, a nanoparticle, or a combination thereof.
[0113] In some embodiments, a cell and / or progeny thereof comprising the system disclosed herein is provided. The system may be transiently or constitutively present in the cell and / or progeny thereof. The cell may be a prokaryotic cell or a eukaryotic cell. The cell may be a plant cell, a non-human animal cell, or a human cell. In some embodiments, an organoid comprising cell and / or progeny thereof is involved. In some embodiments, an organism comprising cell and / or progeny thereof is provided. A. TnpBs and variants thereof
[0114] In some embodiments, a polypeptide comprising a TnpB protein or variant thereof or a polynucleotide encoding the TnpB protein or variant thereof is involved. Non-limiting examples of wild-type TnpBs are shown in Table 1. In some embodiments, the TnpB may be one or more of TnpB 1-93 (Table 1) . In some embodiments, the TnpB may be one or more of TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90 (Table 1) .
[0115] In some embodiments, the TnpB variant may have an amino acid sequence that is at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%identical to the wild-type TnpB. The TnpB variant may have substitution, deletion, and / or insertion of one or more amino acids. In some embodiments, the TnpB variant may comprise one or more of the mutations listed in Table 2. Table 2. TnpB variants
[0116] In some embodiments, the RNA-guided endonuclease activity is enhanced in the TnpB variant compared to its wild-type counterpart. In some embodiments, the activity is increased by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 120%, 140%, 160%, 180%, or 200%. In some embodiments, the TnpB variant lacks the RNA-guided endonuclease activity. In some embodiments, the RNA-guided endonuclease activity is weakened in the TnpB variant compared to its wild-type counterpart. In some embodiments, the TnpB variant has an RNA-guided nickase activity. B. Right end element (reRNA)
[0117] The reRNA ( “guide polynucleotide, ” “guide sequence, ” or “guide RNA” ) disclosed herein refers to a polynucleotide comprising a targeting sequence for specifically binding to the polynucleotide of interest and a scaffold for specifically binding to a TnpB or variant thereof, or a polynucleotide encoding the TnpB or variant thereof. The reRNA comprises one or more stem-loop structures. Examples of predicted stem-loop structures of some reRNAs are shown in Figs. 11A-11T. reRNAs comprising combinations of scaffold and targeting sequences disclosed herein are engineered and cannot be found in nature.
[0118] The targeting sequence comprises a sequence substantially complementary to and thus is capable of specifically binding to the polynucleotide of interest. In some embodiments, the targeting sequence is at least about 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50 nt in length. In some embodiments, the targeting sequence is about 10 to about 50 nt, about 20 to about 40 nt, or about 30 nt in length.
[0119] A non-limiting list of inferred wild-type scaffolds is shown in Table 1. In some embodiments, the scaffold may be a wild-type scaffold (including an inferred wild-type scaffold) or a variant thereof. The scaffold variant may contain deletion, substitution, and / or insertion of one or more nucleic acids of the wild-type scaffold. In some embodiments, the deletion is a 5’ -truncation. In some embodiments, the deletion is within the wild-type scaffold. In some embodiments, the deletion removes one or more stem-loop regions from the wild-type scaffold. In some embodiments, the deletion is a 5’ -truncation of about 20, 40, 60, 80, or 100 nt. In some embodiments, the deletion is inside the wild-type scaffold such as those exemplified in Table 4. The size of the deletion inside the wild-type scaffold may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, or 70 nt.
[0120] In some embodiments, the scaffold is at least about 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, or 250 nt in length, and particularly at least about 100, 120, 125, 130, 135, 140, 160, 180, or 200 nt in length. In some embodiments, the scaffold is about 80 to about 250 nt, about 100 to about 220 nt, or about 120 to 200 nt.
[0121] In some embodiments, using a scaffold variant can increase or decrease the RNA-guided endonuclease activity of a specific TnpB compared to using the corresponding wild-type scaffold. In some embodiments, the activity is increased by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 120%, 140%, 160%, 180%, or 200%. Methods
[0122] The present disclosure provides methods of using the disclosed systems for targeted manipulation of a polynucleotide of interest and / or a molecule associated therewith.
[0123] In some embodiments, the method involves editing of the nucleic acid sequence of the polynucleotide of interest by contacting the system disclosed herein with the polynucleotide of interest. The editing may result in insertion, deletion, and / or substitution of at least one nucleotide in the polynucleotide of interest. In some embodiments, the method involves using the system comprising the donor polynucleotide. In some embodiments, a single-stranded nicking or a double-stranded cleaving of the polynucleotide of interest is involved. In some embodiments, a non-sense mutation is introduced in a gene of interest whereby abolishes the expression of functional protein from the gene, which may be referred to as a “knock out. ” In some embodiments, the polynucleotide of interest, which contains one or more mutations and / or one or more undesirable nucleotides, is corrected to its wild-type and / or desirable version. In some embodiments, a non-homologous end joining (NHEJ) of the cleaved polynucleotide of interest is involved. In some embodiments, a homology-directed repair (HDR) of the cleaved polynucleotide of interest is involved, wherein the repair template is provided by the donor polynucleotide and whereby the donor polynucleotide or a fragment thereof is inserted into the polynucleotide of interest. In some embodiments, the donor polynucleotide comprises one or more genes listed in Tables 3A and 3B or a fragment thereof.
[0124] In some embodiments, the method involves the use of the system comprising a fusion protein comprising a TnpB or variant thereof and an effector or a first polynucleotide encoding the fusion protein. The TnpB variant may have weakened or no endonuclease activity or may have nickase activity. In some embodiments, the method involves promoting or preventing direct or indirect binding of a molecule to the polynucleotide of interest by contacting the system disclosed herein with the polynucleotide of interest. The molecule may already be associated with the polynucleotide of interest. The molecule may be an inorganic molecule or an organic molecule. The molecule may be a nuclease that is not the TnpB or variant thereof, a polymerase, a base editor, a prime editor, an epigenetic modifier, a polymerase, a transposase, a recombinase, a reverse transcriptase, a label, a transcription modulator, a transcription factor, a topoisomerases, a topoisomerases, a helicase, a kinase, a phosphatase, an integrase, a ligase, an antibody, an antigen, and a combination thereof. Such promotion or prevention of direct or indirect binding of the molecule to the polynucleotide of interest may result in changes in transcription, replication, repair, and / or recombination of a polynucleotide of interest. It may also result in modification of the nucleotides of the polynucleotide of interest and / or changes in the secondary or higher structure of the polynucleotide of interest by contacting the system disclosed herein with the polynucleotide of interest. It may further result in modification of a molecule (e.g., a polymerase, a transcription factor, a label, etc. ) already associated with the polynucleotide of interest. In some embodiments, such molecules may be dissociated or modified. For example, an inactive transcription factor bound to the polynucleotide of interest may be specifically phosphorylated and activated by the system comprising a kinase as an effector. Applications
[0125] In some embodiments, cells, tissues, organoids, organs, or organisms comprising the system disclosed herein are generated. In some embodiments, cells, tissues, organoids, organs, or organisms comprising a specific exogenous polynucleotide sequence are generated by using the system disclosed herein.
[0126] In some embodiments, the system disclosed herein is used to treat or prevent a condition or a disease in a subject in need thereof. In some embodiments, the treatment involves inserting or knocking out a gene of interest or a fragment thereof. In some embodiments, the treatment involves correcting one or more undesirable mutations. In some embodiments, the treatment involves increasing or decreasing the expression of a gene. Non-limiting examples of the diseases being treated and the corresponding genes of interest being manipulated are listed in Table 3A. In some embodiments, the system can be used to manipulate one or more genes associated with one or more cellular functions, such as those in Table 3B, whereby treats or prevents a disease. Table 3A. Examples of diseases / disorders and associated genes Table 3B: Examples of cellular functions and associated genes
[0127] In some embodiments, the prevention and / or treatment involves introducing the system to the subject. In some embodiments, the treatment involves introducing one or more cells comprising the system or progeny thereof to the subject. In some embodiments, the treatment involves introducing an organoid comprising such cells or progeny thereof to the subject. In some embodiments, the subject is a mammal and in particular a human.
[0128] In some embodiments, the system disclosed herein is used to detect a polynucleotide of interest in an analyte. In some embodiments, a label (such as a fluorescent, luminescent, and / or a chromogenic protein) is introduced to the specific polynucleotide of interest where the label binds directly or indirectly to the polynucleotide. In some embodiments, detection of the complex comprising the guide RNA and the TnpB or variant thereof indicates the presence of the corresponding polynucleotide of interest. In some embodiments, the system comprising a reporter polynucleotide is used for the detection, wherein the TnpB-guide RNA complex, after specifically binding to the polynucleotide of interest, cleaves the reporter polynucleotide. In some embodiments, cleavage of the reporter polynucleotide relies on the collateral activity of the TnpB-guide RNA complex, wherein the TnpB-guide RNA complex, upon specifically binding to the polynucleotide of interest, exhibits non-specific endonuclease activity (i.e., the collateral activity) and cleaves the reporter polynucleotide. In some embodiments, one or both ends of the reporter polynucleotide is labeled with one or more reporter molecules such as fluorophores, antigens, biotin, etc. In some embodiments, cleavage of the reporter polynucleotide, which indicates the presence of the polynucleotide of interest, can be detected by an immunochemical or immunofluorescent procedure.
[0129] In some embodiments, the detection is for diagnostic purposes wherein the analyte is a sample from a biopsy, a body fluid, feces, etc. of a subject. In some embodiments, the polynucleotide of interest is a genome or fragment thereof or a microbe such as a pathogenic microbe. In some embodiments, the microbe is a virus such as an influenza virus or a SARS-CoV-2 virus. In some embodiments, the polynucleotide of interest is polynucleotide encoding a neoantigen or fragment thereof. EXAMPLES
[0130] The following examples are intended to illustrate, but not limit, the invention. Accordingly, from the above discussion and the Examples, one skilled in the art can ascertain essential characteristics of this disclosure, and without departing from the spirit and scope thereof, can make various changes and modifications to adapt to various uses and conditions.
[0131] Unless defined otherwise herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Example 1. Determination of RNA-guided endonuclease activity in TnpBs A. Plasmid construction
[0132] The RNA-guided endonuclease activity of TnpBs was determined by using one of two different sets of plasmids. One of the plasmid sets was constructed according to Yang et al. (Appl Biochem Biotechnol, 2016, 180: 655-667) and Kong et al. (Nat Commun, 2023, 14: 2046) , which involves a TnpB-reRNA plasmid (Fig. 2A) and an EGFxxFP reporter plasmid (Fig. 2B) . To construct the TnpB-reRNA plasmid, the coding sequence of a TnpB was human codon-optimized and synthesized by General Biol Company and inserted downstream of an EF-1α promoter and a NLS in pUC19. Since reRNA is encoded by the 3’ -sequence of the IS605 (Karvelis et al., 2021) and previous transcriptome profiling data revealed that non-coding RNAs (ncRNAs) expressed in bacteria and archaea are usually about 200 nt in length, the sequence of approximately 200 nt from the 3’ end of IS605 is referred to herein as the “inferred wild-type” scaffold (Table 1) . An reRNA-encoding sequence, including one targeting sequence selected from Target 1 (CCATTACAGTAGGAGCATAC; SEQ ID NO: 334) , Target 2 (GCTCGGAGATCATCATTGCG; SEQ ID NO: 335) ; and Target 3 (GAGTCCCTTGGCGCCC; SEQ ID NO: 336) followed by an reRNA scaffold sequence, was synthesized by General Biol and inserted between the two BsmBI restriction sites downstream of a U6 promoter in pUC19. Unless otherwise specified, the inferred wild-type reRNA scaffold sequences in Table 1 are paired with the corresponding TnpB and variants thereof in the TnpB-reRNA plasmids. The EGFxxFP reporter plasmid contained an mCherry-T2A-EGFxxFP cassette, with a deactivated EGFP (EGFxxFP) harboring a TAM sequence followed by one target sequences (Target 1, 2, or 3) inserted between EGFx CDS and xFP CDS. As the CDSs of EGFx and xFP partially overlapped with each other, TnpB-mediated double-strand cleavage at the target sequence between the two coding sequences would trigger single-strand annealing (SSA) and reinstate EGFP expression.
[0133] The other plasmid set included a TnpB-reRNA plasmid (Fig. 2C) and a RFP-EGFP reporter plasmid (Fig. 2D) . In the TnpB-reRNA plasmid, the transcription of TnpB and reRNA was driven by a CMV promoter and a U6 promoter respectively. In the RFP-EGFP reporter plasmid, a TAM and Target 2 were inserted between the CDSs of RFP and EGFP such that the EGFP CDS was out of frame relative to the start codon of RFP CDS. TnpB-mediated double-strand cleavage would create indels between the two CDSs and thereby reinstate EGFP expression. B. Quantification of RNA-guided endonuclease activity in TnpBs using transfected cells
[0134] HEK 293T cells (American Type Culture Collection) were cultured in complete growth medium including high-glucose Dulbecco’s Modified Eagle Medium (Invitrogen) supplemented with 10%Fetal Bovine Serum (Gemini Bio) and 1%GlutMAXTM (Invitrogen) in an incubator at 37 ℃ with 5%CO2. All experiments were performed in 24-well plates. In each well, 120,000 HEK 293T cells were plated in 0.5 mL complete growth medium, co-transfected with 700 ng of the TnpB-reRNA plasmid and 700 ng of the reporter plasmid using LipofectamineTM 3000 (ThermoFisher Scientific) at 24 hours after plating. GFP and mCherry (or RFP) expression in the transfected cells was determined with a CytoFLEXTM LX Flow Cytometer (Beckman Coulter) at 48 hours after transfection. The co-transfected plasmids shared the same target sequence. Negative controls were prepared for each TnpB by transfecting HEK 293T cultures with only 700 ng of the TnpB-reRNA plasmid. GFP-expressing ( “GFP+” ) cells and the mCherry-expressing ( “mCherry+” ) or RFP-expressing ( “RFP+” ) cells were calculated after removing the background signals in the negative controls. The RNA-guided endonuclease activity in TnpB was quantified as the percentage of GFP+ cells in the total co-transfected cells, i.e., the mCherry+ cells or RFP+ cells. Example 2. Screening for new TnpBs A. Identification of candidate TnpBs by bioinformatic analysis
[0135] Genome sequences of approximately 1,500,000 prokaryotic organisms in GenBank were analyzed for sequences potentially encoding TnpB of IS605 with previously described putative TAM sequences (Xiang et al., Nat Biotechnol, 2023) . B. Quantification of RNA-guided endonuclease activity in candidate TnpBs
[0136] RNA-guided endonuclease activity of candidate TnpB 1-93 was examined as mentioned in Example 1. The percentages of GFP+ cells in the total co-transfected cells in different cultures are reported in Figs. 3A-3H. Representative microscopic images of the cells co-transfected with the reporter plasmid and TnpB 8, 18, 19, and 22, respectively, and their corresponding negative controls are shown in Figs. 4A-4D.Example 3. Impact of reRNA modification on RNA-guided endonuclease activity in TnpBs A. 5’-Truncation of reRNA scaffold
[0137] TnpB-reRNA plasmids encoding an reRNA with a 5’ -truncated scaffold were constructed. The truncated reRNA scaffolds were 100, 120, 125, 130, 135, 140, 160, 180, or 200 nt in length. RNA-guided endonuclease activity in TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90 was tested at the presence of the modified reRNA scaffolds and the inferred wild-type reRNA scaffolds according to Example 1. Target 1, 2, or 3 was used as the targeting sequence.
[0138] The percentages of GFP+ cells in the total co-transfected cells (the mCherry+ or RFP+ cells) in cultures transfected with different combination of TnpB and reRNA scaffold are reported in Figs. 5A-5X. The results indicate that the reRNA scaffold length for achieving maximal RNA-guided endonuclease activity varies in different TnpBs. B. Internal deletion in reRNA scaffold
[0139] TnpB-reRNA plasmids encoding an reRNA scaffold with a deletion of one or more predicted stem-loop structures in the reRNA scaffold were constructed. The sequences of the modified reRNA scaffold are summarized in Table 4. RNA-guided endonuclease activity in TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90 was tested at the presence of the modified reRNA scaffolds and the inferred wild-type reRNA scaffolds according to Example 1. Target 2 was used as the targeting sequence. Table 4. reRNA scaffold with internal deletions in stem region
[0140] The percentages of GFP+ cells in the total co-transfected cells (the mCherry+ or RFP+ cells) in cultures transfected with different combination of TnpB and reRNA scaffold are reported in Figs. 6A-6S.Example 4. RNA-guided endonuclease activity in TnpB variants A. Tnp B variants with a single mutation
[0141] RNA-guided endonuclease activity in variants of TnpB 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90, each of which contained a single mutation, was tested according to Example 1. Target 1, 2, or 3 was used as the targeting sequence.
[0142] The percentages of GFP+ cells in the total transfected cells (the mCherry+ or RFP+ cells) in cultures transfected with different TnpB variants are reported in Figs. 7A-7T. B. TnpB variants with multiple mutations
[0143] RNA-guided endonuclease activity in variants of TnpB 18, 76, and 90, each of which contained multiple mutations, was tested according to Example 1. Target 1 or 2 was used as the targeting sequence.
[0144] The percentages of GFP+ cells in the total transfected cells (the mCherry+ or RFP+ cells) in cultures transfected with different TnpB variants are reported in Figs. 8A-8D.Example 5. TnpB-mediated endogenous gene editing
[0145] TnpB 8, TnpB 18, and TnpB 78 were used to edit endogenous genes of interest. The genes and the corresponding targeting sequences used are listed in Table 5. Table 5. Targeting sequences used in endogenous gene editing
[0146] To test TnpB 8 and 18-mediated endogenous gene editing, the corresponding TnpB-reRNA plasmids were transfected in HEK293T cells. Genomic DNA of the transfected cells was extracted at 48 h after transfection and sequenced at the pre-determined gene editing sites with next-generation sequencing. To test TnpB 78-mediated endogenous gene editing, the TnpB-reRNA plasmids and EGFxxFP reporter plasmids were co-transfected into HEK293T cells. Genomic DNA was extracted from GFP+ co-transfected cells isolated by fluorescence-activated cell sorting (FACS) at 48 h after transfection. The genomic DNA was sequenced at the pre-determined gene editing sites with next-generation sequencing.
[0147] Indels were successfully introduced to most of the target genes. The gene editing efficiencies, indicated by the percentage of indels detected in each target gene, are reported in Figs. 9A-9C.Example 6. TnpB variants (dTnpBs) lacking RNA-guided endonuclease activity
[0148] Variants of TnpB 8 were prepared by introducing a single amino acid mutation of D185A, R258A, or E279A, and their RNA-guided endonuclease activity was tested according to Example 1. Without wishing to be bound by theory, data in Fig. 10 demonstrate that these variants have weakened or almost no RNA-guided endonuclease activity in this experiment.Example 7. Transcription regulation with dTnpB fusion proteins
[0149] A dTnpB-reRNA plasmid is constructed to express a fusion protein comprising a dTnpB and a transcription modulator, e.g., a transcription factor (TF) or a DNA methyltransferase (DNMT) . The plasmid is constructed in a similar way as the TnpB-reRNA plasmid in Example 1 with an addition of the transcription modulator CDS upstream or downstream of the dTnpB CDS. Sequences encoding a compatible reRNA scaffold and a targeting sequence capable of annealing to a specific transcription regulatory sequence, e.g., a promoter of a gene of interest, are also included in the dTnpB-reRNA plasmid.
[0150] The dTnpB-reRNA plasmids are introduced into the host cells, which leads to an increase (e.g., at the presence of dTnpB-TF fusion protein) or a decrease (e.g., at the presence of dTnpB-DNMT fusion protein) in the expression level of the gene of interest as confirmed by qRT-PCR-or sequencing-based methods.Example 8. Base editing with dTnpB fusion proteins
[0151] A dTnpB-reRNA plasmid is constructed to express a fusion protein comprising a dTnpB and a base editor, e.g., a deaminase. The plasmid is constructed in a similar way as the TnpB-reRNA plasmid in Example 1 with an addition of the base editor CDS upstream or downstream of the dTnpB CDS. Sequences encoding a compatible reRNA scaffold and a targeting sequence capable of annealing to a pre-determined base editing site are also included in the dTnpB-reRNA plasmid.
[0152] The dTnpB-reRNA plasmids are introduced into the host cells. Successful editing of the pre-determined site is confirmed by sequencing of the genomic DNA extracted from those cells.
Claims
1.A system for manipulating a polynucleotide of interest, comprising(i) a polypeptide comprising a TnpB or variant thereof or a first polynucleotide encoding the polypeptide; and(ii) a guide polynucleotide comprising a targeting sequence and a scaffold or a second polynucleotide encoding the guide polynucleotide;wherein the polypeptide is capable of specifically binding to the guide polynucleotide via the scaffold to form a complex; andwherein the targeting sequence is capable of specifically binding to the polynucleotide of interest.2.The system of claim 1, wherein(i) the first polypeptide is provided by the first polynucleotide and(ii) the guide polynucleotide is provided by second polynucleotide.3.The system of claim 1 or 2, wherein the TnpB or variant thereof has nickase activity.4.The system of any of claims 1-3, wherein the targeting sequence is about 10 to about 50 nt in length.5.The system of any of claims 1-4, wherein the scaffold is about 50 nt to about 250 nt in length.6.The system of any of claims 1-5, wherein the polypeptide comprises an amino acid sequence sharing at least 80%sequence identity to one or more of SEQ ID NOs: 3, 8, 15, 18, 19, 22, 28, 39, 48, 55, 56, 57, 59, 63, 65, 72, 75, 76, 78, and 90.7.The system of any of claims 1-6, wherein the polypeptide comprises:(i) SEQ ID NO: 3 or a variant of thereof, wherein the variant comprises one or more mutations selected from Q72K, T88R, T112R, L160K, I172R, Q210K, and L351R;(ii) SEQ ID NO: 8 or a variant thereof, wherein the variant comprises one or more mutations selected from S54K and A66K;(iii) SEQ ID NO: 15 or a variant thereof, wherein the variant comprises one or more mutations selected from G56K, E64K, T87K, Q90K, Q97R, N111K, E202R, Q258K, and Q261K;(iv) SEQ ID NO: 18 or a variant thereof, wherein the variant comprises one or more mutations selected from S88K, E97K, S105R, D108K, I110R, I232K, D250R, D316R, and E323R;(v) SEQ ID NO: 19 or a variant thereof, wherein the variant comprises one or more mutations selected from D44H, S88K, S105K, E217K, and I237K;(vi) SEQ ID NO: 22 or a variant thereof, wherein the variant comprises one or more mutations selected from L224R and A326K;(vii) SEQ ID NO: 28 or a variant thereof, wherein the variant comprises one or more mutations selected from D173K and E313K;(viii) SEQ ID NO: 39 or a variant thereof, wherein the variant comprises one or more mutations selected from S83K, S88H, S92K, D100K, G109K, S111R, F120K, Y144R, Q234R, Q248R, A277K, and Q310R;(ix) SEQ ID NO: 48 or a variant thereof, wherein the variant comprises one or more mutations selected from S83K, S88H, S92K, D100K, G109K, S111R, F120K, Y144R, Q234R, Q248R, A277K, and Q310R;(x) SEQ ID NO: 55 or a variant thereof, wherein the variant comprises one or more mutations selected from E40K, G90R, G95K, T121H, S251K, S281K, and I301K;(xi) SEQ ID NO: 56 or a variant thereof, wherein the variant comprises one or more mutations selected from Q90K, S108H, N111K, E202K, L225R, T282K, and E333K;(xii) SEQ ID NO: 57 or a variant thereof, wherein the variant comprises one or more mutations selected from N44K, A59K, Q80K, Q87K, E97K, Q131K, E201H, A209R, A216R, Q221K, Q238K, A274K, E295K, T313K, and V350H;(xiii) SEQ ID NO: 59 or a variant thereof, wherein the variant comprises one or more mutations selected from A135K, D184K, V240R, T254K, and A298K;(xiv) SEQ ID NO: 63 or a variant thereof, wherein the variant comprises a mutation of E291K;(xv) SEQ ID NO: 65 or a variant thereof, wherein the variant comprises one or more mutations selected from E42K, N45K, C62K, A96R, I99R, C214R, N224K, L235R, N245R, I254R, D292R, and T310R;(xvi) SEQ ID NO: 72 or a variant thereof, wherein the variant comprises one or more mutations selected from N93K, N223K, and E306K;(xvii) SEQ ID NO: 76 or a variant thereof, wherein the variant comprises one or more mutations selected from D124G, E144K, L148K, N155K, T197K, A212R, I217K, S272K, N279K, T309K, and Q326K; and / or(xviii) SEQ ID NO: 90 or a variant thereof, wherein the variant comprises one or more mutations selected from S188K, V232K, D233R, N243R, M272K, and T360R.8.The system of any of claims 1-6, wherein the polypeptide comprises:(i) SEQ ID NO: 18 or a variant thereof, wherein the variant comprises a mutation combination of E97K and S88K; E97K and S105R; E97K and D108K; E97K and I110R; E97K and I232K; E97K and D250R; E97K and D316R; E97K and E323R; D108K and S88K; D108K and S105R; D108K and I110R; D108K and I232K; D108K and D250R; D108K and D316R; D108K and E323R; D108K, I232K, and E97K; E97K, I110R, and D108K; E97K, I110R, and D250R; or E97K, I110R, and D316R;(ii) SEQ ID NO: 76 or a variant thereof, wherein the variant comprises a mutation combination of D40K and D97K; D40K and N155K; D40K and T309K; or N155K and T197K; and / or(iii) SEQ ID NO: 90 or a variant thereof, wherein the variant comprises a mutation combination of V232K and N243R; D233R and N243R; V232K, D233R, and N243R, M272K and N243R; or T360R and N243R.9.The system of any of claims 1-8, wherein the scaffold has at least 80%sequence identity to a polynucleotide sequence selected from SEQ ID NOs: 189, 194, 201, 204, 205, 208, 214, 225, 234, 241, 242, 243, 245, 249, 251, 258, 261.262, 264, and 276 or a fragment thereof.10.The system of any of claims 1-9, wherein the scaffold comprises:(i) about 160 nt from 3’-end of SEQ ID NO: 189;(ii) about 120 nt to about 160 nt from 3’-end of SEQ ID NO: 194;(iii) about 125 nt to about 200 nt from 3’-end of SEQ ID NO: 201;(iv) about 125 nt to about 200 nt from 3’-end of SEQ ID NO: 204;(v) about 125 nt to about 200 nt from 3’-end of SEQ ID NO: 205;(vi) about 130 nt to about 200 nt from 3’-end of SEQ ID NO: 208;(vii) about 120 nt to about 140 nt from 3’-end of SEQ ID NO: 214;(viii) about 180 nt from 3’-end of SEQ ID NO: 225;(ix) about 160 nt to about 180 nt from 3’-end of SEQ ID NO: 234;(x) about 120 nt to about 180 nt from 3’-end of SEQ ID NO: 241;(xi) about 140 nt to about 180 nt from 3’-end of SEQ ID NO: 242;(xii) about 120 nt to about 180 nt from 3’-end of SEQ ID NO: 243;(xiii) about 140 nt to about 180 nt from 3’-end of SEQ ID NO: 245;(xiv) about 140 nt to about 180 nt from 3’-end of SEQ ID NO: 249;(xv) about 140 nt to about 180 nt from 3’-end of SEQ ID NO: 251;(xvi) about 140 nt to about 180 nt from 3’-end of SEQ ID NO: 258;(xvii) about 140 nt to about 200 nt from 3’-end of SEQ ID NO: 261;(xviii) about 100 nt to about 200 nt from 3’-end of SEQ ID NO: 262;(xix) about 140 nt to about 200 nt from 3’-end of SEQ ID NO: 264;(xx) about 140 nt to about 200 nt from 3’-end of SEQ ID NO: 276;(xxi) SEQ ID NO: 285;(xxii) SEQ ID NO: 286;(xxiii) SEQ ID NO: 292;(xxiv) SEQ ID NO: 305;(xxv) SEQ ID NO: 306;(xxvi) SEQ ID NO: 310;(xxvii) SEQ ID NO: 311;(xxviii) SEQ ID NO: 312;(xxix) SEQ ID NO: 314;(xxx) SEQ ID NO: 326; and / or(xxxi) SEQ ID NO: 333.11.The system of any of claims 1-10, wherein(i) the polypeptide comprises SEQ ID NO: 3 or a variant thereof and the scaffold comprises SEQ ID NO: 189 or a fragment thereof;(ii) the polypeptide comprises SEQ ID NO: 8 or a variant thereof and the scaffold comprises SEQ ID NO: 194 or a fragment thereof;(iii) the polypeptide comprises SEQ ID NO: 15 or a variant thereof and the scaffold comprises SEQ ID NO: 201 or a fragment thereof;(iv) the polypeptide comprises SEQ ID NO: 18 or a variant thereof and the scaffold comprises SEQ ID NO: 204 or a fragment thereof;(v) the polypeptide comprises SEQ ID NO: 19 or a variant thereof and the scaffold comprises SEQ ID NO: 205 or a fragment thereof;(vi) the polypeptide comprises SEQ ID NO: 22 or a variant thereof and the scaffold comprises SEQ ID NO: 208 or a fragment thereof;(vii) the polypeptide comprises SEQ ID NO: 28 or a variant thereof and the scaffold comprises SEQ ID NO: 214 or a fragment thereof;(viii) the polypeptide comprises SEQ ID NO: 39 or a variant thereof and the scaffold comprises SEQ ID NO: 225 or a fragment thereof;(ix) the polypeptide comprises SEQ ID NO: 48 or a variant thereof and the scaffold comprises SEQ ID NO: 234 or a fragment thereof;(x) the polypeptide comprises SEQ ID NO: 55 or a variant thereof and the scaffold comprises SEQ ID NO: 241 or a fragment thereof;(xi) the polypeptide comprises SEQ ID NO: 56 or a variant thereof and the scaffold comprises SEQ ID NO: 242 or a fragment thereof;(xii) the polypeptide comprises SEQ ID NO: 57 or a variant thereof and the scaffold comprises SEQ ID NO: 243 or a fragment thereof;(xiii) the polypeptide comprises SEQ ID NO: 59 or a variant thereof and the scaffold comprises SEQ ID NO: 245 or a fragment thereof;(xiv) the polypeptide comprises SEQ ID NO: 63 or a variant thereof and the scaffold comprises SEQ ID NO: 249 or a fragment thereof;(xv) the polypeptide comprises SEQ ID NO: 65 or a variant thereof and the scaffold comprises SEQ ID NO: 251 or a fragment thereof;(xvi) the polypeptide comprises SEQ ID NO: 72 or a variant thereof and the scaffold comprises SEQ ID NO: 258 or a fragment thereof;(xvii) the polypeptide comprises SEQ ID NO: 75 or a variant thereof and the scaffold comprises SEQ ID NO: 261 or a fragment thereof;(xviii) the polypeptide comprises SEQ ID NO: 76 or a variant thereof and the scaffold comprises SEQ ID NO: 262 or a fragment thereof;(xix) the polypeptide comprises SEQ ID NO: 78 or a variant thereof and the scaffold comprises SEQ ID NO: 264 or a fragment thereof; and / or(xx) the polypeptide comprises SEQ ID NO: 90 or a variant thereof and the scaffold comprises SEQ ID NO: 276 or a fragment thereof.12.The system of any of claims 1-11, wherein the first polynucleotide and / or the second polynucleotide is operably linked to a regulatory sequence.13.The system of any of claims 1-12, further comprising a donor polynucleotide.14.The system of any of claims 1-13, wherein the system comprises a third polynucleotide comprising the first polynucleotide and the second polynucleotide.15.The system of claim 14, wherein the third polynucleotide further comprises the donor polynucleotide.16.The system of any of claims 1-13, wherein the system comprises a third polynucleotide comprising the first polynucleotide and a fourth polynucleotide comprising the second polynucleotide.17.The system of claim 16, wherein at least one of the third polynucleotide and the fourth polynucleotide further comprise (s) the donor polynucleotide.18.The system of claim 14 or 16, wherein the system comprises a fifth polynucleotide comprising the donor polynucleotide.19.The system of any of claims 14-18, wherein the first polynucleotide, the second polynucleotide, the donor polynucleotide, or a combination thereof is on a vector.20.The system of claim 19, wherein the vector is a viral vector.21.The system of claim 20, wherein the viral vector is a lentiviral vector, a retroviral vector, an adenoviral vector, an adeno-associated viral vector, a herpes simplex viral vector, or a combination thereof.22.The system of claim 20, wherein the vector is a non-viral vector.23.The system of claim 22, wherein the non-viral vector is a plasmid.24.The system of any of claims 1-23, further comprising a reporter polynucleotide capable of being cleaved by the complex that specifically binds to the polynucleotide of interest.25.The system of any of claims 19-24, wherein the vector is attached to or embedded in an auxiliary substance.26.The system of claim 25, wherein the auxiliary substance comprises a viral capsid.27.The system of claim 26, wherein the auxiliary substance comprises a lipid layer, a polymer, an inorganic molecule, a nanoparticle, or a combination thereof.28.The system of any of claim 1-27, wherein the system is used for introducing an insertion, a deletion, and / or a substitution of at least one nucleotide to the polynucleotide of interest.29.The system of any of claims 13-28, wherein the system is used for inserting the donor polynucleotide or fragment thereof to the polynucleotide of interest.30.The system of any of claims 1-29, wherein the system is attached to or embedded in a support.31.A cell and / or progeny thereof, comprising the system of any of claims 1-29.32.The cell and / or progeny thereof of claim 31, wherein the system of any of claims 1-29 is transiently present in the cell and / or progeny thereof.33.The cell and / or progeny thereof of claim 31, wherein the system of any of claims 1-29 is constitutively present in the cell and / or progeny thereof.34.The cell and / or progeny thereof of any of claims 31-33, wherein the cell is a prokaryotic cell.35.The cell and / or progeny thereof of any of claims 31-33, wherein the cell is a eukaryotic cell.36.The cell and / or progeny thereof of any of claims 31-33, wherein the cell is a plant cell.37.The cell and / or progeny thereof of any of claims 31-33, wherein the cell is a non-human animal cell.38.The cell and / or progeny thereof of any of claims 31-33, wherein the cell is a human cell.39.An organoid comprising the cell or progeny thereof of any of claims 31-33.40.An organism comprising the cell or progeny thereof of any of claims 31-33.41.A method for editing a polynucleotide of interest, comprising(i) providing the polynucleotide of interest; and(ii) contacting the system of any of claims 1-30 with the polynucleotide of interest.42.A method for editing a polynucleotide of interest, comprising introducing the system of any of claim 1-29 to a cell comprising the polynucleotide of interest.43.A method for preventing or treating a disease or condition in a subject in need thereof, comprising introducing the system of any of claims 1-29 to at least one cell in the subject.44.A method for preventing or treating a disease or condition in a subject in need thereof, comprising introducing one or more of the cells or progenies thereof of any of claims 31-38 to the subject.45.A method for preventing or treating a disease or condition in a subject in need thereof, comprising introducing the organoid of claim 39 or one or more cells derived therefrom to the subject.46.The method of any of claims 43-45, wherein the subject is a human.47.The method of any of claims 43-46, wherein the disease or condition is one or more of Alzheimer’s disease, myotonic dystrophy type 1, spinal muscular atrophy, Huntington’s disease, sickle cell disease, β-thalassemia, Duchenne muscular dystrophy, and hereditary tyrosinemia.48.The method of any of claims 43-47, wherein the polynucleotide of interest is one or more of APP, Bace1, GSAP, APOE, CD33, GMF, CysLT1R, DMPK, SMN1, SMN2, HTT, BCL11A, ESE, HBB, Dmd, and FAH.49.The method of claim 47 or 48, wherein(i) the disease or condition is Alzheimer’s disease, and the polynucleotide of interest comprises APP, Bace1, GSAP, APOE, CD33, GMF, CysLT1R, EMX1, VEGFA, DNMT1, ALDH1A3, TET1, TET2, RNF2, RUNX1, PCSK9, CXCR4, APOB, DNMT3b, PGK1, MECP2, and / or AGBL1;(ii) the disease or condition is myotonic dystrophy type 1, and the polynucleotide of interest comprises DMPK;(iii) the disease or condition is spinal muscular atrophy, and the polynucleotide of interest comprises SMN1 and / or SMN2;(iv) the disease or condition is Huntington’s disease, and the polynucleotide of interest comprises HTT;(v) the disease or condition is sickle cell disease, and the polynucleotide of interest comprises BCL11A, ESE, and / or HBB;(vi) the disease or condition is β-thalassemia, and the polynucleotide of interest comprises BCL11A, ESE, and / or HBB;(vii) the disease or condition is Duchenne muscular dystrophy, and the polynucleotide of interest comprises Dmd; and / or(viii) the disease or condition is hereditary tyrosinemia, and the polynucleotide of interest comprises FAH.50.A method for detecting a target polynucleotide, comprising contacting the system of any of claims 24-30 with an analyte.51.The method of claim 50, wherein the analyte comprises a sample derived from a biopsy, a body fluid, feces, or a combination thereof from a subject.52.The method of claim 50 or 51, wherein the method is for diagnosis in the subject.53.The method of claim 51 or 52, wherein the subject is a human.54.A system for manipulating a polynucleotide of interest, comprising(i) a fusion protein comprising a TnpB or variant thereof and an effector or a first polynucleotide encoding the fusion protein; and(ii) a guide polynucleotide comprising a targeting sequence and a scaffold or a second polynucleotide encoding the guide polynucleotide;wherein the fusion protein is capable of specifically binding to the guide polynucleotide via the scaffold to form a complex; andwherein the targeting sequence is capable of specifically binding to the polynucleotide of interest.55.The system of claim 54, wherein the TnpB variant has weakened nuclease activity compared to the TnpB.56.The system of claim 55, wherein the TnpB variant comprises the amino acid sequence of SEQ ID NO: 8 and one or more mutations of D185A, R258A, and E279A.57.The system of claim 55 or 56, wherein the effector comprises a nuclear localization signal (NLS) .58.The system of any of claims 54-57, wherein the effector comprises a cell penetrating peptide.59.The system of any of claims 54-58, wherein the system is attached to or embedded in a support.60.The system of any of claims 54-59, wherein the effector comprises one or more of a nuclease that is not the TnpB or variant thereof, a base editor, a prime editor, an epigenetic modifier, a polymerase, a transposase, a recombinase, a reverse transcriptase, a label, a transcription modulator, a transcription factor, a topoisomerases, a helicase, a kinase, a phosphatase, an integrase, a ligase, an antibody, and an antigen.61.The system of claim 60, wherein the base editor is a deaminase, optionally a cytidine deaminase and / or an adenine deaminase.62.The system of claim 60, wherein the label is a fluorescent, luminescent, and / or a chromogenic protein.63.A method for manipulating a polynucleotide of interest and / or one or more molecules associated therewith, comprising contacting the system of any of claims 54-62 with the polynucleotide of interest.64.A method for modifying the polynucleotide of interest and / or one or more molecules associated therewith, comprising contacting the system of any of claims 54-62 with the polynucleotide of interest.65.A method for modulating transcription, replication, repair, and / or recombination of a polynucleotide of interest, comprising contacting the system of any of claims 54-62 with the polynucleotide of interest.66.A method for promoting or preventing direct or indirect binding of a molecule to the polynucleotide of interest, comprising contacting the system of any of claims 54-62 with the polynucleotide of interest.67.A method for labeling a polynucleotide of interest, comprising contacting the system of any of claims 54-62 with the polynucleotide of interest.
Citation Information
Patent Citations
Polynucleotide shuffling method
CN110462046A
Programmable gene editing using guide RNA pair
US20230407280A1
Reprogrammable TNPB polypeptides and use thereof
WO2022159892A1
Novel nucleic acid-guided nucleases and use thereof
WO2023170535A2
Tnpb-based genome editor
WO2024017189A1