Iscb protein and use thereof
Patent Information
- Application Number
- PCT/CN2026/079361
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-08-06
- Filing Date
- 2026-02-13
- Publication Date
- 2026-08-27
Smart Images

Figure PCTCN2026079361-FTAPPB-I100001 
Figure PCTCN2026079361-FTAPPB-I100002 
Figure PCTCN2026079361-FTAPPB-I100003
Abstract
Description
IscB protein and its applications
[0001] This application claims priority to Chinese Patent Application No. 2025101788142, filed on February 18, 2025, and Chinese Patent Application No. 2025110972633, filed on August 6, 2025. The full text of the aforementioned Chinese patent applications is incorporated herein by reference. Technical Field
[0002] This disclosure pertains to the field of CRISPR gene editing, specifically to an IscB protein and its applications. Background Technology
[0003] IscB proteins are structurally similar to Cas9 proteins, and most IscB proteins have a shorter CDS sequence than SpCas9. Summary of the Invention
[0004] This disclosure provides information on the IscB protein and its applications.
[0005] In one aspect, this disclosure provides an IscB protein whose amino acid sequence comprises or is an amino acid sequence having at least 50% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0006] In this disclosure, the IscB protein may be natural or non-natural, for example, it may be modified or engineered.
[0007] The table disclosed lists the IscB protein with different sequences, their corresponding ωRNA backbone sequences, and the TAM motifs they are identified.
[0008] In the specific implementation of this disclosure, the at least 50% identity refers to at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity.
[0009] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identity with any of SEQ ID NO:13-26, 41, 42, 53-93.
[0010] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 80% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0011] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 85% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0012] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 90% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0013] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 95% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0014] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 97% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0015] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 98% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0016] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that has at least 99% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0017] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 99.5% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0018] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 99.7% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0019] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 99.8% identity with any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0020] In the specific embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence that is 100% identical to any one of SEQ ID NO:13-26, 41, 42, 53-93.
[0021] In the specific implementation of this disclosure, the IscB protein retains the function of the protein shown in any of the sequences SEQ ID NO:13-26, 41, 42, 53-93.
[0022] In a specific embodiment disclosed herein, the IscB protein can form a complex with a guide polynucleotide. In a specific embodiment disclosed herein, the IscB protein can specifically bind to the target nucleic acid with the guide polynucleotide.
[0023] In a specific embodiment disclosed herein, the IscB protein can form a complex with a guide polynucleotide, and the complex can specifically bind to the target nucleic acid. In a specific embodiment disclosed herein, the IscB protein can form a complex with a guide polynucleotide, and the complex can specifically bind to the target DNA.
[0024] In specific embodiments disclosed herein, the IscB protein can specifically bind to and cleave target nucleic acids with guide polynucleotides. In specific embodiments disclosed herein, the IscB protein can specifically bind to and cleave target DNA with guide polynucleotides. In specific embodiments disclosed herein, the IscB protein can form a complex with guide polynucleotides, and this complex can specifically bind to and cleave target nucleic acids. In specific embodiments disclosed herein, the IscB protein can form a complex with guide polynucleotides, and this complex can specifically bind to and cleave target DNA.
[0025] In this disclosure, the retention of the protein function shown in any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93 refers to the retention of the ability to form a complex with the guide polynucleotide, the retention of the ability to bind to a target nucleic acid complementary to the guide sequence of the guide polynucleotide, the retention of the ability to target and cleave the target nucleic acid with the guide polynucleotide, and / or the retention of the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0026] In the specific embodiments disclosed herein, the retention of the function of the protein as shown in any of the sequences SEQ ID NO:13-26, 41, 42, 53-93 is to retain the ability to guide the formation of complexes with polynucleotides.
[0027] In the specific embodiments disclosed herein, the function of retaining the protein as shown in any of the sequences SEQ ID NO:13-26, 41, 42, 53-93 is to retain the ability to bind to target nucleic acids complementary to the guide sequence of the guide polynucleotide.
[0028] In the specific implementation of this disclosure, the function of retaining the protein as shown in any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93 is to retain and guide the polynucleotide to target and cleave the target nucleic acid.
[0029] In the specific embodiments disclosed herein, the function of retaining the protein as shown by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93 is to retain the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0030] In a preferred embodiment disclosed herein, the amino acid sequence of the IscB protein comprises or is any of the amino acid sequences shown in SEQ ID NO:13-26, 41, 42, 53-93.
[0031] In the specific embodiments disclosed herein, the TAM sequence (5'→3', target adjacent motif) recognizable by the IscB protein is selected from any one or more of the following:
[0032] A, C, T, G
[0033] TA、TC、GN、AA、AG、TG、AN、GG、CG、TN、NT、NG、GT、NA、CC、AC、GC、AT、CT、GA、TT、CN、NC、CA、
[0034] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG
[0035] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GGCG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC
[0036] AAAAA、AAAAC、AAAAG、AAAAT、AAACC、AAACT、AAAGC、AAAGG、AAAGT、AAATA、AAATT、AACAT、AACTC、AACTG、AAGAG、AAGAT、AAGCC、AAGGA、AAGTA、AATAA、AATA C、AATAG、AATAT、AATCT、AATGA、AATTA、AATTC、AATTG、AATTT、ACAAA、ACAAC、ACAAT、ACACA、ACACC、ACACT、ACATA、ACATC、ACATG、ACATT、ACATT、ACCAC、ACCA CG、ACCTT、ACGAT、ACGCC、ACGGC、ACTGA、ACTGC、ACTTT、AGAAG、AGACA、AGACC、AGAGG、AGATA、AGCAA、AGCAC、AGGAT、AGGTA、AGGTG、AGTAA、AGTCA、AGTCT、A GTGC、AGTGT、AGTTT、ATAAA、ATAAC、ATAAT、ATACA、ATACC、ATAGA、ATATA、ATATC、ATATG、ATCAC、ATCTC、ATGAA、ATGAC、ATGGA、ATGGT、ATGTT、ATTAA、ATTAT、 ATTCA、ATTCT、ATTGA、ATTGC、ATTTA、ATTTT、CAAA、CAAAC、CAACC、CAAGC、CAATA、CAATT、CACAA、CACAC、CACCA、CACCG、CACCT、CACGA、CACTG、CACTT、CAGC A、CAGCT、CATAT、CATCC、CATCT、CATTC、CCAAT、CCACG、CCCAC、CCGAG、CCGCA、CCGCC、CCGTC、CCTAC、CCTAT、CCTCC、CCTCT、CCTGA、CCTGC、CCTTA、CCTTG TT、CGATA、CGATT、CGCCA、CGCCG、CGCCT、CGGCA、CGGGCG、CGTGG、CGTTA、CTATA、CTATC、CTCAG、CTCCC、CTCGG、CTCTC、CTGAA、CTTCC、CTTCT、CTTGA、CTTTA、CTTTG、CTTTT、GAAAA、GAAGA、GAAGG、GAAGT、GACCC、GAGTC、GATAA、GATGC、GATTG、GATTT、GCAAG、GCAAT、GCACG、GCCAT、GCCCG、GCCGA、GCCGC、GCGTA、GCTGA,GCTTC, GCTTG, GCTTT, GGAAG, GGCTT, GGGAA, GGGCA, GGGCT, GGTGC, GGTGG, GGGTGT, GGTTT, GTAAC, GTAAT, GTCCT, GTCTA, GTCTC, GTGCC, GTGCT, GTGTG, GTTAA, GTTAG, GTTTTC, GTTTT, TAAAA, TAAAC, TAAAG, TAAAT, TAAC A. TAACT, TAATA, TAATT, TACAT, TACCA, TACGC, TACTG, TAGAG, TAGGA, TATAA, TATAG, TATAT, TATCA, TATCC, TATCG, TATGA, TATGG, TATGT, TATTA, TATTG, TCAAA, TCAGG, TCATA, TCATC, TCATT, TCCAA, TCCCA, TCCCG, TCC GC, TCCTC, TCCTT, TCGGC, TCTAT, TCTCA, TCTCC, TCTCG, TCTGG, TCTTA, TGAAG, TGACA, TGATA, TGATT, TGCAC, TGCAG, TGCCA, TGCCC, TGCCG, TGCCT, TGCTC, TGCTG, TGCTT, TGGAA, TGGCT, TGGGG, TGGGT, TGGTG, TGTAT, TG TGA, TGTGC, TTGGT, TGTTA, TGTTT, TTAAA, TTAAG, TTAAT, TTACA, TTATA, TTATG, TTATT, TTCAT, TTCCT, TTCTC, TTCTT, TTGCC, TTGGC, TTGTG, TTGTT, TTTAA, TTTAC, TTTAT, TTTGG, TTTGT, TTTTA, TTTTC, TTTTG, TTTTT;,
[0037] N is A, T, C, or G.
[0038] Optionally, the TAM sequence (5'→3') recognizable by the IscB protein may be selected from any one or more of the following:
[0039] WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG, where R=A / G, Y=C / T, M=A / C, K=G / T, S=G / C, W=A / T, H=A / T / C, B=G / T / C, V=G / A / C, D=G / A / T, N=A / T / C / G.
[0040] In the specific implementation scheme disclosed herein, the guiding polynucleotide is ωRNA.
[0041] In the specific implementation scheme disclosed herein, the IscB protein can recognize TAMs with sequence A.
[0042] In the specific implementation of this disclosure, the IscB protein can recognize TAMs with the sequence C.
[0043] In the specific implementation scheme disclosed herein, the IscB protein can recognize TAMs with a sequence of T.
[0044] In the specific implementation scheme disclosed herein, the IscB protein can recognize TAMs with the sequence G.
[0045] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-TA-3'.
[0046] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-TC-3'.
[0047] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-GN-3'.
[0048] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-AA-3'.
[0049] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-AG-3'.
[0050] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-TG-3'.
[0051] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-AN-3'.
[0052] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-GG-3'.
[0053] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-CG-3'.
[0054] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-TN-3'.
[0055] In the specific implementation of this disclosure, the IscB protein can recognize TAM with a sequence of 5'-NT-3'.
[0056] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-NG-3'.
[0057] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-GT-3'.
[0058] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-NA-3'.
[0059] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-CC-3'.
[0060] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-AC-3'.
[0061] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-GC-3'.
[0062] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-AT-3'.
[0063] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-CT-3'.
[0064] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-GA-3'.
[0065] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-TT-3'.
[0066] In the specific implementation of this disclosure, the IscB protein can recognize a TAM with the sequence 5'-CN-3'.
[0067] In the specific implementation of this disclosure, the IscB protein can recognize TAM with the sequence 5'-NC-3'.
[0068] In the specific implementation of this disclosure, the IscB protein can recognize a TAM with the sequence 5'-CA-3'.
[0069] In the specific embodiments disclosed herein, the TAM sequence (5'→3') recognizable by the IscB protein is selected from any one or more of the following:
[0070] WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG, where R=A / G, Y=C / T, M=A / C, K=G / T, S=G / C, W=A / T, H=A / T / C, B=G / T / C, V=G / A / C, D=G / A / T, N=A / T / C / G.
[0071] In the specific implementation scheme disclosed herein, the guiding polynucleotide is ωRNA.
[0072] In the specific implementation disclosed herein, the ωRNA comprises a guide sequence and an ωRNA backbone sequence.
[0073] In some embodiments disclosed herein, the IscB protein is an inactivated IscB variant. In some embodiments disclosed herein, the IscB protein is a nuclease-inactivated variant. In some embodiments disclosed herein, the IscB protein is a dead IscB inactivated variant or a nickase IscB inactivated variant. Optionally, the Ruvc domain of the IscB protein is inactivated.
[0074] In some embodiments of this disclosure, the IscB protein is selected from the active fragments constituting any of the IscB proteins described in this disclosure.
[0075] In some embodiments disclosed herein, the amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of SEQ ID NO:13-26, 41, 42, 53-93.
[0076] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0077] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide containing a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide contains a backbone sequence that can interact with the IscB protein; even further, the backbone sequence contains or is an ωRNA backbone sequence.
[0078] Optionally, the backbone sequence comprises or is a sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any of SEQ ID NO:27-40, 43, 44, 94-134;
[0079] Optionally, the backbone sequence may or may not contain a tracrRNA sequence;
[0080] Optionally, the TAM sequence (5'→3') recognizable by the IscB protein is selected from any one or more of the following: WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG.
[0081] Where: R = A or G, Y = C or T, M = A or C, K = G or T, S = G or C, W = A or T, H = A, T or C, B = G, T or C, V = G, A or C, D = G, A or T, N = A, T, C or G.
[0082] Optionally, the correspondence between the IscB protein, the specific guide polynucleotide backbone sequence associated with the protein and its similar sequences, and the TAM motif is shown in Table 9; the correspondence between the specific IscB protein and its similar proteins, the specific guide polynucleotide backbone sequence associated with the protein and its similar sequences, and the TAM motif in Table 9 is also the same as the correspondence in Table 9.
[0083] Optionally, the IscB protein has an amino acid sequence shown in any one of SEQ ID NO: 13-26, 41, 42, 53-93;
[0084] Optionally, the skeleton sequence is any one of the sequences shown in SEQ ID NO:27-40, 43, 44, 94-134.
[0085] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:13.
[0086] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0087] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:27;
[0088] Optionally, the IscB protein recognition sequence is the TAM motif of WHG.
[0089] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:14.
[0090] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0091] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:28;
[0092] Optionally, the IscB protein recognizes the DHG TAM motif.
[0093] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:15.
[0094] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0095] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:29;
[0096] Optionally, the IscB protein recognizes the ATAAA TAM motif.
[0097] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:16.
[0098] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0099] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:30;
[0100] Optionally, the IscB protein recognition sequence is the TAM motif of ATG.
[0101] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:41.
[0102] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0103] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:43;
[0104] Optionally, the IscB protein can recognize the TAM motif of AYGG.
[0105] In some embodiments disclosed herein, an IscB protein is provided, the amino acid sequence of which comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% identity with any of the sequences in SEQ ID NO:42.
[0106] Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid;
[0107] Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; further, the guide polynucleotide comprises a backbone sequence that can interact with the IscB protein; even further, the backbone sequence comprises or is an ωRNA backbone sequence; optionally, the backbone sequence comprises a sequence having at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, or at least 97% sequence identity with the sequence shown in any one of SEQ ID NO:44;
[0108] Optionally, the IscB protein can recognize the TAM motif of NGG.
[0109] In some embodiments, the IscB protein is a mutant of the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93.
[0110] In some embodiments, the IscB protein is an inactivated variant of the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93.
[0111] In some embodiments, the IscB protein provided herein contains one, two, or more mutations compared to the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof. In some examples, the IscB protein, compared to the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93, contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 3 3, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 7 8, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.In some instances, with SEQ ID The IscB protein represented by any of the sequences NO:13-26, 41, 42, 53-93 is compared to the IscB protein containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide.
[0112] On the other hand, one technical solution provided in this disclosure is: a guide polynucleotide comprising (i) a backbone sequence having at least 50% identity with any one of SEQ ID NO:27-40, 43, 44, 94-134, and (ii) a guide sequence engineered to hybridize with a target nucleic acid; the backbone sequence is linked to the guide sequence, and the guide polynucleotide is capable of forming a complex with an IscB protein and guiding the complex to bind to the target nucleic acid in a sequence-specific manner.
[0113] In some embodiments disclosed herein, the backbone sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0114] In some embodiments disclosed herein, the skeleton sequence has at least 60% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0115] In some embodiments disclosed herein, the skeleton sequence has at least 65% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0116] In some embodiments disclosed herein, the skeleton sequence has at least 70% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0117] In some embodiments disclosed herein, the skeleton sequence has at least 75% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0118] In some embodiments disclosed herein, the skeleton sequence has at least 80% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0119] In some embodiments disclosed herein, the skeleton sequence has at least 85% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0120] In some embodiments disclosed herein, the skeleton sequence has at least 90% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0121] In some embodiments disclosed herein, the skeleton sequence has at least 95% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0122] In some embodiments disclosed herein, the skeleton sequence has at least 96% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0123] In some embodiments disclosed herein, the skeleton sequence has at least 97% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0124] In some embodiments disclosed herein, the skeleton sequence has at least 98% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0125] In some embodiments disclosed herein, the skeleton sequence has 100% sequence identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0126] In the specific embodiments disclosed herein, the backbone sequence comprises or is an ωRNA backbone sequence.
[0127] In a preferred embodiment, the IscB protein is the IscB protein described in this disclosure.
[0128] In specific embodiments disclosed herein, the guide sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-35 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-30 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-22 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-22 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.
[0129] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the guide sequence and the target nucleic acid are 90%-100% complementary.
[0130] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid.
[0131] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the mismatch between the guide sequence and the target nucleic acid does not exceed one nucleotide.
[0132] In specific embodiments disclosed herein, the backbone sequence comprises 15-100 nucleotides. In specific embodiments disclosed herein, the backbone sequence comprises 15-90 nucleotides. In specific embodiments disclosed herein, the backbone sequence comprises 15-80 nucleotides. In specific embodiments disclosed herein, the backbone sequence comprises 15-70 nucleotides. In specific embodiments disclosed herein, the backbone sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-30 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0133] In the specific implementation disclosed herein, the guiding sequence is located at the 3' end of the skeleton sequence.
[0134] In the specific implementation disclosed herein, the guiding sequence is located at the 5' end of the skeleton sequence.
[0135] In a specific embodiment disclosed herein, the guiding polynucleotide further comprises tracrRNA.
[0136] In the specific implementation disclosed herein, the guiding polynucleotide does not contain tracrRNA.
[0137] In specific embodiments disclosed herein, the tracrRNA may be complementary to the backbone sequence. Typically, this complementary pairing is partial base pairing. In specific embodiments disclosed herein, the tracrRNA may interact with the backbone sequence.
[0138] In specific embodiments disclosed herein, the tracrRNA sequence is linked to the backbone sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the backbone sequence are linked via a nucleotide sequence. In specific embodiments disclosed herein, the tracrRNA sequence and the backbone sequence are linked via a nucleotide sequence consisting of 1-10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the backbone sequence are linked via a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the backbone sequence are linked via a nucleotide sequence consisting of 4 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence and the backbone sequence are linked via a 5'-GAAA-3' sequence.
[0139] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 3' end of the backbone sequence.
[0140] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 5' end of the backbone sequence.
[0141] In the specific embodiments disclosed herein, the tracrRNA comprises 10-200 nucleotides. In the specific embodiments disclosed herein, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10- 40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In the specific implementation scheme disclosed herein, the tracrRNA contains 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, and 5 5, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0142] On the other hand, one technical solution provided in this disclosure is: an IscB inactivation variant, wherein the IscB inactivation variant is a nuclease activity inactivation variant of the IscB protein as described in this disclosure.
[0143] In this document, depending on the context, the term "IscB protein" may refer to the inactivated IscB variants. However, given the importance of the inactivated IscB variants (non-limiting examples include their fusion with deaminases for single-base editing, their fusion with transcriptional activation or repression domains for transcriptional regulation, etc.), they will be described separately and with emphasis; this does not mean that the term "IscB protein" necessarily excludes the inactivated IscB variants.
[0144] In the specific embodiments disclosed herein, the IscB inactivating variant is a variant with completely inactivated nuclease activity, i.e., a dead IscB inactivating variant (dIscB). The dIscB can only bind to the target nucleic acid under the mediation of a guide polynucleotide, and has little or no ability to cleave the target nucleic acid. For example, the target nucleic acid cleavage efficiency of the dIscB is ≤20%, ≤15%, ≤10%, ≤5%, ≤4%, ≤3%, ≤2%, or ≤1% of the target nucleic acid cleavage efficiency of the IscB protein before inactivation mutation.
[0145] In the specific embodiments disclosed herein, the IscB inactivating variant is a variant with partially inactivated nuclease activity. Further, the partially inactivated nuclease variant is an IscB nickase (nIscB), which binds to the target nucleic acid under the mediation of a guide polynucleotide and then cleaves one single strand of the double-stranded target nucleic acid without cleaving the other single strand.
[0146] In a preferred embodiment disclosed herein, the inactivating variant of IscB is the inactivation of the Ruvc domain of the IscB protein.
[0147] In a preferred embodiment disclosed herein, the inactivating variant of IscB is the inactivation of the Ruvc-I, Ruvc-II, or Ruvc-III domains of the IscB protein.
[0148] In a preferred embodiment disclosed herein, the IscB inactivating variant is obtained by introducing an inactivating mutation into the Ruvc-I, Ruvc-II, or Ruvc-III domains of the IscB protein.
[0149] In the specific implementation of this disclosure, the TAM sequence recognizable by the inactivated IscB variant is the same as the TAM sequence recognizable by the IscB protein.
[0150] In the specific embodiments disclosed herein, the TAM sequence (5'→3') recognizable by the IscB inactivating variant is selected from any one or more of the following:
[0151] A, C, T, G
[0152] TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA,
[0153] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG
[0154] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,
[0155] N is A, T, C, or G.
[0156] Optionally, the TAM sequence (5'→3') recognizable by the IscB protein variant may be selected from any one or more of the following:
[0157] WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG.
[0158] On the other hand, one technical solution provided by this disclosure is: a fusion protein or conjugate comprising the following elements: (1) an IscB protein as described in this disclosure, or an inactivated variant of IscB as described in this disclosure; and (2) homologous or heterologous functional domains.
[0159] In this document, depending on the context, the term "IscB protein" may refer to the inactive variants of IscB. However, given the importance of the inactive variants of IscB (non-limiting examples include base editing via fusion with deaminases, epigenetic editing via fusion with repressor domains, DNA methylation domains, histone methylation domains, transcription activation domains, or transcription repression domains, etc.), they will be described separately and in detail; this does not mean that the term "IscB protein" necessarily excludes the inactive variants of IscB.
[0160] In some embodiments disclosed herein, a fusion protein is provided, the fusion protein comprising: (1) an IscB protein as described herein, or an inactivated variant of IscB as described herein; and (2) a homologous or heterologous functional domain.
[0161] In some embodiments disclosed herein, a fusion protein is provided, the fusion protein comprising: (1) the IscB protein as described herein; and (2) homologous or heterologous functional domains.
[0162] In a specific embodiment of this disclosure, a conjugate is provided, the conjugate comprising: (1) the IscB protein as described in this disclosure, or the IscB inactivating variant as described in this disclosure; and (2) a homologous or heterologous functional domain.
[0163] In a specific embodiment of this disclosure, a conjugate is provided, the conjugate comprising: (1) the IscB protein as described in this disclosure; and (2) a homologous or heterologous functional domain.
[0164] In some embodiments, the functional domain has enzymatic activity that modifies the target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristylation activity, and / or demyristylation activity.
[0165] In some embodiments, the functional domain is optionally selected from one or more of the following: nucleases (e.g., FokI), methyltransferases, demethylases, DNA repair enzymes, DNA damage enzymes, deaminases, superoxide dismutases, alkylating enzymes, depurinases, oxidases, pyrimidine dimer forming enzymes, integrases, transposases, recombinases, polymerases, ligases, helicases, photolyases, glycosylation enzymes, deglycosylation enzymes, acetyltransferases, deacetylases, kinases, phosphatases, ubiquitin ligases, deubiquitinating enzymes, adenylate acylases, deadenylate acylases, SUMOylating enzymes, deSUMOylating enzymes, myristylases, and / or demyristylases.
[0166] In the specific implementation scheme disclosed herein, the homologous or heterologous functional domains are selected from one, two, three, four or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, affinity tag, reporter tag, affinity domain and reporter domain.
[0167] In some embodiments disclosed herein, the subcellular localization signal may be selected from: nuclear localization signal, nuclear output signal, mitochondrial localization signal, or chloroplast localization signal.
[0168] In the specific embodiments disclosed herein, the fusion protein or conjugate comprises 1, 2, 3, 4, 5, 6, 7, 8, 9 or more of the homologous or heterologous functional domains; the functional domains may be the same or different.
[0169] In some embodiments, the fusion protein or conjugate is arbitrarily connected to 0, 1, 2, 3, 4, 5, 6, 7, 8 or more of the functional domains at the N-terminus and / or C-terminus of the IscB protein.
[0170] In the specific implementation scheme disclosed herein, the fusion protein contains one, two, three, four or more nuclear localization signals.
[0171] In specific embodiments disclosed herein, the fusion protein can be used to achieve base editing, for example, in conjunction with a guide polynucleotide to achieve base editing. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a deaminase domain.
[0172] In the specific embodiments disclosed herein, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and optionally one or two UGI domains. The fusion protein can be used to achieve C→T base editing of target nucleic acids.
[0173] In the specific implementation disclosed herein, the fusion protein includes a nuclear localization signal and an adenosine deaminase domain. The fusion protein can be used to achieve A→G base editing of target nucleic acids.
[0174] In the specific embodiments disclosed herein, the fusion protein comprises a nuclear localization signal, a cytidine deaminase domain, and an adenosine deaminase domain. In the specific embodiments disclosed herein, the fusion protein comprises one, two, or three nuclear localization signals and a deaminase domain. In the specific embodiments disclosed herein, the fusion protein comprises a UGI domain. In the specific embodiments disclosed herein, the fusion protein comprises one, two, or three nuclear localization signals, a deaminase domain, and one or two UGI domains.
[0175] In specific embodiments disclosed herein, the fusion protein can be used to achieve transcriptional activation of a specific target gene, for example, in combination with a guide polynucleotide to achieve transcriptional activation of a specific target gene. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a transcription activation domain.
[0176] In specific embodiments disclosed herein, the fusion protein can be used to achieve transcriptional repression of specific target genes, for example, in combination with a guide polynucleotide to achieve transcriptional repression of specific target genes. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a transcriptional repression domain.
[0177] In specific embodiments disclosed herein, the fusion protein can be used to achieve methylation of a specific target sequence, for example, in conjunction with a guide polynucleotide to achieve methylation of a specific target sequence. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a DNA methylation domain.
[0178] In specific embodiments disclosed herein, the fusion protein can be used to achieve demethylation of specific target sequences, for example, in conjunction with a guide polynucleotide to achieve demethylation of specific target sequences. In specific embodiments disclosed herein, the fusion protein includes a nuclear localization signal and a DNA demethylation domain.
[0179] In a preferred embodiment disclosed herein, the nuclease domain includes a polypeptide having ssDNA cleavage activity and / or a polypeptide having dsDNA cleavage activity.
[0180] In a preferred embodiment disclosed herein, the nuclease domain includes a polypeptide having ssDNA cleavage activity.
[0181] In a preferred embodiment disclosed herein, the nuclease domain comprises a polypeptide having dsDNA cleavage activity.
[0182] In the specific embodiments disclosed herein, the IscB protein or its inactivated variant is directly or indirectly linked to the homologous or heterologous functional domain.
[0183] In the preferred embodiment disclosed herein, the direct connection is a covalent connection, and the indirect connection is a connection via an amino acid linker or a non-amino acid linker.
[0184] In a preferred embodiment disclosed herein, the homologous or heterologous functional domains are fused or conjugated at the N-terminus, C-terminus, or internally relative to the IscB protein or inactivating variant.
[0185] In this disclosure, the fusion protein refers to the connection between the element (1) and the element (2) via peptide segments or direct connection; the conjugate refers to the connection between the element (1) and the element (2) via non-peptide chemical bonds.
[0186] In the specific embodiments disclosed herein, the TAM sequence recognizable by the fusion protein or conjugate is the same as the TAM sequence recognizable by the IscB protein.
[0187] In the specific embodiments disclosed herein, the TAM sequence (5'→3') recognizable by the fusion protein or conjugate is selected from any one or more of the following:
[0188] A, C, T, G
[0189] TA、TC、GN、AA、AG、TG、AN、GG、CG、TN、NT、NG、GT、NA、CC、AC、GC、AT、CT、GA、TT、CN、NC、CA、
[0190] NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT ANG、GNA、GTT、NGA、TAA、GTA、GGN、GNT、NCG、ATT、CCA、CNN、AAA、AAC、ATN、GA G、CTG、ACG、NAA、TAN、NAT、CNA、GCN、GTC、NCN、CTN、CNC、ANT、NNC、CAG、NAN、 ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG
[0191] NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GCGG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC;,
[0192] N is A, T, C, or G.
[0193] Optionally, the TAM sequence (5'→3') recognizable by the fusion protein may be selected from any one or more of the following:
[0194] WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG.
[0195] On the other hand, one technical solution provided in this disclosure is: an isolated nucleic acid, said nucleic acid encoding the IscB protein as described in this disclosure, the IscB inactivating variant as described in this disclosure, or the fusion protein or conjugate as described in this disclosure.
[0196] In some embodiments disclosed herein, the nucleic acid encodes the IscB protein or the fusion protein as described herein.
[0197] In a preferred embodiment disclosed herein, the nucleic acid is codon-optimized for expression in cells.
[0198] In a preferred embodiment disclosed herein, the nucleic acid is codon-optimized for expression in eukaryotes, mammals such as humans or non-human mammals, plants, insects, birds, reptiles, rodents (e.g., mice, rats), fish, worms / nematodes, or yeast.
[0199] On the other hand, one technical solution provided in this disclosure is: an IscB system, the IscB system comprising:
[0200] a. The IscB protein as described in this disclosure, the IscB inactivating variants as described in this disclosure, the fusion protein or conjugate as described in this disclosure, or the isolated nucleic acid as described in this disclosure; and
[0201] b. A guiding polynucleotide, or a polynucleotide sequence encoding the guiding polynucleotide;
[0202] The IscB protein, the inactivated variant of IscB, or the fusion protein or conjugate forms a complex with the guiding polynucleotide; the guiding polynucleotide contains a guiding sequence that is engineered to guide the complex to bind to the target nucleic acid in a sequence-specific manner.
[0203] The isolated nucleic acid encodes an IscB protein as described in this disclosure, an inactivated IscB variant as described in this disclosure, or a fusion protein or conjugate as described in this disclosure.
[0204] In a specific embodiment of this disclosure, the guiding polynucleotide comprises a backbone sequence linked to the guiding sequence. In a specific embodiment of this disclosure, the backbone sequence comprises or is an ωRNA backbone sequence.
[0205] In the specific implementation of this disclosure, the skeleton sequence has at least 50% identity with any one of SEQ ID NO:27-40, 43, 44, 94-134.
[0206] In specific embodiments disclosed herein, the guiding polynucleotide comprises a backbone sequence linked to the guiding sequence. Further, in some embodiments, the backbone sequence has at least 50% identity with any one of SEQ ID NO:27-40, 43, 44, 94-134. In some specific embodiments, the backbone sequence has at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or 100% sequence identity with the sequence shown in any one of SEQ ID NO: 27-40, 43, 44, and 94-134. Furthermore, in some specific embodiments, the backbone sequence comprises or is the sequence shown in any one of SEQ ID NO: 27-40, 43, 44, and 94-134.
[0207] In specific embodiments disclosed herein, the guide sequence comprises 15-60 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-50 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-40 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-35 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-30 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 15-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-25 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 18-22 nucleotides. In specific embodiments disclosed herein, the guide sequence comprises 20-22 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides.
[0208] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the guide sequence and the target nucleic acid are 90%-100% complementary.
[0209] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid.
[0210] In the specific implementation scheme disclosed herein, the guide sequence hybridizes with the target nucleic acid, and the mismatch between the guide sequence and the target nucleic acid does not exceed one nucleotide.
[0211] In the specific implementation scheme disclosed herein, the skeleton sequence comprises at most 30, at most 40, at most 50, at most 60, at most 70, at most 80, at most 90, at most 100, at most 110, at most 120, at most 130, at most 140, at most 150, at most 160, at most 170, at most 180, at most 190, at most 200, and so on. Up to 210, up to 220, up to 230, up to 240, up to 250, up to 260, up to 270, up to 280, up to 290, up to 300, up to 310, up to 320, up to 330, up to 340, up to 350, up to 360, up to 370, up to 380, up to 390, or up to 400 nucleotides.
[0212] In the specific implementation scheme disclosed herein, the backbone sequence comprises at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, at least 220, at least 230, at least 240, at least 250, at least 260, at least 270, at least 280, at least 290, at least 300, at least 310, at least 320, at least 330, at least 340, at least 350, at least 360, at least 370, at least 380, at least 390, or at least 400 nucleotides.
[0213] In the specific implementation scheme disclosed herein, the skeleton sequence includes 15-400, 15-390, 15-380, 15-370, 15-360, 15-350, 15-340, 15-330, 15-320, 15-310, 15-300, 15-290, 15-280, 15-270, 15-260, 15-250, 15- 240, 15-230, 15-220, 15-210, 15-200, 15-190, 15-180, 15-170, 15-160, 15-150, 15-140, 15-130, 15-120, 15-110, 15-100, 15-90, 15-80, 15-70, 15-60 or 15-50 nucleotides.
[0214] In a specific embodiment of this disclosure, the guide sequence comprises 15-50 nucleotides. In a specific embodiment of this disclosure, the guide sequence comprises 15-40 nucleotides. In a specific embodiment of this disclosure, the guide sequence comprises 20-40 nucleotides. In a specific embodiment of this disclosure, the guide sequence comprises 20-30 nucleotides. In the specific implementation scheme disclosed herein, the guiding sequence comprises 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 nucleotides.
[0215] In the specific implementation disclosed herein, the guiding sequence is located at the 3' end of the skeleton sequence.
[0216] In the specific implementation disclosed herein, the guiding sequence is located at the 5' end of the skeleton sequence.
[0217] In a specific embodiment disclosed herein, the guiding polynucleotide further comprises tracrRNA.
[0218] In the specific embodiments disclosed herein, the tracrRNA may be complementary to the ωRNA backbone sequence. Typically, this complementary pairing is a partial base pairing. In the specific embodiments disclosed herein, the tracrRNA may interact with the ωRNA backbone sequence.
[0219] In the specific embodiments disclosed herein, the tracrRNA sequence is linked to the ωRNA backbone sequence. In the specific embodiments disclosed herein, the tracrRNA sequence and the ωRNA backbone sequence are linked via a nucleotide sequence. In the specific embodiments disclosed herein, the tracrRNA sequence and the ωRNA backbone sequence are linked via a nucleotide sequence consisting of 1-10 nucleotides. In the specific embodiments disclosed herein, the tracrRNA sequence and the ωRNA backbone sequence are linked via a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In the specific embodiments disclosed herein, the tracrRNA sequence and the ωRNA backbone sequence are linked via a nucleotide sequence consisting of 4 nucleotides. In the specific embodiments disclosed herein, the tracrRNA sequence and the ωRNA backbone sequence are linked via a 5'-GAAA-3' sequence.
[0220] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 3' end of the ωRNA backbone sequence.
[0221] In the specific implementation disclosed herein, the tracrRNA sequence is located at the 5' end of the ωRNA backbone sequence.
[0222] In the specific embodiments disclosed herein, the tracrRNA comprises 10-200 nucleotides. In the specific embodiments disclosed herein, the tracrRNA comprises 10-190, 10-180, 10-170, 10-160, 10-150, 10-140, 10-130, 10-120, 10-110, 10-100, 10-90, 10-80, 10-70, 10-60, 10-50, 10- 40, 10-30, 10-20, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 10-100, 20-100, 30-100, 40-100, 20-90, 20-80, 20-70, 20-60, 20-50, or 30-50 nucleotides. In the specific implementation scheme disclosed herein, the tracrRNA contains 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, and 5 5, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides.
[0223] In a preferred embodiment of this disclosure, the guiding polynucleotide is the guiding polynucleotide as described in this disclosure.
[0224] In the specific implementation of this disclosure, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA.
[0225] In the preferred embodiment disclosed herein, the DNA is eukaryotic DNA; preferably, the eukaryotic DNA is non-human mammalian DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA.
[0226] In the specific implementation scheme disclosed herein, the target nucleic acid is a disease or symptom-related gene or a gene related to signal transduction biochemical pathways, or the target nucleic acid is a reporter gene; for example, the disease or symptom is a hematological disease or symptom, an ophthalmic disease or symptom, a nervous system disease or symptom, a respiratory disease or symptom, a liver disease or symptom, a metabolic system disease or symptom, cancer, or an infectious disease.
[0227] In some embodiments, the target nucleic acid is a disease or symptom-related gene, wherein the disease or symptom is selected from: hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes mellitus, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Bader-Bieder syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens camp Malnutrition, pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle invasive tumors. Infiltrative bladder cancer, non-alcoholic fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis (ALS). Acute urinary incontinence, acute intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy.Canavan disease, cocaine addiction, Krabby's disease, Krieger-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myotonic dystrophy, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, meat Tumors, breast cancer, Ritter syndrome, triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease I Type A glycogen storage disease, type IIb glycogen storage disease, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal dysplasia, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignancy Osteosclerosis, nutritional epidermolysis bullosa, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0228] In some embodiments, the genes associated with the thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0229] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0230] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0231] The genes related to AATD lung disease include, but are not limited to, AATD;
[0232] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0233] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0234] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0235] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0236] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0237] The genes related to hemophilia B include, but are not limited to, factor IX;
[0238] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0239] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0240] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0241] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0242] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0243] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0244] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0245] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0246] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0247] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0248] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0249] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0250] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0251] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0252] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B.
[0253] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0254] The epilepsy-related genes include, but are not limited to, GAT1.
[0255] On the other hand, one technical solution provided in this disclosure is: a vector system comprising one or more recombinant vectors, wherein the recombinant vectors comprise isolated nucleic acids as described in this disclosure, or the IscB system as described in this disclosure.
[0256] In the specific implementation scheme disclosed herein, the recombinant vector further includes a regulatory sequence.
[0257] In specific embodiments disclosed herein, the vector system comprises one or more recombinant vectors, the recombinant vectors comprising a multinucleotide sequence encoding the IscB protein, IscB inactivating variant, fusion protein, or conjugate disclosed herein, and a multinucleotide sequence encoding the guide multinucleotide.
[0258] In the specific embodiments disclosed herein, the polynucleotide sequence encoding the IscB protein, IscB inactivating variant, fusion protein, or conjugate is operatively linked to regulatory sequence 1.
[0259] In a specific implementation of this disclosure, the polynucleotide sequence encoding the guide polynucleotide is operatively linked to regulatory sequence 2.
[0260] Furthermore, in the specific implementation scheme disclosed herein, the regulatory sequence 1 and the regulatory sequence 2 may be the same or different sequences.
[0261] In a preferred embodiment disclosed herein, the regulatory sequence is selected from one or more of a promoter, enhancer, internal ribosome entry site, and transcription termination signal; the promoter is, for example, a constitutive promoter, an inducible promoter, a broad-spectrum promoter, or a tissue-specific promoter; and / or, the transcription termination signal is, for example, a polyadenylation signal or a polyU sequence.
[0262] In the specific implementation scheme disclosed herein, the backbone of the recombinant vector is an adeno-associated virus vector, a lentiviral vector, a ribonucleoprotein complex, or a virus-like particle.
[0263] In the preferred embodiment disclosed herein:
[0264] When the backbone is an adeno-associated virus vector, the adeno-associated virus vector is a recombinant adeno-associated virus vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74.
[0265] When the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; preferably, the isolated nucleic acid is linked to an aptamer sequence.
[0266] When the backbone is a virus-like particle, the isolated nucleic acid is linked to a gene encoding a gag protein.
[0267] On the other hand, one technical solution provided by this disclosure is: a delivery system comprising: (1) a delivery tool, and (2) an IscB protein as described in this disclosure, a guide polynucleotide as described in this disclosure, an IscB inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, an IscB system as described in this disclosure, or a vector system as described in this disclosure.
[0268] In a preferred embodiment disclosed herein, the delivery tool is a virus, lipid nanoparticles, nanoparticles, liposomes, exosomes, microvesicles, or a gene gun.
[0269] In a preferred embodiment disclosed herein, the delivery tool is a lipid nanoparticle comprising the guiding polynucleotide and mRNA encoding the IscB protein, the inactivating variant of IscB, or the fusion protein or conjugate.
[0270] On the other hand, one technical solution provided in this disclosure is: a cell comprising the IscB protein as described in this disclosure, the guiding polynucleotide as described in this disclosure, the IscB inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the IscB system as described in this disclosure, or the vector system as described in this disclosure.
[0271] In some embodiments disclosed herein, the cells are prokaryotic cells.
[0272] In some embodiments disclosed herein, the cells are eukaryotic cells.
[0273] In some embodiments disclosed herein, the eukaryotic cells are mammalian cells.
[0274] On the other hand, one technical solution provided in this disclosure is: a pharmaceutical composition comprising an IscB protein as described in this disclosure, a guide polynucleotide as described in this disclosure, an IscB inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, an IscB system as described in this disclosure, a carrier system as described in this disclosure, a delivery system as described in this disclosure, or a cell as described in this disclosure.
[0275] Preferably, the pharmaceutical composition comprises pharmaceutically acceptable excipients.
[0276] On the other hand, one technical solution provided in this disclosure is: a kit comprising the IscB protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the IscB inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the IscB system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, or the cell as described in this disclosure.
[0277] In a preferred embodiment disclosed herein, the kit further comprises a cut buffer. The cut buffer may be any buffer known in the art suitable for the cleavage of target nucleic acids by the IscB protein.
[0278] On the other hand, one technical solution provided in this disclosure is the use of the IscB protein, the guide polynucleotide, the IscB inactivating variant, the fusion protein or conjugate, the nucleic acid, the IscB system, the vector system, the delivery system, the cell, the pharmaceutical composition, or the kit as described in this disclosure in the preparation of reagents or drugs for the diagnosis, treatment, and / or prevention of diseases or conditions related to the target nucleic acid.
[0279] In the specific implementation scheme disclosed herein, the disease or condition is a hematological disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, cancer, or an infectious disease; and / or, the reagent or drug is used to: cleave one or more target nucleic acid molecules or create nicks in one or more target nucleic acid molecules, activate or upregulate the expression of one or more target nucleic acid molecules, activate or inhibit the transcription of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label, or detect one or more target nucleic acid molecules, bind one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and mask one or more target nucleic acid molecules.
[0280] In the specific implementation scheme disclosed herein, the diseases or conditions mentioned are selected from: Hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, Hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, Mucopolysaccharidosis type IIIA, Mucopolysaccharidosis type IIIB, Gaucher disease type III, Mucopolysaccharidosis type II, Type II diabetes mellitus, Mucopolysaccharidosis type IV, Gaucher disease type I, Mucopolysaccharidosis type I, Type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bied syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens dystrophy Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic... Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent... Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction, Clabberg disease.Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0281] In the specific implementation scheme disclosed herein, the genes related to thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0282] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0283] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0284] The genes related to AATD lung disease include, but are not limited to, AATD;
[0285] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0286] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0287] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0288] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0289] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0290] The genes related to hemophilia B include, but are not limited to, factor IX;
[0291] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0292] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0293] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0294] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0295] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0296] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0297] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0298] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0299] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0300] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0301] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0302] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0303] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0304] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0305] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B.
[0306] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0307] The epilepsy-related genes include, but are not limited to, GAT1.
[0308] On the other hand, one technical solution provided in this disclosure is: a method for detecting, binding to, or cleaving target nucleic acids, the method comprising contacting the target nucleic acid with an IscB protein as described in this disclosure, a guide polynucleotide as described in this disclosure, an IscB inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, an IscB system as described in this disclosure, a vector system as described in this disclosure, a delivery system as described in this disclosure, a cell as described in this disclosure, a pharmaceutical composition as described in this disclosure, or a kit as described in this disclosure.
[0309] In a preferred embodiment disclosed herein, the method is a method for non-diagnostic and / or therapeutic purposes; and / or the fusion protein or conjugate contains a detectable marker, such as a marker detectable by fluorescence, DNA blotting, or FISH; and / or the method is in vivo or in vitro.
[0310] In a preferred embodiment disclosed herein, when the method involves cleaving the target nucleic acid, the method further includes performing the cleavage reaction using a cut buffer. The cut buffer can be any buffer known in the art suitable for cleaving the target nucleic acid by the IscB protein.
[0311] On the other hand, one technical solution provided in this disclosure is: a method for altering cell state, the method comprising contacting cells with an IscB protein as described in this disclosure, a guide polynucleotide as described in this disclosure, an IscB inactivating variant as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, an IscB system as described in this disclosure, a vector system as described in this disclosure, a delivery system as described in this disclosure, a cell as described in this disclosure, a pharmaceutical composition as described in this disclosure, or a kit as described in this disclosure, thereby altering the cell state.
[0312] In some embodiments disclosed herein, the method results in one or more of the following: increased or decreased expression of a specific gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, promotion and / or inhibition of cell growth in vitro or in vivo, induction of non-responsiveness in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo.
[0313] In a preferred embodiment of this disclosure, the method is a method for non-diagnostic and / or therapeutic purposes; and / or, the method is in vivo or in vitro.
[0314] On the other hand, one technical solution provided in this disclosure is: a method for diagnosing, treating, or preventing diseases or conditions related to target nucleic acids, comprising administering to a sample of a subject in need or to a subject in need an IscB protein as described in this disclosure, a guide polynucleotide as described in this disclosure, an inactivated variant of IscB as described in this disclosure, a fusion protein or conjugate as described in this disclosure, a nucleic acid as described in this disclosure, an IscB system as described in this disclosure, a vector system as described in this disclosure, a delivery system as described in this disclosure, a cell as described in this disclosure, a pharmaceutical composition as described in this disclosure, or a kit as described in this disclosure.
[0315] In the specific implementation scheme disclosed herein, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.
[0316] In the specific implementation scheme disclosed herein, the diseases or conditions mentioned are selected from: Hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, Hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, Mucopolysaccharidosis type IIIA, Mucopolysaccharidosis type IIIB, Gaucher disease type III, Mucopolysaccharidosis type II, Type II diabetes mellitus, Mucopolysaccharidosis type IV, Gaucher disease type I, Mucopolysaccharidosis type I, Type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bied syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens dystrophy Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic... Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent... Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction, Clabberg disease.Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0317] In the specific implementation scheme disclosed herein, the genes related to thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0318] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0319] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0320] The genes related to AATD lung disease include, but are not limited to, AATD;
[0321] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0322] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0323] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0324] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0325] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0326] The genes related to hemophilia B include, but are not limited to, factor IX;
[0327] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0328] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0329] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0330] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0331] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0332] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0333] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0334] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0335] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0336] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0337] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0338] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0339] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0340] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0341] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B.
[0342] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0343] The epilepsy-related genes include, but are not limited to, GAT1.
[0344] On the other hand, one technical solution provided in this disclosure is: the IscB protein as described in this disclosure, the guide polynucleotide as described in this disclosure, the IscB inactivating variant as described in this disclosure, the fusion protein or conjugate as described in this disclosure, the nucleic acid as described in this disclosure, the IscB system as described in this disclosure, the vector system as described in this disclosure, the delivery system as described in this disclosure, the cell as described in this disclosure, the pharmaceutical composition as described in this disclosure, or the kit as described in this disclosure, for the diagnosis, treatment, or prevention of diseases or conditions related to the target nucleic acid.
[0345] In the specific implementation scheme disclosed herein, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.
[0346] In the specific implementation scheme disclosed herein, the diseases or conditions mentioned are selected from: Hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, Hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, Mucopolysaccharidosis type IIIA, Mucopolysaccharidosis type IIIB, Gaucher disease type III, Mucopolysaccharidosis type II, Type II diabetes mellitus, Mucopolysaccharidosis type IV, Gaucher disease type I, Mucopolysaccharidosis type I, Type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bied syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens dystrophy Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic... Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent... Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction, Clabberg disease.Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0347] In the specific implementation scheme disclosed herein, the genes related to thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0348] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0349] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0350] The genes related to AATD lung disease include, but are not limited to, AATD;
[0351] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0352] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0353] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0354] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0355] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0356] The genes related to hemophilia B include, but are not limited to, factor IX;
[0357] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0358] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0359] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0360] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0361] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0362] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0363] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0364] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0365] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0366] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0367] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0368] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0369] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0370] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0371] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B.
[0372] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0373] The epilepsy-related genes include, but are not limited to, GAT1.
[0374] Based on common knowledge in the field, the above-mentioned preferred conditions can be combined arbitrarily to obtain the preferred examples disclosed herein. Attached Figure Description
[0375] Figure 1 shows the TAM motif determined by the CTnB-57 plasmid elimination experiment.
[0376] Figure 2 shows the TAM motif determined by the CTnB-62 plasmid elimination experiment.
[0377] Figure 3 shows the bacterial growth status of TAM-02, TAM-05, TAM-06, and TAM-07 edited by CTnB-57 respectively.
[0378] Figure 4 shows the bacterial growth status of TAM-02, TAM-05, TAM-06, and TAM-07 edited by CTnB-62.
[0379] Figure 5 shows the TAM motif captured by the CTnB-30 plasmid for removal.
[0380] Figure 6 shows the TAM motif captured by the CTnB-50 plasmid for removal.
[0381] Figure 7 shows the bacterial growth status of TAM-05, TAM-06, and TAM-07 edited by CTnB-30.
[0382] Figure 8 shows the bacterial growth on CTnB-30 edited TAM-06 plates with different resistances.
[0383] Figure 9 shows the bacterial growth on plates with different resistance levels on different TAM plasmids edited by CTnB-50.
[0384] Figure 10 shows the editing efficiency of CTnB-30-TTR06.
[0385] Figure 11 shows the editing efficiency of CTnB-30-TTR08.
[0386] Figure 12 shows other TAM motifs identified by IscB.
[0387] Figure 13 shows the TAM motif captured by the CTnB-152 plasmid after removal.
[0388] Figure 14 shows the TAM motif captured by the CTnB-198 plasmid removal process. Detailed Implementation
[0389] In this disclosure, unless otherwise stated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the procedures used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all standard procedures widely used in their respective fields. To better understand this disclosure, definitions and explanations of relevant terms are provided below.
[0390] In this disclosure, "more" means ≥2; "multiple" means two or more.
[0391] In this disclosure, depending on the context, “cutting” may refer to cutting the backbone of a polynucleotide chain; non-limiting examples include cutting a single-stranded DNA completely, cutting one of the single strands of a double-stranded DNA, or cutting both single strands of a double-stranded DNA.
[0392] In this disclosure, depending on the context, "modification" can refer to any form of nucleic acid chain chemical reaction other than "cleavage"; including but not limited to base substitution, addition and / or deletion, as well as methylation and demethylation of nucleic acid chains. Non-limiting examples include base substitution on the target nucleic acid chain through single-base editing (fusion of the IscB disclosed herein with a deaminase domain and in conjunction with gRNA), such as A→G, C→T, T→C, or G→A nucleotide mutations, and other types of nucleotide mutations (e.g., A→T, C→G, T→A, G→C, etc.); base substitution, addition, or deletion can also be achieved through Prime editing technology (fusion of the IscB disclosed herein with a reverse transcriptase and in conjunction with PegRNA), or through HDR homologous recombination (e.g., the IscB disclosed herein in conjunction with gRNA and a donor template); the IscB disclosed herein can also be fused with DNA methyltransferases or DNA demethyltransferases and combined with gRNA for targeted regulation of the methylation level of the target nucleic acid.
[0393] In this disclosure, depending on the context, "regulating the expression of target nucleic acids" can refer to the regulation of the transcription of target nucleic acids; non-limiting examples include enhancing or inhibiting the transcription of target nucleic acids by means of IscB fusion transcriptional activation or repression domains through CRISPRa, CRISPRi technologies.
[0394] In this disclosure, the letters in the amino acid sequence represent single-letter abbreviations of amino acids known in the art, such as those described in J. Biol. Chem, 243, p3558 (1968): Alanine: Ala-A, Arginine: Arg-R, Aspartic acid: Asp-D, Cysteine: Cys-C, Glutamine: Gln-Q, Glutamic acid: Glu-E, Histidine: His-H, Glycine: Gly-G, Asparagine: Asn-N, Tyrosine: Tyr-Y, Proline: Pro-P, Serine: Ser-S, Methionine: Met-M, Lysine: Lys-K, Valine: Val-V, Isoleucine: Ile-I, Phenylalanine: Phe-F, Leucine: Leu-L, Tryptophan: Trp-W, Threonine: Thr-T.
[0395] In this disclosure, "amino acid difference" refers to differences in amino acid residues at specific sites on the amino acid sequence of a protein, including substitutions, additions, or deletions.
[0396] As is known to those skilled in the art, in proteins or peptides, two adjacent amino acids each lose an OH or H atom and undergo dehydration condensation to form a peptide bond. Each amino acid actually exists in the form of an amino acid residue. Therefore, in this disclosure, the terms "amino acid" and "amino acid residue" generally refer to the same thing.
[0397] In this disclosure, if an amino acid is substituted, it means that it is substituted by another amino acid residue that is different from the original amino acid residue. If the original amino acid was originally a positively charged amino acid, and it is substituted to a positively charged amino acid, it means that it is substituted by another positively charged amino acid residue that is different from the original amino acid residue. For example, if the original amino acid residue is R, and it is substituted to a positively charged amino acid, it means that it is substituted to H or K.
[0398] In this article, when referring to RNA sequences, the letter "T" can be used interchangeably with "U". When referring to "guide sequences", the letter "T" can be used interchangeably with "U". When referring to "ωRNA", the letter "T" can be used interchangeably with "U". When referring to "backbone sequences", the letter "T" can be used interchangeably with "U". When referring to "tracrRNA sequences", the letter "T" can be used interchangeably with "U".
[0399] Sequence identity
[0400] As used herein, the term "identity" (or "percent identity") refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are identical at a position when the same base or amino acid monomer subunit is occupied (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine). The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared, multiplied by 100%. For example, if six out of ten positions in two sequences match, then the two sequences have 60% sequence identity. Typically, two sequences are compared to produce the maximum sequence identity. Such alignments can be made using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, CLUSTERalΩ, MAFFT, Probcons, T-Coffee, Probalign, and BLAST, which can be reasonably chosen by those skilled in the art. Those skilled in the art can determine suitable parameters for the alignment of sequences, including any algorithm required to achieve a better or better alignment of the entire length of the sequences being compared, and any algorithm required to achieve a better or better alignment of a local portion of the sequences being compared.
[0401] IscB system
[0402] As used herein, the IscB system typically comprises an IscB protein sequence or its encoded nucleic acid, and a guide polynucleotide or its encoded nucleic acid.
[0403] In some embodiments, the IscB protein described herein refers to a protein whose amino acid sequence includes or has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% sequence identity with any of SEQ ID NO: 13-26, 41, 42, 53-93.
[0404] In this disclosure, the IscB system comprises an IscB protein having at least 50% sequence identity with any one of SEQ ID NO: 13-26, 41, 42, 53-93, or a nucleic acid encoding the IscB protein; and a guiding polynucleotide or a nucleic acid encoding the guiding polynucleotide; the guiding polynucleotide comprising a backbone sequence linked to a guiding sequence, the guiding sequence being engineered to hybridize with a target nucleic acid, the guiding polynucleotide being capable of forming a complex with the IscB protein and guiding the complex to bind sequence-specifically to the target nucleic acid.
[0405] Guided polynucleotides
[0406] As used herein, the term "guide polynucleotide" refers to a molecule in the CRISPR-Cas system that forms a complex with the IscB protein and guides the complex to a target sequence. Typically, a guide polynucleotide contains a backbone sequence linked to a guide sequence that can hybridize with the target sequence. The backbone sequence typically contains an ωRNA backbone sequence and sometimes also contains a tracrRNA sequence. In some embodiments, the guide polynucleotide does not contain a tracrRNA sequence. In some embodiments, the guide polynucleotide contains a tracrRNA sequence.
[0407] In some embodiments, the guide polynucleotide of the IscB system is guide RNA. In some embodiments, the guide polynucleotide is a chemically modified guide polynucleotide. In some embodiments, the guide polynucleotide comprises at least one chemically modified nucleotide.
[0408] In some embodiments, the guiding polynucleotide comprises at least one guide sequence linked to at least one backbone sequence. In some embodiments, the guide sequence is located at the 3' end of the backbone sequence. In some embodiments, the guide sequence is located at the 5' end of the backbone sequence.
[0409] In some implementations, the tracrRNA sequence is linked to the backbone sequence.
[0410] In some embodiments, the tracrRNA sequence is located at the 5' or 3' end of the backbone sequence. In some embodiments, the tracrRNA sequence is located at the 5' end of the backbone sequence. In some embodiments, the tracrRNA sequence is located at the 3' end of the backbone sequence.
[0411] In some implementations, the nucleotide sequence of the guiding polynucleotide comprises, from 5' to 3', the following sequence: tracrRNA, backbone sequence, and guiding sequence.
[0412] In some implementations, the nucleotide sequence of the guiding polynucleotide comprises, from 5' to 3', the following sequence: tracrRNA, a linker sequence, a backbone sequence, and a guiding sequence.
[0413] In some implementations, the nucleotide sequence of the guiding polynucleotide comprises, from 5' to 3', the following sequence: tracrRNA, loop sequence, backbone sequence, and guiding sequence.
[0414] In some implementations, the structure of the guiding polynucleotide is 5'-tracrRNA-loop-backbone sequence-guiding sequence-3'.
[0415] In some implementations, the guide polynucleotide tracrRNA and the backbone sequence are linked by a nucleotide sequence.
[0416] In specific embodiments disclosed herein, the tracrRNA sequence is linked to the backbone sequence via a nucleotide sequence consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence is linked to the backbone sequence via a nucleotide sequence consisting of 4 nucleotides. In specific embodiments disclosed herein, the tracrRNA sequence is linked to the backbone sequence via a 5'-GAAA-3' sequence.
[0417] In some embodiments, the guide sequence is sufficiently complementary to the target nucleic acid sequence to hybridize with the target nucleic acid and guide the IscB complex to bind sequence-specifically to the target nucleic acid. In some embodiments, the guide sequence is 100% complementary to the target nucleic acid, but the guide sequence may have less than 100% complementarity with the target nucleic acid, for example, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or at least 99% complementarity.
[0418] In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with a mismatch of no more than two nucleotides. In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with a mismatch of no more than one nucleotide. In some embodiments, the guide sequence is engineered to hybridize with the target nucleic acid, with or without a mismatch.
[0419] In some embodiments, the IscB system comprises at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different guide polynucleotides. In some embodiments, the guide polynucleotides target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different target nucleic acid molecules, or target at least 2, at least 3, at least 4, at least 5, at least 10, or at least 20 different regions of one or more target nucleic acid molecules.
[0420] In some embodiments, the guiding polynucleotide includes a constant backbone sequence located upstream of a variable guiding sequence. In some embodiments, multiple guiding polynucleotides are part of an array (which may be part of a vector, such as a viral vector or plasmid). For example, a guiding array comprising a sequence backbone-guiding sequence-backbone-guiding sequence-backbone-guiding sequence-...-backbone-guiding sequence may include multiple unique, unprocessed guiding polynucleotides. This allows for multiplexing, such as delivering multiple guiding polynucleotides to a cell or system to target multiple target nucleic acids or multiple regions within a single target nucleic acid.
[0421] The ability of a guide polynucleotide (SNP)-guided complex to bind sequence-specifically to a target nucleic acid can be assessed by any suitable assay. For example, components of an IscB system sufficient to form the complex, including the guide polynucleotide to be tested, can be provided to a host cell containing the corresponding target nucleic acid molecule, for example, by transfection with a vector encoding a component of the complex, and then preferential cleavage within the target sequence can be assessed. Similarly, cleavage of the target nucleic acid sequence can be assessed in vitro by providing the target nucleic acid, components of the complex, including the guide polynucleotide to be tested and a control guide polynucleotide different from the test guide polynucleotide, and comparing the ability of the test and control guide polynucleotides to bind to the target nucleic acid or the rate of cleavage of the target nucleic acid. The ability of the complex to cleave the target nucleic acid or the target nucleic acid itself can also be assessed by the assays described above.
[0422] IscB mutant
[0423] As described herein, when referring to "the corresponding position of the sequence shown in SEQ ID NO:XX" or using similar wording, the corresponding position is determined by amino acid sequence alignment. Typically, two sequences are compared to produce maximum sequence identity. Such alignments can be performed using publicly available and commercially available alignment algorithms and programs, such as, but not limited to, Clustal OMEGA, MAFFT, Probcons, T-Coffee, Probalign, and BLAST, which can be reasonably chosen by those skilled in the art. Those skilled in the art can determine suitable parameters for aligning sequences, including, for example, any algorithm required to achieve a better or optimal alignment of the full length of the compared sequences, and any algorithm required to achieve a better or optimal alignment of a portion of the compared sequences.
[0424] In some embodiments, the IscB protein provided herein contains one or more mutations compared to the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93, such as a single amino acid insertion, a single amino acid deletion, a single amino acid substitution, or a combination thereof. In some examples, the IscB protein, compared to the IscB protein represented by any of the sequences in SEQ ID NO:13-26, 41, 42, 53-93, contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 3 3, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 7 8, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.In some instances, with SEQ ID The IscB protein represented by any of the sequences NO:13-26, 41, 42, 53-93 is compared to the IscB protein containing 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, or 130 amino acid variations (e.g., insertion, deletion, or substitution), but retaining the ability to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide.
[0425] One type of modification or mutation involves replacing amino acid residues with similar biochemical properties; this is known as conserved substitution. Typically, conserved substitutions have little or no effect on the activity of the resulting protein or peptide. For example, a conserved substitution is an amino acid substitution in the IscB protein that essentially does not affect the binding of the IscB protein to target nucleic acid molecules complementary to the guide sequence of the gRNA molecule, and / or the processing of the guide array RNA transcript into gRNA molecules.
[0426] More substantial alterations can be achieved by using less conserved substitutions, for example, by selecting residues that differ more significantly in maintaining the following effects: (a) the structure of the polypeptide backbone in the region where the substitution occurs, for example, as a helical or folded conformation; (b) the charge or hydrophobicity of the region interacting with the target site; or (c) the volume of the side chain. Substitutions that are generally expected to produce the greatest changes in polypeptide function are (a) substitutions between hydrophilic residues (e.g., serine or threonine) and hydrophobic residues (e.g., leucine, isoleucine, phenylalanine, valine, or alanine); (b) substitutions between cysteine or proline and any other residue; (c) substitutions between residues with positively charged side chains (e.g., lysine, arginine, or histidine) and negatively charged residues (e.g., glutamic acid or aspartic acid); or (d) substitutions between residues with large side chains (e.g., phenylalanine) and residues without side chains (e.g., glycine).
[0427] IscB active fragment
[0428] In this disclosure, the IscB protein may contain only the WED-I domain, Helical-I1 domain, PI domain, Helical-I2 domain, Helical-II domain, WED-II domain, Ruvc-I domain, Helical-III domain, BH domain, Ruvc-II domain, Nuc domain and / or Ruvc-III domain.
[0429] The IscB protein disclosed herein may include, in addition to the aforementioned domain, other domains of IscB proteins in the prior art, which together form the complete structure of the IscB protein to achieve the functions of the IscB protein disclosed herein, including but not limited to retaining the ability of the IscB protein to form a complex with gRNA, retaining the ability of the IscB protein to form a complex with gRNA and target target nucleic acids, retaining the ability of the complex formed by the IscB protein and gRNA to target and regulate the expression of target nucleic acids, retaining the ability of the complex formed by the IscB protein and gRNA to target and cleave single-stranded or double-stranded target nucleic acids, retaining the ability of the IscB protein to bind to target nucleic acid molecules complementary to the guide sequence of the guide polynucleotide, and / or retaining the ability to process RNA transcripts containing the guide sequence into guide polynucleotide molecules.
[0430] IscB inactivation variant
[0431] By inactivating the RuvC domain of IscB through point mutation, the IscB protein loses its endonuclease activity, and the resulting dIscB (dead IscB) can only bind to target genes under the mediation of guide polynucleotides, but does not have the function of cutting DNA.
[0432] Alternatively, point mutations can be used to partially deactivate the RuvC domain of IscB, forming IscB nickase (nIscB). This nickase binds to the target gene under the guidance of a guide polynucleotide, cleaving one single strand of the double-stranded nucleic acid without cleaving the other single strand.
[0433] Therefore, dIscB or nIscB can be fused or conjugated with other domains (including but not limited to deaminase domains, transcription activation domains, transcription repression domains, methylation domains, demethylation domains, histone acetylation domains, histone deacetylation domains, and histone domains) to guide polynucleotides to the target sequence of the target nucleic acid, and then perform the corresponding functions with the help of the other domains; for example, C→T conversion of the target nucleic acid bases can be achieved by deamination of cytosine bases, A→G conversion of the target nucleic acid bases can be achieved by deamination of adenine bases, transcriptional repression of the target nucleic acid can be achieved by the transcription repression domain, and transcriptional activation domain can promote the transcription of the target nucleic acid.
[0434] Functional structural domain
[0435] In some embodiments, the IscB protein or IscB inactivating variant is covalently linked or fused with homologous or heterologous functional domains.
[0436] In some embodiments, the functional domain has enzymatic activity that modifies the target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristylation activity, and / or demyristylation activity.
[0437] In some embodiments, the functional domain is optionally selected from one or more of the following: nucleases (e.g., FokI), methyltransferases, demethylases, DNA repair enzymes, DNA damage enzymes, deaminases, superoxide dismutases, alkylating enzymes, depurinases, oxidases, pyrimidine dimer forming enzymes, integrases, transposases, recombinases, polymerases, ligases, helicases, photolyases, glycosylation enzymes, deglycosylation enzymes, acetyltransferases, deacetylases, kinases, phosphatases, ubiquitin ligases, deubiquitinating enzymes, adenylate acylases, deadenylate acylases, SUMOylating enzymes, deSUMOylating enzymes, myristylases, and / or demyristylases.
[0438] In some embodiments, the functional domains are selected from one, two, three, four or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, epitope tag and / or reporter domain.
[0439] In some embodiments disclosed herein, the deaminase domain may be selected from: APOBEC1, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3D, APOBEC3F, activation-induced cytidine deaminase (AID), CDA from lamprey, or a mutant of adenosine deaminase (TadA) engineered to act on DNA.
[0440] In some embodiments, the transcriptional activation domain is optionally selected from: p65, VPR, VP16, VP64, VTR1, VTR2, VTR3, p65, MyoD1, HSF1, RTA, SET7 / 9, and histone acetyltransferase. In some embodiments, the transcriptional activation domain is optionally selected from: sequence fragments from p53 TAD1, p53 TAD2, MLL, E2A, Rtg3, CREB, CREBaB6, Gli3, Gal4, Oaf1, Pip2, Pdr1, and Pdr3.
[0441] In some embodiments, the transcriptional repressor domain is optionally selected from: KOX1, KAP-1, MAD, FKHR, EGR-1, ERD, SID, a tandem of SID (e.g., SID4X), TIEG, v-ERB-A, MBD2, MBD3, TRa, histone methyltransferase, histone deacetylase (HDAC), nuclear hormone receptor (e.g., estrogen receptor or thyroid hormone receptor), DNMT family members (e.g., DNMT1, DNMT3A, DNMT3B), the KRAB domain of MeCP2, ROM2, and AtHD2A.
[0442] In some embodiments, the transcriptional repressor domain is the KRAB domain from the KOX1 protein.
[0443] In some embodiments, the nuclease domain is optionally selected from FokI, a polypeptide with ssDNA cleavage activity, or a polypeptide with dsDNA cleavage activity.
[0444] In some embodiments, the methyltransferase domain is selected from DNA methyltransferases, including but not limited to DNMT1, DNMT3a, and DNMT3b.
[0445] In some embodiments, the demethylase is selected from TET1CD, TET1, ROS1, DME, DML2, and DML3.
[0446] Methylation and demethylation are widely recognized in the field as important mechanisms of epigenetic gene regulation.
[0447] In some embodiments, the homologous or heterologous functional domains are sequence tags useful for the dissolution, purification, or detection of the fusion protein or conjugate. Suitable protein tag sequences are provided herein, including but not limited to biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, multihistidine tags (also known as His tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S-tags, Softag (e.g., Softag 1, Softag 3), strep-tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Other suitable sequences will be readily apparent to those skilled in the art.
[0448] Subcellular localization signals
[0449] In some embodiments, the IscB protein is fused with at least one homologous or heterologous subcellular localization signal. Exemplary subcellular localization signals include organelle localization signals, such as nuclear localization signals (NLS), nuclear export signals (NES), or mitochondrial localization signals.
[0450] Non-limiting examples of NLS include NLS sequences derived from: NLS of the SV40 virus large T antigen; NLS from nucleoplasmic proteins; c-myc NLS; hRNPA1 M9 NLS; sequence fragments from the IBB domain; sequence fragments of the fibroid T protein; sequence fragments of human p53; sequence fragments of mouse c-ablIV; sequence fragments of influenza virus NS1; sequence fragments of hepatitis virus delta antigen; sequence fragments of mouse Mx1 protein; sequence fragments of human PARP enzyme; and sequence fragments of steroid hormone receptors. In some embodiments, the nuclear localization sequence has sufficient strength to drive the fusion protein or conjugate described herein to accumulate in a detectable amount in the nucleus of eukaryotic cells. Generally, the strength of nuclear localization activity can be derived from the number of NLSs, one or more specific NLSs used, or a combination of these factors. Accumulation in the nucleus can be detected by any suitable technique. For example, detectable markers can be fused to IscB proteins to visualize their intracellular location, such as in combination with means for detecting nuclear location (e.g., nuclear-specific dyes such as DAPI). The nucleus can also be isolated from the cell and its contents analyzed using any suitable method for protein detection, such as immunohistochemistry, Western blotting, or enzyme activity assays. Accumulation in the nucleus can also be determined indirectly, such as by measuring the role of nucleic acid targeting complex formation (e.g., measuring DNA or RNA cleavage or mutation at the target sequence, or measuring altered gene expression activity due to the effects of DNA or RNA targeting complex formation and / or DNA or RNA targeting IscB protein activity), compared to a control not exposed to nucleic acid targeting IscB proteins or nucleic acid targeting complexes, or exposed to nucleic acid targeting IscB proteins lacking one or more NLSs.
[0451] Fusion proteins, epigeneity editors, base editors, deaminases
[0452] In some embodiments disclosed herein, the fusion protein is an epigenetic editor fusion protein. In some embodiments disclosed herein, the epigenetic editor comprises an epigenetic editor fusion protein and a guiding polynucleotide.
[0453] In some embodiments disclosed herein, the fusion protein is a base editor fusion protein. In some embodiments disclosed herein, the base editor comprises a base editor fusion protein and a guide polynucleotide.
[0454] In some embodiments disclosed herein, the fusion protein includes a repressor domain.
[0455] In some embodiments disclosed herein, the epigenetic editor fusion protein includes a repressor domain.
[0456] In some embodiments disclosed herein, the homologous or heterologous functional domains are DNA methylation domains, histone methylation domains, and / or repressor domains.
[0457] In some embodiments disclosed herein, the repressor domain specifically binds to an epigenetic effector protein in a cell containing a target gene and directs an epigenetic editor to the target gene to achieve epigenetic modification of nucleotides in the target gene or to achieve epigenetic modification of histones bound to the target gene.
[0458] Patent application CN117136235A describes an apparent edit, the full text of which is incorporated herein by reference.
[0459] In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a DNA methylation domain and / or a histone methylation domain. In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a DNA methylation domain, a histone methylation domain, and / or a repressor domain. In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a DNA methylation domain and / or a repressor domain. In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a histone methylation domain and / or a repressor domain. In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a DNA methylation domain, a histone methylation domain, and a repressor domain. In some embodiments disclosed herein, the epigenetic editor fusion protein comprises a fusion protein, which comprises a DNA methylation domain, a histone methylation domain, and / or a repressor domain.
[0460] In some embodiments disclosed herein, the epigenetic editor fusion protein includes a DNA methylation domain. In some embodiments disclosed herein, the epigenetic editor fusion protein includes DNMT3A, DNMT3B, and / or DNMT3L domains. In some embodiments disclosed herein, the epigenetic editor fusion protein includes a DNMT3A domain. In some embodiments disclosed herein, the epigenetic editor fusion protein includes a DNMT3B domain. In some embodiments disclosed herein, the epigenetic editor fusion protein includes a DNMT3L domain.
[0461] In some embodiments disclosed herein, the epigenetic editor fusion protein includes a histone methylation domain.
[0462] In some embodiments disclosed herein, the fusion protein comprises: the IscB protein as described herein, the IscB inactivating variant as described herein, or its DNA-binding domain.
[0463] In some embodiments disclosed herein, the fusion protein comprises: an IscB protein as described herein, or an inactivating variant of IscB as described herein; and a repressor domain.
[0464] In some embodiments of this disclosure, the epigeneity editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and one, two, three, or four repressor domains. In some embodiments of this disclosure, the epigeneity editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least one repressor domain. In some embodiments of this disclosure, the epigeneity editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least two repressor domains. In some embodiments of this disclosure, the epigeneity editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least three repressor domains. In some embodiments of this disclosure, the epigeneity editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and one repressor domain. In some embodiments of this disclosure, the epigenetic editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and two repressor domains. In some embodiments of this disclosure, the epigenetic editor comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and three repressor domains.
[0465] In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and one, two, three, or four repressor domains. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least one repressor domain. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least two repressor domains. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and at least three repressor domains. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and one repressor domain. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and two repressor domains. In some embodiments of this disclosure, the fusion protein comprises: an IscB protein as described in this disclosure or an inactivating variant of IscB as described in this disclosure, and three repressor domains.
[0466] In some embodiments disclosed herein, one of the repressor domains is a KRAB domain. In some embodiments disclosed herein, the KRAB domain is a KOX1 KRAB domain.
[0467] In some embodiments disclosed herein, the repressor structural domain is selected from at least one of the following: ZIM3, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZNF680, ZNF41, ZNF189, ZNF528, ZNF543, ZNF554, ZNF140, ZNF610, ZNF264, ZNF350, ZNF8, ZNF582, ZNF30, ZNF324, ZNF98, ZNF669, ZNF677, ZNF596, ZNF214, ZNF37A, ZNF34, ZNF250, ZNF547, ZNF273, ZNF354A, Z FP82, ZNF224, ZNF33A, ZNF45, ZNF175, ZNF595, ZNF184, ZNF419, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF566, ZNF729, ZIM2, ZNF254, ZNF764, ZNF785, ZNF10, CBX5, RYBP, YAF2, MGA, CBX1, SCMH1, MPP8, SUMO3, HERC2, BIN1, PCGF2, TOX, FOXA1, FOXA2, IRF2BP1, IRF2BP2, IRF2BPL IRF-2BP1_2N-terminal domain, HOXA13, HOXB13, HOXC13, HOXA11, HOXC11, HOXC10, HOXA10, HOXB9, HOXA9, ZFP28, ZN334, ZN568, ZN37A, ZN181, ZN510, ZN862, ZN140, ZN208, ZN248, ZN571, ZN699, ZN726, ZIK1, ZNF2, Z705F, ZNF14, ZN 471, ZN624, ZNF84, ZNF7, ZN891, ZN337, Z705G, ZN529, ZN729, ZN419, Z705 A. ZNF45, ZN302, ZN486, ZN621, ZN688, ZN33A, ZN554, ZN878, ZN772, ZN22 4. ZN184, ZN544, ZNF57, ZN283, ZN549, ZN211, ZN615, ZN253, ZN226, ZN730 , Z585A, ZN732, ZN681, ZN667, ZN649, ZN470, ZN484, ZN431, ZN382, ZN254 , ZN124, ZN607, ZN317, ZN620, ZN141, ZN584, ZN540, ZN75D, ZN555, ZN658,ZN684、RBAK、ZN829、ZN582、ZN112、ZN716、HKR1、ZN350、ZN480、ZN416、ZNF92、ZN100、ZN736、ZNF74、CBX1、ZN443、ZN195、ZN530、ZN782、ZN791、ZN331、Z354C、ZN157、ZN727、ZN550、ZN793、ZN235、ZNF8、ZN724、ZN573、ZN577、ZN789、ZN718、ZN300、ZN383、ZN429、ZN677、ZN850、ZN454、ZN257、ZN264、ZFP82、ZFP14、ZN485、ZN737、ZNF44、ZN596、ZN565、ZN543、ZFP69、SUMO1、ZNF12、ZN169、ZN433、SUMO3、ZNF98、ZN175、ZN347、ZNF25、ZN519、Z585B、ZIM3、ZN517、ZN846、ZN230、ZNF66、ZFP1、ZN713、ZN816、ZN426、ZN674、ZN627、ZNF20、Z587B、ZN316、ZN233、ZN611、ZN556、ZN234、ZN560、ZNF77、ZN682、ZN614、ZN785、ZN445、ZFP30、ZN225、ZN551、ZN610、ZN528、ZN284、ZN418、MPP8、ZN490、ZN805、Z780B、ZN763、ZN285、ZNF85、ZN223、ZNF90、ZN557、ZN425、ZN229、ZN606、ZN155、ZN222、ZN442、ZNF91、ZN135、ZN778、RYBP、ZN534、ZN586、ZN567、ZN440、ZN583、ZN441、ZNF43、CBX5、ZN589、ZNF10、ZN563、ZN561、ZN136、ZN630、ZN527、ZN333、Z324B、ZN786、ZN709、ZN792、ZN599、ZN613、ZF69B、ZN799、ZN569、ZN564、ZN546、ZFP92、YAF2、ZN723、ZNF34、ZN439、ZFP57、ZNF19、ZN404、ZN274、CBX3、ZNF30、ZN250、ZN570、ZN675、ZN695、ZN548、ZN132、ZN738、ZN420、ZN626、ZN559、ZN460、ZN268、ZN304、ZIM2、ZN605、ZN844、SUMO5、ZN101、ZN783、ZN417、ZN182、ZN823、ZN177、ZN197、ZN717、ZN669、ZN256、ZN251、CBX4、PCGF2、CDY2、CDYL2、HERC2、ZN562、ZN461、Z324A、ZN766、ID2、TOX、ZN274、SCMH1、ZN214、CBX7、ID1、CREM、SCX、ASCL1、ZN764、SCML2、TWST1、CREB1、TERF1、ID3、CBX8、CBX4、GSX1、NKX22、ATF1、TWST2、ZNF17、TOX3、TOX4、ZMYM3、I2BP1、RHXF1、SSX2、I2BPL、ZN680、CBX1、TRI68、HXA13、PHC3、TCF24、CBX3、HXB13、HEY1、PHC2、ZNF81、FIGLA、SAM11、KMT2B、HEY2、JDP2、HXC13、ASCL4、HHEX、HERC2、GSX2、BIN1、ETV7、ASCL3、PHC1、OTP、I2BP2、VGLL2、HXA11、PDLI4、ASCL2、CDX4、ZN860、LMBL4、PDIP3、NKX25、CEBPB、ISL1、CDX2、PROP1、SIN3B、SMBT1、HXC11、HXC10、PRS6A、VSX1、NKX23、MTG16、HMX3、HMX1、KIF22、CSTF2、CEBPE、DLX2、ZMYM3、PPARG、PRIC1、UNC4、BARX2、ALX3、TCF15、TERA、VSX2、HXD12、CDX1、TCF23、ALX1、HXA10、RX、CXXC5、SCML1、NFIL3、DLX6、MTG8、CBX8、CEBPD、SEC13、FIP1、ALX4、LHX3、PRIC2、MAGI3、NELL1、PRRX1、MTG8R、RAX2、DLX3、DLX1、NKX26、NAB1、SAMD7、PITX3、WDR5、MEOX2、NAB2、DHX8、FOXA2、CBX6、EMX2、CPSF6、HXC12、KDM4B、LMBL3、PHX2A、EMX1、NC2B、DLX4、SRY、ZN777、NELL1、ZN398、GATA3、BSH、SF3B4、TEAD1、TEAD3、RGAP1、PHF1、FOXA1、GATA2、FOXO3、ZN212、IRX4、ZBED6、LHX4、SIN3A、RBBP7、NKX61、TRI68、R51A1、MB3L1, DLX5, NOTC1, TERF2, ZN282, RGS12, ZN840, SPI2B, PAX7, NKX62, ASXL2, FOXO1, GATA3, GATA1, ZMYM5, ZN783, SPI2B, LRP1, MIXL1, SGT1, LMCD1, CEBPA, GATA2, SOX14, WTIP, PRP19, CBX6, NKX11, RBBP4, DMRT2, SMCA2, and fragments thereof. In some embodiments disclosed herein, at least one of the repressor domains is selected from: ZIM3, ZNF264, ZN577, ZN793, ZFP28, ZN627, RYBP, TOX, TOX3, TOX4, I2BP1, SCMH1, SCML2, CDYL2, CBX8, CBX5, and CBX1, and fragments thereof.
[0468] In some embodiments of this disclosure, the fusion protein comprises, from its N-terminus to its C-terminus: DNMT3A-DNMT3L-IscB-KOX1KRAB-second repressor domain. In some embodiments of this disclosure, a linker connects the domains of the fusion protein. In some embodiments of this disclosure, the linker is an XTEN linker. In some embodiments of this disclosure, the XTEN linker is selected from XTEN-16, XTEN-18, and XTEN-80. In some embodiments of this disclosure, the fusion protein comprises, from its N-terminus to its C-terminus: DNMT3A-DNMT3L-XTEN80-IscB-XTEN16-KOX1 KRAB-XTEN18-second repressor domain.
[0469] In some embodiments of this disclosure, the epigenesis editor comprises a guiding polynucleotide. In some embodiments of this disclosure, the epigenesis editor comprises a guiding polynucleotide as described in any one of these disclosures.
[0470] In some embodiments of this disclosure, the guiding polynucleotide can form a complex with the epigenetic editor. In some embodiments of this disclosure, the guiding polynucleotide can form a complex with the epigenetic editor and guide the complex sequence to specifically bind to the target nucleic acid.
[0471] In some embodiments of this disclosure, the epigenetic editor comprises a fusion protein, wherein the fusion protein comprises (a) an IscB protein as described in this disclosure, or an IscB inactivating variant as described in this disclosure; (b) a repressor domain; (c) a first catalytic domain selected from the DNMT3A and DNMT3L catalytic domains; and (d) a second catalytic domain selected from the DNMT3A and DNMT3L catalytic domains. In some embodiments of this disclosure, the first catalytic domain has fewer than 380 amino acids, or the second catalytic domain has fewer than 380 amino acids.
[0472] In some embodiments disclosed herein, the fusion protein comprises a DNMT domain (a first DNMT domain).
[0473] In some embodiments disclosed herein, the fusion protein further comprises a second DNMT domain. In some embodiments, the first DNMT domain is selected from the DNMT3A, DNMT3B, DNMT3C, and DNMT3L domains. In some embodiments, the first DNMT domain is the DNMT3A domain. In some embodiments, the first DNMT domain is the DNMT3L domain. In some embodiments, the first DNMT domain is the human DNMT domain. In some embodiments, the human DNMT domain is the human DNMT3A domain. In some embodiments, the human DNMT domain is the human DNMT3L domain. In some embodiments, the first DNMT domain is the mouse DNMT domain. In some embodiments, the mouse DNMT domain is the mouse DNMT3A domain. In some embodiments, the mouse DNMT domain is the mouse DNMT3L domain. In some embodiments, the first DNMT domain is the DNMT3A domain, and the second DNMT domain is the DNMT3L domain. In some embodiments, the first DNMT domain is the human DNMT3A domain, and the second DNMT domain is the human DNMT3L domain. In some embodiments, the first DNMT domain is the human DNMT3A domain, and the second DNMT domain is the mouse DNMT3L domain. In some embodiments, the first DNMT domain is the mouse DNMT3A domain, and the second DNMT domain is the human DNMT3L domain. In some embodiments, the first DNMT domain is the mouse DNMT3A domain, and the second DNMT domain is the mouse DNMT3L domain.
[0474] In some embodiments, the first DNMT domain is the catalytic portion of the DNMT domain. In some embodiments, the second DNMT domain is the catalytic portion of the DNMT domain.
[0475] In some embodiments, the epigenetic editor comprises a fusion protein containing an effector domain. In some embodiments, the effector domain comprises a histone methyltransferase domain. In some embodiments, the effector domain includes a DOT1L domain, a SET domain, an SUV39H1 domain, a G9a / EHMT2 protein domain, an EZH1 domain, an EZH2 domain, a SETDB1 domain, or any combination thereof. In some embodiments, the effector domain comprises a histone-lysine-N-methyltransferase SETDB1 domain.
[0476] In some embodiments, the effector domain comprises a DNA methyltransferase domain or a histone methyltransferase domain. The DNA methyltransferase domain can mediate methylation at a DNA nucleotide (e.g., at any one of the A, T, G, or C nucleotides). In some embodiments, the methylated nucleotide is N6-methyladenosine (m6A). In some embodiments, the methylated nucleotide is 5-methylcytosine (5MC). In some embodiments, methylation occurs at a CG (or CpG) dinucleotide sequence. In some embodiments, methylation occurs at a CHG or CHH sequence, where H is any one of A, T, or C.
[0477] In some embodiments, the effector domain comprises a DNA methyltransferase (DNMT) domain that catalyzes the transfer of methyl groups to cytosine, thereby repressing the expression of a target gene by recruiting repressive regulatory proteins. In some embodiments, the effector domain comprises a DNA methyltransferase (DNMT) family protein domain. In some embodiments, the effector domain comprises a DNMT1 domain. In some embodiments, the effector domain comprises a TRDMT1 domain. In some embodiments, the effector domain comprises a DNMT3 domain. In some embodiments, the effector domain comprises a DNMT3A domain. In some embodiments, the effector domain comprises a DNMT3B domain. In some embodiments, the effector domain comprises a DNMT3C domain. In some embodiments, the effector domain comprises a DNMT3L domain. In some embodiments, the effector domain comprises a fusion of the DNMT3A and DNMT3L domains.
[0478] In some embodiments, the effector domain recruits one or more protein domains that repress the expression of target genes. In some embodiments, the effector domain interacts with a scaffold protein domain that recruits one or more protein domains that repress the expression of target genes. For example, the effector domain may recruit or interact with a scaffold protein domain that recruits PRMT, HDAC, SETDB1, or NuRD protein domains. In some embodiments, the effector domain comprises a Krüppel-associated box (KRAB) repressor domain; a repressor element silencing transcription factor (REST) repressor domain, a KRAB-associated protein 1 (KAP1) domain, a MAD domain, an FKHR (forkhead in the rhabdomyosarcoma gene) repressor domain, an AegR-1 (early growth response gene product-1) repressor domain, an ets2 repressor factor repressor domain (ERD), a MAD smSIN3 interaction domain (SID), a WRPW motif of a hair-associated basic helix-loop-helix (BhlH) repressor protein; an HP1 alpha chromo-shadow repression domain, or any combination thereof. In some embodiments, the effector domain comprises a KRAB domain. In some embodiments, the effector domain comprises a three-part motif containing a 28 (TRIM28, TIF1-β, or KAP1) protein.
[0479] In some embodiments, the effector domain comprises a protein domain that represses the expression of the target gene. For example, the effector domain may comprise a functional domain derived from a zinc finger repressor protein. In some embodiments, the effector domain comprises a functional repressor domain derived from the KOX1 / ZNF10 domain, KOX8 / ZNF708 domain, ZNF43 domain, ZNF184 domain, ZNF91KRAB domain, HPF4 domain, HTF10 domain, or HTF34 domain, or any combination thereof. In some implementations, the effector domains include domains derived from ZIM3 protein structures, such as ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZNF680, ZNF41, ZNF189, ZNF528, ZNF543, ZNF554, ZNF140, ZNF610, ZNF264, ZNF350, ZNF8, ZNF582, ZNF30, ZNF324, ZNF98, ZNF669, ZNF677, ZNF596, ZNF214, and ZNF37A. Functional blocking structural domains of ZNF34, ZNF250, ZNF547, ZNF273, ZNF354A, ZFP82, ZNF224, ZNF33A, ZNF45, ZNF175, ZNF595, ZNF184, ZNF419, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF566, ZNF729, ZIM2, ZNF254, ZNF764, ZNF785, or any combination thereof. In some embodiments, the structural domain is a ZIM3 structural domain, a ZNF554 structural domain, a ZNF264 structural domain, a ZNF324 structural domain, a ZNF354A structural domain, a ZNF189 structural domain, a ZNF543 structural domain, a ZFP82 structural domain, a ZNF669 structural domain, or a ZNF582 structural domain, or any combination thereof. In some embodiments, the structural domain is a ZIM3 structural domain.
[0480] In some implementations, the effect subdomain may be an alternative KRAB domain (e.g.). Alternatively or otherwise, the effect subdomain may be a non-KRAB domain.
[0481] In some implementations, the protein fusion construct may have 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 effector domains.
[0482] In some implementations, the effector domain contains a functional domain that represses or silences gene expression, and this functional domain is part of a larger protein (e.g., a zinc finger repressor). Functional domains capable of regulating gene expression (e.g., repressing or increasing gene expression) can be identified from larger proteins using known methods and the methods provided herein. For example, functional effector domains that can reduce or silence target gene expression can be identified based on the sequence of the repressor or activator protein. The amino acid sequence of a protein with gene expression-regulating function can be obtained from an available genome browser (e.g., the UCSD Genome Browser or the Ensembl Genome Browser). For example, the ZNF10 protein.
[0483] Protein annotation databases such as UniProt or Pfam can be used to identify functional domains in whole protein sequences. Using these tools, a repressor domain can be identified in the ZNF10 protein sequence. In some cases, various functional domains identified from larger proteins can be tested. Databases may differ in specific boundary domains. For example, in some embodiments, the repressor domain of ZNF10 is primarily amino acids 14-85 of the ZNF10 protein. In some embodiments, the repressor domain of ZNF10 is primarily amino acids 14-85 of the ZNF10 protein. In some embodiments, the repressor domain of ZNF10 is primarily amino acids 13-54 of the ZNF10 protein. As a starting point, gene expression regulatory activity can be tested on the largest sequence covering all regions identified by different databases, for example, testing the region of the ZN10 protein containing amino acids 13-85 as a starting point. In another implementation, the starting point region may be truncated at the N-terminus or C-terminus by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more amino acids, and various truncations may be tested to identify the minimum functional unit.
[0484] In some embodiments, the effector domain comprises a histone deacetylase protein domain. In some embodiments, the effector domain comprises an HDAC family protein domain, such as HDAC1, HDAC3, HDAC5, HDAC7, or HDAC9 protein domain. In some embodiments, the effector domain removes an acetyl group. In some embodiments, the effector domain comprises a nucleosome remodeling domain. In some embodiments, the effector domain comprises a nucleosome remodeling and deacetylase complex (NURD) that removes an acetyl group from the histone.
[0485] In some embodiments, the effector domain contains a three-part motif of a 28 (TRIM28, TIF1-β, or KAP1) protein. In some embodiments, the effector domain contains one or more KAP1 proteins. The KAP1 protein in the epigeneity editor can form a complex with one or more other effector domains of the epigeneity editor or with one or more proteins involved in the regulation of gene expression in the cellular environment. For example, KAP1 can be recruited by the KRAB domain of a transcriptional repressor. In some embodiments, KAP1 interacts with or recruits histone deacetylase proteins, histone-lysine methyltransferase proteins (e.g., those depositing a methyl group on lysine 9 [K9][H3K9] at the histone H3 tail), chromatin remodeling proteins, and / or heterochromatin proteins. In some embodiments, the KAP1 protein interacts with or recruits a complex of one or more proteins that reduce or silence gene expression. In some embodiments, the KAP1 protein interacts with or recruits heterochromatin protein 1 (HP1) proteins (e.g., via the chromosome shadow domain of the HP1 protein), SETDB1, etc. The KAP1 protein interacts with or recruits components of the nucleosome remodeling and deacetylation (NuRD) protein complex. In some embodiments, the KAP1 protein recruits the CHD3 subunit of the nucleosome remodeling and deacetylation (NuRD) complex, thereby reducing or silencing the expression of the target gene. In some embodiments, the KAP1 protein recruits the SETDB1 protein (e.g., recruited to the promoter region of the target gene), thereby reducing or silencing the expression of the target gene via H3K9 methylation associated with, for example, the promoter region of the target gene. In some embodiments, the recruitment of the SETDB1 protein results in heterochromatinization of the chromosomal region containing the target gene, thereby reducing or silencing the expression of the target gene. In some embodiments, the KAP1 protein interacts with or recruits the HP1 protein, thereby reducing or silencing the expression of the target gene via reducing the acetylation of H3K9 or H3K14 on the histone tail associated with the target gene. The recruitment of SETDB1 induces heterochromatinization. In some embodiments, the KAP1 protein interacts with or recruits the ZFP90 protein (e.g., isotype 2 of ZFP90) and / or the FOXP3 protein.
[0486] In some embodiments, the effector domain comprises a protein domain that interacts with or is recruited by one or more DNA epigenetic markers. For example, the effector domain may comprise a methyl CpG-binding protein 2 (MECP2) protein that interacts with methylated DNA nucleotides in a target gene. In some embodiments, the MECP2 protein interacts with methylated DNA nucleotides in CpG islands of the target gene. In some embodiments, the MECP2 protein interacts with methylated DNA nucleotides not in CpG islands of the target gene. In some embodiments, the MECP2 protein in the epigenetic editor results in condensed chromatin structure, thereby reducing or silencing the expression of the target gene. In some embodiments, the MECP2 protein in the epigenetic editor interacts with histone deacetylases (e.g., HDAC), thereby repressing or silencing the expression of the target gene. In some embodiments, the MECP2 protein in the epigenetic editor blocks the access of transcription factors or transcription activators to the target gene, thereby repressing or silencing the expression of the target gene.
[0487] In some implementations, the effector domains include a chromoshadow domain, a ubiquitin-2-like Rad60 SUMO-like (Rad60-SLD / SUMO) domain, a chromatin organization modifier (Chromo) domain, a Yaf2 / RYBP C-terminal binding motif domain (YAF2_RYBP), a CBX family C-terminal motif domain (CBX7_C), a zinc finger C3HC4 type (RING finger) domain (zf-C3HC4_2), a cytochrome b5 domain (Cyt-b5), a helical-loop-helical domain (HLH), a high-mobility cassette domain (HMG-cassette), a sterility α motif domain (SAM_1), a basic leucine zipper domain (BziP_1), a Myb_DNA binding domain, a homology domain, a MYM type zinc finger with an FCS sequence domain (zf-FCS), and an interferon regulator 2 binding protein zinc finger domain (IRF-2BP1_). 2) SSX repression domain (SSXRD), B-box type zinc finger domain (zf-B_box), sterility α motif domain (SAM_2), CXXC zinc finger domain (zf-CXXC), chromosome condensation 1 regulatory factor domain (RCC1), SRC homology 3 domain (SH3_9), sterility α motif / tip domain (SAM_PNT), Vestigial / Tondu family domain (Vg_Tdu), LIM domain, RNA recognition motif domain (RRM_1), basic leucine zipper domain (BziP_2), paired amphiphilic helical domain (PAH), proteasome ATPase OB C-terminal domain (Prot_ATP_ID_OB), Neural Homology 2 domain (NHR2), Helical-Hairpin-Helical Motif domain (HHH_3), Hinge domain for cleaving stimulatory factor subunit 2 (CSTF2_hinge), PPARγ N-terminal domain (PPARγ_N), CDC 48 N-terminal domain (CDC48_2), WD40 repeat domain (WD40), Fip1 motif domain (Fip1), PDZ domain (PDZ_6), von Willebrand factor C domain (VWC), NAB conserved region 1 domain (NCD1), S1 RNA-binding domain (S1), HNF3C-terminal domain (HNF_C), Tudor domain (Tudor_2), histone-like transcription factor (CBF / NF-Y) and archaea histological protein domain (CBFD_NFYB_HMF), zinc finger protein domain (DUF3669), EGF-like domain (CegF), GATA zinc finger domain (GATA), TEA / ATTS domain (TEA), phorbol ester / diacylglycerol binding domain (C1-1), multicomb-like MTF2 factor 2 domain (Mtf2_C), FOXO protein family transactivation domain (FOXO-TAD), homeobox KN junction Domains (homeobox_KN), BED zinc finger domain (zf-BED), zinc finger of C3HC4 type RING domain (zf-C3HC4_4), RAD51 interacting motif domain (RAD51_interaction), p55 binding region of methyl-CpG binding domain protein MBD (MBDa), Notch domain, Raf-like Ras binding domain (RBD), Spin / Ssty family domain (Spin-Ssty), PHD finger domain (PHD_3), LDL receptor type A domain (Ldl_recept_a), CS domain, DM A DNA-binding domain or a QLQ domain. In some embodiments, the effector domain is a protein domain comprising a YAF2_RYBP domain or a homologous domain or any combination thereof. In some embodiments, the homologous domain of the YAF2_RYBP domain is a PRD domain, an NKL domain, a HOXL domain, or a LIM domain. In some embodiments, the effector domain comprises a protein domain selected from the SUMO3 domain, the Chromo domain from M-phase phosphoprotein 8 (MPP8), the chromosomeshadow domain from Chromobox 1 (CBX1), and the SAM_1 / SPM domain from Scm polycomb family protein homolog 1 (SCMH1). In some embodiments, the effector domain comprises HNF3. C-terminal domain (HNF_C). In some embodiments, the HNF_C domain is derived from FOXA1 or FOXA2. In some embodiments, the HNF_C domain contains the EH1 (engrailed homology 1) motif. In some embodiments, the effector domain contains the interferon regulator 2 binding protein zinc finger domain (IRF-2BP1_2). In some embodiments, the effector domain contains a gene from the DNA repair factor HERC2.The Cyt-b5 domain of the E3 ligase. In some embodiments, the effector domain comprises a variant SH3 domain (SH3_9) from bridging integrator 1 (BIN1). In some embodiments, the effector domain is an HMG-box domain from transcription factor TOX or a zf-C3HC4_2RING finger domain from the multicomb component PCGF2. In some embodiments, the effector domain comprises a cromo domain-helicase-DNA binding protein 3 (CHD3). In some embodiments, the effector domain comprises a ZNF783 domain. In some embodiments, the effector domain comprises a YAF2_RYBP domain. In some embodiments, the YAF2_RYBP domain comprises a 32-amino acid Yaf2 / RYBP C-terminal binding motif domain (32AA RYBP).
[0488] In some embodiments, the epigenetic editor described herein alters the chemical modifications of a target gene containing a target sequence. For example, an epigenetic editor containing a methyltransferase domain can methylate DNA or histone residues of the target gene at nucleotides (or histones) near the target sequence, or within 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, or 3000 base pairs flanking the target sequence, thereby inhibiting or silencing the expression of the target gene.
[0489] Epigenetic editor-mediated chemical modifications can occur near the target sequence of the target gene. For example, such modifications can occur within 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs flanking the target sequence. In some embodiments, the chemical modification occurs within 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 base pairs upstream of the 5' end of the target sequence.
[0490] Appearance Editor
[0491] This document describes an epigenetic editor for epigenetic modification and expression regulation of target genes. As used herein, an epigenetic editor can be any agent that binds to a target polynucleotide and has epigenetic regulatory activity. In some embodiments, the epigenetic editor uses IscB or a variant thereof to bind to a polynucleotide at a specific sequence. In some embodiments, the epigenetic editor includes an effector domain capable of regulating the epigenetic state of a nucleic acid sequence at or near the target polynucleotide. In some embodiments, the epigenetic editor is capable of depositing epigenetic editing markers at or near a target polynucleotide on a chromatin region, nucleic acid sequence, or histone amino acid residue. For example, the epigenetic editor may be capable of methylating, demethylating, acetylating, deacetylating, ubiquitinizing, or deubiquitinizing a chromatin region, nucleic acid sequence, or histone amino acid residue at or near a target polynucleotide. In some embodiments, the epigenetic editor is capable of recruiting one or more proteins or complexes involved in transcriptional regulation (e.g., transcription factors, transcription activators, transcription repressors, or insulators) to a chromatin region, nucleic acid sequence, or histone amino acid residue at or near a target polynucleotide.
[0492] The epigenetic editor provided herein may include one or more effector subdomains as described. In some embodiments, the epigenetic editor includes multiple effector subdomains. In some embodiments, the epigenetic editor includes one effector subdomain. In some embodiments, the epigenetic editor includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more effector subdomains. In some embodiments, the epigenetic editor includes at least 2 effector subdomains, such as two repressor domains. In some embodiments, the epigenetic editor includes at least 2 effector subdomains. In some embodiments, the epigenetic editor includes two or more effector subdomains. In some embodiments, two or more effector subdomains work synergistically to result in enhanced regulation of the target gene. For example, the epigenetic editor may include two effector subdomains, one of which induces histone deacetylation and the other causes DNA methylation of the target gene.
[0493] In some embodiments, the epigeneity editor comprises a DNA methylation domain and a histone deacetylation domain. In some embodiments, the epigeneity editor comprises a DNA methylation domain and a repression domain, the repression domain recruiting other DNA methylated, histone methylated, or histone deacetylated proteins. In some embodiments, the epigeneity editor comprises a DNA methylation domain and a scaffold protein, the scaffold protein recruiting other DNA methylated, histone methylated, or histone deacetylated proteins. In some embodiments, the epigeneity editor comprises a DNA methylation domain, a histone deacetylation domain, and a scaffold protein, the scaffold protein recruiting other DNA methylated, histone methylated, or histone deacetylated proteins. In some embodiments, the epigeneity editor comprises two or more DNA methylation domains, histone deacetylation domains, and a scaffold protein, the scaffold protein recruiting other DNA methylated, histone methylated, or histone deacetylated proteins. In some embodiments, the epigenetic editor comprises two or more DNA methylation domains, two or more histone deacetylation domains, and / or two or more scaffold proteins that recruit other DNA methylating, histone methylating, or histone deacetylation proteins. In some embodiments, the epigenetic editor comprises a KRAB domain and a DNMT3 domain, which, compared to an epigenetic effector having only one of the two repressor domains, can synergistically reduce or silence the expression of a target gene. In some embodiments, the epigenetic editor comprises a KRAB domain, a Dnmt3A domain, and a Dnmt3L domain. In some embodiments, the epigenetic editor comprises a DNA-binding domain flanked by a configuration of a KRAB domain and a Dnmt3A-Dnmt3L fusion protein domain. In some implementations, the epigenetic editor contains the following configuration: N-[KRAB]-[DNA binding domain]-[Dnmt3A-Dnmt3L]-C, where “]-[” is any nuclear localization signal, any tag sequence, or any adapter as provided herein.
[0494] In some embodiments, the epigenetic editor comprises a DNA demethylation domain and a histone acetylation domain. In some embodiments, the epigenetic editor comprises a DNA demethylation domain and an activation domain, the activation domain recruiting other DNA demethylated or histone acetylated proteins. In some embodiments, the epigenetic editor comprises a DNA demethylation domain, a histone acetylation domain, and a scaffold protein, the scaffold protein recruiting other DNA demethylated or histone acetylated proteins. In some embodiments, the epigenetic editor comprises two or more DNA demethylation domains, two or more histone acetylation domains, and / or two or more scaffold proteins, the scaffold proteins recruiting other DNA demethylated or histone deacetylated proteins.
[0495] The components of an epigenetic editor can be configured in different ways. For example, a DNA-binding domain may be located at the C-terminus, the N-terminus, or between two or more epigenetic effector domains or other domains. In some embodiments, the DNA-binding domain is located at the C-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is located at the N-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is linked to one or more nuclear localization signals. In some embodiments, the DNA-binding domain is linked to two or more nuclear localization signals. In some embodiments, the flanking parts of the DNA-binding domain are epigenetic effector domains or other domains at both ends. In some embodiments, the epigenetic editor comprises a configuration of [N']-[epigenetic effector domain 1]-[DNA-binding domain]-[epigenetic effector domain 2]-[C']. In some embodiments, the epigenetic editor comprises a configuration of [N']-[epigenetic effector domain 1]-[DNA-binding domain]-[epigenetic effector domain 2]-[epigenetic effector domain 3]-[C']. In some embodiments, the epigenetic editor includes a configuration of N']-[epigenetic effector domain 1]-[epigenetic effector domain 2]-[DNA-binding domain]-[epigenetic effector domain 3]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[epigenetic effector domain 1]-[epigenetic effector domain 2]-[DNA-binding domain]-[epigenetic effector domain 3]-[epigenetic effector domain 4]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[KRAB]-[DNA-binding domain]-[Dnmt3A]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[KRAB]-[DNA-binding domain]-[Dnmt3A]-[Dnmt3L]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[SETDB1]-[DNA-binding domain]-[Dnmt3A]-[Dnmt3L]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[SETDB1]-[DNA-binding domain]-[Dnmt3A]-[C'. In some embodiments, the epigenetic editor includes a configuration of N']-[KRAB]-[DNA-binding domain]-[Dnmt3A-Dnmt3L]-[C', wherein Dnmt3A and Dnmt3L are directly fused via peptide bonds.
[0496] In some embodiments, the epigenetic editor includes a configuration of N'-[Dnmt3A]-[DNA-binding domain]-[KRAB]-[C'. In some embodiments, the epigenetic editor includes a configuration of N'-[Dnmt3A]-[Dnmt3L]-[DNA-binding domain]-[KRAB]-[C'. In some embodiments, the epigenetic editor includes a configuration of N'-[Dnmt3A-Dnmt3L]-[DNA-binding domain]-[KRAB]-[C', wherein Dnmt3A and Dnmt3L are directly fused via peptide bonds. In some embodiments, the epigenetic editor includes a configuration of N'-[Dnmt3A]-[DNA-binding domain]-[SETDB1]-[C'. In some embodiments, the epigenetic editor includes a configuration of N'-[Dnmt3A]-[Dnmt3L]-[DNA-binding domain]-[SETDB1]-[C'. In some embodiments, the epigenetic editor comprises a configuration of N']-[Dnmt3A-Dnmt3L]-[DNA-binding domain]-[SETDB1]-[C', wherein Dnmt3A and Dnmt3L are directly fused via peptide bonds. In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a linker, such as a peptide linker. In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a detectable tag. In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a peptide bond. In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a nuclear localization signal. In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a promoter or regulatory sequence. In the epigenetic editor structure, multiple linkers “]-[” can be the same or can each be a different linker, tag, NLS, or peptide bond.
[0497] The DNA-binding domain (DBD) of an epigenetic editor may comprise any of the DNA-binding domains described herein or known to those skilled in the art. In some embodiments, the DBD comprises one or more zinc finger arrays. In some embodiments, the DBD comprises a TALE DNA-binding domain. In some embodiments, the DBD is an RNA-guided programmable DNA-binding domain, such as a CRISPR-Cas protein domain. Suitable Cas proteins have been provided herein, including nuclease-inactive Cas proteins for epigenetic editing purposes without causing target DNA strand breaks. The Cas protein in an epigenetic editor may be nuclease-inactive Cas9 (dCas9), SaCas9d, SpCas9d, dCas9d with modified PAM specificity, high-fidelity dCas9, nuclease-inactive Cpf1 (dCpf1), dCpf1 with modified PAM specificity, high-fidelity dCpf1, dCas12e, dCasY, or any other Cas protein as described herein.
[0498] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD) and an effector domain that represses or silences the expression of a target gene. In some embodiments, the epigenetic editor comprises an N'-[repression domain]-[DBD]-[-C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor comprises an N'-[DBD]-[repression domain]-[-C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0499] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD) and a DNA methyltransferase domain, the DNA methyltransferase domain depositing one or more methylation markers at a target gene to inhibit or silence the expression of the target gene. In some embodiments, the epigenetic editor comprises an N'-[DNA methyltransferase domain]-[DBD]-[-C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor comprises an N'-[DBD]-[DNA methyltransferase domain]-[-C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0500] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD), a DNA methyltransferase domain, and an effector domain that represses or silences the expression of a target gene. In some embodiments, the epigenetic editor comprises a configuration of N'-[DNA methyltransferase domain]-[DBD]-[repression domain]-[C'], wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor comprises a configuration of N'-[repression domain]-[DBD]-[DNA methyltransferase domain]-[C'], wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0501] In some embodiments, the epigenetic editor comprises a configuration of N']-[DNA methyltransferase domain]-[repression domain]-[DBD]-[C', wherein the linker structure [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0502] The repression domain in the epigenetic editor may contain any expression repressor protein known to those skilled in the art and as described herein, or any homolog or combination thereof. In some embodiments, the repression domain contains a histone deacetylase domain. In some embodiments, the repression domain interacts with a scaffold protein domain that recruits one or more protein domains that repress the expression of a target gene. For example, the repression domain may recruit or interact with a scaffold protein domain that recruits PRMT, HDAC, SETDB1, or NuRD protein domains. In some embodiments, the repression domain interacts with a DNA nucleotide of an epigenetic marker in the target gene, thereby repressing or silencing the expression of the target gene. In some embodiments, the repression domain contains a MECP2 domain. In some embodiments, the repression domain contains a KAP1 domain. In some embodiments, the repression domain contains any domain of Table 2 or Table 3, or any combination or homolog thereof.
[0503] The DNA methyltransferase domain in the epigenetic editor may comprise any DNA methyltransferase protein known to those skilled in the art and as described herein, or any homolog or combination thereof. In some embodiments, the effector domain comprises the DNMT3 domain. In some embodiments, the DNA methyltransferase domain comprises the DNMT3A domain. In some embodiments, the DNA methyltransferase domain comprises the DNMT3B domain. In some embodiments, the DNA methyltransferase domain comprises the DNMT3C domain. In some embodiments, the DNA methyltransferase domain comprises the DNMT3L domain. In some embodiments, the DNA methyltransferase domain comprises a fusion of the DNMT3A-DNMT3L domains. As described herein, the DNMT3A-DNMT3L fusion domain can be in any order, such as N-DNMT3A-DNMT3L-C or N-DNMT3L-DNMT3A-C. In some embodiments, the DNA methyltransferase domain comprises any of the domains in Table 1, or any combination or homolog thereof.
[0504] In some embodiments, the epigenetic editor includes a DNA-binding domain (DBD) and an effector domain that increases the expression of a target gene. In some embodiments, the epigenetic editor includes a configuration of N'-[activation domain]-[DBD]-[C', wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor includes a configuration of N'-[DBD]-[activation domain]-[C', wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0505] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD) and a DNA demethylation domain, the DNA demethylation domain removing one or more methylation markers at a target gene, thereby increasing the expression of the target gene. In some embodiments, the epigenetic editor comprises an N'-[DNA demethylase domain]-[DBD]-[C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor comprises an N'-[DBD]-[DNA demethylase domain]-[C' configuration, wherein the linker [-] is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0506] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD), a DNA demethylase domain, and an activation effector domain that increases the expression of a target gene. In some embodiments, the epigenetic editor comprises a configuration of N'-[DNA demethylase domain]-[DBD]-[activation domain]-[C', wherein the linker structure is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein. In some embodiments, the epigenetic editor comprises a configuration of N'-[activation domain]-[DBD]-[DNA demethylase domain]-[C', wherein the linker structure is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0507] In some embodiments, the epigenetic editor comprises a configuration of N']-[DNA demethylase domain]-[activation domain]-[DBD]-[C', wherein the linker structure []-[ is any one of the adapters, detectable tags, affinity domains, peptide bonds, nuclear localization signals, promoters, and / or regulatory sequences as described herein.
[0508] The activation domain in the epigenetic editor may comprise any expression activating protein known to those skilled in the art and as described herein, or any homolog or combination thereof. In some embodiments, the activation domain comprises a histone acetyltransferase domain. In some embodiments, the activation domain interacts with a scaffold protein domain that recruits protein domains that activate the expression of one or more target genes. For example, the activation domain may recruit or interact with a scaffold protein domain that recruits one or more transcription factors or activators. In some embodiments, the activation domain comprises a herpes simplex virus protein 16 (VP16) activation domain. In some embodiments, the activation domain comprises an activation domain containing a tandem repeat sequence of multiple VP16 activation domains. In some embodiments, the activation domain comprises four tandem copies of VP16 (a VP64 activation domain). In some embodiments, the activation domain comprises eight tandem copies of VP16 (a VP128 activation domain). In some embodiments, the activation domain comprises ten tandem copies of VP16 (a VP160 activation domain). In some embodiments, the activation domain comprises the p65 activation domain of NFκB. In some embodiments, the activation domain comprises an Epstein-Barr virus R trans-activator (Rta) activation domain. In some embodiments, the activation domain comprises a fusion of multiple activators, such as a triple activator (VPR activation domain) of VP64, p65, and Rta activation domains. In some embodiments, the activation domain comprises any one of the domains in Table 5 or Table 6, or any homologs or combinations thereof.
[0509] The DNA demethylation domain in the epigenetic editor may comprise any of the DNA demethylating proteins known to those skilled in the art and described herein, or any homologs or combinations thereof. In some embodiments, the DNA demethylation domain comprises a TET family protein domain. In some embodiments, the DNA demethylation domain comprises a TET1, TET2, or TET3 protein domain. In some embodiments, the DNA demethylation domain comprises a TET1 protein domain. In some embodiments, the DNA demethylation domain comprises any of the domains in Table 4, or any homologs or combinations thereof.
[0510] In some embodiments, the epigenetic editor that can reduce or silence the expression of a target gene includes a Dnmt3A-Dnmt3L fusion protein domain. In some embodiments, the epigenetic editor also includes a repression scaffold or recruiting protein domain, such as a KRAB domain, a KAP1 domain, or a MECP2 domain. In some embodiments, the epigenetic editor includes a Dnmt3A-Dnmt3L fusion domain and other repression domains that reduce or silence the expression of a target gene. The repression domain in the epigenetic editor may include any expression repression protein known to those skilled in the art and as described herein, or any homolog or combination thereof. In some embodiments, the repression domain includes a histone deacetylase domain. In some embodiments, the repression domain interacts with a scaffold protein domain that recruits one or more protein domains that repress the expression of a target gene. For example, the repression domain may recruit or interact with a scaffold protein domain that recruits PRMT, HDAC, SETDB1, or NuRD protein domains. In some embodiments, the repressor domain interacts with the DNA nucleotides of an epigenetic marker in the target gene, thereby repressing or silencing the expression of the target gene. In some embodiments, the repressor domain comprises the MECP2 domain. In some embodiments, the repressor domain comprises the KAP1 domain. In some embodiments, the repressor domain comprises any one of the domains in Table 2 or Table 3, or any combination or homolog thereof.
[0511] In some embodiments, the epigenetic editor includes a Dnmt3A-Dnmt3L fusion domain and a KAP1 domain. In some embodiments, the epigenetic editor includes the following configuration: N]-[Dnmt3A-3L]-[KAP1]-[DBD]-[C], where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[KAP1]-[Dnmt3A-3L]-[DBD]-[C], where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[Dnmt3A-3L]-[KAP1]-[C], where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[KAP1]-[Dnmt3A-3L]-[C], where the linker structure []-[ can be any of the adapters provided herein. In some implementations, the epigenetic editor comprises the following configuration: N]-[KAP1]-[DBD]-[Dnmt3A-3L]-[C, wherein the linker structure []-[ can be any of the adapters provided herein.
[0512] In some embodiments, the epigenetic editor includes a Dnmt3A-Dnmt3L fusion domain and a MECP2 domain. In some embodiments, the epigenetic editor includes the following configuration: N]-[Dnmt3A-3L]-[MECP2]-[DBD]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[MECP2]-[Dnmt3A-3L]-[DBD]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[Dnmt3A-3L]-[MECP2]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[MECP2]-[Dnmt3A-3L]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor comprises the following configuration: N]-[MECP2]-[DBD]-[Dnmt3A-3L]-[C, wherein the linker structure []-[ can be any of the adapters provided herein.
[0513] In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnmt3L fusion domain and a heterochromatin protein 1 (HP1) domain. In some embodiments, the epigenetic editor comprises the following configuration: N]-[Dnmt3A-3L]-[HP1]-[DBD]-[C, wherein the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor comprises the following configuration: N]-[HP1]-[Dnmt3A-3L]-[DBD]-[C, wherein the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor comprises the following configuration: N]-[DBD]-[Dnmt3A-3L]-[HP1]-[C, wherein the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[HP1]-[Dnmt3A-3L]-[C, wherein the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[HP1]-[DBD]-[Dnmt3A-3L]-[C, wherein the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[Dnmt3A-3L]-[DBD]-[HP1]-[C, wherein the linker structure []-[ can be any of the adapters provided herein.
[0514] In some embodiments, the epigenetic editor includes a Dnmt3A-Dnmt3L fusion domain and a SETDB1 domain. In some embodiments, the epigenetic editor includes the following configuration: N]-[Dnmt3A-3L]-[SETDB1]-[DBD]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[SETDB1]-[Dnmt3A-3L]-[DBD]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[Dnmt3A-3L]-[SETDB1]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some embodiments, the epigenetic editor includes the following configuration: N]-[DBD]-[SETDB1]-[Dnmt3A-3L]-[C, where the linker structure []-[ can be any of the adapters provided herein. In some implementations, the epigenetic editor includes the following configuration: N]-[SETDB1]-[DBD]-[Dnmt3A-3L]-[C, where the linker structure []-[ can be any of the adapters provided herein.
[0515] In some implementations, the epigenetic editor includes the Dnmt3A-Dnmt3L fusion domain and the SETDB1 domain, KAP1 domain, KRAB domain and / or MECP2 domain, in any order and combination thereof.
[0516] In some embodiments, the epigenetic editor that reduces or silences the expression of a target gene comprises a DBD and an affinity domain that specifically binds to a repression domain. For example, the epigenetic editor may comprise a DBD and an antibody against the repression domain. In some embodiments, the epigenetic editor comprises a DBD and a KAP1 affinity domain. In some embodiments, the epigenetic editor comprises a DBD and a KRAB affinity domain. In some embodiments, the epigenetic editor comprises a DBD and a SETDB1 affinity domain. In some embodiments, the epigenetic editor comprises a DBD and a MECP2 affinity domain. In some embodiments, the epigenetic editor comprises a DNA methyltransferase and a repression domain-binding affinity domain. In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnm3L fusion and a repression domain-binding affinity domain. In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnm3L fusion and a KAP1 affinity domain. In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnm3L fusion and a KRAB affinity domain. In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnm3L fusion and a SETDB1 affinity domain. In some embodiments, the epigenetic editor comprises a Dnmt3A-Dnm3L fusion and a MECP2 affinity domain. As used herein, the affinity domain may be an antibody, a single-chain antibody, a nanobody, or an antigen-binding sequence, an antibody, a nanobody, a functional antibody fragment, a single-chain variable fragment (scFv), a Fab, a single-domain antibody (sdAb), a VH domain, a VL domain, a VNAR domain, a VHH domain, a bispecific antibody, a biantibody, or a functional fragment or a combination thereof.
[0517] In some embodiments, the epigenetic editor for reducing or silencing target gene expression comprises a DBD and an affinity domain that specifically binds to a DNA methyltransferase domain. For example, the epigenetic editor may comprise a DBD and a DNA methyltransferase antibody. In some embodiments, the epigenetic editor comprises a DBD and a Dnmt3A affinity domain. In some embodiments, the epigenetic editor comprises a DBD and a Dnmt3L affinity domain. In some embodiments, the epigenetic editor comprises a repression domain and a DNA methyltransferase binding affinity domain. In some embodiments, the epigenetic editor comprises a repression domain and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editor comprises a repression domain and a Dnmt3L affinity domain. In some embodiments, the epigenetic editor comprises one or more of the KAP1, KRAB, and MECP2 domains and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editor comprises one or more of the KAP1 domains and a Dnmt3A binding affinity domain. In some embodiments, the epigenetic editor comprises one or more of the KAP1, KRAB, and MECP2 domains and a Dnmt3L binding affinity domain. In some embodiments, the epigenetic editor comprises one or more of the KAP1 domains and a Dnmt3L binding affinity domain. The affinity domain can be an antibody, single-chain antibody, nanobody, or antigen-binding sequence, antibody, nanobody, functional antibody fragment, single-chain variable fragment (scFv), Fab, single-domain antibody (sdAb), VH domain, VL domain, VNAR domain, VHH domain, bispecific antibody, biantibody, or functional fragment or a combination thereof.
[0518] In some embodiments, the epigenetic editor for reducing or silencing target gene expression includes a DBD and a first affinity domain that specifically binds to a DNA methyltransferase domain and a second affinity domain that specifically binds to a repression domain. For example, the epigenetic editor may include a DBD and a DNA methyltransferase antibody and a repression domain antibody. In some embodiments, the epigenetic editor includes a DBD, a KAP1 affinity domain, and a Dnmt3A affinity domain. In some embodiments, the epigenetic editor includes a DBD, a KAP1 affinity domain, and a Dnmt3L affinity domain. In some embodiments, the epigenetic editor includes a DBD, a MECP2 affinity domain, and a Dnmt3A affinity domain. In some embodiments, the epigenetic editor includes a DBD, a MECP2 affinity domain, and a Dnmt3L affinity domain. In some embodiments, the epigenetic editor includes a DBD, a KRAB affinity domain, and a Dnmt3A affinity domain. In some implementations, the epigenetic editor includes a DBD, a KRAB affinity domain, and a Dnmt3L affinity domain. The affinity domain can be an antibody, a single-chain antibody, a nanobody, an antigen-binding sequence, a functional antibody fragment, a single-chain variable fragment (scFv), a Fab, a single-domain antibody (sdAb), a VH domain, a VL domain, a VNAR domain, a VHH domain, a bispecific antibody, a biantibody, or a functional fragment or a combination thereof.
[0519] Base Editor
[0520] A base editor (BE) is an agent containing a polypeptide capable of modifying bases (e.g., A, T, C, G, or U) within a nucleic acid sequence (e.g., DNA or RNA). A base editor may comprise a macromolecule or macromolecular complex capable of converting nucleobases in a polynucleotide sequence to another nucleobase (e.g., conversion or transversion) at one or more locations within a base editing window. A base editor may comprise (a) a nucleotide convertase, nucleoside convertase, or nucleobase convertase, and (b) an IscB protein as described herein or an inactivated variant of IscB as described herein. The IscB protein may be catalytically inactivated or impaired such that it does not cleave a single-stranded nucleic acid target or nicks or cleaves at most one strand of a double-stranded nucleic acid target. The IscB may be any of the IscB proteins described herein or its DNA-binding domain.
[0521] A base editor may comprise: an IscB protein as described herein, an inactivating variant of IscB as described herein, or a DNA-binding domain thereof, fused or linked to a domain having base editing activity to produce a base editor fusion protein. The base editor fusion protein may include one or more adapters, such as peptide adapters between domains. In some embodiments, the domain having base editing activity is linked to a guide RNA (e.g., via an RNA-binding motif on the guide RNA and an RNA-binding domain fused to a deaminase).
[0522] In some embodiments, the base editor comprises a base editor fusion protein or its encoded nucleic acid.
[0523] In some embodiments, the base editor fusion protein comprises any of the IscB proteins described in this disclosure, IscB inactivating variants, or their DNA-binding domains.
[0524] In some embodiments, the base editor fusion protein comprises: any of the IscB proteins described herein, an inactivating variant of IscB, or its DNA-binding domain; and a deaminase. In some embodiments, the deaminase is a base deaminase. In some embodiments, the deaminase is a cytidine deaminase. In some embodiments, the deaminase has a cytidine deaminase domain. In some embodiments, the deaminase is an adenosine deaminase. In some embodiments, the deaminase has an adenosine deaminase domain.
[0525] In some embodiments, the base editor comprises: any of the IscB proteins, inactivating variants of IscB, or their DNA-binding domains as described in this disclosure; and a guide polynucleotide.
[0526] In some embodiments, the base editor fusion protein is a modular, programmable protein comprising a deaminase domain fused to any of the IscB proteins, inactivating variants of IscB, or their DNA-binding domains as described herein. Through hydrolysis and subsequent cellular processing without generating double-strand DNA breaks, the adenine base editor (ABE) converts A:T base pairs to G:C base pairs, and the cytosine base editor (CBE) converts C:G base pairs to T:A base pairs. The corresponding deaminases of the base editors are guided to their target sites by guide RNA (gRNA) within the D10A nickase Cas9 (nCas9). The cytidine deaminase of the CBE guides the conversion of cytosine to uridine, thereby achieving C→T (or G→A) substitution. Cytidine deaminases primarily refer to deaminases that act on deoxycytidine, cytidine, or both deoxycytidine and cytidine to convert cytosine to uridine. Cytidine deaminase and cytosine deaminase are used interchangeably herein. In cases where the goal is to disrupt genes in vivo for therapeutic purposes, cytosine base editors (CBEs) can potentially introduce stop codons directly into the coding sequence of a gene by altering specific codons for glutamine (CAG→TAG, CAA→TAA), arginine (CGA→TGA), and tryptophan (TGG→TAG / TAA / TGA, where cytosine on the antisense strand is edited) (nonsense mutations).
[0527] In contrast, the adenosine deaminase of the ABE directs the conversion of adenosine to inosine, thereby achieving an A→G (or T→C) substitution. Adenosine deaminases primarily refer to deaminases that act on deoxyadenosine, adenosine, or both deoxyadenosine and adenosine to convert adenine to hypoxanthine or alternatively, to inosine. Because inosine's structure is similar to guanosine (inosine does not contain the exocyclic amino group of guanosine), inosine tends to behave like guanosine. Through subsequent cellular processing, inosine is eventually replaced by guanosine. Therefore, adenosine deaminase achieves A→G (or T→C) substitution. Adenosine deaminase and adenine deaminase are used interchangeably in this article. The adenine base editor (ABE) cannot directly introduce a stop codon because there is no A→G change that would lead to a nonsense mutation.
[0528] For example, adenine base editors can be used to disrupt gene function by editing start codons (ATG→GTG or ATG→ACG). A second strategy by which adenine base editors can disrupt gene function is by editing splice sites, either splice donors at the 5' end of introns or splice acceptors at the 3' end of introns. Splice site disruption can result in the inclusion of intron sequences in messenger RNA (mRNA) (potentially introducing nonsense, frameshift, or in-frame mutations that lead to premature stop codons or insertions / deletions of amino acids that disrupt protein activity) or the absence of exon sequences (potentially introducing nonsense, frameshift, or in-frame insertion / deletion mutations).
[0529] Base editors may contain nucleic acid-binding proteins such as Cas nickases and cytidine deaminases. Base editors may contain uracil-DNA glycosylation enzymes.
[0530] In some embodiments, the base editor comprises (1) any of the IscB proteins, IscB inactivating variants, or their DNA-binding domains as described herein; (2) a deaminase domain for deaminating bases (e.g., adenosine deaminase or cytidine deaminase); and (3) one or more guide polynucleotides. In some embodiments, the base editor system comprises a base editor fusion protein (or base editor) containing (1) and (2), or a polynucleotide (e.g., mRNA) encoding a base editor containing (1) and (2). In some embodiments, the base editor is an adenine or adenosine base editor (ABE). In some embodiments, the base editor is a cytosine base editor (CBE).
[0531] In some embodiments, the nucleic acid encoding the base editor fusion protein is mRNA. In some embodiments, upon administration, the mRNA is translated into the base editor fusion protein in the targeted cells or subject. In some embodiments, the base editor fusion protein forms a ribonucleoprotein (RNP) complex in the targeted cells or subject.
[0532] It should be appreciated that the fusion proteins of this disclosure may include one or more additional features. For example, in some embodiments, the fusion protein may include a nuclear localization sequence (NLS), a cytoplasmic localization sequence, an export sequence (such as a nuclear export sequence) or other localization sequence, and a sequence tag that can be used to dissolve, purify, or detect the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, multihistidine tags (also known as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1, Softag 3), strep tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Additional suitable sequences will be apparent to those skilled in the art. In some embodiments, the fusion protein includes one or more His tags.
[0533] cytidine deaminase domain
[0534] In some embodiments, the base editor contains a deaminase, which is a cytosine deaminase. In some embodiments, the cytosine deaminase domain is fused to the N-terminus of the napDNAbp domain.
[0535] In some embodiments, the deaminase is an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the deaminase is APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, or APOBEC4 deaminase. In some embodiments, the deaminase is an activation-inducible deaminase (AID). In some embodiments, the deaminase is lamprey CDA1 (pmCDA1) deaminase. In some embodiments, the deaminase is derived from humans, chimpanzees, gorillas, monkeys, cattle, dogs, rats, or mice. In some embodiments, the deaminase is derived from humans. In some embodiments, the deaminase is derived from rats. In some embodiments, the deaminase is human APOBEC1 deaminase. In some embodiments, the deaminase is pmCDA1. In some embodiments, the adenosine deaminase is human APOBEC3G. In some embodiments, the deaminase is a variant of human APOBEC3G. In some embodiments, the deaminase is rat APOBEC1.
[0536] In some implementations, the base editor contains the FERNY or evoFERNY deaminase domain. "FERNY" is reconstructed from ancestral sequences truncated at the N- and C-termini of the APOBEC family phylogenetic tree. Evolved FERNY genotypes also exhibit high GC activity, and despite being shorter proteins, their activity is comparable to that of APOBEC.
[0537] adenosine deaminase domain
[0538] In some embodiments, the adenosine deaminase is a known variant of the adenosine deaminase TadA7.10, which contains the following mutations compared to wild-type ecTadA: W23R, H36L, P48A, R51L, L84F, A106V, D108N, H123Y, S146C, D147Y, R152P, E155V, I156F, and K157N. In some embodiments, the adenosine deaminase is a variant of TadA derived from species other than *Escherichia coli*, such as *Staphylococcus aureus*, *Salmonella typhi*, *Shewanella putrefactive*, *Haemophilus influenzae*, *Strychnos nucleus*, or *Bacillus subtilis*.
[0539] In some embodiments, the adenosine deaminase domain comprises TadA-8e, a variant of E. coli TadA7.10. TadA-8e contains the following substitutions relative to TadA7.10: T111, D119, F149, R26, V88, A109, H122, T166, and D167. In some embodiments, the adenosine deaminase domain comprises TadA-8e (V106W), which contains the V106W substitution relative to TadA-8e.
[0540] In various implementations, adenosine deaminase hydrolyzes and deaminates the target adenosine in the nucleic acid of interest into inosine, which is then read as guanosine (G) by DNA polymerase.
[0541] In some embodiments, the adenosine deaminase domain of the base editor contains a single adenosine deaminase or monomer. In some embodiments, the adenosine deaminase domain contains 2, 3, 4, or 5 adenosine deaminases. In some embodiments, the adenosine deaminase domain contains two adenosine deaminases or a dimer. In some embodiments, the deaminase domain contains a dimer of an engineered (or evolved) deaminase and a wild-type deaminase (such as a wild-type E. coli-derived deaminase).
[0542] In some embodiments, any adenosine deaminase provided herein is capable of deaminating adenine, for example, deaminoing adenine in the deoxyadenosine nucleoside of DNA. Adenosine deaminases can be derived from any suitable organism. In some embodiments, the adenosine deaminase is a naturally occurring adenosine deaminase comprising one or more mutations corresponding to any mutation provided herein (e.g., mutations in ecTadA). Those skilled in the art will be able to identify any homologous protein and the corresponding residues in the corresponding encoding nucleic acid using methods well known in the art, for example, by sequence alignment and determination of homologous residues. In some embodiments, the adenosine deaminase is derived from *Escherichia coli*.
[0543] In some embodiments, the adenosine deaminase domain comprises TadA9 or a variant thereof. TadA9 contains V82S and Q154R substitutions relative to TadA-8e. In some embodiments, the adenosine deaminase domain comprises TadA20 or a variant thereof. TadA20 contains I76Y, V82S, Y123H, Y147R, and Q154R substitutions relative to TadA7.10. TadA20 may be referred to in the art as TadA*8.20.
[0544] carrier system
[0545] Another aspect of this disclosure relates to a vector system comprising the IscB system described herein, the vector system comprising one or more recombinant vectors, the recombinant vectors comprising a polynucleotide sequence encoding the IscB protein and a polynucleotide sequence encoding the guide polynucleotide.
[0546] In some embodiments, the vector system comprises at least one plasmid or viral recombinant vector (e.g., retrovirus, lentivirus, adenovirus, adeno-associated virus, or herpes simplex virus). In some embodiments, the polynucleotide sequence encoding the IscB protein and the polynucleotide sequence encoding the guide polynucleotide are located on the same recombinant vector. In some embodiments, the polynucleotide sequence encoding the IscB protein and the polynucleotide sequence encoding the guide polynucleotide are located on multiple recombinant vectors.
[0547] In some embodiments, the polynucleotide sequence encoding the IscB protein and / or the polynucleotide sequence encoding the guiding polynucleotide is operatively linked to a regulatory sequence (also called a regulatory element). The regulatory element includes promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Regulatory elements include those that constitutively express the nucleotide sequence in many types of host cells, and those that express the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters may be expressed directly primarily in the desired tissue of interest, such as muscle, neurons, bone, skin, blood, a specific organ (e.g., liver, pancreas), or a specific cell type (e.g., lymphocytes). Regulatory elements may also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, and may or may not be tissue- or cell type-specific. In some embodiments, the regulatory element is an enhancer element, such as WPRE, CMV enhancer, R-U5 segment in the LTR of HTLV-1, SV40 enhancer, or intron sequence between exons 2 and 3 of rabbit β-globin.
[0548] In some embodiments, the recombinant vector comprises a pol III promoter (e.g., U6 and H1 promoters), a pol II promoter (e.g., a retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), a cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), an SV40 promoter, a dihydrofolate reductase promoter, a β-actin promoter, a glycerol phosphokinase (PGK) promoter, or an EF1α promoter), or a pol III promoter and a pol II promoter.
[0549] In some embodiments, the promoter is a constitutive promoter, which is continuously active and not regulated by external signals or molecules. Suitable constitutive promoters include, but are not limited to, CMV, RSV, SV40, EF1α, CAG, and β-actin promoters. In some embodiments, the promoter is an inducible promoter regulated by external signals or molecules (e.g., transcription factors).
[0550] In some embodiments, the promoter is a tissue-specific promoter that can be used to drive tissue-specific expression of the IscB protein. Suitable muscle-specific promoters include, but are not limited to, CK8, MHCK7, myoglobin promoter (Mb), desmin promoter, muscle creatine kinase promoter (MCK) and its variants, and SPc5-12 synthesis promoter. Suitable immune cell-specific promoters include, but are not limited to, the B29 promoter (B cells), the CD14 promoter (monocytes), the CD43 promoter (leukocytes and platelets), CD68 (macrophages), and the SV40 / CD43 promoter (leukocytes and platelets). Suitable blood cell-specific promoters include, but are not limited to, the CD43 promoter (leukocytes and platelets), the CD45 promoter (hematopoietic cells), INF-β (hematopoietic cells), the WASP promoter (hematopoietic cells), the SV40 / CD43 promoter (leukocytes and platelets), and the SV40 / CD45 promoter (hematopoietic cells). Suitable pancreas-specific promoters include, but are not limited to, the elastase-1 promoter. Suitable endothelial cell-specific promoters include, but are not limited to, the Fit-1 promoter and the ICAM-2 promoter. Suitable neuronal tissue / cell-specific promoters include, but are not limited to, the GFAP promoter (astrocytes), the SYN1 promoter (neurons), and NSE / RU5' (mature neurons). Suitable kidney-specific promoters include, but are not limited to, the NphsI promoter (podocytes). Suitable bone-specific promoters include, but are not limited to, the OG-2 promoter (osteoblasts, odontoblasts). Suitable lung-specific promoters include, but are not limited to, the SP-B promoter (lung). Suitable liver-specific promoters include, but are not limited to, the SV40 / Alb promoter. Suitable heart-specific promoters include, but are not limited to, α-MHC.
[0551] AAV carrier
[0552] Another aspect of this disclosure relates to an adeno-associated virus (AAV) vector comprising the IscB system described herein, wherein the AAV vector comprises DNA encoding the IscB protein and / or guide polynucleotides described herein.
[0553] In some embodiments of this disclosure, the AAV vector contains a DNA sequence encoding the IscB protein described herein. In some embodiments of this disclosure, the AAV vector contains a DNA sequence encoding the fusion protein described herein. In some embodiments of this disclosure, the AAV vector contains a DNA sequence encoding the guide polynucleotide described herein.
[0554] Delivery of CRISPR-Cas systems via AAV vectors is described in Maeder et al., Nature Medicine 25:229-233 (2019), which is incorporated herein by reference in its entirety. In some embodiments disclosed herein, the AAV vector comprises an ssDNA genome containing an RNA-guided nuclease side-linked to an ITR and a coding sequence for the guide RNA.
[0555] In some embodiments disclosed herein, the IscB protein, guide polynucleotide, IscB inactivating variant, fusion protein or conjugate containing the IscB protein, isolated nucleic acid, and / or IscB system described herein are packaged in AAV vectors, for example, packaged into AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74 capsids.
[0556] In some embodiments disclosed herein, the IscB protein, guide polynucleotide, IscB inactivating variant, fusion protein or conjugate containing the IscB protein, isolated nucleic acid and / or IscB system are packaged into an AAV2, AAV5, AAV6, AAV8, AAV9 or AAVPHP.eB capsid.
[0557] In some implementation schemes disclosed herein, the AAV carriers described herein may be selected from: AAV2 / 2, AAV2 / 3, AAV2 / 4, AAV2 / 5, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2 / 10, AAV2 / 11, AAV2 / 12, AAV2 / 13, AAV2 / PHP.B, AAV2 / PHP.B2, AAV2 / PHP.B3, AAV2 / PHP.A, AAV2 / PHP.eB, AAV2 / PHP.eS, AAV2 / 2.7m8, AAV2 / 8.7m8, AAV2 / ShH10, AAV2 / rh10, and AAV2 / rh74.
[0558] In some implementation schemes disclosed herein, the AAV carriers described herein may be selected from: AAV2 / 2, AAV2 / 5, AAV2 / 6, AAV2 / 8, AAV2 / 9, and AAV2 / PHP.eB.
[0559] In some implementations, the IscB system described herein is packaged in an AAV carrier containing an engineered capsid with tissue tropism, such as an engineered ocular tissue tropism capsid.
[0560] Lipid nanoparticles
[0561] Another aspect of this disclosure relates to lipid nanoparticles (LNPs) comprising the IscB system described herein, wherein the LNP comprises the guiding polynucleotide described herein and mRNA encoding the IscB protein described herein.
[0562] The delivery of LNPs using a CRISPR-Cas system is described in Gillmore et al., N. Engl. J. Med., 385:493-502 (2021). The lipid nanoparticles (LNPs) consist of four lipids, including a proprietary ionizable lipid LP000001; DSPC; cholesterol; and DMG-PEG2k. The LNP suspension was prepared in an aqueous buffer of Tris, NaCl, and sucrose at pH 7.4. The full text of this paper is incorporated herein by reference. In some embodiments, in addition to the RNA payload (IscB mRNA and guide polynucleotide), the lipid nanoparticles (LNPs) also contain four components: cationic or ionizable lipids, cholesterol, cofactor lipids, and PEG-lipids. In some embodiments, the cationic or ionizable lipids include cKK-E12, C12-200, ALC-0315, DLin-MC3-DMA, DLin-KC2-DMA, FTT5, Moderna SM-102, and Intellia LP01. In some embodiments, the PEG-lipids comprise PEG-2000-C-DMG, PEG-2000-DMG, or ALC-0159. In some embodiments, the accessory lipids comprise DSPC. The components of the LNP are described in Paunovska et al., Nature Reviews Genetics 23:265-280 (2022), and the FDA-approved LNP contains variants of four basic components: cationic or ionizable lipids, cholesterol, accessory lipids, and polyethylene glycol (PEG) lipids, which are incorporated herein by reference in their entirety.
[0563] Lentiviral vector
[0564] Another aspect of this disclosure relates to a lentiviral vector comprising the IscB system described herein, wherein the lentiviral vector comprises the guide polynucleotide described herein and mRNA encoding the IscB protein described herein. In some embodiments, the lentiviral vector is pseudotyped with a homologous or heterologous envelope protein such as VSV-G. In some embodiments, the mRNA encoding the IscB protein is linked to an aptamer sequence.
[0565] RNP complex
[0566] Another aspect of this disclosure relates to a ribonucleoprotein complex comprising the IscB system described herein, wherein the ribonucleoprotein complex is formed from the guide polynucleotide and IscB protein described herein. In some embodiments, the ribonucleoprotein complex can be delivered to eukaryotic cells, mammalian cells, or human cells by microinjection or electroporation. In some embodiments, the ribonucleoprotein complex can be packaged in virus-like particles and delivered in vivo to mammalian or human subjects.
[0567] Virus-like particles
[0568] Another aspect of this disclosure relates to virus-like particles (VLPs) comprising the IscB system described herein, wherein the virus-like particles comprise the guide polynucleotide and IscB protein described herein or a ribonucleoprotein complex composed of the guide polynucleotide and IscB protein.
[0569] Banskota et al. Cell 185(2):250-265 (2022) reported the development and application of DNA-free virus-like particles (eVLPs) for efficient packaging and delivery of base editors or Cas9 ribonucleoprotein; Mangeot et al., Nature Communications 10(1):1-15 (2019) induced efficient genome editing in cell lines and primary cells (including human induced pluripotent stem cells, human hematopoietic stem cells, and mouse bone marrow cells) using engineered mouse leukemia virus-like particles (Nanoblades) loaded with Cas9-sgRNA ribonucleoprotein; Campbell, et al., Molecular Therapy 27:151-163 (2019) utilized a specialized extracellular vesicle called a “gesicle” to efficiently but transiently deliver Cas9 targeting HIV long terminal repeats (LTRs) in the form of a ribonucleoprotein. Gesicles are produced by expressing vesicular stomatitis virus glycoproteins and packaging proteins (as its cargo), thus eliminating the need for transgenic delivery and allowing for more precise control over Cas9 expression. Mangeot et al. Molecular Therapy, 19(9):1656-1666 (2011) reported that overexpression of the spike glycoprotein of vesicular stomatitis virus (VSV-G) in human cells induced the release of fusion vesicles called gesicles. Biochemical and functional studies showed that glial cells bind proteins from producing cells and can transport them to recipient cells. This protein transduction method allows for the direct transport of cytoplasmic, nuclear, or surface proteins in target cells. These references all describe engineered VLPs, the full text of which is incorporated herein by reference.
[0570] In some embodiments, the engineered virus-like particles (VLPs) are pseudotyped with homologous or heterologous envelope proteins such as VSV-G. In some embodiments, the IscB protein is fused to a gag protein (e.g., MLVgag) via a cleavable linker, wherein cleavage of the linker in the target cell exposes the NLS located between the linker and the IscB protein. In some embodiments, the fusion protein or conjugate comprises (e.g., from 5' to 3') a gag protein (e.g., MLVgag), one or more NES, a cleavable linker, one or more NLS, and IscB, as described in Banskota et al. Cell 185(2):250-265 (2022).
[0571] In some embodiments, the IscB protein is fused with a first dimerizing domain that is capable of dimerizing or heterodimerizing with a second dimerizing domain fused to a membrane protein, wherein the presence of a ligand promotes the dimerization and enriches the IscB protein or fusion protein or conjugate into the VLP, as described in Campbell, et al., Molecular Therapy 27:151-163 (2019).
[0572] cell
[0573] Another aspect of this disclosure relates to cells comprising the IscB system described herein. Cells (e.g., those that can be used to generate cell-free systems) can be eukaryotic or prokaryotic. Examples of such cells include, but are not limited to, bacterial, archaea, plant, fungal, yeast, insect, and mammalian cells, such as Lactobacillus, Lactococcus, Bacillus (e.g., Bacillus subtilis), Escherichia (e.g., Escherichia coli), Clostridium, Yeast, or Pichia pastoris (e.g., Saccharomyces cerevisiae or Pichia pastoris), Kluyveromyces lactis, Salmonella typhimurium, Drosophila cells, Caenorhabditis elegans cells, Xenopus laevis cells, SF9 cells, C129 cells, 293 cells, Neurospora, and immortalized mammalian cell lines (e.g., HeLa cells, bone marrow cell lines, and lymphoid cell lines).
[0574] In some embodiments, the cells are prokaryotic cells, such as bacterial cells, such as *Escherichia coli*. In some embodiments, the cells are eukaryotic cells, such as mammalian cells or human cells. In some embodiments, the cells are primary eukaryotic cells, stem cells, tumor / cancer cells, circulating tumor cells (CTCs), blood cells (e.g., T cells, B cells, NK cells, Tregs, etc.), hematopoietic stem cells, specialized immune cells (e.g., tumor-infiltrating lymphocytes or tumor suppressor lymphocytes), or stromal cells in the tumor microenvironment (e.g., cancer-associated fibroblasts, etc.). In some embodiments, the cells are brain or neuronal cells of the central or peripheral nervous system (e.g., neurons, astrocytes, microglia, retinal ganglion cells, rod / cone cells, etc.).
[0575] Target nucleic acid or target DNA
[0576] In some embodiments disclosed herein, the target nucleic acid is target DNA.
[0577] The IscB system described herein can be used to target one or more target nucleic acid molecules, such as those present in biological samples, environmental samples (e.g., soil, air, or water samples).
[0578] In some embodiments disclosed herein, the target nucleic acid is a disease or symptom-related gene. In some embodiments disclosed herein, the target nucleic acid is a disease-related gene. In some embodiments disclosed herein, the disease-related gene is a pathogenic gene that directly causes the disease. In some embodiments disclosed herein, the disease-related gene is an abnormal gene that directly causes the disease or a gene whose expression is abnormal. For example, an unfavorable mutation in the gene leads to the occurrence of the disease. As another example, overexpression or underexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, overexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, underexpression of the gene leads to the occurrence of the disease. In some embodiments disclosed herein, overexpression of the gene is associated with the occurrence of the disease. In some embodiments disclosed herein, underexpression of the gene is associated with the occurrence of the disease.
[0579] In some embodiments disclosed herein, the disease or condition is a hematologic disease or condition, an ophthalmic disease or condition, a neurological disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, cancer, or an infectious disease.
[0580] In some embodiments disclosed herein, the diseases or conditions mentioned are selected from: hemophilia A, Best yolk-like macular dystrophy, B-cell acute lymphoblastic leukemia, hemophilia B, CDKL5 deficiency, CLN2 disease, Niemann-Pick disease type C, Dravet syndrome, FOXG1 syndrome, GM1 ganglioside storage disease, GM2 ganglioside storage disease, HIV infection, HSV infection, Usher syndrome type IB, Usher syndrome type IA, mucopolysaccharidosis type IIIA, mucopolysaccharidosis type IIIB, Gaucher disease type III, mucopolysaccharidosis type II, type II diabetes mellitus, mucopolysaccharidosis type IV, Gaucher disease type I, mucopolysaccharidosis type I, type I diabetes mellitus, Usher syndrome type I, KCNQ2 epileptic encephalopathy, Leber hereditary optic neuropathy, Leigh Prader-Willi syndrome, SLC13A5 deficiency, X-linked myotubular myopathy, X-linked retinoschisis, X-linked retinitis pigmentosa, α1-antitrypsin deficiency, α-mannosin storage disease, α-thalassemia, β-thalassemia, Alzheimer's disease, Budd-Bied syndrome, white spot retinal degeneration, leukocyte adhesion defect type I, galactosemia, bladder cancer, overactive bladder, phenylketonuria, nasopharyngeal carcinoma, Bietti lens dystrophy Pyruvate kinase deficiency, erectile dysfunction, autosomal recessive congenital ichthyosis, adult dextran disorders, traumatic arthritis, homozygous familial hypercholesterolemia, fragile X syndrome, thalassemia, hypophospholipase syndrome, epilepsy, multiple myeloma, multiple system atrophy, frontotemporal dementia, catecholamine-sensitive polymorphic ventricular tachycardia, Fabry disease, Fanconi anemia, aromatic amino acid decarboxylase deficiency, radiation-induced xerostomia, non-Hodgkin's lymphoma, non-muscle-invasive bladder cancer, non-alcoholic... Fatty liver disease, non-small cell lung cancer, hypertrophic cardiomyopathy, hypertrophic scarring, obesity, peroneal muscular atrophy type 1A, peroneal muscular atrophy type 2A, pulmonary hypertension, Friedrich's ataxia, peritoneal cancer, liver cancer, hepatocellular carcinoma, dry age-related macular degeneration, Sjögren's syndrome, hyperuricemia, hyperlipidemia, Gaucher disease, autism spectrum disorder, osteoarthritis, bone marrow failure syndrome, citrullinemia type I, coronary heart disease, cystinosis, melanoma, Huntington's disease, amyotrophic lateral sclerosis, urge incontinence, acute intermittent... Intermittent porphyria, acute lymphoblastic leukemia, spinocerebellar ataxia, spinal muscular atrophy with respiratory distress type 1, spinal muscular atrophy, familial amaurotic dementia, methylmalonic acidemia, thyroid cancer, pseudohypertrophic muscular dystrophy, anaplastic astrocytoma, intermittent claudication, borderline epidermolysis bullosa, glioma, glioblastoma, corneal transplant rejection, colorectal cancer, progressive multifocal leukoencephalopathy, progressive familial intrahepatic cholestasis, giant axonal neuropathy, canavan disease, cocaine addiction, Clabberg disease.Kriegler-Najjar syndrome, oral cancer, Happy Puppet syndrome, diffuse endogenous pontine glioma, Lafra disease, rheumatoid arthritis, sickle cell disease, lymphedema, ovarian cancer, chronic lymphocytic leukemia, chronic granulomatous disease, chronic kidney disease with anemia, chronic pain, chronic hepatitis B, Menkes disease, cystic fibrosis, Natherton syndrome, ornithine carbamoyltransferase deficiency, Parkinson's disease, Pompe disease, uveitis, prostate cancer, vestibular schwannoma, myositis, ankylosing spondylitis, castration-resistant prostate cancer, glaucoma, achromatopsia, ischemic heart failure, lysosomal storage disease, sarcoma, breast cancer, Ritter syndrome. Triple-negative breast cancer, Sandhoff's disease, color blindness, heart failure with reduced ejection fraction, neuronal ceroid lipofuscin deposition, adrenoleukodystrophy, renal cell carcinoma, wet age-related macular degeneration, eczema, thrombocytopenia with immunodeficiency syndrome, esophageal cancer, optic neuropathy, optic atrophy, retinal vein occlusion, retinitis pigmentosa, rhodopsin-mediated autosomal dominant retinitis pigmentosa, ependymoma, fallopian tube cancer, bilateral vestibular disease, Sturges' disease, diabetic macular edema, diabetic neuropathy, diabetic retinopathy, diabetic peripheral neuropathy, diabetic foot, glycogen storage disease, glycogen storage disease type Ia, glycogen storage disease Type IIb, atopic dermatitis, hearing loss, hearing impairment, head and neck cancer, squamous cell carcinoma of the head and neck, Wilson's disease, stable angina, Usher syndrome, choroidal agenesis, congenital amaurosis, congenital adrenal hyperplasia, cardiomyopathy, angina pectoris, heart failure, COVID-19 infection, pleural mesothelioma, acne vulgaris, severe combined immunodeficiency, severe limb ischemia, oculopharyngeal muscular dystrophy, pancreatic cancer, graft-versus-host disease, hereditary retinal dystrophy, hereditary angioedema, hepatitis B, metachromatic leukodystrophy, psoriatic arthritis, recessive hereditary dystrophy-type epidermolysis bullosa, infantile malignant osteosclerosis. Nutritional bullous epidermolysis, scleroderma, primary immunodeficiency, heterozygous familial hypercholesterolemia, limb-girdle muscular dystrophy type 2B, limb-girdle muscular dystrophy type 2C, limb-girdle muscular dystrophy type 2D, limb-girdle muscular dystrophy type 2E, limb-girdle muscular dystrophy type 2I, limb-girdle muscular dystrophy type 2L, limb ischemic diseases, lipoprotein lipase deficiency, severe congenital agranulocytosis, wrinkles, stroke, sciatica, schizophrenia, depression, drug addiction, autism, idiopathic pulmonary fibrosis, hyperlipidemia, thyroxine transporter protein (ATTR) amyloidosis, AATD liver disease, and AATD lung disease.
[0581] The genes related to thyroxine transporter protein (ATTR) amyloidosis include, but are not limited to, ATTR;
[0582] The genes associated with Leber hereditary optic neuropathy include, but are not limited to, MT-ND4;
[0583] The genes associated with AATD liver disease include, but are not limited to, AATD.
[0584] The genes related to AATD lung disease include, but are not limited to, AATD;
[0585] The genes related to graft-versus-host disease include, but are not limited to, thymidine kinase genes;
[0586] The genes associated with hereditary retinal dystrophy include, but are not limited to, RPE65;
[0587] The spinal muscular atrophy-related genes mentioned above include, but are not limited to, SMN1;
[0588] The genes related to osteoarthritis include, but are not limited to, TGF-β1;
[0589] The genes related to hemophilia A include, but are not limited to, factor VIII;
[0590] The genes related to hemophilia B include, but are not limited to, factor IX;
[0591] The genes associated with cystic fibrosis include, but are not limited to, CFTR;
[0592] The Parkinson's disease-related genes include, but are not limited to, Gad1, Gad2, PTBP1, KEAP1, RE1, Amigo1, Gprc5c, Let-7a, Pnky, LRRK2, SNCA gene, GBA gene, miR-92b gene, miR-9 gene, miR-124 gene, miR-181 gene, HMGB1, TRIM72, GPNMB, and REST;
[0593] The genes associated with Usher syndrome include, but are not limited to, USH2A;
[0594] The genes related to α-thalassemia, β-thalassemia, and sickle cell disease include, but are not limited to, BCL11A, HBG, HBA, and HBB;
[0595] The genes related to pulmonary hypertension include, but are not limited to, eNOS;
[0596] The genes associated with Sturgeon's disease include, but are not limited to, ABCA4;
[0597] The genes associated with age-related macular degeneration include, but are not limited to, VEGFA, VEGFR, IL17, Kir7.1, LCN-2, IRAK-M, CD59, LTA4H, GPX4, GLS1, PAPP-A, cGAS, STING, mTOR, GCN2, Nrf2, Ang2, CTGF, complement C3, complement C5, CHFR4b, DOCK6, CTSS gene, ELN gene, and FGF2;
[0598] The glaucoma-related genes include, but are not limited to, AQP1, ADRB2, NMNTA2, NRP1, Hrh1, Anxa2, OPA1, Cx43, ANGPTL7, MYOC, ROCK1, ROCK2, TIMP1, TIMP2, TIMP3, TIMP4, carbonic anhydrase CA2, carbonic anhydrase CA4, and carbonic anhydrase CA12.
[0599] The genes associated with idiopathic pulmonary fibrosis include, but are not limited to, CTGF;
[0600] The genes related to high blood lipids include, but are not limited to, PCSK9;
[0601] The genes related to Alzheimer's disease include, but are not limited to, NGF;
[0602] The genes related to coronary heart disease include, but are not limited to, VEGFA and bFGF;
[0603] The genes associated with anemia in chronic kidney disease include, but are not limited to, EPO;
[0604] The genes associated with congenital amaurosis include, but are not limited to, RPE65;
[0605] The genes associated with retinitis pigmentosa include, but are not limited to, PDE6B.
[0606] The genes associated with phenylketonuria include, but are not limited to, PAH; and / or
[0607] The epilepsy-related genes include, but are not limited to, GAT1.
[0608] Non-limiting examples of such target nucleic acids also include those listed in U.S. Provisional Patent Applications 61 / 736,527 and 61 / 748,427, filed December 12, 2012 and January 2, 2013, respectively, and International Application PCT / US2013 / 074667, filed December 12, 2013, all of which are incorporated herein by reference.
[0609] In some embodiments, the target nucleic acid is a reporter gene. Examples of reporter genes include, but are not limited to, glutathione S-transferase (GST), horseradish peroxidase (HRP), chloramphenicol acetyltransferase (CAT), β-galactosidase, β-glucuronidase, luciferase, green fluorescent protein (GFP), HcRed, DsRed, cyan fluorescent protein (CFP), yellow fluorescent protein (YFP), and autofluorescent proteins including blue fluorescent protein (BFP).
[0610] Applications for treating or preventing diseases
[0611] Another aspect of this disclosure relates to a pharmaceutical composition comprising an IscB protein as described herein, a guide polynucleotide as described herein, an inactivating variant of IscB as described herein, a fusion protein or conjugate as described herein, a nucleic acid as described herein, an IscB system as described herein, a vector system as described herein, a delivery system as described herein, or a cell as described herein. The pharmaceutical composition may comprise, for example, an AAV vector encoding the IscB protein or an inactivating variant of IscB as described herein and a guide polynucleotide. The pharmaceutical composition may comprise, for example, lipid nanoparticles containing the guide polynucleotide as described herein and mRNA encoding the IscB protein. The pharmaceutical composition may comprise, for example, a lentiviral vector containing, for example, the guide polynucleotide as described herein and mRNA encoding the IscB protein. The pharmaceutical composition may comprise, for example, virus-like particles containing, for example, the guide polynucleotide as described herein and the IscB protein, or a ribonucleoprotein complex formed from said guide polynucleotide and IscB protein.
[0612] Another aspect of this disclosure relates to the use of IscB proteins, guide polynucleotides, inactivating variants, fusion proteins or conjugates, nucleic acids, IscB systems, vector systems, delivery systems, cells, pharmaceutical compositions, or kits in cutting or editing target nucleic acids in mammalian cells.
[0613] Another aspect of this disclosure relates to the use of IscB proteins, guide polynucleotides, inactivating variants, fusion proteins or conjugates, nucleic acids, IscB systems, vector systems, delivery systems, cells, pharmaceutical compositions, or kits as described herein for any of the following purposes: cleaving or creating nicks in one or more target nucleic acid molecules; activating or upregulating the expression of one or more target nucleic acid molecules; activating or inhibiting the transcription of one or more target nucleic acid molecules; inactivating one or more target nucleic acid molecules; visualizing, labeling, or detecting one or more target nucleic acid molecules; binding one or more target nucleic acid molecules; transporting one or more target nucleic acid molecules; and masking one or more target nucleic acid molecules.
[0614] Another aspect of this disclosure relates to the use of modifying one or more target nucleic acid molecules with IscB proteins, guide polynucleotides, inactivating variants, fusion proteins or conjugates, nucleic acids, IscB systems, vector systems, delivery systems, cells, pharmaceutical compositions, or kits as described in this disclosure, including one or more of the following: nucleic acid base substitution, nucleic acid base deletion, nucleic acid base insertion, target nucleic acid breakage, nucleic acid methylation, and nucleic acid demethylation.
[0615] Another aspect of this disclosure relates to the use of IscB proteins, guide polynucleotides, inactivating variants, fusion proteins or conjugates, nucleic acids, IscB systems, vector systems, delivery systems, cells, pharmaceutical compositions, or kits as described in this disclosure in the diagnosis, treatment, or prevention of diseases or conditions associated with target nucleic acids.
[0616] Another aspect of this disclosure relates to the use of IscB proteins, guide polynucleotides, inactivating variants, fusion proteins or conjugates, nucleic acids, IscB systems, vector systems, delivery systems, cells, pharmaceutical compositions, or kits as described in this disclosure in the preparation of medicaments for the diagnosis, treatment, or prevention of diseases or conditions associated with target nucleic acids.
[0617] In some embodiments, the pharmaceutical composition is delivered in vivo to a human subject. The pharmaceutical composition can be delivered via any effective route. Exemplary routes of administration include, but are not limited to, intravenous infusion, intravenous injection, intraperitoneal injection, intramuscular injection, intratumoral injection, subcutaneous injection, intradermal injection, intraventricular injection, intravascular injection, cerebellar injection, intraocular injection, subretinal injection, intravitreal injection, intra-anterior chamber injection, intratympanic injection, intranasal administration, and inhalation.
[0618] Diagnostic applications
[0619] Another aspect of this disclosure relates to an in vitro composition comprising the IscB system described herein and detector DNA that cannot be used with the markers for guiding polynucleotide hybridization described herein.
[0620] Another aspect of this disclosure relates to the use of the IscB system described herein in detecting target nucleic acids in nucleic acid samples suspected of containing target nucleic acids.
[0621] Another aspect of this disclosure relates to the use of the IscB system described herein in detecting target nucleic acids in nucleic acid samples containing target nucleic acids.
[0622] In some implementations, the target nucleic acid to be detected is the target RNA.
[0623] In some embodiments, the target nucleic acid to be detected is target DNA. In some embodiments, methods for detecting target DNA include an IscB protein fused to a fluorescent protein or other detectable marker, and a guide polynucleotide containing a guide sequence specific to the target DNA. The binding of IscB to the target DNA can be visualized by microscopy or other imaging methods.
[0624] In some implementations, methods for detecting target nucleic acids in cell-free systems result in the generation of detectable markers or enzyme activity. For example, by using an IscB protein, a guide polynucleotide containing a guide sequence specific to the target nucleic acid, and a detectable marker, the target nucleic acid will be recognized by IscB. Binding of IscB to the target nucleic acid triggers its DNase activity, which leads to the cleavage of both the target nucleic acid and the detectable marker.
[0625] In some implementations, the detectable marker is DNA linked to a fluorescent probe and a quencher. The intact detectable DNA is linked to the fluorescent probe and the quencher, suppressing fluorescence. After the detectable DNA is cleaved by IscB, the fluorescent probe is released from the quencher and exhibits fluorescent activity. This method can be used to determine the presence of target DNA in lysed cell samples, lysed tissue samples, blood samples, saliva samples, environmental samples (e.g., water, soil, or air samples), or other lysed cell or cell-free samples. This method can also be used to detect pathogens, such as viruses or bacteria, or to diagnose disease states, such as cancer.
[0626] In some implementations, the detection of target nucleic acids helps diagnose diseases and / or pathological conditions, or the presence of viral or bacterial infections.
[0627] Example
[0628] The present disclosure is further illustrated below by way of examples, but these examples are not intended to limit the scope of the disclosure to the examples described. Experimental methods in the following examples that do not specify specific conditions were performed according to conventional methods and conditions, or as selected in accordance with the product instructions.
[0629] Through bioinformatics analysis and experimental verification, the inventors screened and obtained several new IscB proteins with DNA cleavage capabilities.
[0630] Experimental Example 1: Screening of IscB Protein
[0631] Through complex bioinformatics methods, multiple active IscB proteins were ultimately screened and verified.
[0632] The amino acid sequences of different IscB proteins are as follows:
[0633] >CTnB-57
[0634] >CTnB-62
[0635] The backbone sequence of the ωRNA used in conjunction with the CTnB-57 protein is as follows:
[0636] The backbone sequence of the ωRNA used in conjunction with the CTnB-62 protein is as follows:
[0637] Experiment Example 2: Verifying TAM Composition Based on the In Vivo Editing Activity of IscB Protein in Bacteria
[0638] In this embodiment, a plasmid library containing a 7-nt random sequence was first constructed, and then bacterial expression plasmids encoding different IscB protein sequences were constructed. After the expression plasmids were transformed into bacteria to prepare competent cells, the 7-nt random sequence plasmid library was electroporated. If the plasmid in the 7-nt random sequence library could be recognized and targeted by the IscB protein, it would be removed from the library, and the bacteria would not grow. The specific steps are as follows:
[0639] 1.7nt random sequence plasmid library construction
[0640] The pLVX-EF1a-BSD (SEQ ID NO:1) vector plasmid was double-digested with EcoRV and XhoI, and the linearized vector was recovered by agarose gel electrophoresis. Using the prepared pCDH-CMV-EGFP-reporter3-EF1-Puro (SEQ ID NO:2) plasmid as a template, the DNA fragment containing the coding sequence of the Puro resistance gene was amplified by PCR using primers Puro-PF1 (SEQ ID NO:3) and Puro-PR1 (SEQ ID NO:4). Homologous recombination (NEB, Gibson) was then employed. The Master Mix was inserted into the enzyme-digested pLVX-EF1a-BSD vector to construct the recombinant vector pLVX-7NN-Puro library plasmid (SEQ ID NO:5). The reaction solution was transformed into Stbl3 competent cells, plated on LB agar plates containing ampicillin, and incubated overnight at 37°C. All colonies were scraped off, and plasmids were extracted to prepare a library plasmid containing the 7NN random sequence.
[0641] 2. Synthesis of bacterial expression plasmids for different IscB proteins and preparation of competent cells containing these plasmids
[0642] a. Synthesis of bacterial expression plasmids for different IscB proteins
[0643] After codon optimization, the IscB protein obtained from bioinformatics analysis was sent to a synthesis company for whole-genome synthesis, resulting in a bacterial expression plasmid of the IscB protein. The specific synthetic sequence is as follows:
[0644] >P15A-HDV-CTnB-57(SEQ ID NO:143)
[0645] >P15A-HDV-CTnB-62(SEQ ID NO:144)
[0646] The underlined markers represent the base sequence of the IscB protein, the bold italics represent the ωRNA sequence, and the bold italics with an underline represent the guide sequence.
[0647] b. Preparation of competent cells containing bacterial expression plasmids of different IscB proteins
[0648] Different IscB protein bacterial expression plasmids were transformed into DH5a competent cells, and single clones were selected and inoculated into LB medium containing chloramphenicol and cultured overnight at 37°C.
[0649] c. Plasmid elimination to identify the IscB protein PAM sequence
[0650] 100 ng of pLVX-7NN-Puro library plasmid was used to electroporate competent cells of various IscB protein bacterial expression plasmids and DH5α cells, which were labeled as Lib1 (competent cells electroporated with IscB protein bacterial expression plasmids) and Lib2 (competent cells electroporated with DH5α cells), respectively.
[0651] After electroporation, add 10 ml of LB medium and incubate at 37℃ and 220 rpm for 2 h to recover and culture.
[0652] After resuscitation, the bacterial culture was centrifuged at 4000 rpm for 2 min to collect the bacterial cells. After resuspending in 400 μL LB, the cells were spread on LB plates. The bacterial culture of DH5α was spread on LB plates containing ampicillin resistance, and the competent cells of various IscB protein bacterial expression plasmids were spread on LB plates containing chloramphenicol and ampicillin resistance. The cells were incubated overnight at 37°C.
[0653] Bacterial cells were scraped from the culture plate and plasmid DNA was extracted using the alkaline lysis method.
[0654] 100 ng of each of the two extracted plasmid DNAs was used as PCR templates and PCR amplification was performed using primers SiteSeq-PF1 and SiteSeqPuro-PR. The obtained fragments were used to construct an amplicon library using an NGS library construction kit (SynplSeq DNA Library Prep Kit for Illumina). For detailed library construction procedures, please refer to the kit instructions. The constructed library was then subjected to NGS sequencing.
[0655] The primer sequences for amplicon library construction are as follows:
[0656] SiteSeq-PF1: cagcacaaaaggaaactcaccctaactg (SEQ ID NO: 6).
[0657] SiteSeqPuro-PR: TGTGGCGCGTGGCGGGGTAG (SEQ ID NO: 7).
[0658] By comparing the differences in NGS sequencing of Lib1 and Lib2 cells after transformation with different IscB proteins, the TAM reads showing the differences were obtained and analyzed. The motifs were extracted to form the TAM composition. The TAM motifs of CTnB-57 and CTnB-62 proteins are shown in Figures 1 and 2, and the specific base proportions of each motif are shown in Tables 1 and 2.
[0659] Table 1: Base percentage of TAM Motifs captured by CTnB-57 plasmid elimination
[0660] Table 2: Base percentage of TAM Motifs captured by CTnB-62 plasmid elimination
[0661] Based on the TAM motif and the proportion of each base in the specific motif, the CTnB-57 TAM sequence is obtained as WHG, and the CTnB-62 TAM sequence is obtained as DHG (W = A or T; H = A, C or T; D = A, G or T).
[0662] Experiment Example 3: Verification of the prokaryotic editing activity of CTnB-57 and CTnB-62 proteins
[0663] In this embodiment, a targeting plasmid containing a single TAM sequence was first constructed. Then, the targeting plasmid containing the single TAM sequence was introduced into bacterial competent cells containing IscB protein expression plasmids for detection via electroporation. If the introduced plasmid containing the specific TAM sequence can be recognized and targeted by the IscB protein, it will be removed from the library, and the bacteria will be unable to grow on the antibiotic-protected plate. The specific steps are as follows:
[0664] 1. Construction of a single TAM plasmid
[0665] The pUC57Kan vector plasmid (SEQ ID NO:8) was linearized by NdeI+XhoI restriction enzyme digestion. Primers containing the target site and different TAM sequences were synthesized and annealed. The resulting vector was then annealed with the NdeI+XhoI-digested linearized vector and ligated to obtain the TAM clone. The synthesized primers and TAM information are shown in Table 3.
[0666] Table 3: Primer information for constructing plasmids with different TAM sequences
[0667] The TAM-01-pUC57K plasmid sequence is shown below. The other TAM-related plasmid sequences, except for the TAM sequence, are identical to TAM-01-pUC57K.
[0668] >TAM-01-pUC57K(SEQ ID NO:145)
[0669] The text in bold italics with an underline is the target sequence, and the text in bold italics is the TAM sequence.
[0670] 2. Electroporation of a single TAM plasmid into competent cells of the target protein to verify editing activity.
[0671] Competent cells containing IscB protein prepared in Example 1 were transfected with a single TAM plasmid. Specifically, CTnB-57 competent cells were electroporated with TAM-02-pUC57K, TAM-05-pUC57K, TAM-06-pUC57K, and TAM-07-pUC57K plasmids, respectively. The specific transformation steps were the same as in Example 2. The transformed and recovered products were plated on plates containing kanamycin and / or chloramphenicol and incubated overnight at 37°C. If the TAM plasmid sequence is consistent with the TAM recognized by IscB, the bacteria will not grow on the plate. The editing ability of IscB on bacteria can be determined by observing the bacterial growth on the plate. Specifically, after CTnB-57 cells were transfected with a single TAM plasmid, the transformed and recovered products were divided into three aliquots. Figure 3 shows the first column of plates coated with chloramphenicol and kanamycin, the second column with chloramphenicol, and the third column with kanamycin. Bacterial growth was observed, as shown in Figure 3. Since CTnB-57 recognizes the TAM sequence as WHG, which is consistent with TAM-05 and TAM-07, bacteria do not grow on TAM-05 and TAM-07 plates, as expected.
[0672] Using the same identification method as CTnB-57, TAM-02-pUC57K, TAM-05-pUC57K, TAM-06-pUC57K, and TAM-07-pUC57K plasmids were electroporated into CTnB-62 competent cells. The recovered bacterial culture was divided into four portions and plated on chloramphenicol, kanamycin, and a combination of chloramphenicol and kanamycin antibodies, respectively. The editing ability of the plasmids was assessed by observing bacterial growth. Specific growth data is shown in Figure 4. Based on the bacterial growth, the two DHG-related TAM plasmids did not result in bacterial growth after transformation, consistent with expectations.
[0673] Experiment Example 4: Verification of the Single Base Editing Capabilities of CTnB-57 and CTnB-62
[0674] In this embodiment, the RuvC Domain of CTnB-57 and CTnB-62 proteins was first inactivated to obtain a nickase version of CTnB. Then, a deaminase was fused, and a target site was designed on the E. coli LacZ gene based on the specific recognition TAM, and a single-base editing test was performed.
[0675] 1. Synthesis of CTnB-57 and CTnB-62 single-base editing plasmids
[0676] Bioinformatics analysis confirmed that mutating amino acid D to A at position 58 of the CTnB-57 protein and mutating amino acid D to A at position 58 of the CTnB-62 protein yields nickase versions. Based on the recognition TAM of these two proteins, the sequence GCTCACATTTAATGTTG (SEQ ID NO:166) was selected as the editing target on the lacZ gene. A vector was designed after fusing cytosine deaminase, and a plasmid was synthesized. The specific synthesized sequence is as follows:
[0677] >P15A-CTnB-57-01(SEQ ID NO:146)
[0678] >P15A-CTnB-62-01(SEQ ID NO:147)
[0679] The underlined markers are the base sequences of the nickase version of the IscB protein, the bold italics are the ωRNA sequences, and the bold italics with underlines are the target sequences.
[0680] 2. Verification of single-base editing functions of CTnB-57 and CTnB-62
[0681] Two gene-synthesized plasmids, P15A-CTnB-57-01 and P15A-CTnB-62-01, were transformed into TOP10-F' competent cells after synthesis. The resulting bacterial culture was incubated overnight at 37°C with chloramphenicol added. Primers designed based on the LacZ sequence were used for PCR amplification, followed by Sanger sequencing. The single-base editing efficiency was analyzed using TIDE.
[0682] The specific amplification and sequencing primers are shown in Table 4. Experimental results show that CTnB-62-01 can target sequences in the prokaryotic genome. The fourth base C (bold & underline) is 100% edited to T, indicating single-base editing activity. CTnB-57-01 has a relatively low editing efficiency at this C position, approximately 4%.
[0683] Table 4: Primer information for LacZ target sequence amplification and sequencing
[0684] Experimental Example 5: Screening of IscB Protein
[0685] CTnB-30 and CTnB-50 were tested using essentially the same methods as in Examples 2 and 3 above.
[0686] The amino acid sequences of different IscB proteins are as follows:
[0687] >CTnB-30
[0688] >CTnB-50
[0689] The backbone sequence of the ωRNA used in conjunction with the CTnB-30 protein is as follows:
[0690] The backbone sequence of the ωRNA used in conjunction with the CTnB-50 protein is as follows:
[0691] The bacterial expression plasmid sequence is as follows
[0692] >P15A-HDV-CTnB-30(SEQ ID NO:150):
[0693] >P15A-HDV-CTnB-50(SEQ ID NO:151)
[0694] The underlined markers represent the base sequence of the IscB protein, the bold italics represent the ωRNA sequence, and the bold italics with an underline represent the target sequence.
[0695] By comparing the differences in NGS sequencing of Lib1 and Lib2 cells after transformation with different IscB proteins, the differential TAM reads were obtained and analyzed. The motifs were extracted to form the TAM composition. The TAM motifs of CTnB-30 and CTnB-50 proteins are shown in Figures 5 and 6, and the specific base proportions of each motif are shown in Tables 5 and 6.
[0696] Table 5: Base percentage of TAM Motifs captured in the CTnB-30 plasmid elimination experiment
[0697] Table 6: Percentage of TAM Motif Bases Captured in CTnB-50 Plasmid Elimination Experiment
[0698] Based on the TAM motif image and the proportion of each base in the specific motif, the TAM sequence recognized by CTnB-30 is ATAAA, and the TAM sequence recognized by CTnB-50 is ATG.
[0699] Validation of prokaryotic editing activity of CTnB-30 and CTnB-50 proteins
[0700] Competent cells containing the IscB protein prepared in Example 1 were transfected with a single TAM plasmid. Specifically, TAM-05-pUC57K, TAM-06-pUC57K, and TAM-07-pUC57K plasmids were electroporated into CTnB-30 competent cells, following the same transformation steps as in Example 1. The transformed and revived products were then plated on double-antibody plates containing kanamycin and chloramphenicol and incubated overnight at 37°C. The results are shown in Figure 7. Because CTnB-30 recognizes the TAM sequence ATAAA, which is consistent with TAM-06, bacteria did not grow on the TAM-06 plate, as expected.
[0701] TAM-06 plasmid was used to transform CTnB-30 competent cells. The transformation and recovery products were divided into 3 portions and transformed onto plates coated with chloramphenicol and / or kanamycin, respectively. The bacterial growth was observed, as shown in Figure 8.
[0702] The plasmids TAM-05-pUC57K, TAM-06-pUC57K, and TAM-07-pUC57K were electroporated into CTnB-50 competent cells, respectively. The recovered bacterial cultures were divided into three aliquots and plated on chloramphenicol, kanamycin, and a combination of chloramphenicol and kanamycin, respectively. The specific growth status is shown in Figure 9. Bacteria did not grow on the two TAM plates associated with ATG, as expected.
[0703] Experimental Example 6: Verification of Eukaryotic Editing Activity of CTnB-30
[0704] In this embodiment, the CTnB-30 protein expression framework and ωRNA expression framework were first inserted into the eukaryotic expression plasmid pCDNA3.1(+). Then, based on the TAM sequence corresponding to CTnB-30, a target sequence was designed for the human TTR gene (Gene ID: 7276). A plasmid expressing CTnB-30 protein and a specific target ωRNA was constructed and transfected into HEK293 cells for eukaryotic activity verification.
[0705] 1. Construction of CTnB-30 protein and corresponding ωRNA eukaryotic expression plasmid
[0706] pCDNA3.1(+) was digested with HindIII and XbaI, the plasmid was linearized, and the vector fragment V1 was obtained by gel recovery.
[0707] Gene synthesis of U6-ωRNA expression framework fragment F1.
[0708] 2. Primers ChkSV40NLS-PF2 (SEQ ID NO:10) and CTnBNLS-PR1 (SEQ ID NO:11) were designed and synthesized. CTnB-30 fragment F2 was amplified using P15A-HDV-CTnB-30 as a template. Vector V1 underwent homologous recombination with fragments F1 and F2, and after transformation, CTnB-30 and the ωRNA eukaryotic expression plasmid CTnB-30-pCDNA3.1 (SEQ ID NO:12) were obtained.
[0709] 3. Construction of CTnB-30 targeting hTTR (Gene ID: 7276) gene plasmid
[0710] The CTnB-30-pCDNA3.1 plasmid was digested with BsmBI enzyme, linearized, and gel recovered to obtain vector V2. Primers were designed based on the selected target sequence and annealed to V2 for ligation. The vector was then transformed into E. coli to obtain expression plasmids containing different target sequences. Specific primer and target sequence information is shown in Table 7.
[0711] Table 7: Primer information for constructing plasmids with different target sequences
[0712] Two different target plasmids were transfected into HEK293T cells using PEI. After 48 hours, the cells were collected and lysed using DirectPCR Lysis Reagent (Cell) (VIAGEN: 302-C). Different primers were selected according to different target sites (Table 8) for PCR amplification. Sanger sequencing was performed, and the editing efficiency was analyzed using TIDE (2.3% and 4.3%, respectively), as shown in Figures 10 and 11.
[0713] Table 8: Primer information for amplification and sequencing after transfection of different target sequences
[0714] Experiment 7: Testing of other IscB proteins
[0715] Other IscB proteins were tested using essentially the same methods as in Examples 2 and 3 above, as shown in Table 9 and Figure 12.
[0716] Table 9. TAM of other IscB proteins determined experimentally.
[0717] Experiment Example 8: Identification of TAM based on the bacterial in vivo editing activity of IscB protein
[0718] Vectors for CTnB-152 and CTnB-198 were constructed and tested using methods essentially the same as those used in Experiments 2 and 3 above.
[0719] >CTnB-152
[0720] >CTnB-198
[0721] ωRNA backbone sequence of CTnB-152:
[0722] ωRNA backbone sequence of CTnB-198:
[0723] The bacterial expression plasmid sequences of CTnB-152 and CTnB-198 are as follows:
[0724] >P15A-HDV-CTnB-152
[0725] >P15A-HDV-CTnB-198
[0726] The underlined markers represent the base sequence of the IscB protein, the bold italics represent the ωRNA sequence, and the bold italics with an underline represent the guide sequence.
[0727] Preparation of competent cells containing bacterial expression plasmids of different IscB proteins: bacterial expression plasmids of different IscB proteins synthesized by genes were transformed into DH5a competent cells, and single clones were picked and inoculated into LB medium containing chloramphenicol and cultured overnight at 37°C.
[0728] Plasmid elimination to identify the PAM sequence of the IscB protein
[0729] 100 ng of pLVX-7NN-Puro library plasmid was electroporated into competent cells containing various IscB protein bacterial expression plasmids and DH5α cells, respectively, and labeled as Lib1 (competent cells electroporated with IscB protein bacterial expression plasmids) and Lib2 (competent cells electroporated with DH5α cells). Cells were cultured, and amplicon libraries were constructed.
[0730] By comparing the differences in NGS sequencing of Lib1 and Lib2 cells after transformation with different IscB proteins, the differential TAM reads were obtained and analyzed. The motifs were extracted to form the TAM composition. The TAM motifs of CTnB-152 and CTnB-198 proteins are shown in Figures 13 and 14, and the specific base proportions of each motif are shown in Tables 10 and 11.
[0731] Table 10: Base percentage of TAM Motifs captured by CTnB-152 plasmid elimination
[0732] Table 11: Base percentage of TAM Motifs captured by CTnB-198 plasmid elimination
[0733] The TAM sequence identified by CTnB-152 is AYGG, while the TAM sequence identified by CTnB-198 is NGG.
[0734] Table 12. Primer information for constructing plasmids with different TAM sequences
[0735] The TAM-08-pUC57K plasmid sequence is shown below. The other TAM-related plasmid sequences, except for the TAM sequence, are identical to TAM-08-pUC57K.
[0736] >TAM-08-pUC57K(SEQ ID NO:9)
[0737] Bold italics with underlines indicate the guide sequence, while bold italics indicate the TAM sequence.
[0738] The competent cells containing the IscB protein prepared in this embodiment were transfected with a single TAM plasmid. Specifically, the CTnB-152 competent cells were electroporated with TAM-08-pUC57K and TAM-09-pUC57K plasmids, respectively. The electroporated and revived bacterial cultures were divided into three aliquots and plated with chloramphenicol and kanamycin, respectively, and with both chloramphenicol and kanamycin. CTnB-152 recognized the TAM sequence AYGG, which is consistent with TAM-09-pUC57K. Consequently, no bacteria grew on the corresponding double-antibody plates, as expected.
[0739] TAM-08-pUC57K and TAM-09-pUC57K plasmids were electroporated into CTnB-198 competent cells, respectively. The recovered bacterial cultures were divided into three portions and plated onto chloramphenicol, kanamycin, and chloramphenicol and kanamycin plates, respectively. The editing ability of the plasmids was assessed by observing bacterial growth. The results showed that bacteria transformed with NGG-related TAM plasmids did not grow on the dual-antibody plates, while bacteria transformed with TAM plasmids that did not conform to NGG continued to grow on the dual-antibody plates, consistent with expectations.
[0740] sequence list
Claims
1. An IscB protein, characterized in that, The amino acid sequence of the IscB protein comprises or is an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of SEQ ID NO:13-26, 41, 42, 53-93; Optionally, the IscB protein retains the function of the protein shown in any one of the sequences SEQ ID NO:13-26, 41, 42, 53-93; Optionally, the TAM sequence (5'→3') recognizable by the IscB protein may be selected from any one or more of the following: A, C, T, G TA, TC, GN, AA, AG, TG, AN, GG, CG, TN, NT, NG, GT, NA, CC, AC, GC, AT, CT, GA, TT, CN, NC, CA, NTN, ANN, TTN, ATC, NAC, AGA, TGC, TCT, NGN, CGC, NTC, GCA, TCG, TTT, CCG, GGG, NAG, ACA, CGG, CNG, ACN, GTG, CNT, TTG, TCN, GGT, TNC, CCN, CGT, TGG, CGA, NGG, TCC, AGT, NCA, CAN, TCA, NNG, TAC, CCT, NTG, CGN, TGN, CAT, NGC, GNG, GNC, NNA, GAA, TTC, CTT, ATA, TAT, GCT, NCC, TTA, AGN, GNN, CAA, CAC, AGG, NTT, ANG, GNA, GTT, NGA, TAA, GTA, GGN, GNT, NCG, ATT, CCA, CNN, AAA, AAC, ATN, GAG, CTG, ACG, NAA, TAN, NAT, CNA, GCN, GTC, NCN, CTN, CNC, ANT, NNC, CAG, NAN, ATG, NCT, CCC, AAN, TGT, TNA, ACC, GAT, ACT, AAT, GGA, GAN, ANC, GAC, NNT, CTA, TNN, GCG, GTN, TNT, AAG, TAG, NGT, NTA, ANA, CTC, GCC, TGA, GGC, AGC, TNG, NGAA、GANC、GCNC、NTNT、TGGG、AAGG、AAGN、NTNN、TCGT、CNTG、NTGG、CCGN、AT AT、TGCA、NGGT、TGNT、NNTG、NCCG、ACAT、GNTG、CGCG、GACN、NTCG、TCNG、CTGC、 TNNC, GGTN, CGNN, TCCA, AGCN, TNAG, GGAC, GATC, AANA, NATG, CCAG, NAAT, TCNT, CACT, CGGC, CGAN, CNCA, ATNT, NNNG, NGCT, CTGG, GGAN, NTNC, ATTC, AATG CNTC、TGGN、NATC、GTCG、ACNC、GCNN、GACT、CTNT、NCTT、NAGG、NANC、CTTA、GT CT、ANAG、NGCN、CNNA、TCAG、ACAC、NCGG、TNNT、CAAG、ACCT、CCCA、GTNC、ANTC、 GACC、AACG、TTAA、TCCG、CGCC、NCCN、TTNA、NCNT、NGCA、AGNN、AATC、GGGA、GN AN、NAGA、CGNA、GTAT、GTNA、ATNC、ACNA、GGAA、NTCC、GGCG、AATN、CNNT、AGGC、 GCGN、GTGC、TTGA、AAGC、GAAG、ATNG、TGCT、TACT、CTAN、GGCT、GNGC、GTCN、CG AA、CNAC、GCCT、TAGG、ANGC、TNAA、GANT、NCNA、NCCT、AGAN、GTAA、TTTN、ATGA、 TGNA, CANC, ACGA, CCAC, CCGG, CTNG, CNGN, GGTA, NGNC, GTTT, CTAA, TNCT, CTGN, NGAC, TGTA, TANN, GCNT, GCTC, CNCG, AAAN, CCNT, GANA, CACA, CTNA, ANTN TTNT, CCTG, TNTT, CANA, NTAN, CACG, GGAT, TTTC, GNCG, TACA, GTAC, GAGC, ACNN, ATGG, AANT, ATCC, ACCG, AGNC, TGTT, NCAT, ATTA, GNTT, GAGN, TNAC, GCCG NTNG, GTGG, GNGN, ACCA, NTAA, ACTN, NCTG, NCTA, TTTT, GCNG, NTAG, CAAA, GGNA, CNTN, TTAG, TCTG, NCTN, TATG, GGCGT, TANT, GGGT, NACN, ACTG, CCNG, GNNTCCAT, GNTA, NANT, TACN, TGTN, ATCT, NCAN, TNGG, CNNN, AAGT, ATTN, GGNN, CAGC, CGTN, GCCC, GCTT, CNAT, NANA, CCNN, GNGA, TNGN, GCAG, CGNG, CCTT, NGAG NCNG、AANG、GGTC、ACTC、TGAA、NAGN、NNCA、ACGG、TGAC、TCCN、ANNN、TCGN、TA AN、CAGG、TTAN、NGAN、NTGC、CCNC、TNTN、ATGN、GTGN、GCAT、NNGN、NNCC、CCNA、 CNAG, GNAC, CGNT, TTCN, TAGN, ANCT, NATN, GTGA, TNGT, CTAT, CCCG, TNCA, NGTA, NNGA, CGTG, TAAT, CGCA, NNCG, NGTC, NAGT, GNAT, TNTC, NCGC, NGGN, CATN GTTN, AGTA, GNNG, TTNN, TGNC, NAAA, TNCC, CACC, CTCT, TTGN, GCTA, NTTT, TGAN, TNAN, NGAT, CCTN, GAAT, GTCA, NTCN, GCCA, ANTG, TGGC, CAAC, TTTA, TGTC CGGA, NCGN, AGNT, NCGA, ANCG, ACAA, TAGT, CGAG, NCAA, AATA, AGGG, GNGT, CAGA, AGGT, GGGG, ANAC, TGGT, GTGT, GNCA, GTTA, NGTT, TNNG, NCAG, CACN, GCAN GAAC、NCCA、TTCC、NCNN、GNNN、ANGT、NTNA、CCCT、GNAA、TTNG、GTNN、GGNG、TC TA、NCAC、GANG、TTCG、CCTC、CNGG、ANNA、TCAN、ATCG、NTGA、CGTA、TTAC、GCTN、 GCTG, NGTG, TCCC, CANN, NNNA, TAGA, ACGT, AGAT, GATG, GCCN, TGNG, GCGC, CCGA, GNCN, NTTG, NNAT, TNCG, NANG, GGTG, NCCC, GNCC, CAAT, CGCN, CNGA, NTTC TTCT, NGGA, AGTC, CNNC, NACG, AGTN, NANN, ACAG, GNCT, TACC, CNTA, TGTG, CATC, GACA, TCTT, NTCT, CTGA, AGGA, GATA, TNAT, CCTA, GGAG, ANCC, AANC, GTANGCNA, TGNN, TANC, GNTN, AGCG, CTAG, NNAA, AGTT, CTAC, TACG, TTNC, TNTA, ANTT, ATAC, TCCT, TCAC, NGGC, NTTN, NNTC, CANT, ATAA, TGCC, CTCC, TNNA, GTNG ACGN、GGCA、AAAG、TTGT、NGNA、NAAN、TATN、CGGG、CATA、ATGC、ACGC、ACCN、AT TT、TCNA、TNGC、NACA、NACC、CTCN、GGCC、TANG、AGAA、TNGA、TAGC、CAGN、GGCN、 ANNT, NNNC, TCAT, CATT, TAAA, ATGT, TGAG, CGCT, TCGG, GCAC, GTAG, NTCA, NATT, ANTA, CCCN, ACTA, AAAA, GAAN, TATT, NNAC, TGAT, GGGN, CCAA, GNGG, CCAN GTCC, NNCT, AGNG, CNTT, CNCT, GANN, GGTT, AGCT, CATG, NTAC, TNCN, NNTN, TGGA, GATT, AGCA, TAAG, GCGA, ACTT, ANGN, NTGN, AACN, AACT, TCAA, NTAT, TCGA NCTC, NNNGG, ANGG, NNTT, GTNT, CTNN, CGGN, TAAC, GGNC, GAAA, ACNG, GNAG, TTGG, CTTC, CNGT, TNNN, TNTG, GTTG, TCNN, CGGT, GAGA, CNNG, NCNC, GAGG, AGCC ATNN、NNNT、AGAC、AACC、ANNC、ANNG、ACAN、GTTC、TATA、GNTC、NCGT、NGNT、CG TC、CCGC、CGAC、GACG、ATTG、GNNC、CNAA、TATC、AGNA、CTNC、TTCA、ANCA、ACCC、 AGTG、CCGT、ANAT、CTGT、GGGC、NTTA、NAAG、AANN、CNAN、NNCN、ANAA、ANAN、CT TG、NGNN、AGAG、TANA、TCNC、GCAA、NGNG、NAGC、NATA、ATCN、CGTT、CNGC、GATN、 NNTA, AAGA, CTTT, AAAC, AGGN, ACNT, NTGT, CTTN, ATCA, NACT, NNAG, NGTN, NAAC, TGCG, GGNT, ATAN, TTGC, ANCN, CCCC, ANGA, NGCG, TCTC, CTCG, ATNA, AATTNNAN, NNGT, TCGC, ATAG, CAAN, AACA, TTAT, CAGT, GNNA, TGCN, GGCG, NGGG, CANG, TTTG, GAGT, AAAT, CTCA, CNCN, CNCC, TCTN, CGNC, NGCC, CGAT, NNGC AAAAA、AAAAC、AAAAG、AAAAT、AAACC、AAACT、AAAGC、AAAGG、AAAGT、AAATA、AAATT、AACAT、AACTC、AACTG、AAGAG、AAGAT、AAGCC、AAGGA、AAGTA、AATAA、AATA C、AATAG、AATAT、AATCT、AATGA、AATTA、AATTC、AATTG、AATTT、ACAAA、ACAAC、ACAAT、ACACA、ACACC、ACACT、ACATA、ACATC、ACATG、ACATT、ACATT、ACCAC、ACCA CG、ACCTT、ACGAT、ACGCC、ACGGC、ACTGA、ACTGC、ACTTT、AGAAG、AGACA、AGACC、AGAGG、AGATA、AGCAA、AGCAC、AGGAT、AGGTA、AGGTG、AGTAA、AGTCA、AGTCT、A GTGC、AGTGT、AGTTT、ATAAA、ATAAC、ATAAT、ATACA、ATACC、ATAGA、ATATA、ATATC、ATATG、ATCAC、ATCTC、ATGAA、ATGAC、ATGGA、ATGGT、ATGTT、ATTAA、ATTAT、 ATTCA、ATTCT、ATTGA、ATTGC、ATTTA、ATTTT、CAAA、CAAAC、CAACC、CAAGC、CAATA、CAATT、CACAA、CACAC、CACCA、CACCG、CACCT、CACGA、CACTG、CACTT、CAGC A、CAGCT、CATAT、CATCC、CATCT、CATTC、CCAAT、CCACG、CCCAC、CCGAG、CCGCA、CCGCC、CCGTC、CCTAC、CCTAT、CCTCC、CCTCT、CCTGA、CCTGC、CCTTA、CCTTG TT、CGATA、CGATT、CGCCA、CGCCG、CGCCT、CGGCA、CGGGCG、CGTGG、CGTTA、CTATA、CTATC、CTCAG、CTCCC、CTCGG、CTCTC、CTGAA、CTTCC、CTTCT、CTTGA、CTTTA、CTTTG、CTTTT、GAAAA、GAAGA、GAAGG、GAAGT、GACCC、GAGTC、GATAA、GATGC、GATTG、GATTT、GCAAG、GCAAT、GCACG、GCCAT、GCCCG、GCCGA、GCCGC、GCGTA、GCTGA,GCTTC、GCTTG、GCTTT、GGAAG、GGCTT、GGGAA、GGGCA、GGGCT、GGTGC、GGTGG、GGTGT、GGTTT、GTAAC、GTAAT、GT CCT、GTCTA、GTCTC、GTGCC、GTGCT、GTGTG、GTTAA、GTTAG、GTTTC、GTTTT、TAAAA、TAAAC、TAAAG、TAAAT、TAAC A, TAACT, TAATA, TAATT, TACAT, TACCA, TACGC, TACTG, TAGAG, TAGGA, TATAA, TATAG, TATAT, TATCA, TATCC, TATCG, TATGA, TATGG, TATGT, TATTA, TATTG, TCAAA, TCAGG, TCATA, TCATC, TCATT, TCCAA, TCCCA, TCCCG, TCC GC, TCCTC, TCTT, TCGGC, TCTAT, TTCA, TCTCC, TTCG, TTGG, TCTTA, TGAAG, TGACA, TGATA, TGATT, TGCAC, TGCAG, TGCCA, TGCCC, TGCCG, TGCCT, TGCTC, TGCTG, TGCTT, TGGAA, TGGCT, TGGGG, TGGGT, TGGTG, TGTAT, TG TGA、TGTGC、TGTGT、TGTTA、TGTTT、TTAAA、TTAAG、TTAAT、TTACA、TTATA、TTATG、TTATT、TTCAT、TTCCT、TTCT C. Optionally, the TAM sequence (5'→3') recognizable by the IscB protein may be selected from any one or more of the following: WHG, DHG, ATAAA, ATG, ATGAHD, ATGAA, DTG, DYGG, ATGAW, AYGG, NGG, where R=A / G, Y=C / T, M=A / C, K=G / T, S=G / C, W=A / T, H=A / T / C, B=G / T / C, V=G / A / C, D=G / A / T, N=A / T / C / G; Optionally, the IscB protein can form a complex with a guide polynucleotide; further, the complex can specifically bind to a target nucleic acid; even further, the complex can cleave the target nucleic acid, modify the target nucleic acid, and / or regulate the expression of the target nucleic acid; Optionally, the IscB protein can form a complex with a guide polynucleotide, the guide polynucleotide comprising a guide sequence that is reverse complementary to the target nucleic acid; optionally, the guide polynucleotide comprises a backbone sequence, the backbone sequence being capable of interacting with the IscB protein; furthermore, the backbone sequence comprises or is an ωRNA backbone sequence. Optionally, the backbone sequence may or may not contain a tracrRNA sequence; Optionally, the IscB protein is a nuclease-inactivated variant; alternatively, the IscB protein is a deadIscB-inactivated variant or a nickase IscB-inactivated variant; alternatively, the Ruvc domain of the IscB protein is inactivated.
2. A guide polynucleotide, comprising, It comprises (i) a backbone sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity with any one of SEQ ID NO: 27-40, 43, 44, 94-134; and (ii) a guide sequence engineered to hybridize with a target nucleic acid; the backbone sequence is linked to the guide sequence, the guide polynucleotide being capable of forming a complex with the IscB protein and guiding the complex to bind to the target nucleic acid in a sequence-specific manner. Optionally, the backbone sequence may comprise or be an ωRNA backbone sequence; Optionally, the IscB protein is the IscB protein as described in claim 1; Optionally, the guide sequence comprises 15-60 nucleotides. Optionally, the guide sequence hybridizes with the target nucleic acid, wherein the mismatch between the guide sequence and the target nucleic acid does not exceed one nucleotide; Optionally, the guiding polynucleotide may or may not contain tracrRNA; Optionally, the guiding sequence is located at the 3' end of the skeleton sequence; Optionally, the guiding sequence is located at the 5' end of the skeleton sequence; Optionally, the tracrRNA sequence is located at the 5' or 3' end of the backbone sequence.
3. An IscB inactivation variant, characterized in that, The IscB inactivating variant is the nuclease activity inactivating variant of the IscB protein as described in claim 1; Optionally, the IscB inactivating variant is a dead IscB inactivating variant or a nickase IscB inactivating variant; Optionally, the IscB inactivating variant is the inactivation of the Ruvc domain of the IscB protein.
4. A fusion protein or conjugate, characterized in that, The fusion protein or conjugate comprises the following elements: (1) The IscB protein as described in claim 1, or the IscB inactivating variant as described in claim 3; and (2) Homologous or heterologous functional structural domains; Optionally, the functional domain has enzymatic activity that modifies the target nucleic acid sequence; for example, nuclease activity, methyltransferase activity, demethylase activity, DNA repair activity, DNA damage activity, deamination activity, superoxide dismutase activity, alkylation activity, depurination activity, oxidation activity, pyrimidine dimer formation activity, integrase activity, transposase activity, recombinase activity, polymerase activity, ligase activity, helicase activity, photolyase activity, glycosylation activity, deglycosylation activity, acetyltransferase activity, deacetylase activity, kinase activity, phosphatase activity, ubiquitin ligase activity, deubiquitination activity, adenylation activity, deadenylation activity, SUMOylating activity, deSUMOylating activity, myristylation activity and / or demyristylation activity; Optionally, the homologous or heterologous functional domains may be selected from one or more of the following: subcellular localization signal, DNA binding domain, protease domain, transcription activation domain, transcription repression domain, nuclease domain, deaminase domain, uracil DNA glycosylase domain (UDG), uracil DNA glycosylase repression domain (UGI), methylase, demethylase, transcription release factor, histone acetylase domain, histone deacetylase domain, DNA ligase, affinity tag, reporter tag, affinity domain and reporter domain; Optionally, the nuclease domain includes a polypeptide with ssDNA cleavage activity and / or a polypeptide with dsDNA cleavage activity; Optionally, the IscB protein or its inactivating variant is directly or indirectly linked to the homologous or heterologous functional domain; preferably, the direct link is a covalent link, and the indirect link is a link via an amino acid linker or a non-amino acid linker. Optionally, the homologous or heterologous functional domains are fused or conjugated at the N-terminus, C-terminus, or internally relative to the IscB protein or inactivating variant.
5. An isolated nucleic acid, comprising, The nucleic acid encodes the IscB protein as described in claim 1, the IscB inactivating variant as described in claim 3, or the fusion protein or conjugate as described in claim 4; Optionally, the nucleic acid is codon-optimized for expression in cells; Optionally, the nucleic acid is codon-optimized for expression in prokaryotic cells; Optionally, the nucleic acid is codon-optimized for expression in eukaryotic cells; Optionally, the nucleic acid is codon-optimized for expression in eukaryotes, mammals such as humans or non-human mammals, plants, insects, birds, reptiles, rodents (e.g., mice, rats), fish, worms / nematodes, or yeast.
6. An IscB system characterized by, The IscB system includes: a. The IscB protein of claim 1, the IscB inactivating variant of claim 3, the fusion protein or conjugate of claim 4, or the nucleic acid of claim 5; and b. A guiding polynucleotide, or a polynucleotide sequence encoding the guiding polynucleotide; The IscB protein, IscB inactivating variants, fusion proteins, or conjugates form a complex with the guiding polynucleotide; the guiding polynucleotide contains a guiding sequence that is engineered to guide the complex to bind to the target nucleic acid in a sequence-specific manner. Optionally, the guiding polynucleotide comprises a backbone sequence linked to the guiding sequence; optionally, the backbone sequence has at least 50% identity with any one of SEQ ID NO:27-40, 43, 44, 94-134; optionally, the backbone sequence comprises or is an ωRNA backbone sequence. Optionally, the guide sequence comprises 15-35 nucleotides, and / or the guide sequence hybridizes with the target nucleic acid, wherein the guide sequence and the target nucleic acid are 90%-100% complementary, preferably with a mismatch of no more than one nucleotide; Optionally, the guide sequence comprises 15-60 nucleotides; Optionally, the guide sequence hybridizes with the target nucleic acid; Optionally, the guide sequence may mismatch the target nucleic acid by no more than one nucleotide; Optionally, the guiding polynucleotide further comprises tracrRNA; optionally, the tracrRNA sequence is linked to the backbone sequence; optionally, the tracrRNA comprises 10-200 nucleotides. Optionally, the guiding sequence is located at the 3' end of the skeleton sequence; Optionally, the guiding sequence is located at the 5' end of the skeleton sequence; Optionally, the tracrRNA sequence is located at the 5' or 3' end of the backbone sequence; Optionally, the target nucleic acid is DNA or RNA, preferably dsDNA or ssDNA; Optionally, the DNA is eukaryotic DNA; preferably, the eukaryotic DNA is non-human mammal DNA, non-human primate DNA, human DNA, plant DNA, insect DNA, bird DNA, reptile DNA, rodent DNA, fish DNA, worm / nematode DNA, or yeast DNA. Optionally, the target nucleic acid is a disease or symptom-related gene or a gene related to signal transduction biochemical pathways, or the target nucleic acid is a reporter gene.
7. A vector system characterized by comprising The vector system comprises one or more recombinant vectors, the recombinant vectors comprising the isolated nucleic acid as described in claim 5, or the IscB system as described in claim 6; Optionally, the recombinant vector further comprises a regulatory sequence; Optionally, the polynucleotide sequence encoding the IscB protein, IscB inactivating variant, fusion protein, or conjugate is operatively linked to the regulatory sequence, and / or, the polynucleotide sequence encoding the guiding polynucleotide is operatively linked to the regulatory sequence; more preferably, the regulatory sequence is selected from one or more of: promoters, enhancers, internal ribosome entry sites, and transcription termination signals, wherein the promoters are, for example, constitutive promoters, inducible promoters, broad-spectrum promoters, or tissue-specific promoters, and / or, the transcription termination signals are, for example, polyadenylation signals or polyU sequences; Optionally, the backbone of the recombinant vector is an adeno-associated virus vector, a lentiviral vector, a ribonucleoprotein complex, or a virus-like particle; preferably, when the backbone is an adeno-associated virus vector, the adeno-associated virus vector is a recombinant adeno-associated virus vector of serotype AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV PHP.B, AAV PHP.B2, AAV PHP.B3, AAV PHP.A, AAV PHP.eB, AAV PHP.eS, AAV2.7m8, AAV8.7m8, AAV ShH10, AAVrh10, or AAVrh74; when the backbone is a lentiviral vector, the lentiviral vector is pseudotyped with an envelope protein; Optionally, the isolated nucleic acid is linked to an aptamer sequence; When the backbone of the recombinant vector is a virus-like particle, the isolated nucleic acid is linked to the gene encoding the gag protein.
8. A delivery system characterized by, The delivery system includes: (1) Delivery tools, and (2) The IscB protein as claimed in claim 1, the guiding polynucleotide as claimed in claim 2, the IscB inactivating variant as claimed in claim 3, the fusion protein or conjugate as claimed in claim 4, the nucleic acid as claimed in claim 5, the IscB system as claimed in claim 6, or the vector system as claimed in claim 7; Optionally, the delivery tool may be a virus, lipid nanoparticles, nanoparticles, liposomes, exosomes, microvesicles, or a gene gun; Optionally, the delivery tool is a lipid nanoparticle containing the guiding polynucleotide and mRNA encoding the IscB protein, the inactivated variant of IscB, or the fusion protein or conjugate.
9. A cell, characterized in that, The cell comprises the IscB protein as claimed in claim 1, the guiding polynucleotide as claimed in claim 2, the IscB inactivating variant as claimed in claim 3, the fusion protein or conjugate as claimed in claim 4, the nucleic acid as claimed in claim 5, the IscB system as claimed in claim 6, the vector system as claimed in claim 7, or the delivery system as claimed in claim 8; Optionally, the cells are prokaryotic cells; Optionally, the cells are eukaryotic cells; Optionally, the eukaryotic cell is a mammalian cell.
10. A pharmaceutical composition, characterized in that, The pharmaceutical composition comprises the IscB protein as claimed in claim 1, the guiding polynucleotide as claimed in claim 2, the IscB inactivating variant as claimed in claim 3, the fusion protein or conjugate as claimed in claim 4, the nucleic acid as claimed in claim 5, the IscB system as claimed in claim 6, the carrier system as claimed in claim 7, the delivery system as claimed in claim 8, or the cell as claimed in claim 9; Preferably, the pharmaceutical composition further comprises pharmaceutically acceptable excipients.
11. A reagent kit, characterized in that, The kit comprises the IscB protein as claimed in claim 1, the guide polynucleotide as claimed in claim 2, the IscB inactivating variant as claimed in claim 3, the fusion protein or conjugate as claimed in claim 4, the nucleic acid as claimed in claim 5, the IscB system as claimed in claim 6, the vector system as claimed in claim 7, the delivery system as claimed in claim 8, the cell as claimed in claim 9, or the pharmaceutical composition as claimed in claim 10.
12. Use of the IscB protein of claim 1, the guiding polynucleotide of claim 2, the IscB inactivating variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the IscB system of claim 6, the vector system of claim 7, the delivery system of claim 8, the cell of claim 9, the pharmaceutical composition of claim 10, or the kit of claim 11 in the preparation of reagents or pharmaceuticals for the diagnosis, treatment, and / or prevention of diseases or conditions related to the target nucleic acid; Optionally, the disease or condition is a hematological disease or condition, an ophthalmic disease or condition, a nervous system disease or condition, a respiratory disease or condition, a liver disease or condition, a metabolic disease or condition, cancer, or an infectious disease; and / or, the reagent or drug is used to: cleave one or more target nucleic acid molecules or create nicks in one or more target nucleic acid molecules, activate or upregulate the expression of one or more target nucleic acid molecules, activate or inhibit the transcription of one or more target nucleic acid molecules, inactivate one or more target nucleic acid molecules, visualize, label, or detect one or more target nucleic acid molecules, bind one or more target nucleic acid molecules, transport one or more target nucleic acid molecules, and mask one or more target nucleic acid molecules.
13. A method for detecting, binding to, or cleaving target nucleic acids, characterized in that, The method includes contacting the target nucleic acid with the IscB protein as claimed in claim 1, the guide polynucleotide as claimed in claim 2, the IscB inactivating variant as claimed in claim 3, the fusion protein or conjugate as claimed in claim 4, the nucleic acid as claimed in claim 5, the IscB system as claimed in claim 6, the vector system as claimed in claim 7, the delivery system as claimed in claim 8, the cell as claimed in claim 9, the pharmaceutical composition as claimed in claim 10, or the kit as claimed in claim 11. Optionally, the method is for non-diagnostic and / or therapeutic purposes; and / or, the fusion protein or conjugate contains a detectable marker, such as a marker detectable by fluorescence, DNA blotting, or FISH; and / or, the method is an in vivo or in vitro method.
14. A method for altering cell state, characterized in that, The method includes contacting cells with the IscB protein as described in claim 1, the guiding polynucleotide as described in claim 2, the IscB inactivating variant as described in claim 3, the fusion protein or conjugate as described in claim 4, the nucleic acid as described in claim 5, the IscB system as described in claim 6, the vector system as described in claim 7, the delivery system as described in claim 8, the cell as described in claim 9, the pharmaceutical composition as described in claim 10, or the kit as described in claim 11, thereby altering the cell state. Optionally, the method results in one or more of the following: increased or decreased expression of a specific gene, induction of cell senescence in vitro or in vivo, cell cycle arrest in vitro or in vivo, promotion and / or inhibition of cell growth in vitro or in vivo, induction of non-responsiveness in vitro or in vivo, induction of apoptosis in vitro or in vivo, and induction of necrosis in vitro or in vivo. Optionally, the method is a method for non-diagnostic and / or therapeutic purposes; and / or, the method is an in vivo or in vitro method.
15. A method for diagnosing, treating, or preventing diseases or conditions related to a target nucleic acid, characterized in that, The IscB protein of claim 1, the guide polynucleotide of claim 2, the IscB inactivating variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the IscB system of claim 6, the carrier system of claim 7, the delivery system of claim 8, the cell of claim 9, the pharmaceutical composition of claim 10, or the kit of claim 11 may be administered to a sample of a subject in need or to a subject in need. Optionally, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.
16. The IscB protein of claim 1, the guiding polynucleotide of claim 2, the IscB inactivating variant of claim 3, the fusion protein or conjugate of claim 4, the nucleic acid of claim 5, the IscB system of claim 6, the vector system of claim 7, the delivery system of claim 8, the cell of claim 9, the pharmaceutical composition of claim 10, or the kit of claim 11, for the diagnosis, treatment, or prevention of diseases or conditions related to the target nucleic acid; Optionally, the disease or condition is a blood system disease or condition, an eye disease or condition, a nervous system disease or condition, a respiratory system disease or condition, a liver disease or condition, a metabolic system disease or condition, cancer, or an infectious disease.