IscB variant, base editing system and base editing method
By modifying the structure of the IscB protein and designing fusion proteins, the limitations of loading and delivering large Cas proteins in microbial editing were overcome, enabling efficient and precise microbial genome editing, expanding the editing window and reducing off-target effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-13
AI Technical Summary
Existing Cas9 and Cas12a-based base editing systems have limited loading capacity in delivery tools such as plasmids, bacteriophages, and viral vectors due to the large protein size. Furthermore, their delivery efficiency in in situ editing of microorganisms and in complex environments is limited. In addition, they suffer from problems such as low efficiency in editing natural proteins, limited target recognition sequences, and off-target effects of editing tools.
By studying the structure and function of the IscB protein and engineering it, IscB variants were developed, including mutations of specific amino acid residues, such as changing H339 to A, to form fusion proteins by combining deaminase and deaminase inhibitors. Furthermore, by optimizing the spacer sequence length and linker peptides, an efficient and compact base editing system was constructed.
A miniaturized base editing tool has been developed, which has higher editing efficiency and a wider editing window. It alleviates the editor's proximity bias and reduces off-target effects, making it suitable for precise editing of microbial genomes.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of gene editing technology, and more particularly to IscB variants, base editing systems, and base editing methods. Background Technology
[0002] Existing Cas9 and Cas12a-based base editing systems generally suffer from large protein sizes (Cas9 approximately 1368 amino acids, Cas12a approximately 1300 amino acids), limiting their loading capacity in delivery tools such as plasmids, bacteriophages, and viral vectors, especially in in-situ editing of microorganisms and delivery efficiency in complex environments. IscB is a nuclease whose protein retains the core cleavage domains of Cas9, including the HNH and RuvC nuclease domains, and possesses basic double-stranded DNA cleavage capabilities. Due to its small size and compact structure, IscB naturally possesses superior vector loading and easier expression advantages, making it highly promising for development into a miniature genome editing tool.
[0003] However, the IscB system still faces several technical bottlenecks, such as low efficiency in editing natural proteins, limited target recognition sequences, and off-target effects of editing tools. These need to be overcome through structure-guided engineering and RNA design optimization. Furthermore, developing a robust and broadly adaptable IscB gene editing platform, considering the diversity of genomes and delivery methods in complex microbial environments, remains crucial for achieving precise in-situ genome editing in microorganisms.
[0004] In summary, microbial base editors based on small IscB proteins possess unique technological advantages and broad application prospects. Through in-depth structural-functional research and engineering modifications, efficient, compact, and easily deliverable IscB base editing tools can be developed. This not only overcomes the limitations of traditional large Cas proteins in in-situ microbial editing but also promotes the upgrading of microbial genetic modification technology and the expansion of industrial applications, meeting the urgent need for efficient, precise, and flexible tools in the modern microbial genome editing field. Summary of the Invention
[0005] In view of this, the technical problem to be solved by the present invention is to provide IscB variants and base editing systems and methods.
[0006] The IscB variant provided by this invention has a mutation of H to A at position 339 of wild-type IscB, and at least one amino acid residue at positions 38, 54, 99, 365, 366, 367, 368, 369, 370, 373, 374, 376, 377, 378, 379, 386, 401, and 456 is mutated to arginine.
[0007] In this invention, the IscB variant includes at least one of the following mutations: K38R, P54R, Q99R, Y365R, M366R, V367R, H368R, Q369R, F370R, H373R, D374R, Q376R, A377R, C378R, H379R, S386R, A401R, or S456R.
[0008] In some embodiments, the IscB variant is mutated to arginine at position 368H and / or position 376Q of wild-type IscB.
[0009] In this invention, the amino acid sequence of the wild-type IscB is as shown in SEQ ID NO.1, or is an amino acid sequence that has more than 70% identity with the sequence shown in SEQ ID NO.1 and has IscB activity.
[0010] The present invention also provides a fusion protein comprising a deaminase, IscB, and a deaminase inhibitor;
[0011] IscB is either wild-type or a variant of IscB as described above.
[0012] In this invention, the deaminase is a cytosine deaminase; preferably, the cytosine deaminase is derived from humans, mice, and / or eels; preferably, the cytosine deaminase has an amino acid sequence as shown in any one of SEQ ID NO: 2 to 5, or the amino acid sequence of the cytosine deaminase has more than 70% identity with the sequence shown in any one of SEQ ID NO: 2 to 5, and has cytosine deaminase activity.
[0013] In this invention, the deaminase inhibitor is a cytosine deaminase inhibitor. Preferably, the cytosine deaminase inhibitor is derived from mice. Preferably, the cytosine deaminase inhibitor has the amino acid sequence shown in SEQ ID NO:6; or the amino acid sequence of the cytosine deaminase inhibitor has more than 70% identity with the sequence shown in SEQ ID NO:6 and has the activity of a cytosine deaminase inhibitor.
[0014] This invention does not limit the connection order of deaminase, IscB, and deaminase inhibitor in the fusion protein. This invention does not limit the specific type or number of repetitions of the deaminase in the fusion protein. This invention also does not limit the specific type or number of repetitions of the deaminase inhibitor in the fusion protein. In the embodiments of this invention, the deaminase, IscB, and deaminase inhibitor in the fusion protein are all repeated once.
[0015] In a specific embodiment, the deaminase in the fusion protein is located at the N-terminus of IscB, and the deaminase inhibitor is located at the C-terminus of IscB. The deaminase and IscB can be directly linked or linked via a linker fragment. Alternatively, IscB and the deaminase inhibitor can be directly linked or linked via a linker fragment.
[0016] More specifically, the fusion protein comprises, from N-terminus to C-terminus, the following components in sequence: deaminase, linker1, wild-type IscB, linker2, and deaminase inhibitor.
[0017] Alternatively, the fusion protein may include, from N-terminus to C-terminus, a deaminase, linker1, the previously described IscB variant, linker2, and a deaminase inhibitor.
[0018] In this invention, the lengths of linker1 and / or linker2 are independently selected from 1 to 50 amino acid residues. They can be composed of multiple repetitions of a single amino acid or random repetitions of multiple amino acids.
[0019] In a specific embodiment, linker1 contains 4 to 20 amino acid residues; preferably, linker1 has at least one of the amino acid sequences shown in SEQ ID NO: 9 to 14. Preferably, its amino acid sequence is shown in SEQ ID NO: 12.
[0020] In a specific embodiment, the linker2 contains 4 to 20 amino acid residues; preferably, the linker2 has at least one of the amino acid sequences shown in SEQ ID NO: 9 to 14. Preferably, its amino acid sequence is shown in SEQ ID NO: 12.
[0021] More specifically, the fusion protein, from N-terminus to C-terminus, consists of: cytosine deaminase, linker1, an IscB variant with arginine mutated at position 368H and / or position 376Q of wild-type IscB, linker2, and cytosine deaminase inhibitor.
[0022] Furthermore, the present invention also provides a nucleic acid that encodes the IscB variant as described above; or encodes a fusion protein as described above.
[0023] Furthermore, the present invention also provides the application of the IscB variant, the fusion protein, and / or the nucleic acid in the preparation of a base editing system.
[0024] This invention achieves efficient C-to-T editing across a wide range by screening cytosine deaminases from different sources. By adapting to different spacer lengths, the editing window of CBE is further expanded. Furthermore, by optimizing the linker between the nIscB protein and the deaminase and by modifying the protein structure of the IscB protein through mutation, the editing efficiency of the base editor is effectively improved, and the editor's proximity bias is alleviated. At the same time, the editing specificity of the base editor is verified through editor specificity evaluation.
[0025] Furthermore, the base editing system of the present invention includes:
[0026] I) ωRNA or DNA encoding ωRNA; and
[0027] II) IscB variants as described above, fusion proteins as described above, and / or nucleic acids as described above.
[0028] In this invention, in the base editing system, I) and II) are located on the same plasmid vector, or they can be located on different plasmid vectors. Specifically, if I) and II) are located on different plasmid vectors, the coding nucleic acid of one plasmid IscB variant is responsible for expressing the IscB variant with base editing activity; the other plasmid carries a ωRNA coding sequence, which contains a target sequence complementary to the target gene, and can guide the IscB variant to bind to the target site. If I) and II) are located on the same plasmid vector, the coding nucleic acid of the IscB variant and the coding sequence of the ωRNA are integrated in tandem into the same vector. This vector may also include regulatory elements such as screening marker genes (e.g., resistance genes, fluorescent genes), promoters, and terminators. This invention does not limit whether I) and II) are located on the same plasmid vector; both can achieve good base editing effects. In a specific embodiment, I) and II) are located on the same plasmid vector.
[0029] In this invention, the ωRNA includes a spacer sequence and a scaffold region, wherein the target length of the spacer sequence is 16~30bp.
[0030] In this invention, the ωRNA-mediated base editing window is the 2nd to 18th bases. Specifically, the scaffold region has a nucleic acid sequence as shown in SEQ ID NO:16; restriction endonuclease cleavage sites are set upstream and downstream of the spacer sequence; preferably, the restriction endonuclease is BSAI.
[0031] Furthermore, the present invention also provides a delivery carrier comprising at least one of the following:
[0032] i) ωRNA or nucleic acid encoding ωRNA;
[0033] ii) Nucleic acids encoding the IscB variant as described above;
[0034] iii) Nucleic acid encoding the fusion protein as described above.
[0035] In this invention, the delivery vector is a genetically engineered plasmid or a viral vector;
[0036] The genetically engineered plasmid is a plasmid adapted to prokaryotic cells and / or a plasmid adapted to eukaryotic cells.
[0037] The viral vector is an adeno-associated virus vector, a lentiviral vector, and / or an adenovirus vector.
[0038] If the delivery vector is a plasmid adapted to prokaryotic cells, it further includes at least one of a prokaryotic promoter, a ribosome binding site, a replication initiation site, a replication protein, and / or a selection marker.
[0039] If the delivery vector is a plasmid adapted to eukaryotic cells, it further includes at least one of a eukaryotic promoter, a nuclear localization signal, a replication initiation site, a replication protein, and / or a selection marker.
[0040] If the delivery carrier is a viral carrier, it also includes viral packaging elements.
[0041] In a specific embodiment, the delivery vector is a plasmid vector, which sequentially includes: nucleic acid encoding ωRNA, nucleic acid encoding the fusion protein as described above, a promoter, a rep site, and a selection marker.
[0042] Furthermore, the present invention also provides a host comprising the delivery carrier as described above;
[0043] It may contain nucleic acids encoding the IscB variant as described above or nucleic acids encoding the fusion protein as described above integrated into its genome; or its culture product may contain the delivery vector as described above.
[0044] In a specific embodiment, the host is Escherichia coli.
[0045] Furthermore, the present invention also provides a reagent comprising: the base editing system as described above; and / or the delivery carrier as described above.
[0046] The reagents described in this invention also include acceptable excipients such as buffer solutions and protectants required for storage, transformation, and / or transfection.
[0047] Furthermore, the present invention also provides a base editing method, which includes: treating the sample to be edited with the reagents as described above.
[0048] In this invention, the base editing method specifically includes mixing a reagent containing the base editor as described above with the sample to be edited. The sample to be edited includes, but is not limited to, bacteria. In a specific embodiment, the bacteria is *Escherichia coli*.
[0049] This invention provides a variant of IscB and a small base editor based on this variant. This gene editing tool is characterized by its small size and compact structure, enabling unbiased editing of the preceding base target C, and possessing a wider editing window and higher editing efficiency. This invention has been successfully applied to the genome editing of prokaryotic microorganisms, adding a powerful tool to the microbial gene editing toolbox. Attached Figure Description
[0050] Figure 1 It is a basic component of the nIscB-CBE plasmid;
[0051] Figure 2 The basic workflow of gene editing using microbial base editors;
[0052] Figure 3 The gene editing efficiency of APOBEC-nIscB-UGI at different sites in Escherichia coli NEB-10-beta was measured.
[0053] Figure 4 Editing efficiency and editing window of base editors for cytosine deaminases from different sources, as well as editing effect on target C at different positions under various pre-base conditions;
[0054] Figure 5 Comparison of editing effects and editing windows of hAID-nIscB-UGI in E. coli NEB-10-beta under different spacer target lengths;
[0055] Figure 6 Comparison of the editing effects of hAID-nIscB-UGI in E. coli NEB-10-beta under different types of linker conditions;
[0056] Figure 7 The structure of the nIscB protein and the different selected mutation sites are shown, and the editing efficiency is compared under different single mutation conditions.
[0057] Figure 8 Comparison of the editing effects of hAID-nIscB-UGI-H368R and hAID-nIscB-UGI-WT at different sites on Escherichia coli NEB-10-beta;
[0058] Figure 9The C-to-T editing effects of hAID-nIscB-UGI-H368R and hAID-nIscB-UGI-WT under different pre-base conditions;
[0059] Figure 10 Different mutation sites selected upstream and downstream of the H368 site for the nIscB protein are shown, and their editing efficiency is compared under several different single mutation conditions.
[0060] Figure 11 The editing effects of hAID-nIscB-UGI-H368R-Q376R on different sites in Escherichia coli NEB-10-beta;
[0061] Figure 12 This study compares the editing effects of constructing a ωRNA group with dinucleotide mismatch sites in a 20 bp ωRNA sequence and testing three base editors: nIscB-WT, nIscB-H368R, and nIscB-H368R-Q376R.
[0062] Figure 13 The experimental principle of orthogonal R-loops and the comparison of off-target effects in four types of CBEs were explained.
[0063] Figure 14 To evaluate the off-target effects of base editors nIscB-CBEs across the entire genome. Detailed Implementation
[0064] This invention provides IscB variants, a base editing system, and a base editing method. Those skilled in the art can refer to this document and appropriately modify the process parameters to achieve the desired results. It is particularly important to note that all similar substitutions and modifications are obvious to those skilled in the art and are considered to be included in this invention. The methods and applications of this invention have been described through preferred embodiments. Those skilled in the art can clearly modify or appropriately change and combine the methods and applications described herein without departing from the content, spirit, and scope of this invention to implement and apply the technology of this invention.
[0065] Unless otherwise defined in this invention, the scientific and technical terms associated with this invention shall have the meanings understood by one of ordinary skill in the art.
[0066] The terms “comprising,” “including,” and “having” are used interchangeably to indicate the inclusiveness of a scheme, meaning that the scheme may contain elements other than those listed. It should also be understood that the use of “comprising,” “including,” and “having” herein also provides for schemes “consisting of…”.
[0067] The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural.
[0068] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items.
[0069] In this invention, the term "variant" or "mutant" refers to a derived protein obtained by substituting, deleting, inserting, or modifying amino acid residues based on the original wild-type protein amino acid sequence. By modifying the protein structure of the IscB protein through protein structural mutation, this invention effectively improves the editing efficiency of the base editor and alleviates the editor's near-base bias.
[0070] In this invention, the "fusion protein" refers to a recombinant protein obtained by tandemly expressing genes encoding two or more protein domains with different functions (such as IscB protein, deaminase, etc.) using genetic engineering techniques. This protein is used to construct a miniaturized base editor. In this fusion protein, the fragments can be linked together by a linker to form a complete spatial conformation.
[0071] In this invention, the "nucleic acid" is a biomolecule carrying genetic information, which can be DNA or RNA, and has various conformations such as double-stranded and single-stranded. It is mainly used to encode IscB variants and fusion proteins, and is the core foundation of the base editing system. After being introduced into host cells through a vector, it expresses functional proteins to achieve gene editing. In this invention, the nucleic acid can exist in fragment form, in a plasmid vector, or linked to the host genome; this invention does not limit its presence in this regard.
[0072] In this invention, the "base editing system" uses an IscB protein variant as its core and achieves precise base editing through ωRNA-mediated synthesis. The base editing system provided by this invention is a powerful, miniaturized cytosine base editor capable of meeting various needs of microbial base editing. This invention adapts spacer sequences of different target lengths to the ωRNA of the base editor, forming cytosine base editors with different target lengths. The base editor with a spacer target length of 20 bp further expands its editing window to a wide range of C2-C18, which is larger than the editing window of currently reported base editors based on the nIscB protein. Furthermore, this invention effectively improves the editing efficiency of the cytosine base editor by adapting the linker peptide and introducing corresponding mutations at different positions in the nIscB protein.
[0073] In this invention, the "delivery vector" refers to a tool capable of delivering bioactive molecules (such as nucleic acids, proteins, etc.) to target cells or tissues. In this invention, the delivery vector can accurately transport active ingredients such as ωRNA, nucleic acids encoding IscB variants, and nucleic acids encoding fusion proteins into host cells.
[0074] In this invention, the plasmids adapted to prokaryotic cells contain a prokaryotic promoter that can initiate gene transcription, a ribosome binding site that facilitates protein synthesis, a replication initiation site and replication proteins that ensure autonomous replication of the plasmid, and selection markers that facilitate the screening and identification of cells containing the plasmid. In the plasmids adapted to eukaryotic cells, the eukaryotic promoter can initiate gene expression in eukaryotic cells, and the nuclear localization signal can guide the plasmid into the cell nucleus. In addition, the plasmids adapted to eukaryotic cells may optionally include elements such as replication initiation sites, replication proteins, and selection markers.
[0075] In this invention, the viral vector may be an adeno-associated virus vector, a lentiviral vector, or an adenovirus vector; the adeno-associated virus vector includes an inverted terminal repeat (ITR), a target gene expression cassette (containing a promoter, a target gene coding sequence, and a terminator), and an origin of replication; optionally, it also includes a selection marker gene. The lentiviral vector includes at least one of a long terminal repeat (LTR), a packaging signal (ψ), a Rev response element (RRE), a central polypurine tract (cPPT), a target gene expression cassette, and a selection marker gene; optionally, it also includes a promoter regulatory element. The adenovirus vector includes a viral origin of replication, a late promoter (MLP), a packaging signal (ψ), and a target gene expression cassette; optionally, it also contains E1 and E3 gene deletion regions.
[0076] In this invention, the term "host" refers to an organism or cell capable of having its delivery vector, integrating nucleic acid encoding an IscB variant or fusion protein, introduced into it. The host can be a prokaryote, such as bacteria, or a eukaryote, such as yeast or mammalian cells. The host of this invention can be introduced into the delivery vector using various methods. For prokaryotic hosts, common methods include transformation and transduction. Transformation refers to the direct introduction of exogenous DNA into host cells, while transduction involves using a viral vector to introduce exogenous DNA into host cells. The host described in this invention is used for the preparation, preservation, or amplification of the delivery vector as described above. Furthermore, by introducing the delivery vector of this invention into a suitable organism and expressing components such as the IscB variant, fusion protein, and / or ωRNA in the organism, precise base editing of the organism's genome can be achieved.
[0077] In this invention, the "linker" refers to a polypeptide sequence that connects the functional domains of a fusion protein. It possesses good hydrophilicity and spatial flexibility, effectively eliminating steric hindrance between domains and ensuring that each functional domain independently and completely exerts its biological activity. Furthermore, this linker is non-immunogenic and does not interfere with the overall stability, targeting ability, or other core functions of the fusion protein. The numbers 1 and 2 are used only for linker naming and are not intended to limit the order, quantity, or other relationships.
[0078] In this invention, "identity" refers to the proportion of identical residues (amino acid residues or nucleotides) at the same positions between the target amino acid or nucleotide sequence and the reference sequence after alignment. Sequence alignment employs conventional algorithms in the art (such as BLAST and ClustalW), and reasonable gaps can be introduced as needed to optimize the alignment results. The calculation results directly reflect the degree of homology and similarity between the two sequences. In this invention, "identity of 70% or more" means an identity of no less than 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9%.
[0079] The numerical ranges and parameters involved in this invention have been presented as precisely as possible in the specific embodiments. However, any numerical value inevitably contains standard deviations due to individual test methods. Therefore, unless otherwise expressly stated, it should be understood that all numerical ranges or specific data used in this disclosure may have a reasonable deviation within a certain range, such as ±10%, ±5%, ±1%, or ±0.5%.
[0080] All test materials used in this invention are common commercial products and are readily available in the market. This invention constructs a highly efficient, wide-range, and low-off-target miniaturized CBE base editor, thereby enabling a wide range of base mutations in microorganisms, especially highly efficient and wide-range C-to-T base mutations, thus playing an important role in the exploration of microbial gene function and the enhancement of phenotypic function.
[0081] It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply the order of execution. Some or all steps can be executed in parallel or sequentially. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The present invention is further illustrated below with reference to embodiments:
[0082] Example 1: rAPOBEC-nIscB-UGI for Escherichia coli genome editing
[0083] The rat-derived rAPOBEC cytosine deaminase, IscB protein, and cytosine deaminase inhibitor UGI were synthesized. rAPOBEC was linked to the N-terminus of IscB, and UGI was linked to the C-terminus of IscB. IscB was mutated to alanine A at position H339 to obtain the nicking enzyme nIscB.
[0084] The ωRNA corresponding to nIscB was synthesized, and a BSAI restriction site was designed upstream and downstream of the spacer sequence. After restriction, multiple target sites were quickly replaced by designing complementary oligonucleotide chain pairs: (a) 5'-TAGTNNN...NNN(N22)-3'; (b) 5'-AGCCNNN...NNN(N22 antisense strand)-3'.
[0085] To evaluate the editing activity of rAPOBEC-nIscB-UGI, a cytosine base editor for ωRNA at multiple sites was constructed, and the single plasmid system was transformed into wild-type Escherichia coli NEB-10-beta. After culturing for 48 h, colonies were collected and Sanger sequencing was performed on the target region to verify the C-to-T editing efficiency at different sites.
[0086] Six target sites within the spacer region were evaluated by detecting each cytosine at positions 1 to 16 of the spacer region of *E. coli* NEB-10-beta (position 1 was designated at the 5' end of the spacer region). Sequencing results showed that rAPOBEC-nIscB-UGI exhibited differentiated editing efficiency at different sites. Comprehensive statistical analysis revealed that its editing window in *E. coli* NEB-10-beta reached a 7 bp editing window (C8-C14). Figure 3 ).
[0087] Example 2: Deaminase screening enables effective wide-range gene editing
[0088] Traditional cytosine base editors exhibit severe base bias during genome editing due to the use of rat cytosine deaminase rAPOBEC. Specifically, they favor target Cs with a leading T base, while essentially not editing target Cs with a leading G base.
[0089] This invention screened cytosine deaminases from various sources by evaluating the editing efficiency of all NCs (where N represents any nucleotide and the second nucleotide C is the target nucleotide). First, rAPOBEC-nIscB-UGI was tested, which exhibited a significant base bias: TC > AC / CC > GC. This base bias is detrimental to efficient genome editing. It also suggests that previous work using rAPOBEC-based base editors struggled to efficiently edit all types of bases, prompting this invention to explore deaminases from different sources to mitigate this bias. Therefore, this invention utilizes deaminases from different sources, such as evolutionarily engineered rat APOBEC (evoAPOBEC), human AID (hAID), and an evolved version of eel CDA (evoCDA), to test their respective editing effects and base biases.
[0090] Screening of several cytosine deaminases from different sources revealed that hAID significantly broadened the editing window of the base editor, achieving a 13 bp editing window (C2-C14) in *E. coli* NEB-10-beta, which is 6 bp larger than that of rAPOBEC-nIscB-UGI. Its average editing efficiencies in TC, CC, GC, and AC were 40.1%, 31.3%, 26.7%, and 13.4%, respectively. Overall, hAID-nIscB-UGI exhibited a wider editing window and higher editing efficiency. Figure 4 ).
[0091] The amino acid sequence of the fusion protein with hAID as the core is: SEQ ID NO:7.
[0092] The nucleic acid sequence encoding the fusion protein as described above is SEQ ID NO:8.
[0093] Example 3: Adapting spacer target length to achieve a wider range of gene editing
[0094] In one embodiment of the present invention, the deaminase is AID. To select a suitable spacer targeting length, we first designed four different spacer targeting lengths for the cyoB2 site of E. coli NEB-10-beta (16 bp sequence: CACCCATATCGTTGGT, 20 bp sequence: TTGCCACCCATATCGTTGGT, 25 bp sequence: TCATGTTGCCACCCATATCGTTGGT, and 30 bp sequence: CATCATCATGTTGCCACCCATATCGTTGGT), and then rapidly changed the targeting sequence of different lengths using the BSAI restriction site in the spacer portion.
[0095] Four base editors with different target lengths were constructed and used for editing at the cyoB_2 site. All of them exhibited some C-to-T editing activity, and the base editors with spacer target lengths of 20 bp, 25 bp, and 30 bp showed wider editing windows. The editing effects at three other sites (all located on the cyoB gene, denoted as cyoB_1, nusG_1, and lpd_1) at four different spacer target lengths were further tested.
[0096] Comprehensive analysis of the test results shows that the spacer-targeted base editor with a length of 20 bp exhibits a wider editing window and higher editing efficiency. Tests show that its editing window in E. coli NEB-10-beta reaches a wide range of C2-C18, which is wider than most miniaturized base editors reported to date. Figure 5 ).
[0097] Example 4: Linker adaptation and nIscB protein mutation modification for efficient gene editing
[0098] In one embodiment of the present invention, the linker peptide between the nIscB protein and hAID was first optimized and adapted. Seven different linkers (0 aa, 4 aa, 8 aa, 12 aa, 16 aa, 18 aa, 20 aa) were tested at three sites on E. coli NEB-10-beta (cyoB_1 sequence: TTCGCGCCCGCACCCATCGT, cyoB_2 sequence: TTGCCACCCATATCGTTGGT, lpd_1 sequence: CCACTGACGCGCTGGAACTG). The editing results at the three sites showed that the hAID-nIscB-UGI constructed with the 16 aa linker exhibited higher editing efficiency. Figure 6 ).
[0099] Furthermore, based on the construction of a 16 aa linker base editor, nIscB protein structure was modified by mutation, selecting appropriate sites to mutate to positively charged lysine R to obtain nIscB protein with higher affinity. Specifically, the following sites were included: K38R, P54R, Q99R, H368R, S386R, A401R, and S456R. Figure 7 ).
[0100] Editing tests were conducted on the cyoB_1 and cyoB_2 sites of E. coli NEB-10-beta, and the editing performance of different mutant nIscB base editors was compared. Figure 8 Tests showed that the H368R single mutant can improve the editing efficiency of the editor. Furthermore, the editing effect of the nIscB-H368R single mutant was tested at four other sites. Figure 9 A comprehensive comparison revealed that the nIscB-H368R single mutant can improve editing efficiency and effectively alleviate the problem of low editing efficiency when the target C is preceded by A. Its average editing efficiency in the C2-C18 range can reach 62.7%, which is 16.5% higher than that of the base editor without the H368R single mutation.
[0101] Furthermore, we searched for single mutation sites upstream and downstream of the H368 site in the nIscB protein that could also improve editing efficiency. Specifically, we selected the following sites: Y365R, M366R, V367R, Q369R, F370R, H373R, D374R, Q376R, A377R, C378R, and H379R. We tested the editing effects of these single-mutant base editors on the rcsD_1 (sequence AAGATCTCACCTCCGGATTT) and aroE_1 (sequence GCACCTTTACCACCAGCACT) sites. Figure 10 The results showed that the nIscB-Q376R single mutant could also improve the editing efficiency of the base editor. Furthermore, a nIscB-H368R-Q376R double mutant was constructed based on the nIscB-H368R single mutant. (The amino acid sequence of the resulting fusion protein differs from the amino acid sequence described in Example 2 only at the mutation site, and its encoding nucleic acid sequence also differs only at the mutation site.)
[0102] The editing effect of the nIscB-H368R-Q376R double mutant base editor was tested on six sites (mscL_1, ddlB_1, satP_1, aroE_1, rcsD_1, lpd_1) of E. coli NEB-10-beta. Figure 11Sequencing results showed that the nIscB-H368R-Q376R double mutant base editor improved editing efficiency by 30.1% compared to the nIscB-H368R-WT base editor.
[0103] Example 5: Specificity assessment of nIscB-CBEs
[0104] In one embodiment of the present invention, to investigate the tolerance of nIscB-CBEs to ωRNA mutations, a ωRNA mutant strategy was first adopted, and a set of ωRNAs targeting the aroE_1 and kefB_1 genes on the E. coli NEB-10-beta genome was constructed, containing dinucleotide mismatch sites distributed in the ωRNA sequence (20 bp). The editing effects were then tested against three base editors: nIscB-WT, nIscB-H368R, and nIscB-H368R-Q376R. Figure 12 The test results showed that nIscB's ωRNA exhibited higher tolerance to mutations at the distal TAM end, but lower tolerance to mutations near the TAM end. Compared with nIscB-WT and nIscB-H368R, nIscB-H368R-Q376R had a lower overall off-target editing level and stronger editing specificity.
[0105] Furthermore, to evaluate the non-ωRNA-dependent off-target effects of base editors nIscB-CBEs in the genome, orthogonal R-loop structures were generated at six target sites within the *E. coli* NEB-10-beta genome using CRISPR-dCas12a, thereby assessing the off-target editing activity of different CBEs at these R-loop sites. Figure 13 Next-generation sequencing results showed that among the four tested CBEs (hAID-nIscB-UGI-WT, hAID-nIscB-UGI-H368R, hAID-nIscB-UGI-H368R-Q376R, and hAID-SpCas9n-UGI), the three CBEs represented by the small protein nIscB all exhibited extremely low off-target editing, and were lower than the off-target effect of hAID-SpCas9n-UGI.
[0106] Furthermore, to evaluate the off-target effects of base editors nIscB-CBEs across the entire genome, we transformed several different CBEs (hAID-nIscB-UGI-WT, hAID-nIscB-UGI-H368R, hAID-nIscB-UGI-H368R-Q376R, hAID-SpCas9n-UGI) into *E. coli* NEB-10-beta. Whole-genome sequencing results showed that... Figure 14The main off-target mutations in the genome are of the C>T type, accompanied by extremely low off-target effects and extremely low indel incidence. hAID-SpCas9n-UGI showed the highest single nucleotide variant (SNV) rate, while the off-target editing of hAID-nIscB-UGI-WT, hAID-nIscB-UGI-H368R, and hAID-nIscB-UGI-H368R-Q376R was basically consistent, indicating that mutations in the nIscB protein do not exacerbate off-target editing by the editor.
[0107] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
[0108] Table 1. Fusion protein related sequences
[0109]
[0110]
[0111] Table 2 Related Component Sequence
[0112]
[0113] Table 3 Target primer sequences involved in the examples
[0114]
[0115]
[0116]
[0117] Note:
[0118] The ununderlined portion represents the DNA sequence of the target corresponding to the spacer region.
[0119] The underlined portion indicates the BSAI enzyme ligation site where the corresponding sequence is ligated into the digestion plasmid.
Claims
1. An IscB variant in which the H at position 339 of wild-type IscB is mutated to A, and at least one of the following amino acid residues is mutated to arginine: position 38, position 54, position 99, position 365, position 366, position 367, position 368, position 369, position 370, position 373, position 374, position 376, position 377, position 378, position 379, position 386, position 401, and position 456.
2. The IscB variant according to claim 1, characterized in that, It includes at least one of the following mutations: K38R, P54R, Q99R, Y365R, M366R, V367R, H368R, Q369R, F370R, H373R, D374R, Q376R, A377R, C378R, H379R, S386R, A401R, or S456R.
3. The IscB variant according to claim 1 or 2, characterized in that, It has a mutation at position 368 (H) and / or position 376 (Q) of wild-type IscB to arginine.
4. The IscB variant according to any one of claims 1 to 3, characterized in that, The amino acid sequence of the wild-type IscB is shown in SEQ ID NO.
1.
5. Fusion proteins, including deaminases, IscB, and deaminase inhibitors; Wherein IscB is wild type, or is a variant of IscB as described in any one of claims 1 to 3.
6. The fusion protein according to claim 5, characterized in that, The deaminase is a cytosine deaminase; preferably, the cytosine deaminase is derived from humans, mice, and / or eels; preferably, the cytosine deaminase has an amino acid sequence as shown in any one of SEQ ID NO:2-5.
7. The fusion protein according to claim 5, characterized in that, The deaminase inhibitor is a cytosine deaminase inhibitor. Preferably, the cytosine deaminase inhibitor is derived from mice. Preferably, the cytosine deaminase inhibitor has the amino acid sequence shown in SEQ ID NO:
6.
8. The fusion protein according to any one of claims 5 to 7, characterized in that, It comprises, from N-terminus to C-terminus, a deaminase, linker 1, the IscB variant as described in any one of claims 1 to 3, linker 2, and a deaminase inhibitor.
9. The fusion protein according to claim 8, characterized in that, The linker1 contains 4 to 20 amino acid residues; preferably, the linker1 has at least one of the amino acid sequences shown in SEQ ID NO: 9 to 14. The linker2 contains 4 to 20 amino acid residues; preferably, the linker2 has at least one of the amino acid sequences shown in SEQ ID NO:9 to 14.
10. A nucleic acid encoding the IscB variant as described in any one of claims 1 to 3; or encoding the fusion protein as described in any one of claims 4 to 10.
11. The use of the IscB variant according to any one of claims 1 to 3, the fusion protein according to any one of claims 4 to 9, and / or the nucleic acid according to claim 10 or 11 in the preparation of a base editing system.
12. A base editing system, comprising: I) ωRNA or DNA encoding ωRNA; and II) the IscB variant according to any one of claims 1 to 3, the fusion protein according to any one of claims 4 to 9, and / or the nucleic acid according to claim 10 or 11.
13. The base editing system according to claim 13, characterized in that, The ωRNA includes a spacer sequence and a scaffold region, wherein the spacer sequence has a target length of 16-30 bp.
14. The base editing system according to claim 13 or 14, characterized in that, The ωRNA-mediated base editing window is the 2nd to 18th base position.
15. The base editing system according to any one of claims 13 to 15, characterized in that, The scaffold region has a nucleic acid sequence as shown in SEQ ID NO:16; restriction endonuclease cleavage sites are provided upstream and downstream of the spacer sequence; preferably, the restriction endonuclease is BSAI.
16. A delivery carrier comprising at least one of the following: i) ωRNA or nucleic acid encoding ωRNA; ii) The nucleic acid encoding the IscB variant as described in any one of claims 1 to 3; iii) The nucleic acid encoding the fusion protein according to any one of claims 4 to 9.
17. The delivery carrier according to claim 16, characterized in that, The delivery vector is a genetically engineered plasmid or a viral vector. The genetically engineered plasmid is a plasmid adapted to prokaryotic cells and / or a plasmid adapted to eukaryotic cells. The viral vector is an adeno-associated virus vector, a lentiviral vector, and / or an adenovirus vector.
18. The delivery carrier according to claim 17, characterized in that, If the delivery vector is a plasmid adapted to prokaryotic cells, it further includes at least one of a prokaryotic promoter, a ribosome binding site, a replication initiation site, a replication protein, and / or a selection marker. If the delivery vector is a plasmid adapted to eukaryotic cells, it further includes at least one of a eukaryotic promoter, a nuclear localization signal, a replication initiation site, a replication protein, and / or a selection marker. If the delivery carrier is a viral carrier, it also includes viral packaging elements.
19. The delivery carrier according to any one of claims 16 to 18, characterized in that, It includes, in sequence: nucleic acid encoding ωRNA, nucleic acid encoding the fusion protein according to any one of claims 4 to 9, promoter, rep site and selection marker.
20. Host, It contains the delivery carrier as described in any one of claims 16 to 19; Or, the genome of the device contains nucleic acid encoding the IscB variant of any one of claims 1 to 3 or nucleic acid encoding the fusion protein of any one of claims 4 to 9; Or its culture product contains the delivery vector as described in any one of claims 16 to 19.
21. Reagents, including: The base editing system according to any one of claims 12 to 15; and / or The delivery carrier according to any one of claims 16 to 19.
22. Base editing methods, including: The sample to be edited is treated with the reagent described in claim 21.