Iscb mutant proteins and uses thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- EAST CHINA NORMAL UNIV
- Filing Date
- 2024-07-22
- Publication Date
- 2026-07-31
AI Technical Summary
The existing CRISPR-Cas9 gene editing technology has problems such as low editing activity, low delivery efficiency and high commercialization cost in mammalian cells. In particular, the editing activity of IscB protein in the mammalian genome is no more than 4.4%, which cannot meet the practical application needs.
By mutation of the IscB protein amino acid site, especially replacement mutations to arginine or other non-positively charged amino acids, it improves its interaction with double-stranded DNA, and builds an IscB mutant protein with higher editing activity, and fuses it with deaminase to form an ultra-small base editing tool, using modified ωRNA to improve delivery efficiency and editing accuracy.
It improves the editing activity of IscB protein in mammalian cell lines, reduces commercialization costs, and enhances the application potential of gene editing tools, especially in precision medicine and crop breeding.
Abstract
Description
IscB mutant protein and its application
[0001] This application claims priority from PCT patent application No. PCT / CN2023 / 143184, filed December 29, 2023. This application incorporates the entirety of the PCT patent application. Technical Field
[0002] The present invention belongs to the field of bioengineering technology, and in particular relates to an IscB mutant protein and an application thereof. Background Art
[0003] Gene-editing tools, such as CRISPR-Cas9, are the vehicles of a revolutionary third-generation gene-editing technology, enabling humanity to effectively manipulate the most fundamental genetic information within organisms. Early in its application, this technology significantly shortened the time required to develop animal models for disease. Following significant breakthroughs in clinical research over the past two to three years for intractable genetic diseases such as β-thalassemia and transthyretin amyloidosis, it has recently gained prominence in CAR-T cell therapy. CRISPR-Cas9 technology was awarded the 2020 Nobel Prize in Chemistry for its enormous potential and economic value in numerous fields critical to national well-being, including biomedicine, agricultural breeding, new energy, and synthetic biology. Further promoting the development of various gene-editing tools, particularly non-CRISPR-Cas proteins and derivatives with independent intellectual property rights, is of vital importance.
[0004] The most widely used CRISPR-Cas system currently includes the Cas9 system (SpCas9) from Streptococcus pyogenes. After SpCas9 recognizes a specific nucleic acid combination of several (usually 2-6) bases in double-stranded DNA (dsDNA) (called PAM, protospacer adjacent motif), the crRNA-mediated Cas9 nuclease recognizes and completely matches the target DNA sequence, and then uses the HNH and RuvC nuclease domains to cut the two strands of DNA respectively. The HNH domain cuts the target strand (target strand) that is complementary to the RNA, and the RuvC domain cuts the non-target strand (non-target strand). After the double-strand break (DSB) is formed, the broken DNA double strand is repaired by the non-homologous end joining mechanism (NHEJ) or the homologous recombination repair mechanism (HDR). During the repair process, the purpose of gene editing can be achieved by inserting or deleting (indel) caused by the NHEJ mechanism, or by inserting exogenous target fragments at specific sites in the genome using HDR.
[0005] Despite the groundbreaking success of CRISPR-Cas9 technology, the following key issues remain to be addressed: 1) the damage to cells caused by Cas9 cleaving the genome; and 2) the difficulty in efficiently delivering various editing tools in vivo.
[0006] The safety issue with Cas9 stems from the fact that DSBs formed after genome cutting can cause cell apoptosis. When cells use NHEJ to repair DSBs, indels can lead to unexpected mutations, large deletions, and chromosomal rearrangements, which pose a significant threat to normal cellular function. Using DNA homologous recombination templates to precisely repair DNA via HDR to introduce specific mutations typically has a low editing success rate (approximately 0.1%-5%). To avoid DSBs while performing precise gene editing, David Liu's laboratory, based on the working principles of the CRISPR-Cas9 system, first developed efficient C·G to T·A base conversion tools—the cytosine base editor (CBE) and the adenine base editor (ABE)—by fusing the cytosine deaminase APOBEC or the adenine deaminase TadA with Cas9n / Cas9d (nickase / dead, nickase or fully inactivated forms). Because base editors (BEs) do not generate DSBs and do not rely on the HDR mechanism, they are able to achieve the goal of efficiently and accurately editing single nucleotides. Currently, base editing tools with different characteristics have been developed, such as Td-CGBE, which can precisely edit C to G, and ABE9, which reduces DNA and RNA off-target events to background levels.
[0007] How to efficiently achieve the in vivo delivery of SpCas9 and base editing tools has always been another major problem that has plagued academia and industry. Currently, the most commonly used method is to use viral vectors represented by adeno-associated virus (AAV) for delivery, but AAV has defects such as relatively small vector capacity, high commercial production costs, and transfection differences between different serotypes in different tissues. In particular, the defect that the vector capacity is only about 4.7kb seriously affects the packaging efficiency of SpCas9, and it is even more difficult to accommodate editing tools such as BE or Prime Editing (PE) that are built based on the Cas9 skeleton. This problem can be improved to a certain extent by splitting the Cas9 system into two parts and delivering them separately, but it also brings two new problems: reduced delivery efficiency and significantly increased commercial costs. The problem of achieving efficient delivery can be partially solved by using a more compact gene editing system. Therefore, in 2021, Zhang Feng's laboratory reported the Type II-D Cas9, which consists of approximately 700 amino acids (aa), and other laboratories reported the CRISPR-Cas Type V Cas12f subtype, which has a peptide chain length of 400-600aa. Although these subsequently discovered CRISPR-Cas homologous proteins have a relatively small size, their editing efficiency for genomic dsDNA is not as good as SpCas9. Moreover, since Cas12f only has a RuvC domain and needs to form a dimer to effectively perform its gene editing function, it has not been popularized in various theoretical or practical applications.
[0008] Bioinformatics methods have recently revealed the presence of the IscB protein family within the IS200 / IS605 transposon family, which possesses both RuvC and HNH nuclease domains. Current research suggests that IscB proteins are the evolutionary ancestors of the Cas9 family. In addition to possessing two nuclease domains similar to those of Cas9, IscB proteins also form protein-nucleic acid complexes with single-stranded non-coding RNA (ncRNA) with a conserved secondary structure, known as omega RNA (ωRNA). Recognition of the target double-stranded DNA (dsDNA) by this complex requires the IscB protein to first interact with a target-adjacent motif (TAM) sequence (consisting of 2-6 bases) at one end of the target. Subsequently, the guide sequence on the ωRNA mediates the IscB protein's correct recognition of the dsDNA double strand, and the RuvC and HNH domains then cleave one DNA strand separately, forming a DSB. Compared with SpCas9 of 1053aa and CBE, ABE, and PE of more than 1700aa, the IscB proteins currently discovered through bioinformatics prediction are only about 400aa, which is extremely beneficial for making full use of existing AAV delivery tools to carry out in vivo gene editing.
[0009] However, according to Zhang Feng's 2021 article, the editing activity of IscB protein, a nuclease with gene editing potential, on the mammalian genome does not exceed 4.4% at the tested targets, making it impossible to carry out practical applications of gene editing or to build new editing tools using this protein as a skeleton.
[0010] Summary of the Invention
[0011] Based on this, it is necessary to address the problem of the lack of IscB proteins with better editing activity in the above-mentioned existing technologies and propose an IscB mutant protein, which has a stronger interaction with the double-stranded DNA target and improves the editing activity in mammalian cell lines.
[0012] The present invention provides an IscB mutant protein. Compared with the wild-type IscB protein, the IscB mutant protein has arginine at at least one of the following amino acid positions: position 84, position 96, position 102, position 111, position 159, position 368, and position 386; preferably, the amino acids at positions 401 and 456 are also arginine.
[0013] The wild-type IscB (OgeuIscB) is derived from the human gut metagenome and belongs to the IS200 / IS605 transposon family.
[0014] In some embodiments of the present invention, the IscB mutant protein comprises two substitution mutations or three substitution mutations, one of the two substitution mutations is D96R, and two of the three substitution mutations are E84R and D96R.
[0015] In some embodiments of the present invention, the IscB mutant protein comprises two substitution mutations or three substitution mutations, the two substitution mutations being any two of E84R, D96R, H368R, S386R and S456R, and the three substitution mutations being any three of E84R, D96R, V159R, H368R, S386R, C111R, K92R, A401R, S456R, K118R, M102R and N167R.
[0016] In some embodiments of the present invention, the IscB mutant protein comprises two substitution mutations or three substitution mutations, the two substitution mutations are E84R-D96R, D96R-H368R, D96R-S386R or D96R-S456R, preferably E84R-D96R, D96R-H368R or D96R-S386R, more preferably E84R-D96R, the three substitution mutations are E84R-D96R-V159R, E84R-D96R-H368R, E84R-D96R-S386R, D96R-E84R-C111R, D96R-E84R-K92R, D96R-E84R-A401R, E84R-D96R-S456R, D96R-E84R-K118R, D96R-E84R-M102R or E84R-D96R-N167R, preferably E84R-D96R-V159R, E84R-D96R-H368R, E84R-D96R-S386R or D96R-E84R-C111R, more preferably E84R-D96R-V159R, E84R-D96R-H368R or E84R-D96R-S386R.
[0017] In some embodiments of the present invention, the amino acid sequence of the wild-type IscB protein is shown in SEQ ID NO: 1.
[0018] On the other hand, the present invention also provides an IscB mutant protein, wherein, compared with the wild-type IscB protein or the above-mentioned IscB mutant protein, any one of the amino acids at positions 60, 192, 339 and 342 of the IscB mutant protein is a non-positively charged amino acid, wherein the non-positively charged amino acid is selected from glycine, alanine, valine, leucine, isoleucine, methionine, proline, tryptophan, serine, threonine, tyrosine, cysteine, phenylalanine, asparagine, glutamine, aspartic acid and glutamate, and the non-positively charged amino acid is preferably alanine.
[0019] Understandably, mutating the amino acids at these sites to anything other than the three positively charged amino acids (arginine, lysine, and histidine) could alter their activity, reducing or inactivating the RuvC domain of IscB, leaving it with only the ability to cleave single-stranded DNA, thus generating the nickase IscB (nIscB), which is then used to construct an ultra-small base editing tool. However, mutating the amino acids at these sites to alanine, which has the smallest molecular weight and simplest spatial structure, has a superior effect.
[0020] The above-mentioned nIscB protein, which only has the ability to break single-stranded DNA, can be used to construct base editors and can also be used to perform gene editing on the genome with lower off-target levels. Since this nIscB only produces single-stranded gaps instead of double-strand breaks at the target site, two ωRNAs are required instead of one for gene editing. The two ωRNAs are designed on complementary DNA chains and are in close proximity (the sequences are no more than 20bp apart) to ensure that double-strand breaks in DNA will only occur when both chains are cut by nIscB. This paired nIscB requires two ωRNAs to work together to produce double-strand breaks in DNA, thereby greatly reducing off-target effects and improving the safety of the gene editing system.
[0021] Furthermore, the aforementioned IscB mutant protein can also be used to develop a multiplex orthogonal genome editing system, which utilizes two RNA sequences containing different guide RNAs to recruit cytosine deaminases or adenine deaminases fused to the corresponding binding proteins, and simultaneously achieve CBE and ABE base editing at different target sites. Using the nickase activity of IscB, paired ωRNAs are introduced into the genome to generate a DSB at a third target site, making the multiplex orthogonal genome editing system a system with triple (C-to-T, A-to-G, and knock-out) editing capabilities.
[0022] In some embodiments, the IscB mutant protein (nIscB) comprises any one of the following substitution mutations: D60A, E192A, H339A, D342A, preferably comprises the H339A substitution mutation.
[0023] On the other hand, the present invention also provides an IscB mutant protein, wherein, compared with the wild-type IscB protein or the above-mentioned IscB mutant protein, the amino acids at positions 246 and 342 of the IscB mutant protein are replaced with non-positively charged amino acids, thereby causing the IscB mutant protein to lose its double-stranded DNA cleavage function, wherein the non-positively charged amino acids are selected from glycine, alanine, valine, leucine, isoleucine, methionine, proline, tryptophan, serine, threonine, tyrosine, cysteine, phenylalanine, asparagine, glutamine, aspartic acid and glutamate, and the non-positively charged amino acids are preferably alanine.
[0024] Through the above mutations, the protein loses its ability to cut double-stranded DNA, becoming dead IscB (abbreviated as dIscB).
[0025] Understandably, this dIscB protein that has lost its double-stranded DNA cleavage function can be used to construct a gene activation system in addition to constructing a base editor. For example, by fusing dIscB-ωRNA with the bacteriophage MS2 coat protein and the transcription activator VP64 to construct a transcription activation complex, and then using dIscB's ωRNA to target the promoter sequence of a gene, the complex can recruit transcription factors that regulate mammalian cell transcription to activate or further upregulate gene expression. Similarly, a gene repression system can be constructed by fusing dIscB with a transcription repressor (such as KRAB protein). By targeting specific regions of the dsDNA double strand by ωRNA, gene expression is then inhibited by interference from the transcription repressor. In addition to directly regulating DNA, dIscB can also be fused with epigenetic modification enzymes to form dIscB epigenetic modification tools, which can achieve precise epigenetic modifications such as methylation and demethylation of specific DNA sites, acetylation and demethylation of histones on chromosomes, etc.
[0026] On the other hand, the present invention also provides a fusion protein comprising the aforementioned IscB mutant protein and a functional protein, wherein the functional protein is selected from at least one of an HMG-D protein and a nuclear localization signal protein.
[0027] These functional proteins possess specific functions, thereby improving the performance of fusion proteins. For example, HMG-D protein, a nonspecific double-stranded DNA binding protein, can be used to enhance the editing efficiency of IscB mutant proteins. Nuclear localization signal proteins can interact with nuclear import vectors, enabling protein transport into the cell nucleus. Fusion with the nuclear localization signal domain can further enhance the editing efficiency of IscB mutant proteins.
[0028] In some embodiments of the present invention, the amino acid sequence of the HMG-D protein is shown in SEQ ID NO:11.
[0029] In some embodiments of the present invention, the IscB mutant protein is fused to the functional protein at its amino terminus.
[0030] In some embodiments of the present invention, the IscB mutant protein and the functional protein are connected via amino acid Linker1.
[0031] The amino acid sequence of Linker1 can be selected based on the specific fusion purpose and may be composed of a single amino acid or several amino acids. For example, the amino acid sequence of Linker1 may be as shown in SEQ ID NO: 16: SGGSSGGSSGSETPGTSESATPESSGGSSGGS.
[0032] In some embodiments of the present invention, the functional protein comprises an HMG-D protein and at least two NLS proteins, wherein the HMG-D protein is fused to the amino terminus of the IscB mutant protein, and the NLS proteins are fused to the amino terminus of the HMG-D protein and the carboxyl terminus of the IscB mutant protein, respectively, and the NLS proteins are independently selected from the group consisting of: sv40 NLS, nucleoplasmin NLS, cmyc NLS, and bpNLS.
[0033] The amino acid sequences of the sv40 NLS, nucleoplasmin NLS, cmyc NLS and bpNLS are shown in SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14 and SEQ ID NO: 15, respectively.
[0034] In some embodiments of the present invention, the amino terminus of the HMG-D protein is coupled to an NLS, and the carboxyl terminus of the IscB mutein is coupled to two NLSs. Preferably, an sv40 NLS is coupled to the amino terminus of the HMG-D protein, and an sv40 NLS and a nucleoplasmin NLS (NLP) are coupled sequentially to the carboxyl terminus of the IscB mutein.
[0035] On the other hand, the present invention also provides a fusion protein comprising the aforementioned IscB mutant protein (nIscB) and adenine deaminase or cytosine deaminase, which can be used as an ultra-small base editor.
[0036] In some specific embodiments of the present invention, the IscB mutant protein is fused to the adenine deaminase or cytosine deaminase at its amino terminus.
[0037] In some specific embodiments of the present invention, the adenine deaminase is TadA-8e protein.
[0038] In some specific embodiments of the present invention, the amino acid sequence of the TadA-8e protein is as shown in SEQ ID NO:57.
[0039] In some specific embodiments of the present invention, TadA-8e is directly fused to the amino terminus of the nIscB mutant protein (nickase IscB).
[0040] In some specific embodiments of the present invention, the IscB mutant protein and adenine deaminase or cytosine deaminase are connected via amino acid Linker2, and the length of amino acid Linker2 is 1-48 amino acids, preferably 1-32 amino acids, further preferably 16-32 amino acids, and more preferably 32 amino acids.
[0041] This linker 2 is sufficient to connect the IscB mutant protein and the deaminase while maintaining their activity. However, a 32-amino acid hinge connecting the amino-terminal TadA8e and the carboxy-terminal IscB mutant protein optimizes the editing efficiency of the entire iABE, primarily editing the adenine at position 3 of the target site. It is understood that the numbers "1, 2, and 3" in this linker 2, Linker 1, and Linker 3 mentioned below serve only as linker designations and have no practical meaning.
[0042] The sequence of the amino acid Linker2 is selected from any one of SEQ ID NO: 63 to SEQ ID NO: 68, preferably the sequence shown in SEQ ID NO: 68.
[0043] In some specific embodiments of the present invention, the fusion protein (ultra-small base editor) is also fused with HMG-D protein.
[0044] In some embodiments of the present invention, the HMG-D protein is located between the IscB protein and adenine deaminase or cytosine deaminase.
[0045] In some specific embodiments of the present invention, the HMG-D protein is fused to the fusion protein (a small base editor fused with nickase IscB and adenine deaminase or cytosine deaminase) via amino acid Linker 3.
[0046] In some embodiments of the present invention, the sequence of Linker 3 is selected from SEQ ID NO:16.
[0047] In some embodiments of the invention, HMG-D is fused between TadA-8e and the nickase IscB via Linker3.
[0048] It is understandable that the fusion of the functional fragments in the above fusion protein can be directly connected or connected through a linker, and can be adjusted according to the actual application scenario and common knowledge in the field.
[0049] On the other hand, the present invention also provides an ω RNA comprising a backbone sequence and a guide sequence, wherein the backbone sequence can form a complex with the wild-type Iscb protein, the aforementioned IscB mutant protein or the aforementioned fusion protein.
[0050] The above-mentioned ωRNA can form a complex with the IscB mutant protein through the backbone sequence, guide it to the target DNA through the guide sequence, cut the double-stranded DNA to form DSB, or perform editing, or simply play a role in guiding positioning.
[0051] The present invention modifies the single-stranded RNA (i.e., ωRNA) that performs a guide function, such as replacing part of the ribonucleic acid and shortening the length of the entire chain. This can maintain editing activity while being more conducive to chemical synthesis or delivery system packaging for subsequent commercial applications.
[0052] In some specific embodiments of the present invention, the backbone sequence comprises a basic backbone sequence or a deformed backbone sequence, the sequence of the basic backbone sequence is shown in SEQ ID NO: 3, and the deformed backbone sequence is obtained by replacing or truncating the basic backbone sequence, and the replacement is to replace the incompletely paired bases in the basic backbone sequence with completely paired bases, or to replace AT in the basic backbone sequence with CG.
[0053] In the present invention, an attempt was made to shorten the length of ωRNA without reducing editing efficiency. The secondary structure of ωRNA contains multiple stem-loop structures. The present invention shortens the "stems" in different stem-loop structures, replaces incompletely paired bases or replaces AT bases with CG bases, ultimately obtaining ωRNA with significantly shortened length while maintaining editing efficiency.
[0054] In some specific embodiments of the present invention, the substitution is replacing GAAG in the basic backbone sequence with GAAA or replacing UGCG in the basic backbone sequence with UGAA.
[0055] In some specific embodiments of the present invention, the truncation is to truncate the stem in the stem-loop structure of the basic backbone sequence. For example, the stem portion can be truncated on one side by 1-15 nt.
[0056] In some specific embodiments of the present invention, the backbone sequence of the ωRNA is selected from any one of the sequences shown in SEQ ID NO: 32-SEQ ID NO: 56; preferably any one of the sequences shown in SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 39, SEQ ID NO: 40, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 47, and SEQ ID NO. 56, and more preferably any one of the sequences shown in SEQ ID NO: 32, SEQ ID NO: 33, SEQ ID NO: 34, SEQ ID NO: 39, SEQ ID NO: 44, SEQ ID NO: 45, SEQ ID NO: 47, and SEQ ID NO. 56.
[0057] In some specific embodiments of the present invention, the bases from positions +15 to +20 and +25 to +29 of the P1 stem-loop structure are deleted, and the UGCG in the loop is replaced with UGAA to obtain a deformed backbone sequence, namely P1-del11.
[0058] In some specific embodiments of the present invention, the bases from positions +14 to +20 and +25 to +30 of the P1 stem-loop structure are deleted, and the UGCG in the loop is replaced with UGAA to obtain a deformed backbone sequence, namely P1-del13.
[0059] In some specific embodiments of the present invention, the bases from positions +13 to +20 and +25 to +31 of the P1 stem-loop structure are deleted, and the UGCG in the loop is replaced with UGAA to obtain a deformed backbone sequence, namely P1-del15.
[0060] In some specific embodiments of the present invention, the bases from positions +12 to +20 and +25 to +32 of the P1 stem-loop structure are deleted, and the UGCG in the loop is replaced with UGAA to obtain a deformed backbone sequence, namely P1-del17.
[0061] In some specific embodiments of the present invention, the bases at positions +66 and +71 of the P2 stem-loop structure are deleted, and the GAAG in the loop is replaced with GAAA to obtain a modified backbone sequence, namely P2-del2.
[0062] In some specific embodiments of the present invention, the bases at positions +65-66 and +71-72 of the P2 stem-loop structure are deleted, and the GAAG in the loop is replaced with GAAA to obtain a deformed backbone sequence, namely P2-del4.
[0063] In some specific embodiments of the present invention, the P5 stem-loop structure is deleted from the bases at positions +174 and +178 to obtain a deformed backbone sequence, namely P5-del2.
[0064] In some specific embodiments of the present invention, the P5 stem-loop structure is deleted from bases at positions +173-174 and +178-179 to obtain a deformed backbone sequence, namely P5-del4.
[0065] In some specific embodiments of the present invention, the C at position +103 in the P3 stem-loop structure is replaced by A, destroying the structure of continuous C in the P3 stem-loop, and replacing C with A to obtain a deformed backbone sequence, namely, P3-polyC destruction.
[0066] In some specific embodiments of the present invention, the AU pairing at positions +11 and +33 in the P1 stem-loop structure is replaced by a GC pairing to obtain a deformed backbone sequence, P1-GC.
[0067] At the same time, the above-mentioned better truncation or base replacement methods were combined and experimentally verified.
[0068] In some specific embodiments of the present invention, P1-del15 is combined with P1-GC to obtain a deformed backbone sequence, namely P1-del15-P1-GC.
[0069] In some specific embodiments of the present invention, P1-del15 and P2-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P2-del4.
[0070] In some embodiments of the present invention, P1-del15 is combined with P3-polyC disruption to obtain a deformed backbone sequence, namely P1-del15-P3-polyC disruption.
[0071] In some embodiments of the present invention, the P1-del15, P2del4 and P3-polyC disruptions are combined to obtain a deformed backbone sequence, namely, the P1-del15-P2-del4-P3-polyC disruption.
[0072] In some embodiments of the present invention, the P1-del15, P1-GC and P3-polyC disruptions are combined to obtain a deformed backbone sequence, namely, P1-del15-P1-GC-P3-polyC disruption.
[0073] In some embodiments of the present invention, P1-del15, P3-polyC disruption and P5-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P3-polyC disruption-P5-del4.
[0074] In some specific embodiments of the present invention, P1-del15, P1-GC and P2-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P2-del4.
[0075] In some specific embodiments of the present invention, P1-del15, P1-GC and P5-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P5-del4.
[0076] In some specific embodiments of the present invention, P1-del15, P2-del4, P3-polyC-disruption and P5-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P2-del4-P3-polyC-disruption-P5-del4.
[0077] In some specific embodiments of the present invention, P1-del15, P3-polyC-disruption, P5-del4 and del-loop are combined to obtain a deformed backbone sequence, namely P1-del15-P3-polyC-disruption-P5-del4-del-loop.
[0078] In some embodiments of the present invention, P1-del15, P1-GC, P2-del4 and P3-polyC-disruption are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P2-del4-P3-polyC-disruption.
[0079] In some specific embodiments of the present invention, P1-del15, P1-GC, P3-polyC-disruption and P5-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P3-polyC-disruption-P5-del4.
[0080] In some specific embodiments of the present invention, P1-del15, P1-GC, P3-polyC-destruction and del-loop are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P3-polyC-destruction-del-loop.
[0081] In some specific embodiments of the present invention, P1-del15, P1-GC, P2-del4, P3-polyC-disruption and P5-del4 are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P2-del4-P3-polyC-disruption-P5-del4.
[0082] In some specific embodiments of the present invention, P1-del15, P1-GC, P2-del4, P3-polyC-destruction, P5-del4 and del loop are combined to obtain a deformed backbone sequence, namely P1-del15-P1-GC-P2-del4-P3-polyC-destruction-P5-del4-del-loop.
[0083] On the other hand, the present invention also provides a base editing system comprising the above-mentioned IscB mutant protein and / or the above-mentioned fusion protein.
[0084] In some specific embodiments of the present invention, the base editing system also includes the above-mentioned ωRNA.
[0085] On the other hand, the present invention also provides a polynucleotide encoding the aforementioned IscB mutant protein, or encoding the aforementioned fusion protein, or encoding the aforementioned ω RNA.
[0086] On the other hand, the present invention also provides a recombinant expression vector comprising the polynucleotide as described above.
[0087] On the other hand, the present invention also provides a delivery system, comprising a vector comprising a polynucleotide encoding the aforementioned IscB mutant protein, a polynucleotide encoding the aforementioned fusion protein, and / or a polynucleotide encoding the aforementioned ω RNA.
[0088] In some specific embodiments of the present invention, the delivery system is selected from: an adenoviral delivery system, a lipid nanoparticle delivery system and a lentiviral delivery system, preferably an adenoviral delivery system.
[0089] The IscB mutant protein of the present invention has a sequence length only one-third that of the Cas9 protein, but it also has the ability to induce double-stranded DNA breaks mediated by single-stranded RNA. Therefore, the above base editing system is more conducive to packaging in an AAV virus for in vivo delivery.
[0090] On the other hand, the present invention also provides the use of the above-mentioned IscB mutant protein, the above-mentioned fusion protein, the above-mentioned ωRNA, the above-mentioned base editing system, the above-mentioned polynucleotide, the above-mentioned recombinant expression vector, and the above-mentioned delivery system in the preparation of reagents and / or drugs for gene editing.
[0091] In some embodiments, the drug is a drug for treating a genetic disease, tumor, or autoimmune disease. The genetic diseases include, but are not limited to, β-thalassemia, Pompe disease, phenylketonuria, hereditary deafness, primary hyperoxaluria, hemophilia, hypercholesterolemia, etc.; the tumors include large B-cell lymphoma, small cell lung cancer, etc.; and the autoimmune diseases include systemic lupus erythematosus and scleroderma, etc.
[0092] The present invention also provides a base editing method, wherein the target gene to be edited is contacted with the above-mentioned base editing system; preferably, the base editing method is used for non-diagnostic or therapeutic purposes.
[0093] In some embodiments, the above-mentioned base editing methods are used for non-diagnostic or therapeutic purposes, such as scientific research, preparation of cell models, animal models, drug preparation, etc.
[0094] On the basis of conforming to the common sense in this field, the above-mentioned conditions can be arbitrarily combined to obtain the preferred embodiments of the present invention.
[0095] It is understood that in the present invention, when referring to amino acid replacements, or describing the type of amino acid at a certain position in the IscB mutant protein, or using amino acids as linker fusion proteins, the described amino acids refer to the amino acid residues that constitute the protein.
[0096] It is understood that the stem-loop structure in the ω RNA backbone sequence of the present invention refers to the RNA chain folding back on itself, and the complementary sequences pairing to form a localized double helix, called the stem or arm. The unpaired bases between the two segments remain single-stranded, forming a protruding loop (loop), i.e., a ring. The stem and loop constitute a stem-loop structure (stem-loop), also known as a hairpin structure (hairpin structure).
[0097] The positive progress effect of the present invention is:
[0098] The IscB mutant protein of the present invention is obtained through analysis and experimental exploration of the ternary complex structure formed by IscB, ω RNA, and a double-stranded DNA (dsDNA) substrate. Mutation of the above-mentioned key amino acid sites greatly enhances the interaction between the protein and the double-stranded DNA target, thereby improving the editing activity in mammalian cell lines.
[0099] Furthermore, the present invention also constructs fusion proteins using IscB mutants as a backbone to further enhance their editing activity against double-stranded DNA targets, or integrates deaminases to enable precise single-base editing. This significant technical improvement to the IscB protein will greatly enhance its potential as a next-generation tool for gene editing, promoting its application in precision medicine, animal model development, and crop breeding.
[0100] In particular, compared to the 4104nt nucleic acid sequence of the existing Cas9 protein expression sequence, the IscB mutant protein of the present invention has a DNA sequence of only 1485nt, yet still possesses the ability to induce double-stranded DNA breaks mediated by single-stranded RNA. This makes it more feasible to use this as a framework for constructing various derivative tools, such as single-base editing tools, and to package these tools within an AAV virus for in vivo delivery. This facilitates large-scale commercial production for gene and cell therapies, while significantly reducing commercialization costs and possessing great application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] Figure 1 is a schematic diagram of the crystal structure of the IscB-ωRNA complex bound to substrate DNA (PDB: 7XHT);
[0102] FIG2 is the Sanger sequencing result after amino acid position 96 in the IscB mutant protein plasmid was substituted with arginine (R);
[0103] FIG3 shows the DNA sequence of amino acid position 96 of the wild-type IscB protein;
[0104] Figure 4 shows the editing efficiency of DNMT1-sg2 by IscB mutant proteins with different single-point mutations;
[0105] Figure 5 shows the editing efficiency of DNMT1-sg2 by the double-point mutation IscB mutant protein;
[0106] Figure 6 shows the editing efficiency of DNMT1-sg2 by triple-point-mutated IscB mutant proteins;
[0107] FIG7 shows the effect of the fusion position of HMG-D and IscB mutant proteins on editing efficiency;
[0108] FIG8 shows the effect of different fusion modes of NLS and IscB mutant proteins on editing efficiency;
[0109] Figure 9 is a schematic diagram of the structure of ωRNA;
[0110] Figure 10 shows the editing efficiency of EMX1-sg1 after modification of different sites of ωRNA;
[0111] Figure 11 shows the editing efficiency of EMX1-sg1 by combining different modifications of ωRNA;
[0112] Figure 12 shows the comparison of ABE editing efficiency from A to G on VEGFA-sg11 by selecting different nick mutation sites based on IscB as the backbone;
[0113] FIG13 shows the effects of different hinge lengths on target editing position and editing efficiency at the VEGFA-sg11 target site;
[0114] Figure 14 shows the effect of HMG-D fusion at different positions on adenine catalysis in iABE (N: amino terminus, M: middle, C: carboxyl terminus);
[0115] Figure 15 shows the results of the screening of Tyr targets;
[0116] FIG16 shows the results of screening for PCSK9 targets. DETAILED DESCRIPTION
[0117] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. Preferred embodiments of the present invention are shown in the accompanying drawings. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present disclosure.
[0118] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0119] Unless otherwise specified, the reagents used in the following examples are all commercially available; the methods used in the following examples are all conventional methods unless otherwise specified.
[0120] Example 1
[0121] Design and verification of IscB mutant proteins (single-point mutations).
[0122] 1. Methods
[0123] 1. Design
[0124] Based on the cryo-electron microscopic structure of the ternary complex formed by IscB, ωRNA, and a double-stranded DNA (dsDNA) substrate (Figure 1), we analyzed the protein structure and hypothesized that certain amino acid sites are key regions for IscB protein binding to the dsDNA substrate. We designed a series of single-point mutations targeting these key amino acid sites to affect protein activity by altering the hydrophobicity or polarity of these amino acids. We also designed an endogenous target site, DNMT1-sg2, from a human gene (DNMT1), for screening and evaluation.
[0125] 2. Plasmid construction
[0126] For the plasmid expressing IscB protein, the DNA sequence of the wild-type IscB protein (such as SEQ ID NO: 2) was first synthesized based on the amino acid sequence of the wild-type IscB protein (such as SEQ ID NO: 1). Then, primers were designed according to the amino acid mutation site, and the codon of the mutant amino acid was introduced by PCR amplification. Subsequently, the PCR fragment containing the mutation site was seamlessly cloned and assembled (the kit was Vazyme ClonExpress MultiS One Step Cloning Kit, C113-01).
[0127] For the plasmid expressing ωRNA, the ωRNA backbone sequence contained therein is shown in SEQ ID NO: 3, and the target of the guide sequence is shown in Table 1 below. Two oligos were synthesized, with CACC added to the positive strand and AGCC added to the reverse strand. The synthesized oligos were annealed from 95°C to room temperature and ligated into the U6-ωRNA-EF1α-GFP vector linearized with Bbs1.
[0128] The amino acid sequence of the wild-type IscB protein (OgeuIscB) is as follows:
[0129] The DNA sequence of the wild-type IscB protein (OgeuIscB) is as follows:
[0130] The above ωRNA backbone sequence is as follows:
[0131] The constructed plasmid was subjected to Sanger sequencing to ensure the correctness of the IscB sequence. For example, the Sanger sequencing results for the DNA sequence of amino acid position 96 of the IscB mutant protein, where the wild-type aspartic acid (D) is mutated to arginine (R), are shown in Figures 2-3 , where Figure 2 shows the sequencing result and Figure 3 shows the DNA sequence of amino acid position 96 of the wild-type IscB.
[0132] 3. Cell transfection
[0133] HEK293T cells (commercially available ATCC CRL-3216 cell line) were revived and passaged twice before use in transfection experiments. On day 1, HEK293T cells were cultured in a 10 cm dish. The specific procedures were as follows:
[0134] (1) Add trypsin to the culture dish and place it in a 37°C incubator for 2 minutes to digest the HEK293T cells;
[0135] (2) DMEM medium with serum was added to neutralize trypsin, and the cells were pipetted to suspend the adherent HEK293T cells;
[0136] (3) Collect the cell solution in a 1.5 ml centrifuge tube, centrifuge at 1000 rpm for 3 minutes, and remove the supernatant;
[0137] (4) Add DMEM medium with serum to resuspend the cells and count the cells according to 2×10 5 Cells / well were seeded in 24-well plates.
[0138] Perform transfection on the second day, observe the cell status in each well, and perform transfection after the cell density reaches 70%-90% and the status is normal.
[0139] The amount of plasmid transfection per well was: IscB mutant plasmid:ωRNA plasmid = 750 ng:250 ng. The transfection reagent was PEI (Polyethylenimine), and the amount added was 3 μL PEI for every 1 μg of plasmid. Three wells were set up in each group.
[0140] Table 1. The ωRNA sequences targeting the target sites are as follows:
[0141] 4. Genome Extraction and Amplicon Library Preparation
[0142] 72h after transfection, use QuickExtract TMGenomic DNA was extracted using DNA Extraction Solution (QE09050, Epicentre). Following the Hitom kit's protocol, corresponding identification primers were designed, including a bridging sequence 5'-GGAGTGAGTACGGTGTGC-3' (SEQ ID NO: 7) at the 5' end of the forward identification primer and a bridging sequence 5'-GAGTTGGATGCTGGATGG-3' (SEQ ID NO: 8) at the 5' end of the reverse identification primer. This yielded a single PCR product, which was then used as a template for a second round of PCR amplification. The constructed Hitom sample library was gel-purified and sent to the company for sequencing (sequencing service provider: Jiangxi Hipulos Biotechnology Co., Ltd.).
[0143] Table 2. Target identification primers used
[0144] In the table, F represents the forward identification primer, and R represents the reverse identification primer.
[0145] 5. Analysis and statistics of deep sequencing results
[0146] Bioinformatics software was used to write scripts to batch calculate the insertion / deletion ratio and base substitution ratio in the editing results, and GraphPad Prism 9.1.0 was used for statistical mapping.
[0147] 2. Results
[0148] DNMT1 editing activity
[0149] The following table and Figure 4 exemplify the editing activity of some IscB mutant proteins on the DNMT1 gene. In Figure 4, the horizontal axis is the relative ratio (average) of the editing activity of each mutant working system to the wild-type IscB protein, calculated by dividing the editing efficiency of each mutant working system by the editing efficiency of wild-type IscB (relative efficiency), and the vertical axis is the specific editing efficiency (absolute efficiency). Each point is an IscB mutant protein with a single-point substitution mutation at a different site.
[0150] Table 3. Editing activity of IscB mutant proteins on the DNMT1 gene
[0151] The number is the number of each experimental group. The information under the mutation item is the mutation type. Indel is the editing efficiency, which is calculated by dividing the number of reads with deletion and insertion by the total reads.
[0152] The above only shows the experimental screening results of some IscB mutant proteins. It can be seen from the results that some of the above-mentioned IscB mutant proteins have editing activities comparable to or better than the wild-type IscB protein. The mutated D96R mutant protein increases the editing activity from less than 5% of wild-type IscB to nearly 30%, far exceeding the wild-type IscB.
[0153] The above results show that mutations in nine amino acids, namely E84, D96, M102, C111, V159, H368, A401, S456, and S386, significantly improve the editing efficiency of the protein at the target. Taking into account the structural analysis of the ternary complex formed by IscB, ωRNA, and double-stranded DNA (dsDNA) substrate, it is believed that the three amino acid sites E84, D96, and V159 have greater potential.
[0154] Example 2
[0155] Design and validation of IscB mutant proteins (combination mutations).
[0156] To further improve the editing efficiency of the IscB mutant protein, this example, based on the results of the single-point mutation in Example 1 and according to the structure of IscB, combined screening of amino acid sites that potentially affect the substrate structure was performed.
[0157] 1. Methods
[0158] The plasmid was constructed and tested according to the method of Example 1.
[0159] 2. Results
[0160] 1. Double-point mutation screening
[0161] The results of the double-point mutation screening are shown in the following table and Figure 5.
[0162] Table 4. Editing activity of IscB mutant protein (double point mutation) on DNMT1 gene
[0163] The results showed that the editing efficiency of E84R-D96R, D96R-S456R, D96R-H368R, and D96R-S386R for the DNMT1 target site exceeded 30%, with excellent editing efficiency, especially the E84R-D96R mutant, which exceeded 40%.
[0164] 2. Three-point mutation screening
[0165] The results of the three-point mutation screening are shown in the following table and Figure 6.
[0166] Table 5. Editing activity of IscB mutant protein (three-point mutation) on DNMT1 gene
[0167] The results showed that the triple-point mutation IscB mutant protein can further improve the editing efficiency. Among them, the IscB mutant proteins of D96R-E84R-V159R (hereinafter referred to as zh258), D96R-E84R-C111R, D96R-E84R-S386R and D96R-E84R-H368R have an editing efficiency of more than 50%.
[0168] Example 3
[0169] Fusion protein design and validation.
[0170] To further enhance the editing efficiency of the IscB protein, this example builds on previous modifications by fusing the HMG-D protein to the IscB protein backbone, attempting to leverage its strong binding ability to double-stranded DNA to further enhance IscB's gene editing efficiency. The effect of increasing the number of nuclear localization signals (NLS) on the editing efficiency of the fusion protein was also tested.
[0171] 1. Methods
[0172] 1. Plasmid design and construction
[0173] DNA expression templates containing different HMG-D protein, NLS sequences, and IscB protein linkages were amplified using PCR. Complementary cohesive ends were designed at the 5' and 3' ends of the three fragments, followed by assembly and cloning. The constructed plasmids were confirmed to be completely correct using Sanger sequencing before further manipulation.
[0174] The HMG-D protein sequence of this example is:
[0175] The amino acid sequence of sv40NLS (sv40) is: PKKKRKV (SEQ ID NO: 12).
[0176] The amino acid sequence of nucleoplasmin NLS (NLP) is: KRPAATKKAGQAKKKK (SEQ ID NO: 13).
[0177] The amino acid sequence of cmyc NLS (cmyc) is: PAAKRVKLD (SEQ ID NO: 14).
[0178] The amino acid sequence of bpNLS is: KRTADGSEFEPKKKRKV (SEQ ID NO: 15).
[0179] The linker connecting HMG-D and IscB protein is Linker1: SGGSSGGSSGSETPGTSESATPESSGG SSGGS (SEQ ID NO: 16)
[0180] 2. Cell transfection and amplification sequencing
[0181] Refer to the method of Example 1.
[0182] The test targets are HBG-sg4, VEGFA-sg2 and ALDH1A3-sg1, and the corresponding ωRNAs are as follows:
[0183] Table 6. Targets and sequences used
[0184] In the table, Oligo-up is the forward primer and Oligo-dn is the reverse primer.
[0185] 2. Results
[0186] 1. Fusion HMG-D
[0187] The effect of the relative position of the fusion of HMG-D protein and IscB mutant protein (i.e., fusion at the amino terminus or carboxyl terminus of the IscB protein) on the editing efficiency was tested using different targets.
[0188] The results are shown in the following table and Figure 7, where IscB represents the wild-type IscB group, HMG-D-zh258 represents the HMG-D protein located at the amino terminus of IscB, zh258-HMG-D represents the HMG-D protein located at the carboxyl terminus of IscB, Untreated represents the negative control group, and zh258 represents the IscB mutant protein with the E84R-D96R-V159R amino acid triple mutation.
[0189] Table 7. Editing efficiency of IscB mutant proteins at different fusion positions of HMG-D (target is HBG)
[0190] Table 8. Editing efficiency of IscB mutant proteins at different fusion positions of HMG-D (target is VEGFA)
[0191] Table 9. Editing efficiency of IscB mutant proteins at different fusion positions of HMG-D (target ALDH1A3)
[0192] The results showed that fusing HMG-D protein to the amino terminus of IscB protein showed optimal editing efficiency at three different targets, achieving an editing efficiency of more than 80%.
[0193] 2. Fusion NLS
[0194] The effects of different fusion modes of NLS and IscB mutant proteins (i.e., fusion at the amino terminus or carboxyl terminus of the IscB protein and the number of fusions) on editing efficiency were tested using different targets. The results are shown in the following table and Figure 8.
[0195] In the table, each group represents a different protein fusion method, group 1 is sv40-iscb-NLP, group 2 is sv40-zh258-NLP, group 3 is cymc-zh258-NLP, group 4 is bpNLS-zh258-bpNLS, group 5 is sv40-zh258-sv40-NLP, group 6 is sv40-HMG-zh258-NLP, group 7 is cymc-HMG-zh258-NLP, group 8 is bpNLS-HMG-zh258-bpNLS, and group 9 is sv40-HMG-zh258-sv40-NLP.
[0196] Table 10. Editing efficiency of IscB mutant proteins with different NLS fusion patterns (target is VEGFA)
[0197] Table 11 Editing efficiency of IscB mutant proteins with different NLS fusion patterns (target ALDH1A3)
[0198] Table 12. Editing efficiency of IscB mutant proteins with different NLS fusion patterns (target is HBG)
[0199] The results showed that among different targets, group 9 (sv40-HMG-zh258-sv40-NLP) had better editing efficiency, that is, coupling one NLS at the amino terminus of HMG-D and two NLS at the carboxyl terminus of IscB protein could improve the editing efficiency of HMG-D-IscB fusion protein for each target to varying degrees.
[0200] Example 4
[0201] This example uses truncation or replacement to modify ωRNA to varying degrees without reducing the overall editing efficiency.
[0202] 1. Methods
[0203] 1. Plasmid design and construction
[0204] The inventors analyzed the ωRNA, as shown in Figure 9. According to the stem-loop structure (also known as the backbone sequence), it can be marked as seven different regions, P1-P5 and J1-J2. Based on the structural biology research on the ωRNA and IscB protein complex, the interactions between the ribonucleic acid in different regions and the different amino acids of the protein are considered. For example, the guide region in the ωRNA sequence that pairs with the DNA target interacts with specific amino acids in the RuvC domain, REC domain and Bridge Helix (BH) region of IscB, respectively. Some bases in the J1 region interact with arginine or lysine in the BH region or REC region, respectively. Nucleic acid sequences that do not directly interact with amino acids are also suspected to be indirectly involved in stabilizing the structure of the entire nucleic acid-protein complex.
[0205] Although the specific functions of the nucleic acid sequences in each region require further study, attempts can still be made to replace some of the RNA in these seven regions, such as replacing purine with pyrimidine, or truncating the sequence, in order to further improve or maintain the indel efficiency of the current complex for the tested target while minimizing the sequence length.
[0206] These modified ωRNAs were then combined with wild-type IscB protein and the optimal mutant of IscB protein (zh258: D96R-E84R-V159R), assembled and cloned after PCR. The constructed plasmids were confirmed to be completely correct by Sanger sequencing, and these ωRNA-modified complexes were then tested using the EMX1 target site.
[0207] Based on the structure of ωRNA, base mutations or sequence truncations were performed on DNA templates in different regions of the sequence, and assembly and cloning were performed after PCR. The constructed plasmids were sequenced by Sanger sequencing to ensure complete accuracy.
[0208] 2. Cell transfection and amplification sequencing
[0209] The method of Example 1 was used. The test target was EMX1.
[0210] Table 13. Targets and sequences used
[0211] In the table, Oligo-up is the forward primer and Oligo-dn is the reverse primer.
[0212] Table 14. Target identification primers used
[0213] In the table, F represents the forward identification primer, and R represents the reverse identification primer.
[0214] 2. Results
[0215] 1. The editing efficiency experimental results of combining each ωRNA with the optimal mutant (zh258: D96R-E84R-V159R) are shown in the following table and Figures 10 and 11.
[0216] Table 15. Editing efficiency of each ωRNA (truncated or replaced) (target is EMX1)
[0217] Table 16. Editing efficiency of different ωRNAs (truncated or replaced combinations) (target: EMX1)
[0218] The above results show that the ωRNA (165 nt) shown in SEQ ID NO: 56 can improve editing efficiency. Compared with the wild-type sequence length of 206 nt, it is 41 nt shorter. This modification reduces the difficulty of chemical synthesis of this single-stranded RNA, and some ωRNAs can further improve the system's editing efficiency.
[0219] Example 5
[0220] Construction of an adenine base editor based on the IscB mutant protein.
[0221] The purpose of this embodiment is to transform the IscB protein into a nickase that can only degrade single-stranded DNA or a form that has lost its degradation function (dead IscB, dIscB), thereby constructing an ultra-small base editing tool.
[0222] 1. Methods
[0223] 1. Plasmid design and construction
[0224] Based on the comparison results of the degree of amino acid conservation in the primary structure of the protein, and with reference to the amino acid sites of the Cas9 protein responsible for catalyzing DNA breakage, the four amino acids D60, E192, H339, and D342 were mutated to alanine (A) so that the IscB protein only has the ability to break single-stranded DNA, or the two sites H246 and D342 were mutated at the same time to make the IscB protein completely lose the ability to cut double-stranded DNA. Subsequently, TadA8e was fused on the basis of these mutants to construct an adenine base editor (IscB-ABE, iABE or diABE).
[0225] The TadA8e sequence is:
[0226] DNA expression templates for TadA8e and modified IscB proteins were amplified by PCR, then assembled and cloned. The constructed plasmids were sequenced by Sanger sequencing to ensure complete accuracy.
[0227] The VEGFA target was then used to test the editing activity and editing window width of these five iABEs on adenine in the target site.
[0228] 2. Cell transfection and amplification sequencing
[0229] Refer to the method of Example 1.
[0230] Table 17. Targets and sequences used
[0231] In the table, Oligo-up is the forward primer and Oligo-dn is the reverse primer.
[0232] Table 18. Target identification primers used
[0233] In the table, F represents the forward identification primer, and R represents the reverse identification primer.
[0234] 2. Results
[0235] 1. A->G single-base editing activity
[0236] The results of mutating amino acids D60, E192, H339, and D342 to alanine (A) and simultaneously mutating H246 and D342 to alanine are shown in the following table and Figure 12. In the figure, the horizontal axis represents the adenine sites in the VEGFA target sequence (ACAGGTGTGAAAACAG), such as A1 is the first A in the target sequence, and A3 is the second A in the target sequence (the underline in the above sequence represents each site), and the vertical axis represents the single-base editing efficiency of A->G.
[0237] Table 19. Editing efficiency of different ωRNAs (target: VEGFA)
[0238] Test results show that the H339A mutation has the best A->G single-base editing activity for the target sequence while maintaining the activity of cutting single-stranded DNA, and its main editing site is A3 at the third position.
[0239] 2. Effects of different amino acid hinges on editing activity
[0240] The length of the amino acid hinge (linker) connecting IscB and TadA was changed to investigate its effect on editing activity and position. Group 1 was untreated, i.e., an untreated negative control; Group 2 was TadA-0aa-iscb, without a linker; Group 3 was TadA-5aa-iscb, i.e., a linker with a 5-amino acid hinge; Group 4 was TadA-9aa-iscb, i.e., a linker with a 9-amino acid hinge; Group 5 was TadA-12aa-iscb, i.e., a linker with a 12-amino acid hinge; Group 6 was TadA-15aa-iscb, i.e., a linker with a 15-amino acid hinge; Group 7 was TadA-16aa-iscb, i.e., a linker with a 16-amino acid hinge; Group 8 was TadA-32aa-iscb, i.e., a linker with a 32-amino acid hinge. The fusion proteins in groups 2-8 above were all fusions of the amino terminus of the IscB mutant protein with TadA via a linker.
[0241] Group 9 is iscb-0aa-TadA, i.e., no linker; group 10 iscb-5aa-TadA, i.e., a linker with a 5-amino acid hinge; group 11 iscb-9aa-TadA, i.e., a linker with a 9-amino acid hinge; group 12 iscb-12aa-TadA, i.e., a linker with a 12-amino acid hinge; group 13 iscb-15aa-TadA, i.e., a linker5 with a 15-amino acid hinge; group 14 iscb-16aa-TadA, i.e., a linker with a 16-amino acid hinge; group 15 iscb-32aa-TadA, i.e., a linker with a 32-amino acid hinge; the fusion proteins of groups 9-15 above are all the carboxyl terminus of the IscB mutant protein fused to TadA through a linker.
[0242] The amino acid sequence of the linker of the above 5 amino acid hinge is shown in SEQ ID NO:63: GGGGS. The amino acid sequence of the linker of the above 9 amino acid hinge is shown in SEQ ID NO:64: GGSGGSGGS. The amino acid sequence of the linker of the above 12 amino acid hinge is shown in SEQ ID NO:65: AEAAAKEAAAKA. The amino acid sequence of the linker of the above 15 amino acid hinge is shown in SEQ ID NO:66: EAAAKEAAAKEAAAK. The amino acid sequence of the linker of the above 16 amino acid hinge is shown in SEQ ID NO:67: SGSETPGTSESATPES. The amino acid sequence of the linker of the above 32 amino acid hinge is shown in SEQ ID NO:68: SGGSSGGSSGSETPGTSESATPE SSGGSSGGS.
[0243] The results are shown in the following table and Figure 13.
[0244] Table 20. Editing efficiency Indel% of different hinge lengths and protein fusion positions (target is VEGFA)
[0245] The results showed that a 32-amino acid hinge connecting the amino-terminal TadA and the carboxyl-terminal IscB could optimize the editing efficiency of the entire iABE, and the IscB mutant protein had good editing efficiency when fused to TadA through the amino-terminus.
[0246] 3. Effects of different HMG-D fusion positions on editing activity
[0247] The effects of fusing HMG-D protein to different regions of iABE on editing activity are shown in the following table and Figure 14 .
[0248] Table 21. Editing efficiency of different HMG-D fusion positions (target is VEGFA)
[0249] The “C”, “M” and “N” in the names of each group indicate that HMG-D is fused to the carboxyl terminus, the middle terminus and the amino terminus, respectively.
[0250] The results showed that HMG-D inserted between TadA and IscB had the best editing effect on adenine.
[0251] Example 6
[0252] This example studies the application of an editing system based on IscB mutant protein to explore its potential application in the treatment of albinism and lipid-lowering.
[0253] 1. Methods
[0254] 1. Target screening of Tyr gene and PCSK9 gene in albinism
[0255] A 16nt target was designed based on the TAM sequence recognized by IscB and ligated to the 5' end of the ω RNA backbone plasmid after synthesis. The constructed plasmid was sequenced using Sanger sequencing to ensure complete accuracy before proceeding with further work.
[0256] 2. Cell transfection and amplification sequencing
[0257] The method of Example 1 was used, and the cells used were Neuro-2a (N2A) cells.
[0258] 3. Analysis and statistics of deep sequencing results
[0259] Use indel bioinformatics tools to analyze and select 1-2 targets with the best editing effects in the genome.
[0260] 4. In vitro transcription of IscB protein mRNA and ω RNA
[0261] (1) Screening of Tyr genes
[0262] After screening the Tyr gene targets for albinism, we selected the following suitable targets.
[0263] Table 22. Targets and sequences used
[0264] In the table, Oligo-up is the forward primer and Oligo-dn is the reverse primer.
[0265] Table 23. Target identification primers used
[0266] In the table, F represents the forward identification primer, and R represents the reverse identification primer.
[0267] (2) Screening of PCSK9 gene
[0268] After screening PCSK9 targets that control blood lipid concentrations, we selected the following suitable targets.
[0269] Table 24. Targets and sequences used
[0270] In the table, Oligo-up is the forward primer and Oligo-dn is the reverse primer.
[0271] Table 25. Target identification primers used
[0272] In the table, F represents the forward identification primer, and R represents the reverse identification primer.
[0273] 2. Results
[0274] 1. Editing effect on albinism (Tyr) target
[0275] (1) Target screening results
[0276] The screening results for Tyr targets are shown in the following table and Figure 15.
[0277] Table 25. Target screening results
[0278] The results showed that we obtained the target Tyr-sg21 with better editing efficiency, and its editing efficiency can reach more than 40%.
[0279] 2. Editing results of the lipid-lowering (PCSK9) target
[0280] (1) Target screening results
[0281] The target screening results are shown in the following table and Figure 16.
[0282] Table 28. Target screening results
[0283] The results showed that the editing efficiency of the mPCSK9 gene target site could reach more than 50%.
[0284] Example 7
[0285] In this example, ALDH1A3 and EMX1 were used as targets to test the editing efficiency of various mutant proteins and fusion proteins. The guide sequence and test sequence used for ALDH1A3 were the same as those in Example 3, and the guide sequence and test sequence used for EMX1 were the same as those in Example 4.
[0286] 1. Methods
[0287] The following IscB mutants and fusion proteins were selected for testing: IscB (wild type), D96R, zh258, HMG-D-zh258, A401R-S456R, and Q376R-S456R.
[0288] Among them, A401R-S456R is a mutant protein in which the amino acid at position 401 of the wild-type Iscb is mutated from A to R, and the amino acid at position 456 is mutated from S to R. Q376R-S456R is a mutant protein in which the amino acid at position 376 of the wild-type Iscb is mutated from Q to R, and the amino acid at position 456 is mutated from S to R. They were prepared according to the method of CN116583599A.
[0289] Subsequently, the plasmid was constructed and tested according to the method of Example 1.
[0290] 2. Results
[0291] The editing efficiency of each mutant protein and fusion protein on ALDH1A3 and EMX1 targets is shown in the following table.
[0292] Table 29. Editing efficiency of ALDH1A3 and EMX1 targets
[0293] The above results show that compared with the wild-type protein, the mutant proteins D96R, zh258 and HMG-D-zh258 all exhibited better editing efficiency, and were better than the existing mutant proteins A401R-S456R and Q376R-S456R, especially the HMG-D-zh258 protein, which had the best editing effect.
[0294] Example 8
[0295] This example targets ALDH1A3 and EMX1, testing the editing efficiency of the ω RNA (P1-del15-P1-GC-P2-del4-P3-polyC-break-P5-del4-del loop) described in Example 4, using the backbone sequence set forth in SEQ ID NO:56, in combination with different IscB proteins. The guide and test sequences used for ALDH1A3 are the same as those in Example 3, and those used for EMX1 are the same as those in Example 4.
[0296] 1. Methods
[0297] The following IscB mutant proteins were selected for testing: Iscb (wild type), D96R, and zh258. Plasmids were constructed and tested according to the method of Example 1, except that the backbone sequence of the ωRNA used in this example is shown in SEQ ID NO: 56.
[0298] 2. Results
[0299] The editing efficiency of truncated ω RNA against ALDH1A3 and EMX1 targets is shown in the table below.
[0300] Table 29. Editing efficiency of ALDH1A3 and EMX1 targets Note: NT means not tested.
[0301] The above results show that after selecting this shortened ωRNA, the editing efficiency was significantly improved compared with the 206nt wild-type ωRNA selected in Example 7, indicating that shortening the sequence length of ωRNA can, on the one hand, reduce the difficulty of chemical synthesis of the single-stranded RNA, and on the other hand, improve the system editing efficiency.
[0302] Although the above describes specific embodiments of the present invention, it should be understood by those skilled in the art that these are merely illustrative and that various changes or modifications may be made to these embodiments without departing from the principles and essence of the present invention. Therefore, the scope of protection of the present invention is defined by the appended claims.
Claims
1. An IscB mutant protein, characterized in that, Compared with the wild-type IscB protein, the IscB mutant protein has arginine at at least one of the following amino acid positions: position 84, position 96, position 102, position 111, position 159, position 368, position 386; preferably, it also includes arginine at position 401 and position 456 amino acids.
2. The IscB mutant protein according to claim 1, wherein The IscB mutant protein contains two substitution mutations or three substitution mutations. One of the substitution mutations in the two substitution mutations is D96R, and two of the substitution mutations in the three substitution mutations are E84R and D96R.
3. The IscB mutant protein according to claim 1, characterized in that, The IscB mutant protein contains two substitution mutations or three substitution mutations. The two substitution mutations are any two of E84R, D96R, H368R, S386R, and S456R, and the three substitution mutations are any three of E84R, D96R, V159R, H368R, S386R, C111R, K92R, A401R, S456R, K118R, M102R, and N167R.
4. The IscB mutant protein according to claim 1, wherein, The IscB mutant protein contains two substitution mutations or three substitution mutations. The two substitution mutations are E84R-D96R, D96R-H368R, D96R-S386R, or D96R-S456R, preferably E84R-D96R, D96R-H368R, or D96R-S386R, more preferably E84R-D96R. The three substitution mutations are E84R-D96R-V159R, E84R-D96R-H368R, E84R-D96R-S386R, D96R-E84R-C111R, D96R-E84R-K92R, D96R-E84R-A401R, E84R-D96R-S456R, D96R-E84R-K118R, D96R-E84R-M102R, or E84R-D96R-N167R, preferably E84R-D96R-V159R, E84R-D96R-H368R, E84R-D96R-S386R, or D96R-E84R-C111R, more preferably E84R-D96R-V159R, E84R-D96R-H368R, or E84R-D96R-S386R.
5. The IscB mutant protein according to claim 1, characterized in that, The amino acid sequence of the wild-type IscB protein is as shown in SEQ ID NO:
1.
6. An IscB mutant protein, characterized in that, Compared with the wild-type IscB protein or the IscB mutant protein according to any one of claims 1-5, any one of the amino acids at positions 60, 192, 339, and 342 of the IscB mutant protein is a non-positively charged amino acid; the non-positively charged amino acids are selected from glycine, alanine, valine, leucine, isoleucine, methionine, proline, tryptophan, serine, threonine, tyrosine, cysteine, phenylalanine, asparagine, glutamine, aspartic acid, and glutamic acid, and the non-positively charged amino acid is preferably alanine.
7. The IscB mutant protein according to claim 6, wherein, The IscB mutant protein contains any one of the following substitution mutations: D60A, E192A, H339A, D342A, preferably contains the H339A substitution mutation.
8. An IscB mutant protein, characterized in that, Compared with the wild-type IscB protein or the IscB mutant protein described in any one of claims 1-5, the amino acids at positions 246 and 342 of the IscB mutant protein are substituted with non-positively charged amino acids, resulting in the loss of double-stranded DNA cleavage activity. The non-positively charged amino acids are selected from glycine, alanine, valine, leucine, isoleucine, methionine, proline, tryptophan, serine, threonine, tyrosine, cysteine, phenylalanine, asparagine, glutamine, aspartic acid, and glutamic acid, and the non-positively charged amino acid is preferably alanine.
9. A fusion protein, characterized in that, The fusion protein contains the IscB mutant protein described in any one of claims 1-8 and a functional protein, and the functional protein is selected from at least one of HMG-D protein and nuclear localization signal protein.
10. The fusion protein according to claim 9, wherein, The IscB mutant protein fuses with the functional protein at its amino terminus.
11. The fusion protein according to claim 9, wherein, The IscB mutant protein and the functional protein are linked by an amino acid Linker1.
12. The fusion protein according to claim 9, wherein The functional protein includes an HMG-D protein and at least two segments of NLS protein. The HMG-D protein is fused to the amino terminus of the IscB mutant protein, and the NLS proteins are respectively fused to the amino terminus of the HMG-D protein and the carboxyl terminus of the IscB mutant protein. The NLS proteins are independently selected from: sv40 NLS, nucleoplasmin NLS, cmyc NLS, and bpNLS.
13. The fusion protein according to claim 12, wherein, One NLS is coupled to the amino terminus of the HMG-D protein, and two NLSs are coupled to the carboxyl terminus of the IscB mutant protein. Preferably, one sv40 NLS is coupled to the amino terminus of the HMG-D protein, and one sv40 NLS and one nucleoplasmin NLS are sequentially coupled to the carboxyl terminus of the IscB mutant protein.
14. A fusion protein, characterized in that, Contains the IscB mutant protein described in any one of claims 6-8 and adenine deaminase or cytosine deaminase.
15. The fusion protein according to claim 14, wherein, The IscB mutant protein fuses with the adenine deaminase or cytosine deaminase at its amino terminus.
16. The fusion protein according to claim 14, wherein, The adenine deaminase is the TadA-8e protein.
17. The fusion protein according to claim 14, characterized in that, The IscB mutant protein and the adenine deaminase or cytosine deaminase are linked by an amino acid Linker2. The length of the amino acid Linker2 is 1-48 amino acids, preferably 1-32 amino acids, further preferably 16-32 amino acids, and more preferably 32 amino acids.
18. The fusion protein according to claim 17, wherein, The sequence of the amino acid Linker2 is selected from the sequences shown in any one of SEQ ID NO:63-SEQ ID NO:68, preferably the sequence shown in SEQ ID NO:
68.
19. The fusion protein according to claim 16, wherein, An HMG-D protein is also fused.
20. The fusion protein according to claim 19, wherein, The HMG-D protein is located between the IscB protein and the adenine deaminase or cytosine deaminase.
21. The fusion protein according to claim 19, wherein The HMG-D protein is fused to the fusion protein through an amino acid Linker3.
22. An ωRNA, characterized in that, It contains a scaffold sequence and a guide sequence, and the scaffold sequence can form a complex with a wild-type Iscb protein, an IscB mutant protein according to any one of claims 1-8, or a fusion protein according to any one of claims 9-21.
23. The ωRNA according to claim 22, wherein The scaffold sequence contains a basic scaffold sequence or a modified scaffold sequence. The sequence of the basic scaffold sequence is as shown in SEQ ID NO:31; the modified scaffold sequence is obtained by substitution or truncation of the basic scaffold sequence. The substitution is to replace the unpaired bases in the basic scaffold sequence with completely paired bases, or to replace A-T in the basic scaffold sequence with C-G.
24. The ωRNA according to claim 23, wherein The substitution is to replace GAAG in the basic scaffold sequence with GAAA or to replace UGCG in the basic scaffold sequence with UGAA.
25. The ωRNA according to claim 23, wherein, The truncation is to truncate the stem in the stem-loop structure of the basic scaffold sequence.
26. The ωRNA according to claim 22, wherein, The scaffold sequence is selected from any one of the sequences shown in SEQ ID NO:32 - SEQ ID NO.56; preferably any one of the sequences shown in SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:39, SEQ ID NO:40, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:47, SEQ ID NO.56, and more preferably any one of the sequences shown in SEQ ID NO:32, SEQ ID NO:33, SEQ ID NO:34, SEQ ID NO:39, SEQ ID NO:44, SEQ ID NO:45, SEQ ID NO:47, SEQ ID NO.
56.
27. A base editing system, characterized in that, It contains: the IscB mutant protein according to claims 1-8 and / or the fusion protein according to any one of claims 9-21.
28. The base editing system according to claim 27, wherein It further includes the ωRNA according to any one of claims 22-26.
29. A polynucleotide, characterized in that, The polynucleotide encodes the IscB mutant protein according to any one of claims 1-8, or encodes the fusion protein according to any one of claims 9-21, or encodes the ωRNA according to any one of claims 22-26.
30. A recombinant expression vector, characterized in that, The recombinant expression vector contains the polynucleotide according to claim 29.
31. A delivery system, characterized in that, It includes a vector, and the vector contains a polynucleotide encoding the IscB mutant protein according to any one of claims 1-8, a polynucleotide encoding the fusion protein according to any one of claims 9-21, and / or a polynucleotide encoding the ωRNA according to any one of claims 22-26.
32. The delivery system according to claim 31, wherein The delivery system is selected from: an adenovirus delivery system, a lipid nanoparticle delivery system, and a lentivirus delivery system, preferably an adenovirus delivery system. Use of the IscB mutant protein according to any one of claims 1-8, the fusion protein according to any one of claims 9-21, the ωRNA according to any one of claims 22-26, the base editing system according to any one of claims 27-28, the polynucleotide according to claim 29, the recombinant expression vector according to claim 30, and the delivery system according to any one of claims 31-32 in the preparation of a reagent and / or drug for gene editing.
34. The application according to claim 33, wherein The drug is a drug for treating genetic diseases, tumors or autoimmune diseases.
35. A base editing method, characterized in that, Contact the target gene to be edited with the base editing system according to any one of claims 27-28.