A fusion protein, a highly efficient and specific base editing system containing it, and its applications.

By designing a base editing system for fusion proteins, and utilizing SaCas9 nickase and adenine deaminase ecABE8e for specific base editing, the off-target editing problem of existing base editors is solved, achieving efficient A to G base substitution and high safety.

CN116239703BActive Publication Date: 2026-05-26ZHUHAI JIKANG TECHNOLOGY LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHUHAI JIKANG TECHNOLOGY LTD
Filing Date
2023-03-01
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing base editors suffer from off-target editing events when performing A to G base substitutions, leading to safety issues in clinical applications and making it difficult to achieve efficient and specific base substitutions.

Method used

A fusion protein was designed, comprising SaCas9 nickase, adenine deaminase ecABE8e, a flexible linker sequence, and a nuclear localization signal BPNLS sequence, forming a base editing system. Guided by sgRNA, it specifically recognizes and cuts the bases to achieve A to G base editing.

Benefits of technology

This system achieves a high efficiency of A to G base substitution within cells, reaching 90%, with virtually no off-target DNA and RNA editing events, thus improving safety and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116239703B_ABST
    Figure CN116239703B_ABST
Patent Text Reader

Abstract

This application belongs to the field of biotechnology and relates to an efficient and specific editing system, method, and application for A-G base substitution. The fusion protein described in this application comprises, from the N-terminus to the C-terminus, a first SaCas9 nickase fragment, a chimeric deaminase fragment, and a second SaCas9 nickase fragment. The deaminase is selected from adenosine deaminase or a variant thereof, specifically ecTadA8e. The amino acid sequence of ecTadA8e is shown in SEQ ID NO. 20, or has more than 80% sequence identity with the amino acid sequence shown in SEQ ID NO. 20, and possesses the function or activity of ecTadA8e. The fusion protein provided in this application, combined with the corresponding guide RNA, can efficiently and specifically replace the base A with G at the target site, providing an effective tool for repairing pathogenic mutations, studying gene function, and improving cell function, and has promising application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of biotechnology, specifically to a highly efficient and specific base editing system and its applications. Background Technology

[0002] CRISPR / Cas9-mediated gene editing technology, characterized by its simplicity, efficiency, and versatility, has become the most widely used tool in gene editing research. The CRISPR / Cas9 system works by using the Cas9 protein, which has DNA double-strand cutting activity, to target specific sites and perform double-strand DNA cleavage under the guidance of a specific gRNA. This is followed by cellular repair mechanisms, introducing mutations such as insertion, deletion, or substitution of fragments or bases. Based on this, base editors have been developed as a novel gene editing tool. By fusing inactivated or partially inactivated Cas proteins with deaminases or reverse transcriptases, they can efficiently catalyze base conversion without causing DNA double-strand breaks or requiring a donor DNA template. Therefore, they hold great promise for germplasm improvement and gene therapy. Currently, common base editors include the cytosine base editor (CBE), the adenine base editor (ABE), the Guaninebase editor (GBE), and the prime editor (PE). The first three can perform CT base substitution, AG base substitution, and CG base transversion, respectively, while the latter can perform specific base conversions and the insertion or deletion of small sequence fragments.

[0003] Base editing technology was first developed by David R. Liu's team at Harvard University. They developed the first-generation base editors, CBE and ABE, by fusing a mutant Cas9 protein (nickase Cas9) with either cytosine deaminase or adenine deaminase. Due to the activation of the base excision repair pathway, CBE, in addition to inducing C-to-T base conversion, also produces non-T byproducts and indels (insertion and deletion events) to some extent. Subsequent studies reported severe random off-target effects in CBE-treated cells and embryos. Unlike CBE, first-generation ABEs (such as ABE7.10) do not induce significant indels, and ABE7.10 rarely induces non-Cas9-dependent DNA off-target editing. These excellent characteristics give ABEs many advantages for future clinical applications. Some studies have obtained the more active ABE8e (using the deaminase ecTadA8e) through molecular evolution, but the problems of Cas9-independent DNA and RNA off-target effects are serious. Studies have revealed that the reason ABE8e is prone to off-target effects is that the deaminase ecTadA8e linked to Cas9 is constantly activated, causing Cas9 to continuously bind to DNA and perform its cleavage function before finding the intended target. Reported methods include mutagenesis of the deaminase ecTadA8e, regulation of deaminase expression, and optimization of the editor's delivery tool, but off-target events cannot be completely avoided. This poses a significant safety risk to the clinical application of this tool.

[0004] In summary, developing a highly efficient and specific base editor that can induce base substitutions in target sequences without triggering off-target editing is one of the urgent problems to be solved in the field of gene editing. Summary of the Invention

[0005] Based on this, the purpose of this invention is to provide a fusion protein, a base editing system containing the fusion protein, and its applications. This invention utilizes the SaCas9 nickase, adenine deaminase ecABE8e, a flexible linker sequence, and a nuclear localization signal BPNLS sequence to design a fusion protein. The fusion protein can form a complex with guide RNA (sgRNA) to constitute a base editing system. The sgRNA guides the fusion protein to specifically recognize and cleave the target sequence, resulting in adenine (A) to guanine (G) base editing. Using the base editing system of this invention for cell gene editing not only significantly improves the efficiency of A to G base substitution but also minimizes DNA and RNA off-target editing events, making it a highly efficient and specific adenine base editor. This promotes the safe clinical use of base editors for gene and cell therapy.

[0006] This application provides a fusion protein comprising, from the N-terminus to the C-terminus, a first SaCas9 nickase fragment, a chimeric deaminase fragment, and a second SaCas9 nickase fragment; the deaminase is selected from adenosine deaminase or a variant thereof, the adenosine deaminase is selected from ecTadA8e, the amino acid sequence of ecTadA8e is shown in SEQ ID NO.20, or has more than 80% sequence identity with the amino acid sequence shown in SEQ ID NO.20, and possesses the function or activity of ecTadA8e.

[0007] This application also provides an isolated polynucleotide that encodes the aforementioned fusion protein.

[0008] This application also provides an expression vector containing the isolated polynucleotides described above.

[0009] This application also provides an expression system containing the above-described expression vector or a genome in which exogenous polynucleotides described above are integrated.

[0010] This application also provides a base editing system comprising the above-described fusion protein or its encoded polynucleotide.

[0011] This application also provides the use of the above-mentioned fusion proteins, isolated polynucleotides, expression vectors, expression systems or base editing systems in gene editing.

[0012] This application also provides a gene editing method, comprising: performing base editing on a target sequence using the above-described fusion protein, the above-described isolated polynucleotide, the above-described expression vector, or the above-described expression system or the above-described base editing system.

[0013] This application also provides a base-edited cell, obtained by using the above-described gene editing method to edit the target sequence in the cell by A to G or T to C genes.

[0014] This application also provides a reporting system comprising a nucleotide sequence as shown in SEQ ID NO.15.

[0015] This application also provides the use of the above-described reporting system for detecting the AG editing efficiency of the above-described fusion protein, isolated polynucleotide, expression vector or expression system or base editing system.

[0016] This application also provides a method for detecting the AG editing efficiency of a base-editing product, using the aforementioned reporting system to detect the AG editing efficiency of the product under test.

[0017] The beneficial effects of the embodiments in this specification include, but are not limited to: (1) This application provides a new fusion protein with gene editing function and a base editing system containing it, characterized by high editing efficiency and high specificity in vivo and in vitro; (2) The fusion protein and base editing system containing it provided in this application are used to achieve at least one of the following: correction of pathogenic sites, gene function research, enhancement of cell function, and cell therapy; (3) The base editing system provided in this application can achieve AG editing in mammalian cells with a base substitution efficiency of up to 90%, and can also achieve efficient AG editing in mammalian cells with a base substitution efficiency of up to 63%. Attached Figure Description

[0018] This application will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, wherein:

[0019] Figure 1 This is a schematic diagram of a fluorescence reporting system according to some embodiments of this application;

[0020] Figure 2 The bar chart shows the proportion of BFP fluorescence at different time points after cells transfected with saABE8e, as illustrated in some embodiments of this application.

[0021] Figure 3 This is a flowchart illustrating the random insertion of ecTadA8e into the SaCas9 nickase (nSaCas9) protein using a transposase, as shown in some embodiments of this application.

[0022] Figure 4 A bar chart showing the editing efficiency of the reporting system for editors with different chimeric sites, as illustrated in some embodiments of this application;

[0023] Figure 5A This is a schematic diagram of the structure of the SaABE8e base-editing protein according to some embodiments of this application;

[0024] Figure 5B This is a schematic diagram of the structure of a chimeric CE-saABE8e base-editing protein according to some embodiments of this application;

[0025] Figure 6 To analyze the editing efficiency of the CE-saABE8e base editing protein at five sites in HEK293T cells according to some embodiments of this application;

[0026] Figure 7A Off-target analysis of the CE-saABE8e base editing protein at the DNA level, as shown in some embodiments of this application;

[0027] Figure 7B This document describes the off-target effects of the CE-saABE8e base-editing protein at the RNA level, as shown in some embodiments of this application.

[0028] Figure 8 This study analyzes the editing efficiency of the CE-saABE8e base editing protein on five endogenous gene loci in mice, as shown in some embodiments of this application. Detailed Implementation

[0029] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0030] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0031] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0032] This application provides a fusion protein comprising, from the N-terminus to the C-terminus, a first SaCas9 nickase fragment, a chimeric deaminase fragment, and a second SaCas9 nickase fragment; the deaminase is selected from adenosine deaminase or a variant thereof, the adenosine deaminase is selected from ecTadA8e, the amino acid sequence of ecTadA8e is shown in SEQ ID NO.20, or has more than 80% sequence identity with the amino acid sequence shown in SEQ ID NO.20, and possesses the function or activity of ecTadA8e.

[0033] In some embodiments, the fusion protein may be a chimeric protein of a SaCas9 nickase and a deaminase fragment. In some embodiments, the chimeric site of the deaminase fragment may be selected from positions 730 to 744 of the SaCas9 nickase amino acid sequence. In some embodiments, preferably, the chimeric site of the deaminase fragment may be selected from positions 733, 736, 739, or 744 of the SaCas9 nickase amino acid sequence.

[0034] The term "sequence" in this article should generally be understood to include both the relevant amino acid sequence and the nucleic acid or nucleotide sequence encoding the amino acid sequence, unless a more specific interpretation is required in this article.

[0035] The “sequence identity” between two polypeptide sequences indicates the percentage of identical amino acids between the sequences. Sequence identity also indicates the percentage of identical or conserved amino acid substitutions. Methods for evaluating the degree of sequence identity between amino acids or nucleotides are known to those skilled in the art. For example, amino acid sequence identity is typically measured using sequence analysis software. For instance, the BLAST procedure from the NCBI database can be used to determine identity.

[0036] As used herein, the terms “polynucleotide,” “nucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably to refer to nucleic acids including DNA, RNA, their derivatives, or combinations thereof.

[0037] In some embodiments, the fusion protein may further include a nuclear localization signal fragment located at the N-terminus and / or C-terminus of the fusion protein. In some embodiments, the nuclear localization signal fragment may be an optimized nuclear localization signal (BPNLS) or a variant thereof. In some embodiments, the amino acid sequence of the BPNLS may be as shown in SEQ ID NO. 17. In some embodiments, the amino acid sequence of the variant may have more than 80% sequence identity with the BPNLS and possess the function of the BPNLS.

[0038] The term "nuclear localization signal" (NLS) refers to an amino acid sequence that induces the transport of molecules including or linked to such sequences into the nucleus of eukaryotic cells. The NLS can form part of the molecule to be transported. In some embodiments, the NLS can be linked to the rest of the molecule via covalent bonds, hydrogen bonds, or ionic interactions. In some embodiments, the NLS can facilitate the entry of fusion proteins into the cell nucleus.

[0039] In some embodiments, the fusion protein may further include a first flexible linker peptide fragment and a second flexible linker peptide fragment. In some embodiments, the first flexible linker peptide fragment may be located between a first SaCas9 nickase fragment and a chimeric deaminase fragment. In some embodiments, the second flexible linker peptide fragment may be located between the chimeric deaminase fragment and the second SaCas9 nickase fragment.

[0040] In some embodiments, the amino acid sequence of the first flexible linker peptide fragment may be as shown in SEQ ID NO. 18. In some embodiments, the amino acid sequence of the first flexible linker peptide fragment may have more than 80% sequence identity with the amino acid sequence shown in SEQ ID NO. 18, and possess the function or activity of the first flexible linker peptide fragment.

[0041] In some embodiments, the amino acid sequence of the second flexible linker peptide fragment may be as shown in SEQ ID NO.19. In some embodiments, the amino acid sequence of the second flexible linker peptide fragment may have more than 80% sequence identity with the amino acid sequence shown in SEQ ID NO.19 and possess the function or activity of the second flexible linker peptide fragment.

[0042] In some embodiments, the amino acid sequence of the fusion protein may include any of the segments shown in SEQ ID NO. 10 to 13, or an amino acid sequence that has more than 80% sequence identity with one of the amino acid sequences shown in SEQ ID NO. 10 to 13 and has the function of the amino acid sequence defined in SEQ ID NO. 10 to 13.

[0043] This application also provides an isolated polynucleotide that encodes the aforementioned fusion protein.

[0044] Polynucleotides are polymers of nucleotides typically linked from one deoxyribose or ribose to another, and depending on the context, they refer to both DNA and RNA. The term "polynucleotide" in this application does not contain any size limitation and also includes polynucleotides containing modifications, particularly modified nucleotides. In some embodiments, the polynucleotide may be RNA, DNA, or cDNA, etc. Methods for providing the isolated polynucleotides should be known to those skilled in the art; for example, they can be prepared by automated DNA synthesis and / or recombinant DNA techniques, or isolated from suitable natural sources.

[0045] This application also provides an expression vector, which may contain the isolated polynucleotides described above.

[0046] As used herein, "vector" refers to a polynucleotide capable of carrying at least one polynucleotide fragment. A vector can deliver fragments of nucleic acids, or individual polynucleotides, into a host cell. It may contain at least one expression cassette containing a regulatory sequence for the proper expression of the polynucleotide incorporated therein. The polynucleotide to be introduced into the cell (e.g., a polynucleotide encoding a target product or a selectable marker) can be inserted into the expression cassette of the vector for expression therefrom. Vectors according to this application may be in circular or linear (linearized) form and also include vector fragments. The term "vector" also includes artificial chromosomes or similar individual polynucleotides that allow the transfer of exogenous nucleic acid fragments.

[0047] This application also provides an expression system containing the above-described expression vector or a genome in which exogenous polynucleotides described above are integrated.

[0048] The expression system can be a host cell, which can be a prokaryotic cell, such as a bacterial cell; a lower eukaryotic cell, such as a yeast cell or a filamentous fungal cell; or a higher eukaryotic cell, such as a mammalian cell. Representative examples include: *Escherichia coli*, *Streptomyces* spp.; bacterial cells of *Salmonella typhimurium*; fungal cells such as yeast, filamentous fungi, and plant cells; insect cells of *Drosophila* S2 or Sf9; and animal cells such as CHO, COS, 293 cells, or Bowes melanoma cells. Methods for introducing the expression vector into host cells should be known to those skilled in the art, such as microinjection, gene gun methods, electroporation, virus-mediated transformation, electron bombardment, and calcium phosphate precipitation. The choice of expression system depends on various factors, including cell growth characteristics, expression level, intracellular and extracellular expression, post-translational modifications and biological cleanliness of the target protein, as well as regulatory issues and economic considerations in the production of therapeutic proteins.

[0049] In some embodiments, the host cell of the expression system may be selected from eukaryotic cells or prokaryotic cells. In some embodiments, preferably, the host cell may be selected from mouse cells or human cells.

[0050] This application also provides a base editing system comprising the above-described fusion protein or its encoded polynucleotide.

[0051] In some embodiments, the base editing system may further include guide RNA. In some embodiments, the guide RNA may target a target sequence.

[0052] The term "guide RNA" refers to an RNA molecule that can instruct CRISPR effectors with nuclease activity to target and cleave a specified target nucleic acid.

[0053] In some embodiments, the base editing system may comprise one or more vectors. In some embodiments, the one or more vectors may comprise a first regulatory element and a second regulatory element. In some embodiments, the first regulatory element is operatively linked to a polynucleotide encoding the fusion protein. In some embodiments, the second regulatory element is operatively linked to a polynucleotide encoding the guide RNA nucleotide sequence. In some embodiments, the first and second regulatory elements may be located on the same or different vectors.

[0054] In some embodiments, the base editing system may comprise (i) a fusion protein and (ii) a guide RNA or a vector comprising the guide RNA encoding a polynucleotide.

[0055] This application also provides the use of the above-mentioned fusion proteins, isolated polynucleotides, expression vectors, expression systems or base editing systems in gene editing.

[0056] In some embodiments, the gene editing can achieve base substitution. In some embodiments, the gene editing can achieve substitution of A to G or T to C. In some embodiments, the gene editing can be used to achieve at least one of the following: correction of pathogenic sites, gene function research, enhancement of cell function, and cell therapy. In some embodiments, the fusion protein, isolated polynucleotide, expression vector, expression system, or base editing system can be used in combination with other drugs or reagents. In some embodiments, the disease caused by the pathogenic site can be selected from at least one of the following: autoimmune diseases, tumors, viral infectious diseases, and bacterial infectious diseases.

[0057] In some embodiments, the autoimmune diseases may include, but are not limited to: systemic lupus erythematosus, rheumatoid arthritis, systemic vasculitis, scleroderma, pemphigus, dermatomyositis, mixed connective tissue disease, autoimmune hemolytic anemia, etc.

[0058] In some embodiments, the tumor may include, but is not limited to: Ewing sarcoma, neuroendocrine tumor, glioblastoma, neuroblastoma, melanoma, skin cancer, breast cancer, colon cancer, rectal cancer, prostate cancer, liver cancer, kidney cancer, pancreatic cancer, lung cancer, biliary tract cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, medullary thyroid carcinoma, ovarian cancer, glioma, lymphoma, leukemia, myeloma, acute lymphoblastic leukemia, acute myeloid leukemia, chronic lymphocytic leukemia, chronic myeloid leukemia, Hodgkin's lymphoma, non-Hodgkin's lymphoma, or urobladder cancer, etc.

[0059] In some embodiments, the viral infectious diseases may include, but are not limited to: measles, rubella, mumps, chickenpox, HIV / AIDS, condyloma acuminata, viral hepatitis, etc. In some embodiments, the bacterial infectious diseases may include, but are not limited to: tuberculosis, acute tonsillitis, bacterial dysentery, purulent meningitis, scarlet fever, acute pharyngitis, etc.

[0060] This application also provides a gene editing method, comprising: performing base editing on a target sequence using the aforementioned fusion protein, isolated polynucleotide, expression vector or expression system or base editing system.

[0061] In some embodiments, the method may be performed in vitro. In some embodiments, preferably, the method may be performed in cultured cells. In some embodiments, the method may be performed in vivo; preferably, the method may be performed in mammals. In some embodiments, more preferably, the method may be performed in rodents or primates. In some embodiments, even more preferably, the method may be performed in mice or humans.

[0062] This application also provides a base-edited cell, obtained by using the above-described gene editing method to edit the target sequence in the cell by A to G or T to C genes.

[0063] This application also provides a reporting system that may contain a nucleotide sequence as shown in SEQ ID NO.15.

[0064] In some embodiments, the reporting system may comprise a nucleotide sequence as shown in SEQ ID NO. 15. In some embodiments, the reporting system may display green fluorescence. In some embodiments, the reporting system may display blue fluorescence when the nucleotide sequence shown in SEQ ID NO. 15 is mutated to SEQ ID NO. 16. In some embodiments, the reporting system may not display fluorescence when the nucleotide sequence shown in SEQ ID NO. 15 is mutated to SEQ ID NO. 14.

[0065] In some embodiments, the reporter system may comprise a plasmid. In some embodiments, the plasmid may comprise a nucleotide sequence as shown in SEQ ID NO. 15. In some embodiments, the reporter system may further comprise a guide RNA or a vector comprising a polynucleotide sequence encoded by the guide RNA. In some embodiments, the guide RNA may target the codon sequence of amino acid 66 encoded by SEQ ID NO. 15. In some embodiments, preferably, the target sequence of the guide RNA may be as shown in SEQ ID NO. 5.

[0066] This application also provides the use of the above-described reporting system for detecting the AG editing efficiency of the above-described fusion protein, isolated polynucleotide, expression vector or expression system or base editing system.

[0067] This application also provides a method for detecting the AG editing efficiency of a base-editing product, using the aforementioned reporting system to detect the AG editing efficiency of the product under test.

[0068] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent companies. All quantitative experiments in the following examples were performed in triplicate, and the results were averaged.

[0069] Example 1: A fluorescence reporting system for testing AG editing efficiency

[0070] In mammalian cell lines, a base substitution editor was used to edit the acquired reporter system, and the presence and efficiency of the editing were determined by fluorescence signals. The specific implementation is as follows:

[0071] 1. Construction of the reporting system.

[0072] The reporting system described in this invention (operating mode as follows) Figure 1 The reporter protein expressed by the system is shown in SEQ ID NO. 15, and its amino acid sequence is shown in SEQ ID NO. 2. When the codon at amino acid position 66 of the coding strand of the reporter protein is TAC (corresponding to ATG on the non-coding strand), tyrosine is expressed, and the reporter protein encoded by the reporter system exhibits green fluorescence. When an A to G mutation is introduced using a base editing system to target the non-coding strand, mutating ATG on the non-coding strand to GTG, and correspondingly mutating the codon at amino acid position 66 to CAC, histidine is expressed, and the reporter system exhibits blue fluorescence. At this time, the coding nucleotide sequence of the reporter protein becomes as shown in SEQ ID NO. 16, and the amino acid sequence of the reporter protein becomes as shown in SEQ ID NO. 1. When a mutation is introduced using base editing, mutating the codon at amino acid position 66 to TAG, a stop codon is expressed, and the reporter system does not exhibit fluorescence. At this time, the coding nucleotide sequence of the reporter protein becomes as shown in SEQ ID NO. 14, and the amino acid sequence of the reporter protein becomes as shown in SEQ ID NO. 3. The substitution ratio of AG bases can be calculated by analyzing the proportion of blue fluorescence using flow cytometry. The nucleotide sequence of the vector plasmid (reporter system) containing this reporter protein constructed in this embodiment is shown in SEQ ID NO.39, and the corresponding reporter system is named the BFP-AG reporter system.

[0073] SEQ ID NO.15:

[0074] ATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGG

[0075] ACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCAC

[0076] CTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAACTGCCCGTGCCCTGGC

[0077] CCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCAC

[0078] ATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCA

[0079] CCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGG

[0080] CGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAAC

[0081] ATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGA

[0082] CAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGC

[0083] AGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGC

[0084] TGCTGCCCGACAACCACTACCTGAGCACCCAGTCCAAGCTGAGCAAAGACCCCAACGA

[0085] GAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCA

[0086] TGGACGAGCTGTACAAGTGA

[0087] SEQ ID NO.2:

[0088] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTL

[0089] VTTLTYGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLV

[0090] NRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADH

[0091] YQQNTPIGDGPVLLPDNHYLSTQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYKSEQ IDNO.16:

[0092] ATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGG

[0093] ACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCAC

[0094] CTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAACTGCCCGTGCCCTGGC

[0095] CCACCCTCGTGACCACCCTGACCCACGGCGTGCAGTGCTTCAGCCGCTACCCCGACCAC

[0096] ATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCA

[0097] CCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGG

[0098] CGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAAC

[0099] ATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGA

[0100] CAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGC

[0101] AGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGC

[0102] TGCTGCCCGACAACCACTACCTGAGCACCCAGTCCAAGCTGAGCAAAGACCCCAACGA

[0103] GAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCA

[0104] TGGACGAGCTGTACAAGTGA

[0105] SEQ ID NO.1:

[0106] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTL

[0107] VTTLTHGVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLV

[0108] NRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADH

[0109] YQQNTPIGDGPVLLPDNHYLSTQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYKSEQ IDNO.14:

[0110] ATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCTGG

[0111] ACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCCAC

[0112] CTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAACTGCCCGTGCCCTGGC

[0113] CCACCCTCGTGACCACCCTGACCTAGGGCGTGCAGTGCTTCAGCCGCTACCCCGACCAC

[0114] ATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCGCA

[0115] CCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAGGG

[0116] CGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCAAC

[0117] ATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCCGA

[0118] CAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACGGC

[0119] AGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGTGC

[0120] TGCTGCCCGACAACCACTACCTGAGCACCCAGTCCAAGCTGAGCAAAGACCCCAACGA

[0121] GAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGGCA

[0122] TGGACGAGCTGTACAAGTGA

[0123] SEQ ID NO.3:

[0124] MVSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTL VTTLT*GVQCFSRYPDHMKQHDFFKSAMPEGYVQERTIFFKDDGNYKTRAEVKFEGDTLV NRIELKGIDFKEDGNILGHKLEYNYNSHNVYIMADKQKNGIKVNFKIRHNIEDGSVQLADH YQQNTPIGDGPVLLPDNHYLSTQSKLSKDPNEKRDHMVLLEFVTAAGITLGMDELYK(* represents the stop codon)

[0125] SEQ ID NO.39:

[0126] GACGGATCGGGAGATCTCCCGATCCCCTATGGTGCACTCTCAGTACAATCTGCTCTGATG

[0127] CCGCATAGTTAAGCCAGTATCTGCTCCCTGCTTGTGTGTTGGAGGTCGCTGAGTAGTGCG

[0128] CGAGCAAAATTTAAGCTACAACAAGGCAAGGCTTGACCGACAATTGCATGAAGAATCTG

[0129] CTTAGGGTTAGGCGTTTTGCGCTGCTTCGCGATGTACGGGCCAGATATACGCGTTGACAT

[0130] TGATTATTGACTAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATAT

[0131] GGAGTTCCGCGTTACATAACTTACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACC

[0132] CCCGCCCATTGACGTCAATAATGACGTATGTTCCCATAGTAACGCCAATAGGGACTTTCC

[0133] ATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAGTACATCAAGTGT

[0134] ATCATATGCCAAGTACGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATT

[0135] ATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCAT

[0136] CGCTATTACCATGGTGATGCGGTTTTGGCAGTACATCAATGGGCGTGGATAGCGGTTTGA

[0137] CTCACGGGGATTTCCAAGTCTCCACCCCATTGACGTCAATGGGAGTTTGTTTTGGCACC

[0138] AAAATCAACGGGACTTTCCAAAATGTCGTAACAACTCCGCCCCATTGACGCAAATGGGC

[0139] GGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTCTCTGGCTAACTAGAGAACC

[0140] CACTGCTTACTGGCTTATCGAAATTAATACGACTCACTATAGGGAGACCCAAGCTGGCTA

[0141] GCATGGTGAGCAAGGGCGAGGAGCTGTTCACCGGGGTGGTGCCCATCCTGGTCGAGCT

[0142] GGACGGCGACGTAAACGGCCACAAGTTCAGCGTGTCCGGCGAGGGCGAGGGCGATGCC

[0143] ACCTACGGCAAGCTGACCCTGAAGTTCATCTGCACCACCGGCAAACTGCCCGTGCCCTG

[0144] GCCCACCCTCGTGACCACCCTGACCTACGGCGTGCAGTGCTTCAGCCGCTACCCCGACC

[0145] ACATGAAGCAGCACGACTTCTTCAAGTCCGCCATGCCCGAAGGCTACGTCCAGGAGCG

[0146] CACCATCTTCTTCAAGGACGACGGCAACTACAAGACCCGCGCCGAGGTGAAGTTCGAG

[0147] GGCGACACCCTGGTGAACCGCATCGAGCTGAAGGGCATCGACTTCAAGGAGGACGGCA

[0148] ACATCCTGGGGCACAAGCTGGAGTACAACTACAACAGCCACAACGTCTATATCATGGCC

[0149] GACAAGCAGAAGAACGGCATCAAGGTGAACTTCAAGATCCGCCACAACATCGAGGACG

[0150] GCAGCGTGCAGCTCGCCGACCACTACCAGCAGAACACCCCCATCGGCGACGGCCCCGT

[0151] GCTGCTGCCCGACAACCACTACCTGAGCACCCAGTCCAAGCTGAGCAAAGACCCCAAC

[0152] GAGAAGCGCGATCACATGGTCCTGCTGGAGTTCGTGACCGCCGCCGGGATCACTCTCGG

[0153] CATGGACGAGCTGTACAAGTGAAAGCTTGGTACCGAGCTCGGATCCACTAGTCCAGTGT

[0154] GGTGGAATTCTGCAGATATCCAGCACAGTGGCGGCCGCTCGAGTCTAGAGGGCCCGTTT

[0155] AAACCCGCTGATCAGCCTCGACTGTGCCTTCTAGTTGCCAGCCATCTGTTGTTTGCCCCT

[0156] CCCCCGTGCCTTCCTTGACCCTGGAAGGTGCCACTCCCACTGTCCTTTCCTAATAAAATG

[0157] AGGAAATTGCATCGCATTGTCTGAGTAGGTGTCATTCTATTCTGGGGGGTGGGGTGGGG

[0158] CAGGACAGCAAGGGGGAGGATTGGGAAGACAATAGCAGGCATGCTGGGGATGCGGTGG

[0159] GCTCTATGGCTTCTGAGGCGGAAAGAACCAGCTGGGGCTCTAGGGGGTATCCCCACGCG

[0160] CCCTGTAGCGGCGCATTAAGCGCGGCGGGTGTGGTGGTTACGCGCAGCGTGACCGCTAC

[0161] ACTTGCCAGCGCCCTAGCGCCCGCTCCTTTCGCTTTCTTCCCTTCCTTTCTCGCCACGTT

[0162] CGCCGGCTTTCCCCGTCAAGCTCTAAATCGGGGGCTCCCTTTAGGGTTCCGATTTAGTGC

[0163] TTTACGGCACCTCGACCCCAAAAAACTTGATTAGGGTGATGGTTCACGTAGTGGGCCAT

[0164] CGCCCTGATAGACGGTTTTTCGCCCTTTGACGTTGGAGTCCACGTTCTTTAATAGTGGAC

[0165] TCTTGTTCCAAACTGGAACAACACTCAACCCTATCTCGGTCTATTCTTTTGATTTATAAGG

[0166] GATTTTGCCGATTTCGGCCTATTGGTTAAAAAATGAGCTGATTTAACAAAAATTTAACGC

[0167] GAATTAATTCTGTGGAATGTGTGTCAGTTAGGGTGTGGAAAGTCCCCAGGCTCCCCAGC

[0168] AGGCAGAAGTATGCAAAGCATGCATCTCAATTAGTCAGCAACCAGGTGTGGAAAGTCCC

[0169] CAGGCTCCCCAGCAGGCAGAAGTATGCAAAGCATGCATCTCAATTAGTCAGCAACCATA

[0170] GTCCCGCCCCTAACTCCGCCCATCCCGCCCCTAACTCCGCCCAGTTCCGCCCATTCTCCG

[0171] CCCCATGGCTGACTAATTTTTTTTATTTATGCAGAGGCCGAGGCCGCCTCTGCCTCTGAG

[0172] CTATTCCAGAAGTAGTGAGGAGGCTTTTTTGGAGGCCTAGGCTTTTGCAAAAAGCTCCC

[0173] GGGAGCTTGTATATCCATTTTCGGATCTGATCAAGAGACAGGATGAGGATCGTTTCGCAT

[0174] GATTGAACAAGATGGATTGCACGCAGGTTCTCCGGCCGCTTGGGTGGAGAGGCTATTCG

[0175] GCTATGACTGGGCACAACAGACAATCGGCTGCTCTGATGCCGCCGTGTTCCGGCTGTCA

[0176] GCGCAGGGGCGCCCGGTTCTTTTTGTCAAGACCGACCTGTCCGGTGCCCTGAATGAACT

[0177] GCAGGACGAGGCAGCGCGGCTATCGTGGCTGGCCACGACGGGCGTTCCTTGCGCAGCT

[0178] GTGCTCGACGTTGTCACTGAAGCGGGAAGGGACTGGCTGCTATTGGGCGAAGTGCCGG

[0179] GGCAGGATCTCCTGTCATCTCACCTTGCTCCTGCCGAGAAAGTATCCATCATGGCTGATG

[0180] CAATGCGGCGGCTGCATACGCTTGATCCGGCTACCTGCCCATTCGACCACCAAGCGAAA

[0181] CATCGCATCGAGCGAGCACGTACTCGGATGGAAGCCGGTCTTGTCGATCAGGATGATCT

[0182] GGACGAAGAGCATCAGGGGCTCGCGCCAGCCGAACTGTTCGCCAGGCTCAAGGCGCGC

[0183] ATGCCCGACGGCGAGGATCTCGTCGTGACCCATGGCGATGCCTGCTTGCCGAATATCATG

[0184] GTGGAAAATGGCCGCTTTTCTGGATTCATCGACTGTGGCCGGCTGGGTGTGGCGGACCG

[0185] CTATCAGGACATAGCGTTGGCTACCCGTGATATTGCTGAAGAGCTTGGCGGCGAATGGG

[0186] CTGACCGCTTCCTCGTGCTTTACGGTATCGCCGCTCCCGATTCGCAGCGCATCGCCTTCTA

[0187] TCGCCTTCTTGACGAGTTCTTCTGAGCGGGACTCTGGGGTTCGAAATGACCGACCAAGC

[0188] GACGCCCAACCTGCCATCACGAGATTTCGATTCCACCGCCGCCTTCTATGAAAGGTTGG

[0189] GCTTCGGAATCGTTTTCCGGGACGCCGGCTGGATGATCCTCCAGCGCGGGGATCTCATG

[0190] CTGGAGTTCTTCGCCCACCCCAACTTGTTTATTGCAGCTTATAATGGTTACAAATAAAGC

[0191] AATAGCATCACAAATTTCACAAATAAAGCATTTTTTTCACTGCATTCTAGTTGTGGTTTGT

[0192] CCAAACTCATCAATGTATCTTATCATGTCTGTATACCGTCGACCTCTAGCTAGAGCTTGGC

[0193] GTAATCATGGTCATAGCTGTTTCCTGTGTGAAATTGTTATCCGCTCACAATTCCACACAAC

[0194] ATACGAGCCGGAAGCATAAAGTGTAAAGCCTGGGGTGCCTAATGAGTGAGCTAACTCAC

[0195] ATTAATTGCGTTGCGCTCACTGCCCGCTTTCCAGTCGGGAAACCTGTCGTGCCAGCTGCA

[0196] TTAATGAATCGGCCAACGCGCGGGGAGAGGCGGTTTGCGTATTGGGCGCTCTTCCGCTT

[0197] CCTCGCTCACTGACTCGCTGCGCTCGGTCGTTCGGCTGCGGCGAGCGGTATCAGCTCAC

[0198] TCAAAGGCGGTAATACGGTTATCCACAGAATCAGGGGATAACGCAGGAAAGAACATGTG

[0199] AGCAAAAGGCCAGCAAAAGGCCAGGAACCGTAAAAAGGCCGCGTTGCTGGCGTTTTTC

[0200] CATAGGCTCCGCCCCCCTGACGAGCATCACAAAAATCGACGCTCAAGTCAGAGGTGGC

[0201] GAAACCCGACAGGACTATAAAGATACCAGGCGTTTCCCCCTGGAAGCTCCCTCGTGCGC

[0202] TCTCCTGTTCCGACCCTGCCGCTTACCGGATACCTGTCCGCCTTTCTCCCTTCGGGAAGC

[0203] GTGGCGCTTTCTCATAGCTCACGCTGTAGGTATCTCAGTTCGGTGTAGGTCGTTCGCTCC

[0204] AAGCTGGGCTGTGTGCACGAACCCCCCGTTCAGCCCGACCGCTGCGCCTTATCCGGTAA

[0205] CTATCGTCTTGAGTCCAACCCGGTAAGACACGACTTATCGCCACTGGCAGCAGCCACTG

[0206] GTAACAGGATTAGCAGAGCGAGGTATGTAGGCGGTGCTACAGAGTTCTTGAAGTGGTGG

[0207] CCTAACTACGGCTACACTAGAAGAACAGTATTTGGTATCTGCGCTCTGCTGAAGCCAGTT

[0208] ACCTTCGGAAAAAGAGTTGGTAGCTCTTGATCCGGCAAACAAACCACCGCTGGTAGCG

[0209] GTTTTTTTGTTTGCAAGCAGCAGATTACGCGCAGAAAAAAAGGATCTCAAGAAGATCCT

[0210] TTGATCTTTTCTACGGGGTCTGACGCTCAGTGGAACGAAAACTCACGTTAAGGGATTTT

[0211] GGTCATGAGATTATCAAAAAGGATCTTCACCTAGATCCTTTTAAATTAAAAATGAAGTTTT

[0212] AAATCAATCTAAAGTATATATGAGTAAACTTGGTCTGACAGTTACCAATGCTTAATCAGTG

[0213] AGGCACCTATCTCAGCGATCTGTCTATTTCGTTCATCCATAGTTGCCTGACTCCCCGTCGT

[0214] GTAGATAACTACGATACGGGAGGGCTTACCATCTGGCCCCAGTGCTGCAATGATACCGCG

[0215] AGACCCACGCTCACCGGCTCCAGATTTATCAGCAATAAACCAGCCAGCCGGAAGGGCC

[0216] GAGCGCAGAAGTGGTCCTGCAACTTTATCCGCCTCCATCCAGTCTATTAATTGTTGCCGG

[0217] GAAGCTAGAGTAAGTAGTTCGCCAGTTAATAGTTTGCGCAACGTTGTTGCCATTGCTACA

[0218] GGCATCGTGGTGTCACGCTCGTCGTTTGGTATGGCTTCATTCAGCTCCGGTTCCCAACGA

[0219] TCAAGGCGAGTTACATGATCCCCCATGTTGTGCAAAAAAGCGGTTAGCTCCTTCGGTCC

[0220] TCCGATCGTTGTCAGAAGTAAGTTGGCCGCAGTGTTATCACTCATGGTTATGGCAGCACT

[0221] GCATAATTCTCTTACTGTCATGCCATCCGTAAGATGCTTTTCTGTGACTGGTGAGTACTCA

[0222] ACCAAGTCATTCTGAGAATAGTGTATGCGGCGACCGAGTTGCTCTTGCCCGGCGTCAATA

[0223] CGGGATAATACCGCGCCACATAGCAGAACTTTAAAAGTGCTCATCATTGGAAAACGTTCT

[0224] TCGGGGCGAAAACTCTCAAGGATCTTACCGCTGTTGAGATCCAGTTCGATGTAACCCAC

[0225] TCGTGCACCCAACTGATCTTCAGCATCTTTACTTTCACCAGCGTTTCTGGGGTGAGCAAA

[0226] AACAGGAAGGCAAAATGCCGCAAAAAAGGGAATAAGGGCGACACGGAAATGTTGAAT

[0227] ACTCATACTCTTCCTTTTTCAATATTATTGAAGCATTTTATCAGGGTTATTGTCTCATGAGCG

[0228] GATACATATTTGAATGTATTTAGAAAAATAAACAAATAGGGGTTCCGCGCACATTTCCCC

[0229] GAAAAGTGCCACCTGACGTC

[0230] 2. Construction of sgRNA expression vector

[0231] Based on the AAV-EFS-SaABE8e-bGH-U6-sgRNA-BsmBI vector (addgene: 189922), a specific sgRNA vector targeting the reporter protein was constructed. This vector contains the reported SaABE8e protein (protein structure shown in Figure 1). Figure 5A (as shown) and its guide RNA expression element. For convenience, the BsmBI restriction site on the vector was replaced with the commonly used BsaI restriction site sequence to obtain the AAV-EFS-SaABE8e-bGH-U6-sgRNA-BsaI vector.

[0232] Based on the SaCas9 design principle, targeting the editing target, namely the 66th amino acid codon of the aforementioned reporter protein, a 22nt targeting sequence was designed as SEQ ID NO.5: CGCCGTGGGTCAGGGTGGTCAC. The corresponding sgRNA expression vector was constructed as follows:

[0233] Oligonucleotide pairs with sticky ends were synthesized according to the target site sequence, as shown in SEQ ID NO.6: CACCCGCCGTGGGTCAGGGTGGTCAC and SEQ ID NO.7: AAACCCGTAGGTCAGGGTGGTCAC. These were dissolved and diluted to a concentration of 10 μM in sterile water. The primers were annealed to form oligoduplexes and ligated into the BsaI-digested linearized AAV-EFS-SaABE8e-bGH-U6-sgRNA-BsaI vector to construct a specifically targeted sgRNA vector, which was named AAV-SaABE8e-gRNA-BFP in this invention. The experimental procedure is as follows:

[0234] 2.1 Annealing forms oligo double chains

[0235] The annealing reaction system is as follows:

[0236] Upstream primer (10 μM) 20μL Downstream primer (10 μM) 20μL

[0237] The annealing procedure is as follows: 95℃ for 5 min, 95-85℃-2℃ / s, 85-25℃-0.1℃ / s, 4℃.

[0238] 2.2 Obtaining linearized vectors by BsaI digestion

[0239] The enzyme digestion reaction system is shown below:

[0240] AAV-EFS-SaABE8e-bGH-U6-sgRNA-BsaI plasmid 2μg 10×Cutsmart buffer 5μL BsaI enzyme (NEB: R0539L) 5μL Add water until 50μL

[0241] After the above system was prepared, it was reacted at 37℃ for 3 hours. 2 μL was taken and run on agarose gel electrophoresis to verify that the vector was completely linearized. The completely linearized enzyme digestion product of the vector was recovered using the AxyPrep PCR recovery kit (Axygen, Code No.: AP-PCR-250G, hereinafter the same) to obtain the linearized vector.

[0242] 2.3 Linkage of annealed products with linearized carrier

[0243] Using the DNA Ligation Kit Ver. 2.1 (Takara, Code No.: 6022Q), the annealed product obtained in step 2.1 and the linearized vector obtained in step 2.2 were ligated. The reaction system is shown below:

[0244]

[0245]

[0246] After the above system was prepared, it was incubated at 16℃ for 30 minutes, then transformed into DH5α competent cells. After recovery on a shaker at 37℃ for 30 minutes, the cells were plated on ampicillin-resistant LB agar plates and incubated overnight at 37℃. The next day, single clones were selected for first-generation sequencing verification. After successful ligation and sequence accuracy, plasmid extraction was performed.

[0247] 3. Mammalian cell line transfection AG editing system and acquisition reporter system.

[0248] HEK293T cells were seeded and cultured in DMEM medium containing 10% FBS (HyClone, Code No.: SH30022.01B, hereinafter the same). Cells were separated into 24-well plates one day before transfection. The following day, transfection was performed when the cell density reached 70%-80%. The medium was replaced with fresh medium two hours before transfection. Following the EZTrans cell transfection solution (Liji Biotechnology Co., Ltd., Code No.: AC04L091, hereinafter the same), 900 ng of the AAV-SaABE8e-gRNA-BFP base editing vector plasmid and 100 ng of the BFP-AG reporter system expression vector plasmid were mixed and co-transfected into the cells. The medium was changed after 6-8 hours, and fluorescence detection was performed by flow cytometry at 24, 48, 72, and 96 hours.

[0249] 4. Fluorescence reporter system analysis of base substitution efficiency

[0250] Using FlowJO software to analyze the flow cytometry results, this invention discovered that AAV-SaABE8e-gRNA-BFP can achieve the conversion of GFP to BFP fluorescence. Figure 2 The fluorescence intensity was highest at 48 hours, and subsequent experiments will use 48 hours as the time point to detect the fluorescence intensity.

[0251] The above results indicate that the reporting system constructed in this invention can accurately reflect AG editing efficiency and can be used as a method for testing the base editing efficiency of adenine base editors.

[0252] Example 2: Construction of chimeric CE-SaABE8e

[0253] Previous studies have shown that SaABE8e, with its N-terminal fused deaminase, is prone to random off-target effects on DNA and RNA, posing a significant safety threat to the in vivo application of base editors. Therefore, this invention proposes to insert the highly efficient ecTadA8e deaminase into the protein domain of the SaCas9 nickase, which can effectively avoid random deamination by the deaminase and reduce off-target effects.

[0254] 1. Construction of the pET-nSaCas9-SagRNA-AmpR(W163X)-KanR vector

[0255] use II. The One Step Cloning Kit (Novaza, Code No.: C112-02, hereinafter the same) was used to construct the pET-nSaCas9-SagRNA-AmpR(W163X)-KanR vector, the nucleotide sequence of which is shown in SEQ ID NO.8. The ampicillin resistance gene on this vector contains a stop codon TAG (in bold in the sequence of SEQ ID NO.8) at amino acid position 163. Ampicillin resistance is only effective when TAG is targeted and edited to TGG, and the corresponding bacteria can grow on ampicillin-treated plates.

[0256] SEQ ID NO.8:

[0257]

[0258] 2. Constructing randomly inserted recombinant vector plasmids using MuA transposase

[0259] The codon-optimized ecTadA8e gene fragment (nucleotide sequence as shown in SEQ ID NO. 9) of *E. coli* was synthesized at Sangon Biotech (Shanghai) Co., Ltd. (hereinafter referred to as Sangon). The ecTadA8e fragment and the pET-nSaCas9-SagRNA-AmpR(W163X)-KanR plasmid were used to construct a recombinant vector of ecTadA8e at different positions in vitro using MuA transposase (Thermo Fisher, F-701). The specific reaction system is as follows:

[0260] ecTadA8e fragment 250ng pET-nSaCas9-SagRNA-AmpR(W163X-KanR plasmid 500ng MuA transposase 1μL 5×Reaction Buffer for MuA Transposase 4μL Add water until 20μL

[0261] The prepared reaction solution was incubated at 30°C for 1 hour to achieve random insertion, and then incubated at 75°C for 10 minutes to inactivate MuA transposase. The DNA was then purified by isopropanol precipitation and resuspended in 5 μL of deionized water, and then transformed into 100 μL of BL21(DE3) competent cells.

[0262] SEQ ID NO.9:

[0263] TCTGAAGTAGAATTTTCCCACGAATACTGGATGCGCCATGCACTGACCCTGGCAAAACGCGCCCGCGACGAACGTGAAGTTCCAGTTGGTGCGGTGCTGGTACTGAACAACCGTGTAATCGGCG AAGGCTGGAATCGTGCGATCGGTCTGCACGATCCGACTGCACACGCAGAAATCATGGCTCTGCGTCAGGGTGGCCTGGTGATGCAAAATTACCGCCTGATCGATGCGACTCTGTATGTTACCTTC GAACCGTGCGTAATGTGTGCAGGTGCTATGATCCACTCCCGTATTGGTCGCGTCGTGTTTGGTGTTCGCAACTCCAAGCGTGGTGCTGCAGGCTCTCTGATGAACGTGCTGAACTACCCGGGCA TGAACCATCGTGTTGAGATCACGGAAGGCATCCTGGCTGACGAATGTGCTGCCCTGCTGTGTGACTTCTACCGTATGCCGCGCCAGGTATTCAACGCCCAGAAGAAGGCGCAGAGCAGCATCAAC

[0264] 3. Screening expression plasmids for functionally embedded fusion proteins in Escherichia coli

[0265] Only bacteria that can express ecTadA8e normally can grow on ampicillin-resistant LB agar plates.

[0266] Therefore, the transformed bacteria were revived in SOC medium for 1 hour, then plated on three LB agar plates containing 10 μg / mL kanamycin, and incubated at 37°C for 16 hours. Since the vector pET-nSaCas9-SagRNA-AmpR(W163X)-KanR carries the normally expressed KanR gene, both the original and recombinant vectors exhibit kanamycin resistance, resulting in abundant colony growth. Colonies from all the plates were scraped off and resuspended in 100 mL LB agar containing 500 μM IPTG. The cultures were incubated for 10–12 hours to induce the expression of the functional intercalation editing fusion protein and repair the mutation in the AmpR(W163X) gene on the vector. Then, reduced cell volumes (5 mL, 1 mL, 500 μL, 100 μL) were inoculated onto 15 cm LB agar plates containing ampicillin (10 μg / mL) and kanamycin (10 μg / mL). After incubation at 37°C overnight, colonies were selected and Sanger sequencing was performed to assess base editing on AmpR(W163X) and determine the ecTadA8e insertion site.

[0267] Based on the Sanger sequencing analysis, the amino acid positions at which ecTadA8e inserts into the SaCas9 nickase were identified as amino acids 123, 128, 460, 665, 723, 730, 731, 732, 733, 734, 735, 736, 738, 739, 740, 741, 742, 743, 744, 755, 832, 901, 911, 912, 913, and 953. In this application, these amino acid positions are referred to simply as "insertion sites".

[0268] 4. Construction of expression vectors carrying chimeric fusion proteins and base editors

[0269] Based on the insertion sites determined through preliminary screening, this invention constructs a recombinant expression vector for a base editor expressed in mammals. This vector carries not only the coding sequence of the chimeric fusion protein and its promoter and other regulatory element sequences, but also guide RNA (gRNA) and its promoter sequence. The construction method is as follows:

[0270] First, based on different insertion sites, gene fragments of different chimeric fusion proteins and their regulatory elements optimized for mammalian codons (hereinafter referred to as "CE fragments") were synthesized at Sangon Biotech. The N-terminus to C-terminus of the CE fragments are as follows: EFS promoter (nucleotide sequence as shown in SEQ ID NO.4), N-terminal BPNLS nuclear localization signal (amino acid sequence as shown in SEQ ID NO.17), first SaCas9 nickase fragment (26 sequences were designed according to different insertion sites), first flexible linker sequence (amino acid sequence as shown in SEQ ID NO.18), ecTadA8e deaminase fragment (amino acid sequence as shown in SEQ ID NO.20), second flexible linker sequence (amino acid sequence as shown in SEQ ID NO.19), second SaCas9 nickase fragment (26 sequences were designed according to different insertion sites), C-terminal BPNLS sequence (amino acid sequence as shown in SEQ ID NO.21), and bGH-polyA sequence (well known in the art).

[0271] Simultaneously, a gene fragment of the optimized gRNA (nucleotide sequence as shown in SEQ ID NO.22) and its human U6 promoter sequence (well known in the art) was synthesized (hereinafter referred to as the "gRNA fragment").

[0272] SEQ ID NO.4:

[0273] GAATTCGCTAGCTAGGTCTTGAAAGGAGTGGGAATTGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGGGTCGGCAATTGATCCGGTGCCTAGAGAAGGTGGCGCGG GGTAAACTGGGAAAGTGATGTCGTGTACTGGCTCCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAACGGGTTTGCCGCCAGAACACAGGACCGGTGCCACC

[0274] SEQ ID NO.17:

[0275] MKRTADGSEFESPKKKRKV

[0276] SEQ ID NO.18:

[0277] SGSETPGTSESATPESGS

[0278] SEQ ID NO.19:

[0279] SGSGSETPGTSESATPES

[0280] SEQ ID NO.20:

[0281] SEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN

[0282] SEQ ID NO.21:

[0283] KRTADGSEFEPKKKRKV

[0284] SEQ ID NO.22:

[0285] GTTTTAGTACTCTGTAATGAAAATTACAGAATCTACTAAAACAAGGCAAAATGCCGTGTTTATCTCGTCAACTTGTTGGCGAGA

[0286] 4.1 Fragment Amplification and Recovery

[0287] Design upstream and downstream primers, and amplify the synthesized genes (CE fragments and gRNA fragments with 26 insertion sites) as DNA templates to obtain two insertion fragments with overlapping sequences at both ends; then amplify the vector backbone fragment using pX601-AAV (Addgene: #61591) as the basic vector (primers and corresponding templates are shown in the table below).

[0288]

[0289] Perform PCR amplification according to the following PCR reaction system:

[0290] 2×FastD Pfu Supermix (Full Metal, Code No.: AS231) 25μL Forward primer (10 μM) 2.5μL Reverse primer (10 μM) 2.5μL DNA template 5ng Add water until 50μL

[0291] The reaction conditions were: denaturation at 95℃ for 3 min; denaturation at 95℃ for 30 s, annealing at 60℃ for 20 s, extension at 72℃ for 4 min, amplification for 32 cycles; holding at 72℃ for 5 min; and holding at 4℃.

[0292] The PCR products were recovered using the AxyPrep PCR recovery kit, and 26 different CE fragment recovery products 1, gRNA fragment recovery product 2, and vector backbone recovery product 3 were obtained. The concentrations were then determined.

[0293] 4.3 Fragment Reassembly

[0294] Using a recombinant kit, 26 recovered products 1 were recombined with recovered products 2 and 3, respectively. The recombinant reaction system is shown below:

[0295] Recycled Product 1 25ng Recycled Product 2 25ng Recycled Product 3 50ng 5×CE II Buffer 2μL Exnase II 1μL Add water until 10μL

[0296] After the above system was prepared, it was incubated at 16℃ for 30 minutes, then transformed into DH5α competent cells. After recovery on a shaker at 37℃ for 30 minutes, the cells were plated on ampicillin-resistant LB agar plates and incubated overnight at 37℃. The next day, single clones were selected for first-generation sequencing verification. After successful ligation and sequence accuracy, plasmid extraction was performed.

[0297] This application uniformly names this base editor and its corresponding base editing fusion protein as CE-SaABE8e. The editor vectors constructed from 26 different insertion sites are distinguished by suffixes. For example, the base editing fusion protein with insertion site 736 and its editor are named CE-SaABE8e-736.

[0298] 5. Editing of the reporting system using the hybrid CE-SaABE8e

[0299] The sgRNA target sequence specifically targeting the non-coding strand containing the 66th amino acid codon of the reporter protein was designed, annealed to form an oligo double strand, and ligated into 26 linearized CE-SaABE8e expression vectors recovered by BsaI restriction enzyme digestion, thus obtaining 26 sgRNA vector plasmids specifically targeting the target site. The specific sgRNA vector construction method is as described in "2. Construction of sgRNA expression vectors" in Case 1.

[0300] HEK293T cells were seeded and cultured in DMEM medium containing 10% FBS at 37°C and 5% CO2. Cells were aliquoted into 24-well plates the day before transfection. The following day, transfection was performed when the cell density reached 70%-80%. Following the EZTrans cell transfection protocol, 900 ng of different chimeric CE-SaABE8e vector plasmids were mixed with 100 ng of the BFP-AG reporter system and co-transfected into the cells. The medium was changed after 6-8 hours, and BFP fluorescence signals were analyzed after 48 hours.

[0301] The results of flow cytometry were analyzed using FlowJO software. The effectiveness of AG base substitution was determined by comparing BFP efficiency. The results showed that the base editing vectors that underwent AG editing were CE-SaABE8e-730, CE-SaABE8e-731, CE-SaABE8e-732, CE-SaABE8e-733, CE-SaABE8e-734, CE-SaABE8e-735, CE-SaABE8e-736, CE-SaABE8e-737, CE-SaABE8e-738, CE-SaABE8e-739, CE-SaABE8e-740, CE-SaABE8e-741, CE-SaABE8e-742, CE-SaABE8e-743, and CE-SaABE8e-744. Figure 4 Among the selected proteins, the four with the highest editing efficiency are CE-SaABE8e-733, CE-SaABE8e-736, CE-SaABE8e-739, and CE-SaABE8e-744, whose corresponding chimeric fusion protein amino acid sequences are shown in SEQ ID NO. 10–13. This application will subsequently use CE-SaABE8e-736 as a representative to conduct base editing experiments on mammalian cell sites and mouse endogenous gene sites.

[0302] SEQ ID NO.10:

[0303]

[0304] SEQ ID NO.11:

[0305]

[0306] SEQ ID NO.12:

[0307]

[0308] SEQ ID NO.13:

[0309]

[0310] Example 3: Efficient AG mutation of CE-SaABE8e in HEK293T cells

[0311] To further investigate the characteristics and efficiency of the AG base substitution editor, this invention edited five sites in HEK293T cells. The specific implementation process is as follows:

[0312] 1. Target site selection and construction of corresponding sgRNA expression vectors.

[0313] The five sites were selected as follows:

[0314] Site1: GTGGTAGACAGCATGTGTCCTAAAGGGT (SEQ ID NO.29);

[0315] Site2:ATTTACAGCCTGGCCTTTGGGGTCGGGT(SEQ ID NO.30);

[0316] Site3: GGAGAGAAAGAGAAGTTGATTGATGGGT (SEQ ID NO.31);

[0317] Site4: GTGTCAGGTAATGTGCTAAACAGAGAGT (SEQ ID NO.32);

[0318] Site5: ATGCATTAACTGAAAATGGTCAAGGAGT (SEQ ID NO.33);

[0319] Design appropriate sgRNA primers, and anneal the upstream and downstream sequences using a programmed sequence (95℃, 5 min; 95℃-85℃ at -2℃ / s; 85℃-25℃ at -0.1℃ / s; hold at 4℃) before ligating them into the BsaI-linearized CE-SaABE8e-736 vector. Positive clones were subjected to culture to extract plasmids (Axygene, Code No.: AP-MN-P-250G) and their concentration was determined for later use. The specific sgRNA vector construction method is as described in Section 2, "Construction of sgRNA Expression Vector," in Case 1.

[0320] 2. Editing of CE-SaABE8e in HEK293T cells

[0321] HEK293T cells were seeded and cultured in DMEM medium containing 10% FBS at 37°C and 5% CO2. Cells were aliquoted into 24-well plates the day before transfection. The following day, transfection was performed when the cell density reached 70%-80%. Following the EZTrans cell transfection protocol, 900 ng of sgRNA vector plasmid and 450 ng of GFP fluorescent expression plasmid were mixed and co-transfected into the cells. The medium was changed after 6-8 hours, and 10,000 GFP-positive cells were sorted by flow cytometry after 48 hours. The cell pellet was lysed and genotypes were identified. The lysis buffer consisted of: 50 mM KCl, 1.5 mM MgCl2, 10 mM Tris (pH 8.0), 0.5% Nonidet P-40, 0.5% Tween 20, and 100 μg / ml Protease K.

[0322] 3. Analysis of the editing effect of CE-SaABE8e in HEK293T cells

[0323] Using first-generation Sanger sequencing, this invention analyzed the aforementioned five sites and statistically determined the corresponding editing efficiencies. Figure 6 The results showed that CE-SaABE8e can achieve efficient AG editing, with a base substitution efficiency ranging from a maximum of 90% to a minimum of 26%.

[0324] Example 4: CE-SaABE8e can significantly reduce off-target effects on DNA and RNA.

[0325] Base editors have been reported to cause off-target effects at both the DNA and RNA levels. This invention will analyze the off-target effects at both the DNA and RNA levels following intracellular site editing by CE-SaABE8e.

[0326] HEK293T cells were seeded and cultured in DMEM medium containing 10% FBS at 37°C and 5% CO2. Cells were separated into 6-well plates the day before transfection. The following day, transfection was performed when the cell density reached 70%-80%. Following the EZTrans cell transfection protocol, 4 μg of CE-SaABE8e or control SaABE8e plasmid was mixed with 2 μg of GFP plasmid and co-transfected into the corresponding wells. The medium was changed after 6-8 hours, and 500,000 GFP-positive cells were sorted by flow cytometry after 48 hours.

[0327] The sorted cells were centrifuged to collect the precipitate, and DNA and RNA were extracted for whole-genome sequencing and RNA-seq sequencing. Comparison with empty-transfected negative control cells revealed no significant off-target effects of CE-SaABE8e at the DNA and RNA levels compared to the reference genome. Figures 7A-7BThis demonstrates that the chimeric base editor developed in this application achieves efficient and specific base editing within mammalian cells.

[0328] Example 5: Packaging and preparation of AAV virus from CE-SaABE8e recombinant expression vector

[0329] Adeno-associated viruses (AAVs) have been used to deliver genes encoding many therapeutic proteins in animal models of human diseases, clinical trials, and FDA-approved drugs. Due to their advantages such as clinical validation, ability to target various clinically relevant tissues, high safety profile, and extensive research, AAVs have become a popular in vivo delivery method. The CE-SaABE8e fusion protein developed in this application uses an AAV recombinant vector (see Example 2 for details) for expression. This vector can be further packaged into an AAV virus and delivered to mammals via injection for base editing.

[0330] To conduct in vivo editing experiments in mice, this invention prepared AAV virus loaded with CE-SaABE8e. The virus packaging and preparation process is as follows:

[0331] 1. Cell transfection

[0332] HEK293T cells were seeded and cultured in DMEM medium containing 10% FBS at 37°C and 5% CO2. When conjugation reached 90%, the cells were transferred to trays at a ratio of 1:3 (approximately 2.5 × 10⁶ cells per tray). 6 Continue culturing. Two hours before transfection, switch to serum-free medium. When cell confluence reaches 80%-90%, begin transfection. Add 9 mL of DMEM, 5.7 μg of CE-SaABE8e-736 plasmid, 11.4 μg of pHelper, 22.8 μg of rep-cap plasmid, and 1 mL of PEI sequentially to a 15 mL centrifuge tube. Vortex to mix, incubate at room temperature for 30 min, and then add dropwise to a 10 cm dish. On day 1 post-transfection, replace the cell culture medium with DMEM containing 10% FBS.

[0333] 2. Virus collection and purification

[0334] Virus was harvested on day 4 post-transfection. Cells and culture medium were collected together using a rubber cell scraper into 50 ml centrifuge tubes and centrifuged at 2,000 g for 10 min. The cell pellet and culture supernatant were collected separately. Each cell pellet was resuspended in 500 μl of hypertonic lysis buffer (40 mM Tris-base, 500 mM NaCl, 2 mM MgCl2, and 100 U / mL). -1Cells were lysed by incubating the culture medium supernatant with PEG-8000 NaCl solution (40% PEG-8000, 2.5 mM NaCl) at 37°C for 1 h. The supernatant was filtered through a 0.45 μm filter, and 5×PEG-8000 NaCl solution (40% PEG-8000, 2.5 mM NaCl) was added to bring the final concentration to 1×(8% PEG, 500 mM NaCl). After incubation for 2 h, the mixture was centrifuged at 3200 g for 30 min, and the precipitate was resuspended in 500 μL of hypertonic lysis buffer. The crude lysates obtained from the above two steps were combined and incubated at 4°C overnight or immediately for ultracentrifugation. The cell lysates were centrifuged at 2000 g for 10 min, and virus purification was performed using iodixanol density gradient centrifugation.

[0335] 3. Virus concentration and titration

[0336] The solution from the previous step was exchanged for cold PBS containing 0.001% F-68 and concentrated using a PES 100kDMWCO column (Thermo Fisher, Pierce 88533). The concentrated AAV solution was sterilely filtered through a 0.22 μm filter, and the AAV virus was titrated by qPCR using the AAVpro Titration Kit Version 2 (Clontech). The solution was then stored at 4°C until use.

[0337] Example 6: Efficient endogenous gene AG mutation of CE-SaABE8e in mice

[0338] Gene editing offers the potential for clinical treatment of various genetic diseases, but most gene editing research and treatment for genetic diseases need to be conducted in vivo. To investigate the characteristics and efficiency of the CE-SaABE8e base editor developed in this application in vivo, AG base substitution editing was performed at five endogenous gene loci in mice. The specific implementation is as follows:

[0339] 1. Target site selection and construction of corresponding sgRNA expression vectors.

[0340] The five sites were selected as follows:

[0341] PCSK9_exon1: GCCACCGCAGCCACGCAGAGCAGTGGGT (SEQ ID NO.34);

[0342] PCSK9_exon5:GCGTGCTTACCTGTCTGTGGAAGCGGGT(SEQ ID NO.35)

[0343] PCSK9_exon8:GCCATCCTGCTCACCTGTCTCATGGGT(SEQ ID NO.36)

[0344] PCSK9_exon9:GCCATCCTGCTTACCTGCCCCATGGGT(SEQ ID NO.37)

[0345] Angptl3_exon4: GTGTTCCATGGGTTTACCTGATTGGGT (SEQ ID NO.38);

[0346] Design appropriate sgRNA primers, and anneal the upstream and downstream sequences using a programmed sequence (95℃, 5 min; 95℃-85℃ at -2℃ / s; 85℃-25℃ at -0.1℃ / s; hold at 4℃) before ligating them into the BsaI-linearized CE-SaABE8e-736 vector. Positive clones were subjected to culture for plasmid extraction, and the concentration was determined for later use. The specific sgRNA vector construction method is as described in Section 2, "Construction of sgRNA Expression Vector," in Case 1.

[0347] 2. Preparation of AAV virus carrying sgRNA expression vector

[0348] The five sgRNA expression vectors constructed in the previous step were transfected into HEK293T cells according to the method in Case 5, and then collected, purified and concentrated to prepare AAV virus.

[0349] 2. Injection and editing in mice.

[0350] The C57BL / 6J mice used in this application were purchased from The Jackson Laboratory. Humanized PCSK9 mice have been previously reported in studies. All mice were housed in a room under a 12-hour light-dark cycle and provided with standard rodent food and water. No immunosuppression or other treatments were administered to the mice before injection or during the experiment, except for pre-bleeding fasting as described below. Multiple bleeding was performed before tail vein delivery of the AAV vector or control to collect pre-injection samples and to acclimate the animals to handling during the procedure.

[0351] Mice were randomly assigned to five experimental groups and one control group. Prior to injection, 4 × 10⁴ AAV viruses were drawn from each experimental group. 10 The dose of vitamin G was diluted to 100 μL with 0.9% sterile phosphate-buffered saline (PBS, pH 7.4). Mice were induced anesthetized with 2–4% isoflurane. After induction, if there was no response to pressure on the toes, the skin was gently pressed to cause the right eye to protrude. The AAV solution was then slowly injected into the retrobulbar sinus using an insulin syringe. One drop of prupacaine hydrochloride ophthalmic solution was then applied to the eye as an analgesic.

[0352] Four weeks after AAV injection, liver samples were collected: To collect liver tissue and a larger volume of serum for testing, mice were euthanized by inhaling carbon dioxide. A portion of the shredded liver tissue was collected for genomic DNA extraction, and another portion was rapidly cryopreserved in liquid nitrogen for RNA extraction. DNA and RNA extracted from the livers of both the experimental and control groups were subjected to whole-genome sequencing and transcriptome sequencing, respectively.

[0353] 3. Editing analysis of CE-SaABE8e in mice

[0354] Using high-throughput sequencing, this invention analyzed the aforementioned five sites and statistically determined the corresponding editing efficiencies. Figure 8 The results showed that CE-SaABE8e can achieve efficient AG editing in mammals, with a base substitution efficiency ranging from a maximum of 63% to a minimum of 35%.

[0355] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0356] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0357] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0358] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A fusion protein, characterized in that, The amino acid sequence of the fusion protein is shown in SEQ ID NO.

10.

2. An isolated polynucleotide, characterized in that, The isolated polynucleotide encodes the fusion protein as described in claim 1.

3. An expression carrier, characterized in that, The expression vector contains the isolated polynucleotide as described in claim 2.

4. An expression system comprising an expression vector as described in claim 3 or a genome in which an exogenous polynucleotide as described in claim 2 is integrated.

5. The expression system as described in claim 4, characterized in that, The host cells of the expression system are selected from eukaryotic cells or prokaryotic cells.

6. The expression system according to claim 5, characterized in that, The host cells are selected from mouse cells and human cells.

7. A base editing system comprising the fusion protein of claim 1 or its encoded polynucleotide.

8. The base editing system as described in claim 7, characterized in that, The base editing system also includes a guide RNA that can guide the fusion protein to target a specific target.

9. The base editing system as described in claim 7, characterized in that, Includes at least one of the following: 1) The base editing system comprises one or more vectors; the one or more vectors comprise (i) a first regulatory element operatively linked to a polynucleotide encoding the fusion protein; and (ii) a second regulatory element operatively linked to a coding polynucleotide of the guide RNA nucleotide sequence; (i) and (ii) are located on the same or different carriers; 2) The base editing system comprises (i) a fusion protein and (ii) a guide RNA or a vector containing the guide RNA encoding a polynucleotide.

10. Use of the fusion protein of claim 1, and / or the isolated polynucleotide of claim 2, and / or the expression vector of claim 3, and / or the expression system of any one of claims 4 to 6, and / or the base editing system of any one of claims 7 to 9 in gene editing for non-disease diagnosis and treatment purposes, wherein the gene editing achieves the substitution of A to G.

11. A gene editing method for purposes other than disease diagnosis and treatment, comprising: Base editing of target sequences is performed using the fusion protein as described in claim 1, the isolated polynucleotide as described in claim 2, the expression vector as described in claim 3, the expression system as described in any one of claims 4 to 6, or the base editing system as described in any one of claims 7 to 9.

12. The method as described in claim 11, characterized in that, The method is performed in vitro.

13. The method as described in claim 11, characterized in that, The method is performed in cultured cells.