Base editing system based on TraC effect protein and application thereof
By developing the fusion protein composition eTraC-ABE/eTraC-CBE based on eTraC effector protein, the shortcomings of the existing base editing system in terms of editing efficiency, editing window and Index frequency are solved, and more efficient and accurate gene editing effects are achieved.
Patent Information
- Application Number
- CN202411698901.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-23
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-23
AI Technical Summary
The existing base editing system has room for optimization in terms of editing efficiency, editing windows and Index frequency, and the off-target effect is relatively obvious.
Based on the functional variant of the TraC effector protein, a series of new fusion protein compositions, eTraC-ABE/eTraC-CBE, was developed to achieve adenine/cytosine base editing. These fusion proteins include DNA binding protein domain, deaminase domain and uracil glycosylase inhibitor domain, which improves editing efficiency and editing accuracy by optimizing the arrangement of amino acid sequences and domains.
The new base editing system eTraC-ABE/eTraC-CBE significantly improves editing efficiency, expands the editing window, and reduces the Indel frequency and off-target effect, providing a wider range of gene editing application scenarios.
Smart Images

Figure CN120026006A_ABST
Abstract
Description
[0001] Priority and related applications
[0002] This application claims priority to Chinese patent application 202311576149.X, filed on November 23, 2023, entitled “A base editor based on TraC effector protein and its application”. The entire contents of the above application, including the appendix, are incorporated into this application by reference. Technical Field
[0003] The present invention belongs to the field of gene editing, and relates to a novel fusion protein with cytosine / adenine base editing function, an editing system, and nucleic acid molecules or constructs encoding them. The present invention also relates to a complex and a kit for gene editing, which comprises the fusion protein of the present invention or nucleic acid molecules encoding them. Background Art
[0004] Base editing refers to the process of replacing nucleotides at specific DNA sites through genetic engineering. Base editing is a type of genome editing technology. In plants, many excellent agronomic traits are produced by single nucleotide variations (SNVs); in humans and animals, many genetic diseases are caused by point mutations in functional genes. Base editing technology can be used to modify and transform the genomic DNA of animals, plants, and microorganisms to create new genotypes and obtain target traits that are beneficial to production applications. It can also be used to correct some serious congenital genetic variations in the treatment of genetic diseases. Therefore, base editing has important application prospects in the creation of excellent germplasm and disease treatment.
[0005] Base editing is achieved with the help of a base editor system. According to the type of bases acted on by the base editor system, the base editor system can be divided into two categories: 1) cytosine base editor system (CBE) and 2) adenine base editor system (ABE). Among them, CBE can act on the cytosine base on the cytosine deoxyribonucleotide (C), while ABE can act on the adenine base on the adenine deoxyribonucleotide (A). CBE editing can convert the CG base pair on the double-stranded DNA into the TA base pair; ABE editing can convert the AT base pair on the double-stranded DNA into the GC base pair.
[0006] The core components of CBE and ABE include sequence-specific DNA binding proteins and deaminases. Among them, the sequence-specific DNA binding protein is responsible for guiding the entire editor to the target DNA site and unwinding the DNA double helix; the deaminase fused with it can bind to the unwound ssDNA substrate and deaminize; depending on the deaminase, the cytosine base (cytidine deaminase) or adenine base (adenosine deaminase) can be deaminated. At present, the sequence-specific DNA binding proteins used in the CBE system and the ABE system are mainly CRISPR-Cas9 proteins and their variants. In addition, there are also reports on base editing systems based on other types of CRISPR-Cas proteins, including Cas12a, Cas12f, Cas12n, etc.
[0007] Base editing technology can be used to more accurately target and modify the genomic DNA of animals, plants, and microorganisms, and therefore has shown great application value in basic research in life sciences, agriculture, biomedicine, etc. However, the existing base editing system still needs to be continuously optimized and improved in terms of editing efficiency, editing window, Indel frequency caused by base editing, off-target effects, etc. Summary of the invention
[0008] In order to enrich the toolbox of base editing systems and broaden their application scenarios, the present invention is based on the functional variant eTraC effector protein between the transposon and the CRISPR-Cas12 intermediate (TraC) effector protein, and provides a series of new fusion protein compositions eTraC-ABE / eTraC-CBE with cytidine / adenosine base editing function. The fusion protein composition can be used to achieve adenine / cytosine base editing.
[0009] In one aspect of the present invention, a DNA-binding protein is provided, characterized in that the DNA-binding protein is a nuclease that completely or partially loses the DNA double-stranded cleavage activity or only retains the single-stranded DNA cleavage activity, and the nuclease is a functional variant of the effector protein between the transposon and the CRISPR-Cas12 intermediate (TraC), and the TraC effector protein refers to the disclosure in PCT / CN2023 / 097783 (publication number WO / 2023 / 232109) and is incorporated herein by reference. The functional variant of the TraC effector protein in the present invention is TraC-5M-7 in the above-mentioned patent, which is renamed as eTraC effector protein in the present invention.
[0010] In some embodiments, the DNA binding protein is a nuclease that completely or partially loses DNA double-strand cleavage activity, and the nuclease is a mutant of the eTraC effector protein;
[0011] In some embodiments, the eTraC mutant comprises one or more mutations at amino acid positions 253, 381, and 464 of SEQ ID NO:1, and the amino acid sequence of the eTraC mutant has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to SEQ ID NO:1.
[0012] In some embodiments, the DNA binding protein comprises an amino acid sequence in which any one or more mutation types of D253A, E381A, and D464A exist in the amino acid sequence provided in SEQ ID NO:1.
[0013] In some embodiments, the DNA binding protein comprises an amino acid sequence of double mutations of D253A and E381A in the amino acid sequence provided in SEQ ID NO:1.
[0014] In some embodiments, the DNA binding protein comprises three mutated amino acid sequences of D253A, E381A, and D464A in the amino acid sequence provided in SEQ ID NO:1.
[0015] In some embodiments, the DNA binding protein comprises the amino acid sequence shown in any one of SEQ ID NO:2, SEQ ID NO:26, SEQ ID NO:33, SEQ ID NO:40 or SEQ ID NO:47.
[0016] Another aspect of the present invention provides a fusion protein, characterized in that the fusion protein comprises:
[0017] i) a DNA binding protein domain, wherein the DNA binding protein domain comprises the DNA binding protein described above;
[0018] ii) deaminase domain;
[0019] Herein, "deaminase" refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase, ie, an enzyme that can remove the amino group of a cytidine molecule or an adenosine molecule.
[0020] In some embodiments, the cytidine deaminase is selected from an APOBEC family deaminase, a SCP1.201 family deaminase, or homologs thereof, or other known cytidine deaminases that can be used in a base editing system.
[0021] In some embodiments, the APOBEC family deaminase is selected from AID deaminase, APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase and APOBEC3H deaminase. The SCP1.201 family deaminase is selected from Sdd2, Sdd3, Sdd4 deaminase, Sdd6 deaminase, Sdd7 deaminase, Sdd9 deaminase, Sdd10 deaminase, Sdd59 deaminase. The homolog of the APOBEC family deaminase is cytidine deaminase 1 (pmCDA1) from sea lamprey (Petromyzonmarinus). In some embodiments, the cytidine deaminase domain comprises the amino acid sequence of SEQ ID NO:3, SEQ ID NO:51-54.
[0022] In some embodiments, the adenosine deaminase comprises TadA deaminase, QB34 deaminase or a functional variant thereof, or other adenosine deaminase that can be used in a base editing system. In some embodiments, the adenosine deaminase is TadA-8e deaminase. In some embodiments, the adenosine deaminase is an amino acid sequence shown in SEQ ID NO: 21.
[0023] In some embodiments, the fusion protein further comprises: iii) a uracil glycosylase inhibitor (UGI) domain.
[0024] As used herein, the term "uracil glycosylase inhibitor" or "UGI" refers to a protein that is capable of inhibiting uracil-DNA glycosylase base excision repair enzymes. In some embodiments, the UGI domain comprises an amino acid sequence as shown in SEQ ID NO:4.
[0025] In some embodiments, the fusion protein further comprises: iv) a nuclear localization sequence (NLS).
[0026] Herein, "nuclear localization sequence (NLS)" or "nuclear localization signal" are used interchangeably and refer to a domain of a protein, usually a short amino acid sequence, which can interact with a nuclear import carrier so that the protein can be transported into the nucleus. In general, one or more NLSs in the fusion protein should have sufficient strength to drive the fusion protein to accumulate in the nucleus of the cell in an amount that can achieve its base editing function. In general, the intensity of nuclear localization activity is determined by the number, position, one or more specific NLSs used in the fusion protein, or a combination of these factors. In some embodiments of the present invention, the NLS of the fusion protein of the present invention may be located at the N-terminus and / or the C-terminus. In some embodiments of the present invention, the NLS of the fusion protein of the present invention may be located between the adenosine deaminase domain, the cytidine deaminase domain, the DNA binding protein domain and / or the UGI. In some embodiments, the fusion protein comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLSs. In some embodiments, the fusion protein is included in or close to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at the N-terminus. In some embodiments, the fusion protein is included in or close to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more NLS at the C-terminus. In some embodiments, the polypeptide comprises a combination of these, such as one or more NLSs included at the N-terminus and one or more NLSs at the C-terminus. When there is more than one NLS, each can be selected to be independent of other NLSs. In general, NLS is composed of one or more short sequences of positively charged lysine or arginine exposed on the protein surface, but other types of NLS are also known. Non-limiting examples of NLS include amino acid sequences SEQ ID NO: 5-8.
[0027] In some embodiments, the fusion protein optionally includes the following structure:
[0028] NH2-[Cytidine deaminase domain]-[UGI domain]-[DNA binding protein domain]-COOH
[0029] NH2-[Cytidine deaminase domain]-[DNA binding protein domain]-[UGI domain]-COOH
[0030] NH2-[DNA binding protein domain]-[cytidine deaminase domain]-[UGI domain]-COOH
[0031] NH2-[DNA binding protein domain]-[UGI domain]-[cytidine deaminase domain]-COOH
[0032] NH2-[UGI domain]-[DNA binding protein domain]-[cytidine deaminase domain]-COOH
[0033] NH2-[UGI domain]-[Cytidine deaminase domain]-[DNA binding protein domain]-COOH
[0034] NH2-[Adenosine deaminase domain]-[DNA binding protein domain]-COOH
[0035] NH2-[DNA binding protein domain]-[adenosine deaminase domain]-COOH
[0036] NH2-[Adenosine deaminase domain]-[DNA binding protein domain]-[Adenosine deaminase domain]-COOH
[0037] Among them, "NH2-" and "-COOH" are only used to indicate the N-terminus and C-terminus of the fusion protein, and "-" indicates that the protein domains are directly connected or connected through an optional linker. The NLS sequence can be selectively connected to the end of the fusion protein or between the domains.
[0038] In some embodiments, an exemplary fusion protein includes one of the following structures:
[0039] NH2-[NLS]-[Cytidine deaminase domain]-[DNA binding protein domain]-[UGI domain]-[UGI domain]-[NLS]-COOH, or
[0040] NH2-[NLS]-[adenosine deaminase domain]-[DNA binding protein domain]-[NLS]-COOH.
[0041] In some embodiments, the fusion protein comprises the amino acid sequence shown in any one of SEQ ID NOs: 16-20, 22, 27-32, 34-39, 41-46, 48-50.
[0042] "Linker" herein refers to connecting two molecules or parts, such as two domains of a fusion protein. Typically, a linker is located between or flanking two groups, molecules or other parts and is connected to each by a covalent bond, thereby connecting the two. In some embodiments, a linker is an organic molecule, group, polymer or chemical part.
[0043] In some embodiments, the linker is an amino acid or multiple amino acids (e.g., a peptide or protein). In some embodiments, the linker comprises an amino acid sequence (GGGS)n(SEQ ID NO:9), (GGGGS)n(SEQ ID NO:10), (G)n, (EAAAK)n(SEQ ID NO:11), (GGS)n(SEQ ID NO:12), (SGGS)n(SEQ ID NO:13), SGSETPGTSESATPES(SEQ ID NO:14), or (XP)n motif or a combination thereof, wherein n is independently an integer of 1-30, and wherein X is any amino acid. In some embodiments, the domains of the fusion protein are connected by a linker.
[0044] Another aspect of the present invention provides a complex, characterized in that the complex comprises the fusion protein of the present invention and a guide RNA bound to the DNA binding protein domain of the fusion protein.
[0045] Herein, the term "complex" refers to a combination of two or more different substances. In some embodiments, the complex refers to a combination formed by binding or associating a DNA binding protein domain of a fusion protein with one or more guide RNAs.
[0046] In some embodiments, the base editing system comprises the fusion protein of any one of the present invention, and / or an expression construct containing a nucleotide sequence encoding the fusion protein; or the complex of the present invention.
[0047] Herein, "base editing system" refers to a system that achieves precise editing of DNA bases based on gene editing technology. Herein, the base editing system refers to the fusion protein and guide RNA of the present invention.
[0048] In some embodiments, the base editing system can be at least one of the following components:
[0049] a) a fusion protein of the present invention, and a guide RNA;
[0050] b) an expression construct comprising the nucleotide sequence of the fusion protein of the present invention, and a guide RNA;
[0051] c) a fusion protein of the present invention, and an expression construct comprising a nucleotide sequence encoding a guide RNA;
[0052] d) an expression construct comprising a nucleotide sequence encoding a fusion protein of the present invention, and an expression construct comprising a nucleotide sequence encoding a guide RNA.
[0053] e) The fusion protein of the present invention and a guide RNA bound to the DNA binding protein domain of the fusion protein of the present invention.
[0054] Another aspect of the present invention provides a host cell, characterized in that the host cell contains the fusion protein of the present invention, and / or an expression construct containing a nucleotide sequence encoding the fusion protein; or the complex of the present invention; or the base editing system of the present invention.
[0055] Another aspect of the present invention provides a base editing method, characterized in that the base editing method comprises contacting the fusion protein of the present invention with a guide RNA, wherein the guide RNA is complementary to a sequence of at least 10 consecutive nucleotides in a target sequence in the genome of an organism.
[0056] In some embodiments, wherein the target sequence is in the genome of an organism. In some embodiments, wherein the organism is a prokaryote. In some embodiments, wherein the prokaryote is a bacterium. In some embodiments, wherein the organism is a eukaryote. In some embodiments, wherein the organism is a plant or a fungus. In some embodiments, wherein the organism is a vertebrate. In some embodiments, wherein the vertebrate is a mammal. In some embodiments, wherein the mammal is a mouse, a rat, or a human. In some embodiments, wherein the organism is a cell. In some embodiments, wherein the cell is a mouse cell, a rat cell, or a human cell. In some embodiments, wherein the cell is a HEK-293 cell.
[0057] In some embodiments, after the base editing system of the present invention is introduced into the cell, the fusion protein of the present invention and the guide RNA are able to form a complex, and the complex specifically targets the target sequence under the mediation of the guide RNA, and causes one or more C to be replaced by T and / or one or more A to be replaced by G in the target sequence or a sequence complementary to the target sequence.
[0058] In some embodiments, the at least one guide RNA may be directed to a target sequence on a sense strand (e.g., a protein coding strand) and / or an antisense strand located in a genomic target nucleic acid region. When the guide RNA targets the sense strand (e.g., a protein coding strand), the base editing system of the present invention may cause one or more Cs in an editing sequence complementary to the target sequence on the antisense strand (e.g., a protein coding strand) to be replaced by T and / or one or more A to be replaced by G. When the guide RNA targets the antisense strand, the base editing system of the present invention may cause one or more Gs in an editing sequence complementary to the target sequence on the sense strand (e.g., a protein coding strand) to be replaced by A and / or one or more T to be replaced by C. In some embodiments, when the guide RNA targets the target sequence, the base editing system of the present invention may cause one or more Gs on a double-stranded sequence (e.g., a protein coding strand) to be replaced by A and / or one or more T to be replaced by C.
[0059] In some embodiments, the "editing sequence" described in the present invention is complementary to the "target sequence" complementary to the sgRNA. In some embodiments, when the target sequence targeted by the guide RNA is the sense strand, the editing sequence is the antisense strand complementary thereto. In some embodiments, when the target sequence targeted by the guide RNA is the antisense strand, the editing sequence is the sense strand complementary thereto. In some embodiments, when the target sequence targeted by the guide RNA is the antisense strand, the editing sequence is the sense strand complementary thereto. In some embodiments, when the target sequence targeted by the guide RNA is the double-stranded DNA targeted by the guide RNA.
[0060] In order to obtain efficient expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the base editing fusion protein is codon optimized for the organism whose genome is to be modified.
[0061] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with codons that are more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit specific preferences for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the efficiency of translation of messenger RNA (mRNA), which in turn is believed to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example at www.kazusa.orjp / codon / The codon usage tables are tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28: 292 (2000).
[0062] In some embodiments, wherein the target sequence is in the genome of an organism. In some embodiments, wherein the organism is a prokaryote. In some embodiments, wherein the prokaryote is a bacterium. In some embodiments, wherein the organism is a eukaryote. In some embodiments, wherein the organism is a plant or a fungus. In some embodiments, wherein the organism is a vertebrate. In some embodiments, wherein the vertebrate is a mammal. In some embodiments, wherein the mammal is a mouse, a rat, or a human. In some embodiments, wherein the organism is a cell. In some embodiments, wherein the cell is a mouse cell, a rat cell, or a human cell. In some embodiments, wherein the cell is a HEK-293 cell.
[0063] In some embodiments, after the base editing system of the present invention is introduced into the cell, the fusion protein of the present invention and the guide RNA are able to form a complex, and the complex specifically targets the target sequence under the mediation of the guide RNA, and causes one or more C to be replaced by T and / or one or more A to be replaced by G in the editing sequence complementary to the target sequence.
[0064] In order to obtain efficient expression in cells, in some embodiments of the present invention, the nucleotide sequence encoding the base editing fusion protein is codon optimized for the organism whose genome is to be modified.
[0065] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more codons) with codons that are more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit specific preferences for certain codons for specific amino acids. Codon bias (differences in codon usage between organisms) is often correlated with the efficiency of translation of messenger RNA (mRNA), which in turn is believed to depend on the nature of the codons being translated and the availability of specific transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell generally reflects the codons most frequently used for peptide synthesis. Therefore, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, for example at www.kazusa.orjp / codon / The codon usage tables are tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28: 292 (2000).
[0066] Another aspect of the present invention provides a method for producing at least one genetically modified cell, characterized in that the method comprises introducing the base editing system of the present invention into at least one of the cells, thereby causing one or more nucleotides in the target nucleic acid region in the at least one cell to be replaced.
[0067] In some embodiments, the method further comprises the step of screening the at least one cell for cells having the desired one or more nucleotide substitutions.
[0068] In some embodiments, the methods of the invention are performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.
[0069] In the present invention, the target nucleic acid region to be modified can be located at any position of the genome, for example, in a functional gene such as a protein coding gene, or, for example, in a gene expression regulatory region such as a promoter region or an enhancer region, thereby achieving modification of the gene function or modification of gene expression. In some embodiments, the desired nucleotide substitution results in a desired gene function modification or gene expression modification.
[0070] In some embodiments, the target nucleic acid region is related to the proterties of the cell or organism. In some embodiments, the mutation in the target nucleic acid region causes a change in the proterties of the cell or organism. In some embodiments, the target nucleic acid region is located in the coding region of protein. In some embodiments, the function-related motif or domain of the target nucleic acid region encoding protein. In some preferred embodiments, one or more nucleotides in the target nucleic acid region replace the amino acid replacement in the amino acid sequence of the protein. In some embodiments, the one or more nucleotides replace the change in the function of the protein.
[0071] In the method of the present invention, the base editing system can be introduced into cells by various methods well known to those skilled in the art.
[0072] Methods that can be used to introduce the base editing system of the present invention into cells include, but are not limited to, calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus and other viruses), gene gun technique, PEG-mediated protoplast transformation, and soil Agrobacterium-mediated transformation.
[0073] In some embodiments, the cell is from a prokaryotic organism such as a bacterium; a eukaryotic organism such as a plant, a fungus, or a vertebrate.
[0074] In some embodiments, the vertebrate is a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; the plant is a crop plant, such as wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava or potato.
[0075] In another aspect, the present invention provides a method for producing a genetically modified plant, comprising introducing the base editing system of the present invention into at least one of the plants, thereby causing one or more nucleotide substitutions within a target nucleic acid region in the genome of the at least one plant.
[0076] In some embodiments, the method further comprises screening the at least one plant for plants having the desired one or more nucleotide substitutions.
[0077] In the method of the present invention, the base editing system can be introduced into plants by various methods known to those skilled in the art. Methods that can be used to introduce the base editing system of the present invention into plants include, but are not limited to, gene gun method, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube channel method, and ovary injection method.
[0078] In the method of the present invention, the modification of the target sequence can be achieved by simply introducing or producing the complex of the base editing fusion protein and the guide RNA in the plant cell, and the modification can be stably inherited without the need to stably transform the exogenous polynucleotide encoding the components of the base editing system into the plant. This avoids the potential off-target effects of the stably existing (continuously produced) base editing system and also avoids the integration of the exogenous nucleotide sequence in the plant genome, thereby having higher biosafety.
[0079] In some preferred embodiments, the introduction is performed in the absence of selection pressure, thereby avoiding integration of the exogenous nucleotide sequence into the plant genome.
[0080] In some embodiments, the introduction includes transforming the base editing system of the present invention into isolated plant cells or tissues, and then regenerating the transformed plant cells or tissues into complete plants. Preferably, the regeneration is carried out in the absence of selection pressure, that is, no selection agent for the selection gene carried on the expression vector is used during tissue culture. Not using a selection agent can improve the regeneration efficiency of the plant and obtain a modified plant that does not contain exogenous nucleotide sequences.
[0081] In other embodiments, the base editing system of the present invention can be transformed into specific parts of the whole plant, such as leaves, stem tips, pollen tubes, panicles or hypocotyls. This is particularly suitable for the transformation of plants that are difficult to regenerate through tissue culture.
[0082] In some embodiments of the invention, the in vitro expressed protein and / or in vitro transcribed RNA molecule (e.g., the expression construct is an in vitro transcribed RNA molecule) is directly transformed into the plant. The protein and / or RNA molecule can achieve base editing in plant cells and then be degraded by the cells, avoiding the integration of exogenous nucleotide sequences in the plant genome.
[0083] Therefore, in some embodiments, genetic modification and breeding of plants using the methods of the present invention can obtain plants whose genomes have no exogenous polynucleotides integrated, ie, non-transgenic (transgene-free) modified plants.
[0084] In some embodiments of the invention, the modified target nucleic acid region is associated with a plant trait such as an agronomic trait, whereby the one or more nucleotide substitutions result in the plant having an altered (preferably improved) trait, such as an agronomic trait, relative to a wild-type plant.
[0085] In some embodiments, the method further comprises the step of screening plants for desired one or more nucleotide substitutions and / or desired traits, such as agronomic traits.
[0086] In some embodiments of the present invention, the method further comprises obtaining offspring of the genetically modified plant. Preferably, the genetically modified plant or its offspring has a desired one or more nucleotide substitutions and / or a desired trait such as an agronomic trait.
[0087] On the other hand, the present invention also provides a genetically modified plant or its offspring or its part, wherein the plant is obtained by the method of the present invention. In some embodiments, the genetically modified plant or its offspring or its part is non-transgenic. Preferably, the genetically modified plant or its offspring has a desired genetic modification and / or a desired trait such as an agronomic trait.
[0088] In another aspect, the present invention also provides a plant breeding method, comprising hybridizing a first genetically modified plant comprising one or more nucleotide substitutions in a target nucleic acid region obtained by the method of the present invention with a second plant not containing the one or more nucleotide substitutions, thereby introducing the one or more nucleotide substitutions into the second plant. Preferably, the first genetically modified plant has a desired trait such as an agronomic trait.
[0089] The present invention also covers the use of the base editing system of the present invention in disease treatment.
[0090] By modifying the disease-related genes through the base editing system of the present invention, it is possible to achieve upregulation, downregulation, inactivation, activation or mutation correction of the disease-related genes, thereby achieving disease prevention and / or treatment. For example, the target nucleic acid region described in the present invention may be located in the protein coding region of the disease-related gene, or, for example, may be located in a gene expression regulatory region such as a promoter region or an enhancer region, so that the disease-related gene function modification or the disease-related gene expression modification can be achieved. Therefore, the modified disease-related genes described herein include modification of the disease-related genes themselves (e.g., protein coding regions), and also include modification of their expression regulatory regions (e.g., promoters, enhancers, introns, etc.).
[0091] "Disease-related" gene refers to any gene that produces a transcription or translation product at an abnormal level or in an abnormal form in cells derived from tissues affected by the disease, compared to tissues or cells of non-disease controls. In the case where the altered expression is related to the appearance and / or progression of the disease, it can be a gene expressed at an abnormally high level; it can be a gene expressed at an abnormally low level. Disease-related genes also refer to genes with one or more mutations or genetic variations that are directly responsible or unbalanced with one or more genes responsible for the etiology of the disease. The mutation or genetic variation is, for example, a single nucleotide variation (SNV). The transcribed or translated product can be known or unknown, and can be at normal or abnormal levels.
[0092] Therefore, the present invention also provides a method for treating a disease in a subject in need thereof, comprising delivering an effective amount of a base editing system of the present invention to the subject to modify a gene associated with the disease (e.g., deaminating mitochondrial DNA by a fusion protein or multiple fusion proteins). Use of a base editing system in the preparation of a pharmaceutical composition for treating a disease in a subject in need thereof, wherein the base editing system is used to modify a gene associated with the disease. The present invention also provides a drug for treating a disease in a subject in need thereof, comprising a base editing system of the present invention, and an optional pharmaceutically acceptable carrier, wherein the base editing system is used to modify a gene associated with the disease.
[0093] In some embodiments, the fusion protein or base editing system described in the present invention is used to introduce point mutations into nucleic acids by deaminating target nuclear bases (e.g., C residues). In some embodiments, the deamination of target nuclear bases leads to the correction of genetic defects, such as in point mutations that cause loss of function in gene products. In some embodiments, genetic defects are associated with diseases or disorders (e.g., lysosomal storage diseases or metabolic diseases, such as type I diabetes). In some embodiments, the methods provided herein can be used to introduce inactivating point mutations into genes or alleles encoding gene products associated with diseases or disorders.
[0094] In some embodiments, the goal of the protocols described herein is to restore the function of dysfunctional genes via genome editing. Nucleobase editing proteins provided herein are useful for in vitro gene editing of human cells, such as correcting disease-associated mutations in human cell culture.
[0095] In some embodiments, the purpose of the scheme described in the present invention is to treat diseases associated with or caused by point mutations, which can be corrected by the DNA base editing fusion protein provided herein. In some embodiments, the disease is a proliferative disease. In some embodiments, the disease is a genetic disease. In some embodiments, the disease is a neoplastic disease. In some embodiments, the disease is a metabolic disease. In some embodiments, the disease is a lysosomal storage disease.
[0096] In some embodiments, the invention describes a method for treating mitochondrial diseases or disorders. As used herein, "mitochondrial diseases" refer to diseases caused by abnormal mitochondria, such as mitochondrial gene mutations, enzyme pathways, etc. Examples of diseases include, but are not limited to: neurological diseases, loss of motor control, muscle weakness and pain, gastrointestinal diseases and dysphagia, poor growth, heart disease, liver disease, diabetes, respiratory complications, epilepsy, vision / hearing problems, lactic acidosis, developmental delays, and susceptibility to infection.
[0097] Examples of diseases described in the present invention include, but are not limited to, genetic diseases, circulatory system diseases, muscle diseases, brain, central nervous and immune system diseases, Alzheimer's disease, secretase disorders, amyotrophic lateral sclerosis (ALS), autism, trinucleotide repeat expansion disorders, hearing diseases, gene-targeted therapy of non-dividing cells (neurons, muscles), liver and kidney diseases, epithelial cell and lung diseases, cancer, Usher syndrome or retinitis pigmentosa-39, cystic fibrosis, HIV and AIDS, beta thalassemia, sickle cell disease, herpes simplex virus, autism, drug addiction, age-related macular degeneration, schizophrenia. Other diseases treated by correcting point mutations or introducing inactivating mutations into disease-related genes are known to those skilled in the art, and the present disclosure is not limited in this regard. In addition to the diseases exemplarily described in the present invention, other related diseases can also be treated with the strategies and fusion proteins provided by the present invention, and the application is obvious to those skilled in the art. The diseases or targets applicable to the present invention refer to WO2015089465A1
[0098] (PCT / US2014 / 070135), WO2016205711A1 (PCT / US2016 / 038181), WO2018141835A1 (PCT / EP2018 / 052491), WO2020191234A1
[0099] (PCT / US2020 / 023713), WO2020191233A1 (PCT / US2020 / 023712), WO2019079347A1 (PCT / US2018 / 056146), WO2021155065A1
[0100] Related diseases for which the base editing system listed in (PCT / US2021 / 015580) is applicable.
[0101] Another aspect of the present invention provides a kit, characterized in that the kit comprises the fusion protein of the present invention, and / or an expression construct containing a nucleotide sequence encoding the fusion protein; or the complex of the present invention; or the base editing system of the present invention; or the host cell of the present invention.
[0102] In some embodiments, the kit further comprises an expression construct encoding a guide RNA backbone, wherein the construct comprises a cloning site that allows a nucleic acid sequence identical or complementary to a target sequence to be cloned into the guide RNA backbone.
[0103] In some embodiments, the kit also includes instructions for use. The kit generally includes a label indicating the intended use and / or method of use of the contents of the kit. The term label includes any written or recorded material provided on or with the kit or otherwise provided with the kit.
[0104] Technical Effects
[0105] (1) Compared with the existing Cas9-based CBE cytosine / ABE adenine base editing system, the novel base editing system eTraC mutant-CBE / eTraC mutant-ABE described in the present invention can recognize sequences of 5'-
[0106] TTN-3' PAM sequence, so it can act on targets other than 5'-NGG-3' PAM;
[0107] (2) The eTraC mutant described in the present invention has higher activity, and the editing efficiency of the new base editing system it constitutes is significantly improved.
[0108] (3) The base editing system of the present invention has a wider editing window, which can expand the application of gene editing tools in genome editing.
[0109] Application scenarios and scope in the editor;
[0110] definition
[0111] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. In addition, the protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology related terms and laboratory operation procedures used herein are terms and routine procedures widely used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present invention are well known to those skilled in the art and are more fully described in the following documents: Sambrook, J., Fritsch, EF and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). At the same time, in order to better understand the present invention, the definitions and explanations of the relevant terms are provided below.
[0112] As used herein, the term "and / or" encompasses all combinations of items connected by the term, and each combination should be considered to have been listed separately herein. For example, "A and / or B" encompasses "A," "A and B," and "B." For example, "A, B, and / or C" encompasses "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."
[0113] When the term "comprising" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present invention. In addition, it is clear to those skilled in the art that the methionine encoded by the start codon at the N-terminus of the polypeptide may be retained in certain practical situations (for example, when expressed in a specific expression system), but it does not substantially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of this application, although it may not contain a methionine encoded by a start codon at the N-terminus, a sequence containing the methionine is also covered at this time, and accordingly, its encoding nucleotide sequence may also contain a start codon; and vice versa.
[0114] "Gene", "genome" as used herein encompasses not only chromosomal DNA present in the cell nucleus, but also organelle DNA present in subcellular components of the cell (eg, mitochondria, plastids).
[0115] As used herein, "organism" includes any organism suitable for genome editing, preferably a eukaryotic organism. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants including monocots and dicots, such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, Arabidopsis, etc.
[0116] "Genetically modified organism" or "genetically modified cell" means an organism or cell that contains an exogenous polynucleotide or a modified gene or expression control sequence in its genome. For example, the exogenous polynucleotide can be stably integrated into the genome of the organism or cell and inherited for consecutive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant DNA construct. The modified gene or expression control sequence is a sequence in the genome of the organism or cell that contains single or multiple deoxynucleotide substitutions, deletions and additions.
[0117] "Polynucleotide", "nucleic acid sequence", "nucleotide sequence" or "nucleic acid fragment" are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers that optionally may contain synthetic, non-natural or altered nucleotide bases. Nucleotides are referred to by their single letter names as follows: "A" is adenosine or deoxyadenosine (RNA or DNA, respectively), "C" represents cytidine or deoxycytidine, "G" represents guanosine or deoxyguanosine, "U" represents uridine, "T" represents deoxythymidine, "R" represents purine (A or G), "Y" represents pyrimidine (C or T), "K" represents G or T, "H" represents A or C or T, "I" represents inosine, and "N" represents any nucleotide.
[0118] "Polypeptide", "peptide", and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide", "peptide", "amino acid sequence" and "protein" may also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma carboxylation of glutamic acid residues, hydroxylation and ADP-ribosylation.
[0119] Sequence "identity" has a meaning recognized in the art, and the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions can be calculated using published techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, for example: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, AM, and Griffin, HG, eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). Although there are many methods to measure the identity between two polynucleotides or polypeptides, the term "identity" is well known to those of skill (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).
[0120] In peptides or proteins, suitable conservative amino acid substitutions are known to those skilled in the art, and generally can be carried out without changing the biological activity of the resulting molecule. Generally, those skilled in the art recognize that single amino acid substitutions in non-essential regions of a polypeptide do not substantially change biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub.co., p. 224).
[0121] As used herein, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence may refer to the transcription of the nucleotide sequence (such as transcription to generate mRNA or functional RNA) and / or translation of RNA into a precursor or mature protein.
[0122] The "expression construct" of the present invention can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be a translatable RNA (such as mRNA).
[0123] An "expression construct" of the present invention may comprise regulatory sequences and a nucleotide sequence of interest from different sources, or regulatory sequences and a nucleotide sequence of interest from the same source but arranged in a manner different from that normally found in nature.
[0124] The term "deaminase" refers to an enzyme that catalyzes a deamination reaction. In some embodiments, the deaminase is a cytidine deaminase or an adenosine deaminase, i.e., an enzyme that can remove the amino group of a cytidine molecule or an adenosine molecule. For example, an enzyme having the same amino acid sequence as shown in any of SEQ ID NOs: 3, 21, 51-54, or having identity and still retaining deamination activity. For example, variants or mutants having a certain degree (e.g., 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%) of sequence identity, these variants or mutants still retain deamination activity.
[0125] The term "mutation" refers to the substitution of a residue within a sequence (e.g., a nucleic acid or amino acid sequence) with another residue or the deletion or insertion of one or more residues within a sequence. Mutations are usually described herein by identifying the original residue, followed by the position of the residue within the sequence and the identity of the newly substituted residue.
[0126] The terms "nucleobase editing system (NBE)" or "base editing system (BE)" can be used interchangeably to refer to a system that achieves precise editing of DNA bases based on gene editing technology. The base editing system described herein comprises a fusion protein and a guide RNA. In some embodiments, the fusion protein comprises an eTraC protein that completely or partially loses cleavage activity fused to a deaminase. In some embodiments, the base editing system can be at least one of the following components: a) a fusion protein of the present invention, and a guide RNA; b) an expression construct comprising a nucleotide sequence of the fusion protein of the present invention, and a guide RNA; c) a fusion protein of the present invention, and an expression construct comprising a nucleotide sequence encoding a guide RNA; d) an expression construct comprising a nucleotide sequence encoding the fusion protein of the present invention, and an expression construct comprising a nucleotide sequence encoding a guide RNA. e) A fusion protein of the present invention and a guide RNA bound to the DNA binding protein domain of the fusion protein of the present invention.
[0127] The terms "DNA binding protein" or "sequence-specific DNA binding protein" can be used interchangeably to refer to a nuclease that completely or partially loses cleavage activity herein. In some embodiments, the nuclease of the present invention is selected from a TraC effector protein, preferably an eTraC protein, whose amino acid sequence is shown in the amino acid sequence of SEQ ID NO:1. The completely or partially lost nuclease can be selected as an eTraC protein that completely or partially loses cleavage activity. In some embodiments, the eTraC protein that completely or partially loses cleavage activity comprises 1-3 corresponding mutations in the three amino acids D253, E381, and D464 of SEQ ID NO:1, which completely or partially loses the cleavage activity of the eTraC protein. In some embodiments, the fusion protein comprises three mutations D253, E381, and D464.
[0128] "Promoter" refers to a nucleic acid fragment that can control the transcription of another nucleic acid fragment. In some embodiments of the present invention, a promoter is a promoter that can control the transcription of a gene in a cell, whether or not it is derived from the cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally regulated promoter or an inducible promoter.
[0129] As used herein, the term "operably linked" refers to the connection of a regulatory element (e.g., but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (e.g., a coding sequence or an open reading frame) such that transcription of the nucleotide sequence is controlled and regulated by the transcription regulatory element. Techniques for operably linking a regulatory element region to a nucleic acid molecule are known in the art.
[0130] "Introducing" a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, RNA, etc.) or a protein into an organism refers to transforming an organism cell with the nucleic acid or protein so that the nucleic acid or protein can function in the cell. "Transformation" as used in the present invention includes stable transformation and transient transformation.
[0131] "Transient transformation" refers to the introduction of a nucleic acid molecule or protein into a cell to perform its function without the stable inheritance of the exogenous nucleotide sequence. In transient transformation, the exogenous nucleic acid sequence is not integrated into the genome.
[0132] "Trait" refers to a physiological, morphological, biochemical, or physical characteristic of a cell or organism.
[0133] "Agronomic traits" specifically refer to measurable indicator parameters of crop plants, including but not limited to: leaf green, grain yield, growth rate, total biomass or accumulation rate, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, plant total nitrogen content, fruit nitrogen content, seed nitrogen content, plant vegetative tissue nitrogen content, plant total free amino acid content, fruit free amino acid content, seed free amino acid content, plant vegetative tissue free amino acid content, plant total protein content, fruit protein content, seed protein content, plant vegetative tissue protein content, herbicide resistance and drought resistance, nitrogen absorption, root lodging, harvest index, stem lodging, plant height, ear height, ear length, disease resistance, cold resistance, salt resistance and tillering number, etc.
[0134] Organisms that can be subjected to genome modification by the base editing system of the present invention include any organisms suitable for base editing, preferably eukaryotic organisms. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants, including monocots and dicots, for example, the plants are crop plants, including but not limited to wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava and potatoes. BRIEF DESCRIPTION OF THE DRAWINGS
[0136] Figure 1 , Schematic diagram of the expression construct of the eTraC-CBE base editing system.
[0137] Figure 2 , Schematic diagram of the expression construct of the eTraC-ABE base editing system.
[0138] Figure 3 , eTraC and its mutants constructed CBE base editing efficiency at different sites.
[0139] Figure 4 , eTraC and its mutants constructed base editing efficiency at different sites.
[0140] Figure 5 , the base editing efficiency of CBE / ABE constructed by the preferred eTraC mutant at different sites.
[0141] Figure 6 , the editing window of CBE / ABE constructed by the preferred eTraC mutant at site21.
[0142] Figure 7 , the editing window of CBE / ABE constructed by the preferred eTraC mutant at site22.
[0143] Figure 8, the editing window of CBE / ABE constructed by the preferred eTraC mutant at site23.
[0144] Fig. 9 , the editing window of CBE / ABE constructed by the preferred eTraC mutant at site24. DETAILED DESCRIPTION
[0145] Example 1 Construction of eTraC-based base editing system fusion protein
[0146] The schematic diagram of the fusion protein expression construct of the base editing system of this embodiment is shown as an example. Figure 1-2 shown.
[0147] Among them, the fusion protein of the cytosine base editing system includes the following structure: NH2-[NLS]-[cytidine deaminase domain]-[DNA binding protein domain]-[UGI domain]-[UGI domain]-[NLS]-COOH, where "-" represents a direct connection between the domains or an optional linker connection. More specifically, the fusion protein structure of this embodiment is: NLS-deaminase-32 amino acid linker-DNA binding protein domain-32 amino acid linker-UGI domain-32 amino acid linker-UGI domain-NLS, where the sequence of the 32 amino acid linker is
[0148] SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 15).
[0149] The fusion protein of the adenine base editing system includes the following structure: NH2-[NLS]-[adenosine deaminase domain]-[DNA binding protein domain]-[NLS]-COOH, where "-" indicates a direct connection between the domains or an optional linker connection. More specifically, the sequence of the linker in the embodiment of the present invention is
[0150] SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 15).
[0151] Example 2 Optimization of eTraC nuclease
[0152] In this embodiment, the sites near the active cleavage region of the eTraC nuclease were screened through structural prediction, and five mutation types were screened at five sites of the eTraC nuclease, namely R150D, D253A, V254A, E381A, and D464A, and single-point mutations were performed on the above five mutation types in the eTraC nuclease. The mutated eTraC nuclease was constructed into a fusion protein as shown in Example 1: TadA-8e adenosine deaminase was connected to its N-terminus to form eTraC-ABE, and the DNA sequence to be edited was base edited from A to G; mini-Sdd6 cytidine deaminase was connected to its N-terminus to form eTraC-CBE (mini-Sdd6), and the DNA sequence to be edited was base edited from C to T. At the same time, three sgRNAs were designed for four different editing sites (site21-24) of the HEK293T cell genome (the editing sequences targeted by sgRNA are shown in SEQ ID NOs:59-62, and the sgRNA sequences are shown in SEQ ID NOs:55-58) to detect the two different types of base editors to detect whether the amino acid mutations in the eTraC nuclease active cleavage region will have an optimization effect on the base editing activity of the eTraC nuclease.
[0153] HEK293T cells were co-transfected with multiple eTraC mutant editing system fusion protein encoding plasmids and corresponding sgRNA plasmids. After 72 hours of transfection, cell DNA was extracted, and the base editing efficiency of the editing sequence was detected by next-generation sequencing (NGS). The detection primers are shown in Table 1. For specific experimental steps, please refer to (Huang, Jiaying et al. "Discovery of deaminase functions by structure-based protein clustering." Cell vol. 186, 15 (2023): 3182-3195. e14.).
[0154] The expression constructs and base editing results of CBE are shown in Figure 2. Figure 3 The expression construct and base editing results of ABE are shown in Figure 4As shown. Among them, the three single mutants of eTraC-D253A, eTraC-E381A and eTraC-D464A showed significantly improved base editing efficiency in two different types of base editing systems, ABE and CBE, especially eTraC-D253A, which had the highest editing efficiency. The editing efficiency of eTraC-R150D and eTraC-V254A was low, and even slightly lower than that of the unmutated eTraC. It is worth noting that the two mutants, eTraC-D253A and eTraC-V254A, differed in amino acid position by only one position, but showed completely different editing abilities, which indicates that the amino acid mutation position has a significant impact on the activity of the eTraC nuclease.
[0155] The three preferred single mutation sites D253A, E381A and D464A were mutated simultaneously to form an eTraC triple mutant, which was named as the eTraC protein with lost DNA cleavage activity (dead eTraC) in the present invention. Dead eTraC-ABE and dead eTraC-CBE (mini-Sdd6) were constructed as described above. The results are as follows Figure 3 and Figure 4 As shown, its editing efficiency is greatly improved compared with the unmutated eTraC nuclease in two different types of base editing systems, ABE and CBE. At the same time, the editing efficiency is also greatly improved compared with the single mutants of eTraC-E381A and eTraC-D464A. Even at site 22 of the CBE editing system, the editing efficiency of dead eTraC is also improved compared with the single mutant with the highest editing efficiency, eTraC-D253A.
[0156] Table 1
[0157]
[0158] Example 3 Verification of the editing effect of eTraC mutant base editing system fusion protein in cells
[0159] In this implementation, the eTraC mutant was used as the DNA binding protein domain, and different cytidine / adenosine deaminases were fused to construct different types of base editing system fusion proteins to verify the editing effects of different base editing systems. The fusion proteins are shown in Table 2 below, and the experimental and detection methods are the same as in Example 2.
[0160] Table 2
[0161]
[0162] In this example, four sgRNAs were designed for four different editing sites (site 21-24) of the HEK293T cell genome (the editing sequences targeted by the sgRNAs are shown in SEQ ID NOs: 59-62, and the sgRNA sequences are shown in SEQ ID NOs: 55-58), and the detection primers are shown in Table 1.
[0163] The experimental results are as follows Figure 5 As shown, all eTraC mutants-CBE and eTraC mutants-ABE connected to different deaminases in the present invention showed the ability to perform base editing on the DNA sequence to be edited. Among them, eTraC-D253A connected to various deaminases showed high editing efficiency at almost all sites; at the same time, the eTraC-D253A E381A double mutant also showed high editing efficiency when connected to hA3A deaminase; in addition, dead eTraC also showed high editing efficiency at site22 when connected to mini-Sdd6 or mini-Sdd9 deaminase.
[0164] The above results show that various base editing systems composed of different eTraC mutants in this embodiment have good editing effects and have a wide range of application scenarios.
[0165] Example 4 Editing window of eTraC mutant base editing system fusion protein in cells
[0166] In this example, the editing windows of the eTraC mutant base editing system at different sites were further analyzed, wherein the first base window close to the PAM on the editing site sequence was defined as 1. The NGS results are shown in Figure 2. Figure 6-9 As shown (wherein the editing efficiency occurring in the non-sgRNA binding chain is represented by a positive number, and the editing efficiency occurring in the sgRNA binding chain is represented by a negative number).
[0167] The results showed that all CBEs and ABEs constructed by eTraC mutants connected to different deaminases had a wide editing window, with the editing window of CBE covering C3-C18; the editing window of ABE covering A2-A14. Among them, the CBE constructed by the eTraC mutant can be edited on both strands of the DNA targeted by the sgRNA. Taking the CBE with mini-Sdd7 as the deaminase as an example, the CBE constructed by the eTraC mutant in the present invention showed obvious double-strand editing activity at site 21-24, while in the CBE base editing system with Cas9 as the DNA binding protein, when mini-Sdd7 was used as the deaminase, no double-strand editing activity was shown (see Huang J, Lin Q, Fei H, et al. Discovery of deaminase functions by structure-based protein clustering [J]. Cell, 2023, 186 (15): 3182-3195. e14.)
[0168] In summary, the eTraC mutant base editing system expands the editing window, enriches the editing types, and helps to achieve precise editing of target sites of various sequences in the genome, thereby further expanding the application scenarios of base editing.
[0169] Although the present invention has been described with reference to specific embodiments thereof, it will be understood by those skilled in the art that various changes may be made and equivalents may be substituted without departing from the true spirit and scope of the present invention. In addition, many modifications may be made to adapt specific circumstances, materials, compositions of matter, processes, process steps or steps to the purpose, concept and scope of the present invention. All such modifications are intended to be within the scope of the appended claims. Sequence Listing:
[0170] >SEQ ID NO:1eTraC
[0171]
[0172] >SEQ ID NO:2dead eTraC(D253A,E381A,D464A)
[0173]
[0174] >SEQ ID NO:3hAPOBEC3A
[0175]
[0176] >SEQ ID NO:4UGI
[0177]
[0178] >SEQ ID NO:5 NLS
[0179]
[0180] >SEQ ID NO:6 NLS
[0181]
[0182] >SEQ ID NO:7 NLS
[0183]
[0184] >SEQ ID NO:8 NLS
[0185]
[0186] >SEQ ID NO:9 linker
[0187]
[0188] >SEQ ID NO:10 linker
[0189]
[0190] >SEQ ID NO:11 linker
[0191]
[0192] >SEQ ID NO:12 linker
[0193]
[0194] >SEQ ID NO:13 linker
[0195]
[0196] >SEQ ID NO:14 linker
[0197]
[0198] >SEQ ID NO:15 linker
[0199]
[0200] >SEQ ID NO:16 dead eTraC-CBE(mini-Sdd3)
[0201]
[0202] >SEQ ID NO:17 dead eTraC-CBE(mini-Sdd6)
[0203]
[0204] >SEQ ID NO:18 dead eTraC-CBE(mini-Sdd7)
[0205]
[0206] >SEQ ID NO:19 dead eTraC-CBE(mini-Sdd9)
[0207]
[0208]
[0209] >SEQ ID NO:20 dead eTraC-CBE(A3A)
[0210]
[0211] >SEQ ID NO:21 TadA*ABE8e(TadA-8e):
[0212] >SEQ ID NO:22 dead eTraC-ABE
[0213]
[0214]
[0215] >SEQ ID NO:23 eTraC-V254A
[0216]
[0217] >SEQ ID NO:24 eTraC-V254A-CBE(mini-Sdd6)
[0218]
[0219]
[0220] >SEQ ID NO:25 eTraC-V254A-ABE
[0221]
[0222] >SEQ ID NO:26 eTraC-D253A
[0223]
[0224] >SEQ ID NO:27 eTraC-D253A-CBE(mini-Sdd3)
[0225]
[0226] >SEQ ID NO:28 eTraC-D253A-CBE(mini-Sdd6)
[0227]
[0228]
[0229] >SEQ ID NO:29 eTraC-D253A-CBE(mini-Sdd7)
[0230]
[0231] >SEQ ID NO:30 eTraC-D253A-CBE(mini-Sdd9)
[0232]
[0233]
[0234] >SEQ ID NO:31 eTraC-D253A-CBE(A3A)
[0235]
[0236] >SEQ ID NO:32 eTraC-D253A-ABE
[0237]
[0238]
[0239] >SEQ ID NO:33 eTraC-E381A
[0240]
[0241] >SEQ ID NO:34 eTraC-E381A-CBE(mini-Sdd3)
[0242]
[0243]
[0244] >SEQ ID NO:35 eTraC-E381A-CBE(mini-Sdd6)
[0245]
[0246] >SEQ ID NO:36 eTraC-E381A-CBE(mini-Sdd7)
[0247]
[0248]
[0249] >SEQ ID NO:37 eTraC-E381A-CBE(mini-Sdd9)
[0250]
[0251] >SEQ ID NO:38 eTraC-E381A-CBE(A3A)
[0252]
[0253]
[0254] >SEQ ID NO:39 eTraC-E381A-ABE
[0255]
[0256] >SEQ ID NO:40 eTraC-D464A
[0257]
[0258]
[0259] >SEQ ID NO:41 eTraC-D464A-CBE(mini-Sdd3)
[0260]
[0261] >SEQ ID NO:42 eTraC-D464A-CBE(mini-Sdd6)
[0262]
[0263]
[0264] >SEQ ID NO:43 eTraC-D464A-CBE(mini-Sdd7)
[0265]
[0266] >SEQ ID NO:44 eTraC-D464A-CBE(mini-Sdd9)
[0267]
[0268]
[0269] >SEQ ID NO:45 eTraC-D464A-CBE(A3A)
[0270]
[0271] >SEQ ID NO:46 eTraC-D464A-ABE
[0272]
[0273]
[0274] >SEQ ID NO:47 eTraC-D253A E381A
[0275]
[0276] >SEQ ID NO:48 eTraC-D253A E381A-CBE(mini-Sdd6)
[0277]
[0278]
[0279] >SEQ ID NO:49 eTraC-D253A E381A-CBE(A3A)
[0280]
[0281] >SEQ ID NO:50 eTraC-D253A E381A-ABE
[0282]
[0283]
[0284] >SEQ ID NO:51 mini-Sdd3
[0285]
[0286] >SEQ ID NO:52 mini-Sdd6
[0287]
[0288] >SEQ ID NO:53 mini-Sdd7
[0289]
[0290] >SEQ ID NO:54 mini-Sdd9
[0291]
[0292] >SEQ ID NO:55 site21-sgRNA
[0293]
[0294] >SEQ ID NO:56 site22-sgRNA
[0295]
[0296] >SEQ ID NO:57 site23-sgRNA
[0297]
[0298] >SEQ ID NO:58site24-sgRNA
[0299]
[0300] >SEQ ID NO:59site21-20nt
[0301]
[0302] >SEQ ID NO:60site22-20nt
[0303]
[0304] >SEQ ID NO:61 site23-20nt
[0305]
[0306] >SEQ ID NO:62site24-20nt
[0307]
[0308] >SEQ ID NO:63site21-detection primer F
[0309]
[0310] >SEQ ID NO:64site21-detection primer R
[0311]
[0312] >SEQ ID NO:65site22-detection primer F
[0313]
[0314] >SEQ ID NO:66site22-detection primer R
[0315]
[0316] >SEQ ID NO:67site23-detection primer F
[0317]
[0318] >SEQ ID NO:68 site23-detection primer R
[0319]
[0320] >SEQ ID NO:69site24-detection primer F
[0321]
[0322] >SEQ ID NO:70 site24-detection primer R
[0323]
[0324] >SEQ ID NO:71eTraC-R150D
[0325]
[0326] >SEQ ID NO:72eTraC-R150D-CBE(mini-Sdd6)
[0327]
[0328]
[0329] >SEQ ID NO:73 eTraC-R150D-ABE
[0330]
Claims
1. A DNA binding protein, characterized in that The DNA binding protein is a nuclease that completely or partially loses DNA double-strand cleavage activity, and the nuclease is an eTraC effector protein mutant; The eTraC effector protein mutant comprises one or more mutations in amino acids 253, 381, and 464 of SEQ ID NO: 1, and the amino acid sequence of the eTraC effector protein mutant has at least 70%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity compared to SEQ ID NO:
1.
2. The DNA binding protein according to claim 1, characterized in that The DNA binding protein comprises an amino acid sequence in which any one or more mutation types of D253A, E381A, and D464A exist in the amino acid sequence provided in SEQ ID NO:
1.
3. The DNA binding protein according to any one of claims 1 or 2, characterized in that The DNA binding protein comprises an amino acid sequence of double mutations of D253A and E381A in the amino acid sequence provided in SEQ ID NO:
1.
4. The DNA binding protein according to any one of claims 1 or 2, characterized in that The DNA binding protein comprises three mutated amino acid sequences of D253A, E381A and D464A in the amino acid sequence provided in SEQ ID NO:
1.
5. The DNA binding protein according to any one of claims 1 to 4, characterized in that The DNA binding protein has an amino acid sequence as shown in any one of SEQ ID NO:2, SEQ ID NO:26, SEQ ID NO:33, SEQ ID NO:40 or SEQ ID NO:
47.
6. A fusion protein, characterized in that The fusion protein comprises: i) a DNA binding protein domain, the DNA binding protein domain comprising at least one DNA binding protein according to any one of claims 1 to 5; ii) Deaminase domain.
7. The fusion protein according to claim 6, characterized in that The deaminase domain is a cytidine deaminase domain or an adenosine deaminase domain.
8. The fusion protein according to claim 7, characterized in that The cytidine deaminase is selected from the group consisting of APOBEC family deaminase, SCP1.201 family deaminase or homologs thereof.
9. The fusion protein according to claim 8, characterized in that The APOBEC family deaminase is selected from AID deaminase, APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase and APOBEC3H deaminase.
10. The fusion protein according to claim 8, characterized in that The SCP1.201 family deaminase is selected from Sdd2, Sdd3, Sdd4 deaminase, mini-Sdd3 deaminase, Sdd6 deaminase, mini-Sdd6 deaminase, Sdd7 deaminase, mini-Sdd7 deaminase, Sdd9 deaminase, mini-Sdd9 deaminase, Sdd10 deaminase, and Sdd59 deaminase.
11. The fusion protein according to claim 8, characterized in that The homolog of the APOBEC family deaminase is cytidine deaminase 1 (pmCDA1) from sea lamprey (Petromyzon marinus).
12. The fusion protein according to claim 7, characterized in that The cytidine deaminase comprises the amino acid sequence shown in any one of SEQ ID NO: 3 and SEQ ID NO: 51-54.
13. The fusion protein according to claim 7, characterized in that The adenosine deaminase comprises TadA deaminase, QB34 deaminase or a functional variant thereof.
14. The fusion protein according to claim 13, characterized in that The adenosine deaminase is TadA-8e deaminase. The fusion protein according to claim 14 , wherein the adenosine deaminase has an amino acid sequence as shown in SEQ ID NO:
21.
16. The fusion protein according to claim 7, characterized in that The fusion protein further comprises: iii) Uracil glycosylase inhibitor (UGI) domain.
17. The fusion protein according to claim 16, characterized in that The UGI domain comprises the amino acid sequence shown in SEQ ID NO:
4.
18. The fusion protein according to any one of claims 6 or 16, characterized in that The fusion protein further comprises: iv) Nuclear localization sequence (NLS).
19. The fusion protein according to claim 18, characterized in that The NLS comprises the amino acid sequence shown in any one of SEQ ID NOs: 5-8.
20. The fusion protein according to claim 18, characterized in that The fusion protein comprises at least one of the following structural formulas (1) or (2): (1) NH2-[NLS]-[cytidine deaminase domain]-[DNA binding protein domain]-[UGI domain]-[UGI domain]-[NLS]-COOH; (2) NH2-[NLS]-[adenosine deaminase domain]-[DNA binding protein domain]-[NLS]-COOH; Wherein "-" indicates direct connection or connection through an optional linker.
21. The fusion protein according to claim 20, characterized in that The linker comprises the amino acid sequence (GGGS)n(SEQ ID NO:9), (GGGGS)n(SEQ ID NO:10), (G)n, (EAAAK)n(SEQ ID NO:11), (GGS)n(SEQ ID NO:12), (SGGS)n(SEQ ID NO:13), SGSETPGTSESATPES (SEQ ID NO: 14), or (XP)n motif, or a combination thereof, wherein n is independently an integer from 1-30, and wherein X is any amino acid.
22. The fusion protein of claim 21, wherein the fusion protein comprises the amino acid sequence shown in any one of SEQ ID NOs: 16-20, 22, 27-32, 34-39, 41-46, 48-50.
23. A composite, characterized in that The complex comprises the fusion protein of any one of claims 6 to 22 and a guide RNA bound to the DNA binding protein domain of the fusion protein.
24. A base editing system for modifying a target nucleic acid region, characterized in that: The base editing system comprises at least one of the following components a) to e): a) the fusion protein of any one of claims 6 to 22, and a guide RNA; b) an expression construct comprising a nucleotide sequence encoding the fusion protein of any one of claims 6 to 22, and a guide RNA; c) the fusion protein of any one of claims 6 to 22, and an expression construct comprising a nucleotide sequence encoding a guide RNA; d) an expression construct comprising a nucleotide sequence encoding the fusion protein of any one of claims 6 to 22, and an expression construct comprising a nucleotide sequence encoding a guide RNA; or e) The complex according to claim 23.
25. A host cell, characterized in that The host cell comprises the fusion protein of any one of claims 6-22, and / or an expression construct containing a nucleotide sequence encoding the fusion protein; or the complex of claim 23; or the base editing system of claim 24.
26. A base editing method, characterized in that: The base editing method comprises contacting the fusion protein of any one of claims 6 to 22 with a guide RNA that is complementary to a sequence of at least 10 consecutive nucleotides in a target sequence in the genome of an organism.
27. The base editing method of claim 26, characterized in that The organism is a prokaryotic organism such as a bacterium; a eukaryotic organism such as a plant, a fungus or a vertebrate.
28. The base editing method of claim 27, characterized in that The vertebrate is a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, or cat.
29. The base editing method of claim 28, characterized in that The plant is a crop plant, for example wheat, rice, corn, soybean, sunflower, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugar cane, tomato, tobacco, cassava or potato.
30. A method for producing at least one genetically modified cell, characterized in that The method comprises introducing the base editing system of claim 24 into at least one of the cells, thereby causing one or more nucleotides in the target nucleic acid region to be replaced in the at least one cell.
31. The method of claim 30, wherein: The method further comprises the step of screening the at least one cell for cells having the desired one or more nucleotide substitutions.
32. The method according to any one of claims 30 to 31, characterized in that The cells are from mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, cats; poultry such as chickens, ducks, geese; plants, preferably crop plants, such as wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, potatoes.
33. A kit, characterized in that The kit comprises the fusion protein of any one of claims 6-22, and / or an expression construct containing a nucleotide sequence encoding the fusion protein; or the complex of claim 23; or the base editing system of claim 24; or the host cell of claim 25.
34. The kit according to claim 33, characterized in that The kit also includes an expression construct encoding a guide RNA backbone, wherein the construct comprises a cloning site that allows a nucleic acid sequence identical or complementary to a target sequence to be cloned into the guide RNA backbone.
Citation Information
Patent Citations
Novel crispr enzymes and systems
WO2016205711A1
Compounds, compositions and methods for cancer treatment
WO2018141835A1
Uses of adenosine base editors
WO2019079347A1
Methods and compositions for editing nucleotide sequences
WO2020191233A1
Base editors, compositions, and methods for modifying the mitochondrial genome
WO2021155065A1
Cited By
Genome editing system based on large serine recombinase and application thereof
CN121065132A
Bxb1 recombinase with improved activity and application thereof
CN121065133A
Genome editing system based on large serine recombinase mutant and application thereof
CN121294392A