Genome editing system based on large serine recombinase and use thereof
By developing a novel large serine recombinase and its genome editing system, the problems of low insertion efficiency and poor specificity of existing recombinases have been solved, enabling efficient editing of large DNA fragments in the genome and expanding the application scenarios of gene editing tools.
Patent Information
- Application Number
- PCT/CN2025/112204
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-05
- Filing Date
- 2025-08-01
- Publication Date
- 2026-02-05
AI Technical Summary
Existing recombinases suffer from low insertion efficiency, poor specificity, and cytotoxicity in genome editing. In particular, macroserine recombinases have limited recombination efficiency in animal and plant cells, making it difficult to meet the needs of efficient and precise editing.
To develop a novel large serine recombinase and its genome editing system, which contains specific amino acid and nucleotide sequences, and combines CRISPR effector proteins and reverse transcriptase to achieve efficient insertion, deletion, and flipping operations on the genome.
Novel large serine recombinases exhibit highly efficient enzymatic activity in plant and animal genomes, significantly improving the integration efficiency of large DNA fragments, enriching the gene editing toolkit, and supporting precise manipulation of target DNA sequences.
Smart Images

Figure CN2025112204_05022026_PF_FP_ABST
Abstract
Description
Genome editing system based on megasericine recombinase and application thereof
[0001] Priority and related applications
[0002] The present application claims priority to Chinese Patent Application No. 202411061014.4, filed on August 2, 2024, entitled “Genome editing system based on megasericine recombinase and application thereof” and Chinese Patent Application No. 202411067822.1, filed on August 5, 2024, entitled “Genome editing system based on megasericine recombinase and application thereof”. The entire contents of the above-cited patent applications are hereby incorporated by reference in their entirety. TECHNICAL FIELD
[0003] The present application belongs to the field of genetic engineering. Specifically, the present application relates to a genome editing system based on megasericine recombinase and application thereof. More specifically, the present application provides megasericine recombinase that can act on the genome, a genome editing system based thereon and application thereof. Methods for editing the genome of an organism using the genome editing system, and genetically modified organisms and their offspring produced by the methods. BACKGROUND
[0004] Site-specific recombinases catalyze the specific recombination of fragments between two specific DNA sequences, mediate DNA fragment integration, excision or inversion, and play a key role in the life cycle of many microorganisms, including bacteria and bacteriophages.
[0005] According to the difference of the residues mediating catalysis, recombinases are divided into tyrosine recombinases and serine recombinases. According to the directionality of the recombination reaction, tyrosine recombinases are further divided into reversible and irreversible tyrosine recombinases; while serine recombinases, due to the recognition of DNA sequences that are not completely inverted palindromic sequences, mediate irreversible recombination reactions, and are divided into megasericine recombinases and small sericin recombinases according to the size and function of the protein molecules. Due to the characteristics of not introducing DNA double-strand breaks in the process of editing the genome and being able to edit large fragments of chromosomes, recombinases can be used as an ideal gene editing tool for DNA large fragment insertion, deletion and inversion, for the treatment of diseases in animals and plants, trait improvement, etc.
[0006] At present, the recombinases in the prior art have problems of low insertion efficiency and poor specificity. For example, the recombination activity of tyrosine recombinase is reversible, resulting in low editing efficiency, and the tyrosine family recombinase has few choices, the cargo fragment that can be effectively and accurately edited is small, and in addition, the tyrosine recombinase Cre has cytotoxicity after overexpression. The large serine recombinase (LSR) has the characteristics of irreversible recombination, making it a potential genome editing tool. However, LSR still has problems of low editing efficiency and single variety. So far, only a few LSRs have been discovered, including Bxb1 and PhiC31 recombinases, but their recombination efficiency in animal and plant cells is very limited. Therefore, finding a recombinase with high activity, high recombination efficiency and wide application scenarios is of great significance for expanding the existing DNA large fragment editing system and developing a gene editing tool library that can accurately manipulate target DNA sequences. SUMMARY
[0007] Problems to be solved by the invention
[0008] The present inventors have unexpectedly found a new type of large serine recombinase that can act on large fragments of DNA in the genome based on a large number of experiments and explorations. Based on this discovery, the inventors have developed a new type of genome editing system based on large serine recombinase, which enriches the application scenarios and choices of recombinase gene editing.
[0009] Solutions to the problems
[0010] The first aspect of the present application provides a large serine recombinase, wherein the large serine recombinase comprises an amino acid sequence shown in any one of SEQ ID NO: 2, 1 and 3-5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity with the amino acid sequence shown in any one of SEQ ID NO: 2, 1 and 3-5.
[0011] In some embodiments, the large serine recombinase corresponds to a recombinase recognition site sequence (RS) selected from an attP site and an attB site.
[0012] In some embodiments, the attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 6-12, and / or the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 13-19.
[0013] In some embodiments, the megasynthase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to the amino acid sequence set forth in SEQ ID NO: 2.
[0014] In some embodiments, the attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 7-8, and / or, the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 14-15.
[0015] The second aspect of the present application provides a genome editing system, wherein the genome editing system comprises:
[0016] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site;
[0017] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase, the recombinase comprising the megasynthase recombinase of the first aspect of the present application.
[0018] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.
[0019] In some embodiments, one or more of the recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.
[0020] In some embodiments, the recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
[0021] In some embodiments, the CRISPR effector protein functional variant is a CRISPR nuclease with full / partial loss of cleavage activity, preferably, the CRISPR nuclease functional variant is a CRISPR nickase, for example, Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase or TraC nickase.
[0022] In some embodiments, the guide RNA comprises a scaffold sequence and a primer binding sequence, and an integration template sequence of the recombinase recognition site sequence (RS).
[0023] In some embodiments, the recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase.
[0024] In some embodiments, the CRISPR nuclease of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nuclease and targets a desired location in the genome, wherein the CRISPR nuclease makes a cut in a strand of the genome, and the reverse transcriptase incorporates a RS integration template sequence in the guide RNA into the cut site, thereby inserting at least one recombinase recognition site sequence (RS) recognizable by the recombinase at the desired location in the genome.
[0025] In some embodiments, the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii Marathonase RT (Marathon RT).
[0026] In some embodiments, the recombinase recognition site integration unit is not covalently linked to a recombinase.
[0027] In some embodiments, the recombinase recognition site integration unit is covalently linked to a recombinase.
[0028] In some embodiments, the genome editing system further comprises a donor comprising: 3) an exogenous nucleotide sequence to be inserted into the genome, optionally, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer.
[0029] In some embodiments, the exogenous nucleotide sequence further comprises one or more recombinase recognition site sequences (RS) recognizable by a recombinase, optionally, wherein the exogenous nucleotide sequence is flanked by one or two recombinase recognition site sequences (RS).
[0030] In some embodiments, the recombinase recognition site sequence (RS) is selected from the group consisting of an attP site, an attB site, optionally, wherein the attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 6-12, and the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 13-19.
[0031] A third aspect of the present application provides a fusion protein comprising the genome editing system of the second aspect of the present application, wherein the CRISPR nuclease is linked to a reverse transcriptase at the C-terminus, and the reverse transcriptase is linked to a recombinase via a linker.
[0032] In a fourth aspect, the present application provides a polynucleotide comprising a nucleotide sequence encoding the megaserine recombinase of the first aspect, the genome editing system of the second aspect, or the fusion protein of the third aspect of the present application, optionally, the polynucleotide is RNA, e.g., mRNA.
[0033] In a fifth aspect, the present application provides an expression construct comprising the polynucleotide of the fourth aspect of the present application.
[0034] In a sixth aspect, the present application provides a cell comprising the megaserine recombinase of the first aspect, the genome editing system of the second aspect, the fusion protein of the third aspect, the polynucleotide of the fourth aspect, or the expression construct of the fifth aspect of the present application.
[0035] In a seventh aspect, the present application provides a kit comprising the megaserine recombinase of the first aspect, the genome editing system of the second aspect, the fusion protein of the third aspect, the polynucleotide of the fourth aspect, the expression construct of the fifth aspect, or the cell of the sixth aspect of the present application.
[0036] In a seventh aspect, the present application provides a method for performing gene editing in an organism or a cell of an organism, wherein the method comprises introducing into the organism or the cell of the organism the megaserine recombinase of the first aspect, the polynucleotide of the fourth aspect, the expression construct of the fifth aspect, the genome editing system of the second aspect, or the fusion protein of the third aspect of the present application.
[0037] In some embodiments, components 1), 2), and optionally 3) of the genome editing system are introduced into the organism or the cell of the organism simultaneously.
[0038] In some embodiments, components 1), 2), or optionally 3) of the genome editing system are introduced into the organism or the cell of the organism stepwise, respectively.
[0039] In some embodiments, component 2) of the genome editing system is introduced into the organism or the cell of the organism alone.
[0040] In some embodiments, the component 1) inserts the RS into the donor construct of the genomic or exogenous nucleotide sequence in the same direction or in the opposite direction.
[0041] In some embodiments, the method comprises recombining DNA of the genome of the organism or the cell of the organism.
[0042] In some embodiments, the recombination of DNA of the genome of the organism or organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome.
[0043] In some embodiments, the megaserine recombinase, the genome editing system, the polynucleotide, or the expression construct is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
[0044] In some embodiments, the organism or organism cell is from a mammal such as a human, a mouse, a rat, a monkey, a dog, a pig, a sheep, a cow, a cat; a poultry such as a chicken, a duck, a goose; a plant, preferably a crop plant, for example, wheat, rice, corn, soybean, sunflower, kiwifruit, leafy vegetable, lettuce, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava, and potato.
[0045] Effects of the invention
[0046] The present application develops a new type of megaserine recombinase and its genome editing system. The new type of megaserine recombinase exhibits beneficial enzyme activity in the genome of plants and animals for large exogenous DNA fragments, and its integration efficiency is significantly improved compared with the widely used high-efficiency recombinase Bxb1 in the prior art. The mining of this high-activity new type of megaserine recombinase is of great significance for expanding the existing DNA large fragment editing system and developing a gene editing tool library for precise manipulation of target DNA sequences. BRIEF DESCRIPTION OF DRAWINGS
[0047] FIG. 1A and FIG. 1B, Sequence alignment of QBRMYA, QBRMYS and Bxb1.
[0048] FIG. 2, Expression construct structure diagram of the new type of megaserine recombinase editing activity verification system in tobacco leaves.
[0049] FIG. 3, New type of megaserine recombinase editing activity verification results.
[0050] FIG. 4, Expression construct structure diagram of the new type of megaserine recombinase inversion activity quantitative system.
[0051] FIG. 5, New type of megaserine recombinase inversion activity quantitative experiment results.
[0052] FIG. 6, Schematic diagram of double pegRNA insertion RS method in the genome.
[0053] Figure 7, schematic diagram of expression construct for recombination enzyme integration activity verification system.
[0054] Figure 8, schematic diagram of recombination enzyme integration activity detection method.
[0055] Figure 9, results of new type large serine recombinase integration activity verification in rice cells.
[0056] Figure 10, results of new type large serine recombinase integration activity verification in human cells. DETAILED DESCRIPTION
[0057] I. Definitions
[0058] In the present application, unless otherwise indicated, the scientific and technical terms used herein have the meanings that would be generally understood by one of ordinary skill in the art. Also, the terms related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology, and laboratory procedures steps used herein are terms and procedures widely used in the corresponding fields. For example, the standard recombinant DNA and molecular cloning techniques used in the present application are well known to those skilled in the art and are described more fully in Sambrook, J., Fritsch, E.F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter "Sambrook"). Also, for better understanding of the present application, the definitions and explanations of the related terms are provided below.
[0059] As used herein, the term "and / or" encompasses all combinations of the items linked by the term, and should be read as if each combination was individually listed. For example, "A and / or B" covers "A," "B," and "A and B." For example, "A, B, and / or C" covers "A," "B," "C," "A and B," "A and C," "B and C," and "A and B and C."
[0060] The term "comprise" as used herein when describing the sequence of a protein or nucleic acid means that the protein or nucleic acid can consist of the sequence, or can have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, but still have the activity described in the present application. Furthermore, it is clear to the skilled person that the methionine encoded by the start codon at the N-terminus of a polypeptide is in some practical cases (e.g. when expressed in a particular expression system) retained, but does not materially affect the function of the polypeptide. Therefore, the specification and claims herein when describing a specific amino acid sequence of a polypeptide, although it can not comprise the methionine encoded by the start codon at the N-terminus, also encompass sequences comprising the methionine, and accordingly, the nucleotide sequence encoding it can comprise the start codon; vice versa.
[0061] "Genome" as used herein encompasses not only chromosomal DNA present in the nucleus of a cell, but also organellar DNA present in subcellular components of a cell (such as mitochondria, plastids).
[0062] "Organism" includes any organism suitable for genome editing, preferably a eukaryote. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cows, cats; poultry such as chickens, ducks, geese; plants including monocotyledons and dicotyledons, for example wheat, rice, maize, soybean, sunflower, kiwifruit, leafy greens, lettuce, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugar cane, tomato, kiwifruit, tobacco, cassava, potato, and the like.
[0063] "Foreign" means a sequence from a foreign species, or, if from the same species, which has been significantly altered from its native form by deliberate human intervention through a modification in composition and / or locus.
[0064] "Polynucleotide", "nucleic acid sequence", "nucleotide sequence", or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally can contain synthetic, non-natural, or altered nucleotide bases. Nucleotides are referred to by their single letter designation: "A" for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" for cytidine or deoxycytidine, "G" for guanosine or deoxyguanosine, "U" for uridine, "T" for deoxythymidine, and "N" for any nucleotide.
[0065] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residues are artificial chemical analogs of naturally occurring amino acids, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.
[0066] As used herein, the term "amino acid" can include natural amino acids, unnatural amino acids, amino acid analogs, and all their D and L stereoisomers. Amino acids and abbreviations and English names in the present disclosure are shown as follows:
[0067] Histidine (His, H); Serine (Ser, S); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Threonine (Thr, T); Phenylalanine (Phe, F); Aspartic acid (Asp, D); Tyrosine (Tyr, Y); Leucine (Leu, L); Isoleucine (lie, I); Arginine (Arg, R); Alanine (Ala, A); Valine (Val, V); Tryptophan (Trp, W); Methionine (Met, M); Asparagine (Asn, N); Cysteine (Cys, C); Lysine (Lys, K); Proline (Pro, P).
[0068] Sequence "identity" has the meaning commonly understood in the art and can be calculated using published techniques to determine the percent sequence identity between two nucleic acid or polypeptide molecules or regions. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule (see, e.g., Computational Molecular Biology, Lesk, A.M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D.W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A.M., and Griffin, H.G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are a number of methods for measuring identity between two polynucleotides or polypeptides, the term "identity" is well known to one of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48:1073 (1988)).
[0069] In peptides or proteins, suitable conservative amino acid substitutions are known to those of skill in the art and can generally be made without altering the biological activity of the resulting molecule. In general, those of skill in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide will not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0070] As used herein, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or a functional RNA) and / or translation of the RNA into a precursor or mature protein.
[0071] An "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (e.g., mRNA).
[0072] An "expression construct" of the application can comprise regulatory sequences and a nucleotide sequence of interest of different origin, or regulatory sequences and a nucleotide sequence of interest of the same origin arranged in a manner different from that normally existing in nature.
[0073] In the present specification, the term "safe harbor site" or "safe harbor locus (SHL)" is well known in the art. A SHL is a genomic locus at which a gene or other genetic element can be safely inserted and expressed without altering the physiological state of the cell. A SHL is further described as a genomic location at which a new gene or genetic element can be introduced without interfering with the expression or regulation of neighboring genes.
[0074] II. Dimeric serine recombinases
[0075] In one aspect, the present application provides a dimeric serine recombinase comprising an amino acid sequence set forth in any one of SEQ ID NOs: 1-5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to an amino acid sequence set forth in any one of SEQ ID NOs: 1-5.
[0076] As used herein, the term "recombinase" and the like refers to a site-specific enzyme that mediates DNA recombination between sequences recognized by the recombinase. During the reaction, an attack on the DNA phosphate backbone is launched by a tyrosine or serine located in the catalytic active center of the recombinase, resulting in a DNA strand break. Recombinases can be divided into two different families: serine recombinases (e.g., resolvases and invertases) and tyrosine recombinases (e.g., integrases). Some examples of serine recombinases include, but are not limited to, Bxbl, Ilin, Gin, Tn3, beta-six, CinH, ParA, gamma delta, Bxbl, TP901, TGl, Rl, R2, R3, R4, R5, MRll, Al 18, U153, and gp29.
[0077] The term "recombination" refers to the excision, integration, inversion, or exchange (e.g., translocation) of a DNA fragment between sequences recognized by a recombinase enzyme. In some embodiments, recombinase enzyme recombination activity includes integration activity or inversion activity of a recombinase enzyme.
[0078] As used herein, the terms "large serine recombinase" and "LSR" are used interchangeably and refer to a class of serine recombinases. Serine recombinases are integrases carried by bacteriophages that integrate large fragments of DNA sequences into the genome of a bacterium without the need for any cellular cofactors by recognizing specific sequences on the bacterial genome and the phage DNA fragment. Many serine recombinases function to resolve transposition intermediates or to regulate gene expression by inverting regulatory sequences. During the catalytic reaction of a serine recombinase, a complex is formed from DNA and the recombinase, then a serine in the recombinase domain attacks the DNA phosphate backbone causing the DNA to form a nick with a 3'-OH double-stranded break end and a 5'-phosphoserine covalently linked to the DNA, inversion of the complex occurs, the double-stranded DNA is re-ligated, and a recombination product is formed. Most serine recombinases have a 150 amino acid catalytic domain at their amino terminus, followed by a small HTH (Helix-Turn-Helix)-DNA binding domain. Large serine recombinases have a similar amino-terminal catalytic domain, but have a larger carboxy-terminal region that varies in size from 300 residues in the bacteriophage R4 and A118 recombinases to 550 residues in the TnpX transposase, and are referred to as large serine recombinases because of the larger carboxy-terminal region of this class of serine recombinases.
[0079] In some embodiments, the large serine recombinase corresponds to a recombinase recognition site sequence (RS) selected from the group consisting of an attP site, an attB site.
[0080] As used herein, the terms "recombinase recognition site sequence," "recombination site," "recognition site," "RS site," "RS sequence," and "RS" generally refer to a nucleic acid (e.g., DNA) sequence that is recognized (e.g., can be bound by) a recombinase polypeptide, which can be naturally occurring, or engineered or optimized.
[0081] As used herein, the terms "attP," "attB" refer to bacteriophage attachment sites (attP) and bacterial attachment sites (attB), respectively, and are collectively referred to herein as "recombinase recognition site sequences" or "recombination sites."
[0082] In some cases, the recombinase recognition site sequence comprises an attB site (SEQ ID NOs: 13-19), an attP site (SEQ ID NOs: 6-12), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the nucleic acid sequences set forth in SEQ ID NOs: 6-19.
[0083] In some embodiments, the attP site comprises a nucleotide sequence selected from the group consisting of the nucleotide sequences set forth in SEQ ID NOs: 6-12, and the attB site comprises a nucleotide sequence selected from the group consisting of the nucleotide sequences set forth in SEQ ID NOs: 13-19.
[0084] In some embodiments, the megaserrnase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the attP site comprises an amino acid sequence selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 7-8, and the attB site comprises an amino acid sequence selected from the group consisting of the amino acid sequences set forth in SEQ ID NOs: 14-15. In some specific embodiments, the attP site comprises the amino acid sequence set forth in SEQ ID NO: 7, and the attB site comprises the amino acid sequence set forth in SEQ ID NO: 14. In some specific embodiments, the attP site comprises the amino acid sequence set forth in SEQ ID NO: 8, and the attB site comprises the amino acid sequence set forth in SEQ ID NO: 15.
[0085] In some embodiments, the megaserrnase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 1, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the attP site comprises the amino acid sequence set forth in SEQ ID NO: 6, and the attB site comprises an amino acid sequence selected from the group consisting of the amino acid sequences set forth in SEQ ID NO: 13.
[0086] In some embodiments, the megasynthase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 3, or an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identical to the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the attP site comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 9-10, and the attB site comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 16-17. In some specific embodiments, the attP site comprises an amino acid sequence set forth in SEQ ID NO: 9, and the attB site comprises an amino acid sequence set forth in SEQ ID NO: 16. In some specific embodiments, the attP site comprises an amino acid sequence set forth in SEQ ID NO: 10, and the attB site comprises an amino acid sequence set forth in SEQ ID NO: 17.
[0087] In some embodiments, the megasynthase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 4, or an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identical to the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, the attP site comprises an amino acid sequence set forth in SEQ ID NO: 11, and the attB site comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 18.
[0088] In some embodiments, the megasynthase recombinase comprises an amino acid sequence set forth in SEQ ID NO: 5, or an amino acid sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identical to the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the attP site comprises an amino acid sequence set forth in SEQ ID NO: 12, and the attB site comprises an amino acid sequence selected from the group consisting of SEQ ID NO: 19.
[0089] III. Genome editing system
[0090] In one aspect, the present application provides a genome editing system, wherein the genome editing system comprises:
[0091] 1) a recombinase recognition site sequence (RS) comprising an attP or attB site (hereinafter also referred to as component 1);
[0092] 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase (hereinafter also referred to as component 2), the recombinase comprising a megasynthase recombinase comprising an amino acid sequence as set forth in any one of SEQ ID NOs: 1-5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to an amino acid sequence as set forth in any one of SEQ ID NOs: 1-5.
[0093] In some embodiments, the recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.
[0094] In some embodiments, the one or more recombinase recognition site sequences (RS) are inserted into the genome at a desired location by a recombinase recognition site integration unit.
[0095] A person skilled in the art can select a suitable gene editing tool to insert the RS into the genome at a desired location.
[0096] In some embodiments, the recombinase recognition site integration unit comprises: a sequence-specific DNA cleavage protein and a RS integration template, and / or an expression construct encoding the sequence-specific DNA cleavage protein and the RS integration template. In some embodiments, the recombinase recognition site integration unit specifically recognizes a desired location in the genome and cleaves DNA, and then inserts a recombinase recognition site sequence into the desired location in the genome under the mediation of the RS integration template through the endogenous or exogenous repair mechanism of the cell.
[0097] In some embodiments, the RS integration template is a donor DNA template or a reverse transcription RNA template, wherein the donor DNA template comprises a desired recombinase recognition site sequence or a complementary sequence thereof, and the reverse transcription RNA template comprises a transcription sequence of a desired recombinase recognition site sequence or a complementary sequence thereof.
[0098] In some embodiments, the RS integration template comprises a DNA synthetic template encoding a RS that can be recognized by a recombinase.
[0099] In some embodiments, the sequence-specific DNA cleavage protein is selected from one or more of a meganuclease (MGN), a zinc finger nuclease (ZFN), a transcription activator-like effector nuclease (TALEN), and a CRISPR effector protein.
[0100] In some embodiments, the recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or a functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
[0101] In some embodiments, the CRISPR effector protein or a functional variant thereof comprises a CRISPR nuclease and functional variants thereof.
[0102] As used herein, the term "CRISPR effector protein" generally refers to a nuclease or a functional variant thereof that is present in a naturally occurring CRISPR system. The term encompasses any effector protein based on a CRISPR system that is capable of achieving sequence-specific targeting within a cell.
[0103] As used herein, the term "CRISPR nuclease" can be derived from a Cas9 nuclease, including a Cas9 nuclease or a functional variant thereof. The Cas9 nuclease can be a Cas9 nuclease from different species, such as spCas9 from S. pyogenes or SaCas9 derived from S. aureus. "Cas9 nuclease" and "Cas9" are used interchangeably herein to refer to an RNA-guided nuclease comprising a Cas9 protein or a fragment thereof (e.g., a protein comprising the active DNA cleavage domain of Cas9 and / or the gRNA binding domain of Cas9). Cas9 is a component of the CRISPR / Cas (clustered regularly interspaced short palindromic repeat and associated systems) genome editing system that can target and cleave a DNA target sequence under the guidance of a guide RNA to form a DNA double-strand break (DSB).
[0104] A "CRISPR nuclease" can also be derived from a Cpf1 nuclease, including a Cpf1 nuclease or a functional variant thereof. The Cpf1 nuclease can be a Cpf1 nuclease from different species, such as Cpf1 nuclease from Francisella novicida U112, Acidaminococcus sp. BV3L6, and Lachnospiraceae bacterium ND2006.
[0105] As used herein, a "functional variant" with respect to a CRISPR nuclease means that it retains at least the ability to target a sequence specifically mediated by a guide RNA. Preferably, the functional variant is a nuclease-inactivated variant, i.e., it lacks the double-stranded nucleic acid cleavage activity. However, a CRISPR nuclease that lacks the double-stranded nucleic acid cleavage activity also encompasses a nickase, which forms a nick in a double-stranded nucleic acid molecule, but does not completely cleave the double-stranded nucleic acid.
[0106] In some preferred embodiments of the present application, the CRISPR effector protein of the present application has nickase activity (or referred to as CRISPR nickase). In some embodiments, the functional variant recognizes a different PAM (protospacer adjacent motif) sequence relative to the wild-type nuclease.
[0107] As used herein, the term "CRISPR nickase" refers to a nuclease-inactivated Cas9 that can be derived from Cas9 of different species, e.g., derived from S. pyogenes Cas9 (SpCas9), or derived from S. aureus Cas9 (SaCas9). Simultaneous mutation of the HNH nuclease subdomain and the RuvC subdomain of Cas9 (e.g., comprising mutations D10A and H840A) renders the nuclease activity of Cas9 inactive, becoming a nuclease-inactivated Cas9 (dCas9). Mutation inactivation of one of the subdomains can render Cas9 to have nickase activity, i.e., obtain Cas9 nickase (nCas9), e.g., nCas9 with only mutation D10A (or referred to as Cas9-D10A), and nCas9 with only mutation H840A (or referred to as Cas9-H840A). In some embodiments, the CRISPR nickase can be derived from Cas12 of different species, e.g., Cas12a, Cas12b, Cas12i. In some embodiments, the CRISPR nickase is TraC protein based on the TraC effector protein of the intermediate transposon and CRISPR-Cas12 intermediate TraC, which only retains the DNA single-strand cleavage activity (wherein the TraC nuclease is referred to the disclosure in PCT / CN2023 / 097783 (Publication No. WO / 2023 / 232109), which is incorporated herein by reference).
[0108] In some embodiments, the functional variant of the CRISPR nuclease is a CRISPR nickase, e.g., Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, and / or TraC nickase (see patents CN202310646033.2, CN202411698901.2).
[0109] In some embodiments, the recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the reverse transcriptase. In these embodiments, the recombinase recognition site integration unit can be based on prime editors and iterations thereof, which have shown insertion, deletion, and base conversion with modest editing efficiency (as disclosed in WO 2020 / 191234 Al, incorporated herein by reference), as well as plant prime editors (PPEs), which are disclosed in Lin, Qiupeng et al. “Prime genome editing in rice and wheat.” Nature biotechnology vol. 38, 5 (2020): 582-585. doi: 10.1038 / s41587-020-0455-x; Lin, Qiupeng et al. “High-efficiency prime editing with optimized, paired pegRNAs in plants.” Nature biotechnology vol. 39, 8 (2021): 923-927. doi: 10.1038 / s41587-021-00868-w; and the like; dual-ePPEs, which are disclosed in Sun, Chao et al. “Precise integration of large DNA sequences in plant genomes using PrimeRoot editors.” Nature biotechnology vol. 42, 2 (2024): 316-327. doi: 10.1038 / s41587-023-01769-w; and the like; TwinPEs, which are disclosed in Anzalone, Andrew V et al. “Programmable deletion, replacement, integration and inversion of large DNA sequences with twin prime editing.” Nature biotechnology vol. 40, 5 (2022): 731-740. doi: 10.1038 / s41587-021-01133-w; and the like, incorporated herein by reference.
[0110] In some embodiments, the recombinase recognition site integration unit further comprises all the technologies in the prior art that can insert specific DNA sequences into DNA fragments at a specific site, including but not limited to: technologies based on CRISPR and transposon systems, such as homologous directed repair (HDR); INTEGRATE system, which is disclosed in Strecker, Jonathan et al. “RNA-guided DNA insertion with CRISPR-associated transposases.” Science (New York, N.Y.) vol. 365, 6448 (2019): 48-53. doi:10.1126 / science.aax9181, and the like; CRISPR-associated transposase (CAST) technology, which is disclosed in Lampe, George D et al. “Targeted DNA integration in human cells without double-strand breaks using CRISPR RNA-guided transposases.” bioRxiv: the preprint server for biology 2023.03.17.533036. 18 Mar. 2023, doi:10.1101 / 2023.03.17.533036. Preprint. and the like, which are incorporated herein by reference.
[0111] As used herein, the term “reverse transcriptase” is an enzyme that directs the synthesis of deoxyribonucleotide triphosphates into complementary DNA (cDNA) using RNA as a template.
[0112] In some embodiments, wherein the reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii mature enzyme RT (Marathon RT).
[0113] In some embodiments, the guide RNA comprises a scaffold sequence, an RS integration template sequence, and a primer binding sequence, the RS integration template sequence being linked to the primer binding sequence.
[0114] As used herein, the terms "guide RNA" and "gRNA" are used interchangeably to refer to an RNA molecule capable of forming a complex with a CRISPR effector protein and capable of targeting the complex to a target sequence due to a certain identity with the target sequence. The guide RNA targets the target sequence through base pairing between the complementary strands of the target sequence. The guide RNA functions in conjunction with the CRISPR effector protein to collectively be referred to as a "scaffold sequence", for example, the scaffold sequence employed by a Cas9 nuclease or a functional variant thereof is typically composed of a crRNA and a tracrRNA that form a complex in part. A "primer binding sequence" refers to a sequence of at least 10, at least 15, or at least 20 contiguous nucleotides of the guide RNA that is complementary to the target sequence. In some embodiments, the recombinase recognition site sequence is linked to the primer binding sequence. In some embodiments, the "guide RNA" is a single guide RNA (sgRNA), a guide RNA for prime editor (pegRNA). Designing a suitable guide RNA based on the CRISPR nuclease used and the target sequence to be edited is within the ability of one skilled in the art.
[0115] In some embodiments, the CRISPR nickase of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with and targets the CRISPR nickase to a desired location in the genome, wherein the CRISPR nickase makes a nick in a strand of the genome, and the reverse transcriptase incorporates a RS integration template sequence in the guide RNA into the nicked site, thereby inserting at least one RS recognizable by the recombinase at the desired location in the genome.
[0116] In some embodiments, the recombination site (RS) is selected from an attP site, an attB site. In some embodiments, the attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 6-12, and the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 13-19.
[0117] The original function of integrase is to make recombination between a short sequence of DNA between a bacteriophage attachment site (attP) and a bacterial attachment site (attB), so it is applied in the process of gene editing. The recognition site of the target DNA in the editing process is called "attB", and the recognition site of the DNA sequence to be edited is called "attP".
[0118] In some cases, the recognition site comprises an attB site (SEQ ID NOs: 13-19), an attP site (SEQ ID NOs: 6-12), or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to the nucleotide sequence set forth in SEQ ID NOs: 6-19. The RS can be inserted into the genome or a fragment thereof of a cell using a nuclease, a gRNA, and / or an integrase, wherein the RS is carried on a guide RNA. The guide RNA can target any site known in the art. The complementary RS can be operably linked to a gene or nucleic acid sequence of interest of an exogenous DNA or RNA. In some embodiments, one RS is added to the target genome. In some embodiments, more than one RS is added to the target genome.
[0119] In some embodiments, the recombinase recognition site integration unit is not covalently linked to the recombinase.
[0120] In some embodiments, the recombinase recognition site integration unit is covalently linked to the recombinase.
[0121] In some embodiments, the genome editing system further comprises a nucleic acid molecule comprising:
[0122] 3) a donor of an exogenous nucleotide sequence to be inserted into the genome (hereinafter also referred to simply as component 3).
[0123] In some embodiments, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer, such as at least 50 bp, at least 100 bp, at least 300 bp, at least 500 bp, at least 1 kb, at least 1.5 kb, at least 2 kb, at least 3 kb, at least 4 kb, at least 5 kb, at least 6 kb, at least 7 kb, at least 8 kb, at least 9 kb, at least 10 kb, at least 20 kb, at least 50 kb.
[0124] In some embodiments, the exogenous nucleotide sequence further comprises one or more RSs that can be recognized by the recombinase.
[0125] In some embodiments, the exogenous nucleotide sequence is flanked by one or two RSs.
[0126] IV. Fusion proteins
[0127] In another aspect of the present application, a fusion / chimeric protein comprising the genome editing system of any one of the above or a portion thereof is provided, wherein the sequence-specific DNA cleavage protein, the reverse transcriptase, and the recombinase are directly linked or linked via a linker.
[0128] In some embodiments, the sequence-specific DNA cleavage protein, e.g., CRISPR nickase, is linked at the C-terminus to a reverse transcriptase, which is linked via a linker to a recombinase.
[0129] As used herein, the term "linker" can be a long 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, non-functional amino acid sequence without secondary structure above. For example, the linker can be a flexible linker.
[0130] V. Polynucleotides
[0131] In one aspect, the present application provides a polynucleotide comprising a nucleotide sequence encoding the aforementioned large serine recombinase, the aforementioned genome editing system, or the aforementioned fusion protein.
[0132] The polynucleotide of the present application can be in the form of DNA or RNA. The DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. The DNA can be single-stranded or double-stranded. The DNA can be a coding strand or a non-coding strand. The RNA form includes mRNA.
[0133] The polynucleotide encoding the large serine recombinase of the present application includes: only a coding sequence encoding the large serine recombinase; a coding sequence of the large serine recombinase and various additional coding sequences; a coding sequence of the large serine recombinase (and optional additional coding sequences) and non-coding sequences.
[0134] VI. Expression Constructs
[0135] In one aspect, the present application provides an expression construct comprising the aforementioned polynucleotide.
[0136] In some embodiments, the expression construct is selected from the group consisting of viral, bacterial, yeast, plant, mammalian cell expression constructs.
[0137] VII. Cells
[0138] In one aspect, the present application provides a cell comprising the aforementioned large serine recombinase, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned polynucleotide, or the aforementioned expression construct.
[0139] VIII. Kits
[0140] In one aspect, the present application provides a kit comprising the aforementioned large serine recombinase, the aforementioned polynucleotide, the aforementioned genome editing system, the aforementioned fusion protein, the aforementioned expression construct, or the aforementioned cell.
[0141] IX. Methods of genome editing
[0142] In another aspect of the present application, a method of genome editing in an organism or an organism cell is provided, selected from: i) introducing into the organism or the organism cell the aforementioned meganuclease and a donor comprising an exogenous nucleotide sequence to be inserted into the genome, wherein the target genome of the organism or the organism cell contains a sequence of a recombinase recognition site (RS) corresponding to the meganuclease; or ii) introducing into the organism or the organism cell the aforementioned genome editing system.
[0143] In some embodiments, the method of genome editing can result in one or more nucleotide substitutions, or one or more insertions of nucleotides. In some embodiments, the substitution of nucleotides is at least 2 bp, at least 10 bp, at least 100 bp, at least 1 kbp, at least 10 kbp, at least 20 kbp, at least 50 kbp in length.
[0144] In some embodiments, components 1), 2) and optionally 3) of the genome editing system are introduced into the organism or the organism cell simultaneously.
[0145] In some embodiments, components 1), 2) or optionally 3) of the genome editing system are introduced into the organism or the organism cell stepwise, respectively.
[0146] In some embodiments, component 2) of the genome editing system is introduced into the organism or the organism cell alone.
[0147] A person skilled in the art can select the RS inserted into the genome and / or the exogenous nucleotide sequence, and the direction of insertion into the genome or the exogenous nucleotide sequence, according to the purpose of DNA recombination, such as deletion, inversion, integration, etc.
[0148] In some embodiments, component 1) inserts the RS into the genome or the donor construct of the exogenous nucleotide sequence in the same direction or in the opposite direction.
[0149] In some embodiments, the method comprises recombining DNA of the genome of the organism or the organism cell.
[0150] In some specific embodiments, the recombining DNA of the genome of the organism or the organism cell comprises deleting DNA, inverting DNA in the genome, and / or integrating exogenous DNA into the genome.
[0151] In the method of the present application, the megaserrnase recombinase or genome editing system can be introduced into the cell by various methods well known to those skilled in the art. In some embodiments, the method of introducing the megaserrnase recombinase or genome editing system into the organism or biological cell is selected from the group consisting of: calcium phosphate transfection, protoplast fusion, electroporation, lipofection, microinjection, viral infection (e.g., baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus, or other viruses), biolistics, N-acetylgalactosamine (GalNAc)-mediated, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
[0152] In some embodiments, the organism or organism cell is from a mammal such as a human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; plant, preferably a crop plant, e.g., wheat, rice, corn, soybean, sunflower, kiwifruit, leafy greens, lettuce, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomato, kiwifruit, tobacco, cassava, potato.
[0153] In the present application, the target nucleic acid region to be edited can be located anywhere in the genome, e.g., in a safe harbor of the genome, within a functional gene such as a protein-coding gene, or e.g., in a gene expression regulatory region such as a promoter region or an enhancer region, thereby effecting a modification of the function of the gene or a modification of the gene expression. In some embodiments, the desired nucleotide sequence substitution results in a desired modification of the function of the gene or a modification of the gene expression.
[0154] Examples
[0155] Embodiments of the present application will be described in detail with reference to the following examples, but the present application is not limited to the following examples. The following examples are merely illustrative of the present application and should not be viewed as limiting the scope of the present application. Unless otherwise indicated, the conditions in the examples are conventional or those recommended by the manufacturer. Unless otherwise indicated, the reagents or instruments used in the examples are commercially available conventional products.
[0156] Example 1: Mining of novel megaserrnase recombinases
[0157] The inventors apply the homologous characteristics in evolutionary taxonomy to the characterization of protein three-dimensional structure through structural analysis and observation of large serine recombinases, and propose a prediction method based on protein structure classification (prediction method, see CN202411131015.1). First, new large serine recombinases are obtained through structural prediction, and then filtered. In the prediction of attP and attB recognition sites, based on the protein electron microscope structure of the interaction between large serine recombinases and DNA, the region of large serine recombinases interacting with DNA is divided, and the amino acid arrangement characteristics of the interaction region are searched in the existing large serine recombinase database to obtain the attP and attB recognition sites of the new large serine recombinase.
[0158] Finally, 8178 new large serine recombinases and the corresponding attP and attB recognition sites of each large serine recombinase are identified. The blast results show that the amino acid sequence similarity of the new recombinase protein and the reported large serine recombinase Bxb1 (SEQ ID NO: 6) is relatively low. For example, QBRMYA and QBRMYS, the amino acid sequence identity with Bxb1 is 47.48% and 58.23% respectively (Fig. 1A and Fig. 1B).
[0159] Example 2: Verification of editing activity of new large serine recombinases
[0160] In this embodiment, the editing activity of the novel large serine recombinase is verified by using a recombinase editing activity verification system. Specifically, the recombinase editing activity verification system in tobacco leaves includes three components: as shown in FIG. 2, the expression construct Recombinase (abbreviated as R) is used to express the recombinase (LSR) whose editing activity is to be verified, and the expression is initiated by the cauliflower mosaic virus CaMV 35S promoter (P-35S) and terminated by the Agrobacterium tumefaciens terminator NOS (T-NOS); the expression construct Rep conditionally expresses the replication protein Rep of the geminivirus, and the expression is initiated by the CaMV 35S promoter (P-35S), followed by two recombinase candidate recognition site RS sequences (attB and attP) in the same direction, a terminator sequence (T-HSP and T-E9) is placed between the two RSs, and the expression is terminated by the Agrobacterium tumefaciens terminator NOS (T-NOS); the expression construct GFP expresses the reporter gene green fluorescent protein GFP. Only when the recombinase whose editing activity is to be verified in the expression construct R is expressed, the two RS sites in the expression construct Rep are recognized, the terminator loop between the two RS sites is removed, and the expression of the Rep protein in the expression construct Rep is initiated. Finally, the Rep protein expressed in the expression construct Rep recognizes the long intergenic region (LTR) in the expression construct GFP, drives the rolling circle replication of the expression construct GFP, turns on the expression of the reporter gene GFP, and determines whether the recombinase whose editing activity is to be verified can recognize the candidate RS site and has activity by detecting the GFP signal. The construction of the recombinase editing activity verification system is shown in FIG. 2.
[0161] The novel megasynthase recombinases screened in Example 1 were subjected to editing activity verification by the above-mentioned recombinase editing activity verification system, and five novel megasynthase recombinases with recombinase activity, QBRMYA, QBRMYS, QBRMyi, QBRStrl and QBRStr2, were obtained, and the results are shown in Figure 3. The RS sequences on the expression construct Rep in combinations 1, 2, 7 and 8 were Bxbl-attP (SEQ ID NO: 21) and Bxbl-attB (SEQ ID NO: 22), respectively; the RS sequences on the expression construct Rep in combinations 3 and 4 were QBRMYA-attP (SEQ ID NO: 6) and QBRMYA-attB (SEQ ID NO: 13), respectively; the RS sequences on the expression construct Rep in combinations 5 and 6 were QBRMYS-attPl (SEQ ID NO: 7) and QBRMYS-attBl (SEQ ID NO: 14), respectively; the RS sequences on the expression construct Rep in combinations 9 and 10 were QBRMYS-attP2 (SEQ ID NO: 8) and QBRMYS-attB2 (SEQ ID NO: 15), respectively. The RS sequences on the expression construct Rep in combination 12 were QBRMyi-attP (SEQ ID NO: 11) and QBRMyi-attB (SEQ ID NO: 18), respectively; the RS sequences on the expression construct Rep in combination 13 were QBRStr2-attP (SEQ ID NO: 12) and QBRStr2-attB (SEQ ID NO: 19), respectively; the RS sequences on the expression construct Rep in combination 14 were QBRStrl-attPl (SEQ ID NO: 9) and QBRStrl-attBl (SEQ ID NO: 16), respectively. Among them, combinations 2, 4, 6, 8, 10 served as corresponding negative controls and did not express recombinases.
[0162] The results are shown in Figure 3. In combination 1, Bxb1 as a positive control, green fluorescence signal can be detected. In combination 2, Bxb1 is not expressed, and no fluorescence signal is detected, indicating that the recombinase editing activity verification system can accurately verify the recombinase activity. Compared with the negative control combination 4, combination 3 can detect green fluorescence signal, indicating that QBRMYA can recognize QBRMYA-attP and QBRMYA-attB sequences, and has recombinase activity. Compared with the negative control combination 6, combination 5 can detect green fluorescence signal, indicating that QBRMYS can recognize QBRMYS-attP1 and QBRMYS-attB1, and has recombinase activity. Compared with the negative control combination 10, combination 9 can detect green fluorescence signal, indicating that QBRMYS can recognize QBRMYS-attP2 and QBRMYS-attB2, and has recombinase activity. The green fluorescence signal of combinations 12, 13 and 14 is weak, but compared with the control (no fluorescence), the fluorescence signal is still obvious, so QBRMyi, QBRStr1 and QBRStr2 have recombinase recognition activity.
[0163] In summary, the present application verifies that QBRMYA, QBRMYS, QBRMyi, QBRStr1 and QBRStr2 are five new recombinase active megaserine recombinases from the recombinases predicted in Example 1.
[0164] Example 3: Activity quantitative experiment of new megaserine recombinase
[0165] The recombinase flip-flop activity quantitative system was used to verify the activity of the new recombinase in this embodiment. The recombinase flip-flop activity quantitative system comprises two components. The R plasmid expresses the recombinase to be verified for flip-flop activity, and the expression of the recombinase is initiated by the cauliflower virus CaMV 35S promoter (P-CaMV 35S) and terminated by the Agrobacterium tumefaciens terminator NOS (T-NOS). The RS plasmid conditionally expresses the firefly luciferase, and the expression is initiated by the cauliflower virus CaMV 35S promoter (P-CaMV 35S). The attB site and the attP site are placed in opposite directions. The reverse open reading frame (ORF) sequence of the firefly luciferase is inserted between the attB site and the attP site. A segment from the 3' untranslated region sequence of Agrobacterium tumefaciens terminates the expression of the firefly luciferase. The cassava vein mosaic virus CSVMV promoter (P-CSVMV) initiates the expression of the Renilla luciferase, and the CaMV 35S terminator terminates the expression of the Renilla luciferase. Only when the recombinase to be verified for flip-flop activity is expressed and recognizes the attB and attP sequences, the ORF sequence of the firefly luciferase between the att sites is flipped, and the firefly luciferase can be expressed. Finally, the activity of the firefly luciferase is detected by using an enzyme marker, representing the activity of the recombinase. The activity of the Renilla luciferase constantly expressed on the RS plasmid is used as an internal reference to correct the expression of the firefly luciferase. The schematic diagram of the expression construct of the recombinase flip-flop activity quantitative system is shown in FIG. 4.
[0166] The luciferase detection kit used in this quantitative verification experiment is Transdetect Double-Luciferase Reporter Assay Kit (Transgen, FR201). The R plasmid and the RS plasmid are transformed into Agrobacterium, and tobacco is infected simultaneously as the experimental group. Only the RS plasmid infects tobacco as the negative control of each recombinase. Bxb1 is used as a positive control, and three biological replicates are set for each experiment.
[0167] The experimental results are shown in Figure 5, and the annotations in the abscissa take Bxb1 as an example. Bxb1+att represents the expression of Bxb1 recombinase, att_Bxb1 represents the expression of the negative control without Bxb1 recombinase, and so on. The firefly luciferase activity of Bxb1 and its RS combination is 0.3 times that of Bxb1, the turnover activity of QBRStr1 and its att1 (QBRStr1-attB1+QBRStr1-attP1) and att2 (QBRStr1-attB2+QBRStr1-attP2) is 0.3 times that of Bxb1, the turnover activity of QBRMyi and its att (QBRMyi-attP+QBRMyi-attB) is 1.2 times that of Bxb1, the turnover activity of QBRStr2 and its att (QBRStr2-attP+QBRStr2-attB) is 2.2 times that of Bxb1, the activity of QBRMYA and its att (QBRMYA-attP+QBRMYA-attB) is 4.9 times that of Bxb1, the turnover activity of QBRMYS and its att1 (QBRMYS-attP1+QBRMYS-attB1) combination is 5.13 times that of Bxb1, the turnover activity of QBRMYS and its att2 (QBRMYS-attP2+QBRMYS-attB2) combination is 7.44 times that of Bxb1, and the negative control att_Bxb1 has almost no firefly luciferase activity.
[0168] As can be seen, the five new large serine recombinases QBRStr1, QBRMyi, QBRStr2, QBRMYA and QBRMYS have significantly higher recombinase turnover activity, especially QBRStr2, QBRMYA and QBRMYS have significantly improved recombination efficiency compared with Bxb1.
[0169] Example 4: Large fragment DNA integration effect of new large serine recombinase in plant and animal genomes
[0170] A recognition site of the recombinase is inserted into a safe harbor site of the genome using a guide editing tool (Prime Editor, PE) and double pegRNA, and the method principle of inserting the RS is shown in Figure 6, and another recognition site of the recombinase is placed near the donor large fragment DNA. When the recombinase is expressed, the recombinase recognizes the site on the genome and the site on the donor, respectively, forms a tetramer, activates the recombinase activity, and integrates the donor large fragment DNA into the genome.
[0171] The endogenous integration activity of the novel large serine recombinase was verified in rice and human cells, respectively, and the expression construct of the recombinase is shown in FIG. 7 (a in FIG. 7 is for rice cells, and b in FIG. 7 is for human cells). The selected site in rice is a safe harbor site GSH1 (kitaake, Chr1:7660637-7661671) on the genome, and a 5.7 kb donor vector is to be integrated, which carries the recognition site of the recombinase, and the PE tool used for inserting RS in the rice genome is ePPE; the TRAC site is selected for human cells HEK293T, and the integration activity of the recombinase on a 5.6 kb donor vector is verified, and the PE tool used is PE2. The names and sequences of the vectors and double pegRNAs used in the above experiments are shown in Table 1.
[0172] Table 1
[0173] The activity detection method of the recombinase for integrating the donor large fragment DNA into the genome is shown in FIG. 8. In order to facilitate the detection of integration efficiency, a R primer on the genome is pre-set on the donor large fragment DNA vector. Therefore, when a pair of F and R primers are used to amplify the genomic DNA, three sequences can be amplified, one is the wild type sequence, the second is the sequence with attB insertion after PE work, and the third is the sequence with donor fragment after recombinase work. By using high-throughput second-generation sequencing, the number of reads of the three amplification products is counted, respectively, wherein the reads number of the second plus the third / the total reads number = the efficiency of PE, and the reads number of the third / (the reads number of the second plus the third) = the recombination efficiency of the recombinase. The specific binding sequences of the F and R primers used for amplification in rice cells are SEQ ID NO: 46 and SEQ ID NO: 47, respectively; and the specific binding sequences of the F and R primers used for amplification in human cells are SEQ ID NO: 23 and SEQ ID NO: 24, respectively.
[0174] The integration efficiency of Cre, Bxbl and QBRMYS on a 5.7 kb size donor large fragment DNA at the safe harbor site GSH1 on the rice genome was detected, respectively, and PE represents the efficiency of inserting the recognition site of the recombinase Cre-Lox66 or Bxbl-attB at the GSH1 site on the genome, and IN represents the integration efficiency of the recombinase on the donor large fragment DNA. The results are shown in FIG. 9, which shows that the integration efficiency of the novel recombinase QBRMYS is significantly higher than that of Bxbl and Cre, which is 2.1 times that of Bxbl and 19.5 times that of Cre.
[0175] The large fragment DNA insertion experiment of the new megaserine recombinase was performed at the safe harbor TRAC site of the genome of HEK293T cells, PE represents the efficiency of inserting the recognition site Bxb1-attB or QBRMYS-attB of the recombinase at the genomic TRAC site, IN represents the integration efficiency of the donor large fragment DNA by the recombinase. The results are shown in Figure 10, which shows that the screened new megaserine recombinase QBRMYS with recombination activity can accurately insert the donor large fragment DNA at the expected site, and the integration efficiency is significantly higher than that of Bxb1 (represented as WT_Bxb1 in Figure 10) and eeBxb1 (a Bxb1 variant with increased activity reported in the prior art, represented as ee_Bxb1 in Figure 10), which is 2.9 times that of Bxb1 and 1.4 times that of eeBxb1. It can be seen that the new megaserine recombinase QBRMYS has very good integration ability for exogenous large fragment DNA in plant and animal genomes, and the integration efficiency is significantly improved compared with the recombinases widely used in the prior art and with high efficiency. The mining of the high-activity new megaserine recombinase is of great significance for expanding the existing DNA large fragment editing system and developing a gene editing tool library for precise manipulation of target DNA sequences.
[0176] It should be noted that although the technical solutions of the present application are described with specific examples, those skilled in the art can understand that the present application should not be limited thereto.
[0177] The above has described various embodiments of the present application, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles, practical applications, or technical improvements in the market of the embodiments, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
[0178] New megaserine recombinase and reference sequence of recognition site in embodiments
[0179] >SEQ ID NO:20 Bxb1
[0180] >SEQ ID NO:21 Bxb1-attP
[0181] >SEQ ID NO:22 Bxb1-attB
[0182] >SEQ ID NO: 23 Specific binding sequence of F primer used for detection of efficiency in human cells
[0183] >SEQ ID NO: 24 R primer binding sequence inserted on donor vector used in human cells
[0184] >SEQ ID NO: 25 eeBxb1
[0185] >SEQ ID NO: 26 Cre amino acid sequence
[0186] >SEQ ID NO: 27 Lox66
[0187] >SEQ ID NO: 28 Lox71
[0188] >SEQ ID NO: 29 ePPE vector, for rice
[0189] >SEQ ID NO: 30 epegRNA1, Cre-Lox66 inserted in rice genome
[0190] >SEQ ID NO: 31 epegRNA2, Cre-Lox66 inserted in rice genome
[0191] >SEQ ID NO: 32 epegRNA1, Bxb1-attB inserted in rice genome
[0192] >SEQ ID NO: 33 epegRNA2, Bxb1-attB inserted in rice genome
[0193] >SEQ ID NO: 34 epegRNA1, QBRMYS-attB inserted in rice genome
[0194] >SEQ ID NO: 35 epegRNA2, QBRMYS-attB inserted in rice genome
[0195] >SEQ ID NO:36 5.7kb donor, donor with Cre-Lox71 for transformation of rice
[0196] >SEQ ID NO:37 5.7kb donor, donor with Bxbl-attP for transformation of rice
[0197] >SEQ ID NO:38 5.7kb donor, donor with QBRMYS-attP for transformation of rice
[0198] >SEQ ID NO:39 PE2 vector, for human cells
[0199] >SEQ ID NO:40 pegRNA1, for insertion of Bxbl-attB into human genome
[0200] >SEQ ID NO:41 pegRNA2, for insertion of Bxbl-attB into human genome
[0201] >SEQ ID NO:42 pegRNA1, for insertion of QBRMYS-attB into human genome
[0202] >SEQ ID NO:43 pegRNA2, for insertion of QBRMYS-attB into human genome
[0203] >SEQ ID NO:44 5.6kb donor, donor with Bxbl-attP for transformation of human cells
[0204] >SEQ ID NO:45 5.6kb donor, donor with QBRMYS-attP for transformation of human cells
[0205] >SEQ ID NO:46 Specific binding sequence for F primer used to detect efficiency in rice
[0206] >SEQ ID NO: 47 R primer binding sequence inserted on the donor vector used in rice
[0207] >SEQ ID NO: 48 CaMV 35S
[0208] >SEQ ID NO: 49 T-NOS
[0209] >SEQ ID NO: 50 Firefly luciferase
[0210] >SEQ ID NO: 51 T-At ORF1 3'UTR
[0211] >SEQ ID NO: 52 P-CSVMV
[0212] >SEQ ID NO: 53 Renilla luciferase
[0213] >SEQ ID NO: 54 T-CaMV35S
Claims
1. A megasynthase recombinant enzyme, wherein, The megaserine recombinase comprises an amino acid sequence as set forth in any one of SEQ ID NOs: 2, 1, and 3-5, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to an amino acid sequence as set forth in any one of SEQ ID NOs: 2, 1, and 3-5.
2. The megasynthase recombinase according to claim 1, wherein, The megaserine recombinase corresponds to a recombinase recognition site sequence (RS) selected from an attP site, an attB site.
3. The megasynthase recombinase according to claim 2, wherein, The attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 6-12, and / or the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 13-19.
4. The megasynthase recombinase according to any one of claims 1 to 3, wherein, The megaserine recombinase comprises an amino acid sequence as set forth in SEQ ID NO: 2, or an amino acid sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 99% identity to an amino acid sequence as set forth in SEQ ID NO:
2.
5. The megasynthase recombinase according to claim 4, wherein, The attP site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 7-8, and / or the attB site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 14-15.
6. A genome editing system, wherein, The genome editing system comprises: 1) a recombinase recognition site sequence (RS) comprising an attP or attB site; 2) a recombinase and / or an expression construct comprising a nucleotide sequence encoding the recombinase, the recombinase comprising a megaserine recombinase as claimed in any one of claims 1-5.
7. The genome editing system of claim 6, wherein, The recombinase recognition site sequence (RS) is naturally occurring, or engineered or optimized.
8. The genome editing system of claim 7, wherein one or more of the recombinase recognition site sequence (RS) is inserted into a desired location in the genome by a recombinase recognition site integration unit.
9. The genome editing system of claim 8, wherein, The recombinase recognition site integration unit comprises: a CRISPR effector protein or a functional variant thereof and / or an expression construct comprising a nucleotide sequence encoding the CRISPR effector protein or the functional variant thereof, and at least one guide RNA and / or at least one expression construct comprising a nucleotide sequence encoding the at least one guide RNA.
10. The genome editing system of claim 9, wherein, The CRISPR effector protein functional variant is a CRISPR nuclease with full / partial loss of cleavage activity, preferably the CRISPR nuclease functional variant is a CRISPR nickase, for example, Cas9-D10A, Cas9-H840A, Cas12a nickase, Cas12b nickase, or TraC nickase.
11. The genome editing system according to claim 9, wherein, The guide RNA comprises a scaffold sequence and a primer binding sequence, and an integration template sequence of a recombinase recognition site sequence (RS).
12. The genome editing system of any one of claims 8-11, wherein, The recombinase recognition site integration unit further comprises a reverse transcriptase and / or an expression construct comprising a nucleotide sequence encoding the reverse transcriptase.
13. The genome editing system of claim 12, wherein, The CRISPR nuclease of the recombinase recognition site integration unit is linked to a reverse transcriptase, the guide RNA interacts with the CRISPR nuclease and targets a desired location in the genome, wherein the CRISPR nuclease makes a cut in a strand of the genome, and the reverse transcriptase incorporates a RS integration template sequence in the guide RNA into the cut site, thereby inserting at least one recombinase recognition site sequence (RS) recognizable by the recombinase at the desired location in the genome.
14. The genome editing system of claim 12 or 13, wherein, The reverse transcriptase is selected from the group consisting of Moloney murine leukemia virus (M-MLV) reverse transcriptase, transcribing heteropolymerase (RTX), avian myeloblastosis virus reverse transcriptase (AMV-RT), and Faecalibacterium prausnitzii Marathonase RT (Marathon RT).
15. The genome editing system of any one of claims 8-14, wherein, The recombinase recognition site integration unit is not covalently linked to a recombinase.
16. The genome editing system of any one of claims 8-14, wherein, The recombinase recognition site integration unit is covalently linked to a recombinase.
17. The genome editing system of any one of claims 6-16, wherein, The genome editing system further comprises a donor comprising: 3) a donor of an exogenous nucleotide sequence to be inserted into the genome, Optionally, the exogenous nucleotide sequence can be about 1 bp to about 50 kb or longer.
18. The genome editing system of claim 17, wherein, The exogenous nucleotide sequence further comprises one or more recombinase recognition site sequences (RS) recognizable by a recombinase, optionally wherein the exogenous nucleotide sequence is flanked by one or two recombinase recognition site sequences (RS).
19. The genome editing system of any one of claims 6-18, wherein, The recombinase recognition site sequence (RS) is selected from the group consisting of an attP site comprising a nucleotide sequence selected from the group consisting of SEQ ID NOs: 6-12, and an attB site comprising a nucleotide sequence selected from the group consisting of SEQ ID NOs: 13-19.
20. A fusion protein comprising the genome editing system of any one of claims 6-19, wherein, The CRISPR nuclease is linked at the C-terminus to a reverse transcriptase, which is linked via a linker to a recombinase.
21. A polynucleotide comprising a nucleotide sequence encoding the megaserine recombinase of any one of claims 1-5, the genome editing system of any one of claims 6-19, or the fusion protein of claim 20, optionally wherein the polynucleotide is RNA, e.g., mRNA.
22. An expression construct comprising the polynucleotide of claim 21.
23. A cell comprising the megaserine recombinase of any one of claims 1-5, the genome editing system of any one of claims 6-19, the fusion protein of claim 20, the polynucleotide of claim 21, or the expression construct of claim 22.
24. A kit comprising the megaserine recombinase of any one of claims 1-5, the genome editing system of any one of claims 6-19, the fusion protein of claim 20, the polynucleotide of claim 21, the expression construct of claim 22, or the cell of claim 23.
25. A method of performing gene editing in an organism or cell of an organism, wherein, The method comprises introducing the megasynthase recombinase of any one of claims 1-5, the polynucleotide of claim 21, the expression construct of claim 22, the genome editing system of any one of claims 6-19, or the fusion protein of claim 20 into an organism or an organism cell.
26. The method of claim 25, wherein, The component 1), the component 2), and the optional component 3) of the genome editing system are introduced into the organism or the organism cell simultaneously.
27. The method of claim 25, wherein, The component 1), the component 2), or the optional component 3) of the genome editing system are introduced into the organism or the organism cell step by step respectively.
28. The method of claim 25, wherein, The component 2) of the genome editing system is introduced into the organism or the organism cell alone.
29. The method of claims 25-28, wherein, The component 1) inserts the RS into the donor construct of the genome or the exogenous nucleotide sequence in the same direction or in the opposite direction.
30. The method of any one of claims 25-29, wherein, The method comprises recombining the DNA of the genome of the organism or the organism cell.
31. The method of claim 30, wherein, The recombining the DNA of the genome of the organism or the organism cell comprises deleting the DNA, inverting the DNA in the genome, and / or integrating the exogenous DNA into the genome.
32. The method of any one of claims 25-31, wherein, The megasynthase recombinase, the genome editing system, the polynucleotide, or the expression construct is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistic method, N-acetylgalactosamine (GalNAc) mediation, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation.
33. The method according to any one of claims 25-32, wherein, The organism or the organism cell is from a mammal such as human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; plant, preferably crop plant, for example wheat, rice, corn, soybean, sunflower, kiwifruit, leafy vegetable, lettuce, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato. The method comprises introducing the megasynthase recombinase of any one of claims 1-5, the polynucleotide of claim 21, the expression construct of claim 22, the genome editing system of any one of claims 6-19, or the fusion protein of claim 20 into an organism or an organism cell. The component 1), the component 2), and the optional component 3) of the genome editing system are introduced into the organism or the organism cell simultaneously. The component 1), the component 2), or the optional component 3) of the genome editing system are introduced into the organism or the organism cell step by step respectively. The component 2) of the genome editing system is introduced into the organism or the organism cell alone. The component 1) inserts the RS into the donor construct of the genome or the exogenous nucleotide sequence in the same direction or in the opposite direction. The method comprises recombining the DNA of the genome of the organism or the organism cell. The recombining the DNA of the genome of the organism or the organism cell comprises deleting the DNA, inverting the DNA in the genome, and / or integrating the exogenous DNA into the genome. The megasynthase recombinase, the genome editing system, the polynucleotide, or the expression construct is introduced into the cell by a method selected from the group consisting of calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), biolistic method, N-acetylgalactosamine (GalNAc) mediation, PEG-mediated protoplast transformation, Agrobacterium tumefaciens-mediated transformation. The organism or the organism cell is from a mammal such as human, mouse, rat, monkey, dog, pig, sheep, cow, cat; poultry such as chicken, duck, goose; plant, preferably crop plant, for example wheat, rice, corn, soybean, sunflower, kiwifruit, leafy vegetable, lettuce, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, tomato, tobacco, cassava and potato.
Citation Information
Patent Citations
Site-specific serine recombinases and methods of their use
CN101194018A
Method for integrating genome with exogenous sequence
CN113355345A
Serine recombinases mediating stable integration into plant genomes
CN113614227A
Systems, methods, and compositions for site-specific genetic modification using programmable addition (PASTE) implemented with site-specific targeting elements
CN116419975A
Method for inserting exogenous sequence in genome at fixed point
CN117126876A