Optimized recombinase recognition site and application thereof
By optimizing the nucleic acid molecule of the attB recognition site of the recombinase, the problem of low attB site activity of the macroserine recombinase Bxb1 in the prior art has been solved, achieving high-efficiency recombination activity and improved gene editing performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies struggle to efficiently screen for the attB recognition site of the large serine recombinase Bxb1 with high recombinant activity using traditional methods, and there is a lack of systematic design or creation of saturated mutant libraries to enhance recombinant activity.
By rationally designing and saturating mutations, the nucleic acid molecule at the recognition site of the recombinase attB is optimized to form a specific base sequence [motif1-L][motif2-L]GACGGCGGWCTCC[motif3][motif2-R][motif1-R], which is then combined with the recognition site of the recombinase attP to construct a site-specific recombination system for application in eukaryotic cells.
This improved the recombination activity of recombinases and the performance of gene editing tools, achieving high flipping and integration activity.
Smart Images

Figure CN121759451A_ABST
Abstract
Description
[0001] Priority and related applications
[0002] This invention claims priority to Chinese Patent Application No. 202510324622.8, filed on March 19, 2025, entitled "Optimized Recombinase Recognition Site and Its Application Thereof". The entire contents of the above-cited patent application are incorporated herein by reference. Technical Field
[0003] This invention belongs to the field of genetic engineering technology and relates to optimized recombinase recognition sites and their applications. Specifically, it relates to a method for improving the recombination activity of the attB recognition site of the macroserine recombinase Bxb1 obtained through rational design and saturation mutation, and its application in gene editing systems. Background Technology
[0004] Large serine recombinases (such as Bxb1 and phiC31) mediate DNA recombination by recognizing a pair of distinct DNA sequences (attB and attP), known as recombinase recognition sites. Specifically, large serine recombinases recognize attB and attP to form a tetramer, thereby activating recombinase activity. This leads to DNA breakage at the central dibases of attB and attP, resulting in the attL and attR sequences. Depending on the relative orientation of the attB and attP sequences, functions such as loop excision (excision), inversion (flipping), integration, or chromosomal translocation can be achieved.
[0005] Optimization of recombinase activity primarily revolves around optimizing the recombinase protein and its recognition site sequence. Regarding recognition site optimization, the long natural recognition site of large serine recombinases makes it difficult to screen for highly active variants using traditional saturation mutagenesis methods. Existing techniques involve modifying specific recognition regions of the recombinase protein through directed evolution to obtain attP variant sequence libraries with altered recombinant activity, but these methods have limited efficiency and lack systematic design. Currently, there are no reports of rationally designing or creating saturation mutant libraries targeting the attB site and then screening them to enhance recombinant activity. Summary of the Invention
[0006] To address the aforementioned technical problems, this application provides a method for obtaining a library of recombinase attB recognition sites through rational design and saturation mutation, and then obtaining recombinase recognition sites with high flipping and integration activity in eukaryotic cells through high-throughput screening.
[0007] The technical solution provided by this invention is as follows:
[0008] This invention provides a nucleic acid molecule with a non-natural recombinase attB recognition site, wherein the non-natural attB recognition site contains the base sequence: [motif1-L][motif2-L]GACGGCGGWCTCC[motif3] [motif2-R][motif1-R].
[0009] In some embodiments, [motif1-L] is selected from at least one of SEQ ID NO:34-37; [motif2-L] is selected from at least one of SEQ ID NO:38-40; [motif3] is selected from at least one of SEQ ID NO:41-46; [motif2-R] is selected from at least one of SEQ ID NO:47-51; [motif1-R] is selected from at least one of SEQ ID NO:52-55; and W is a base A or T.
[0010] In some embodiments, the recombinase is selected from Bxb1, Bxb1 variants, BxZ2, PhiC31, Peaches, Veracruz, Rebeuca, Theia, Benedict, PattyP, Trouble, KSSJEB, Lockley, Scowl, Switzer, Bob3, Abrogate, Doom, ConceptII, Anglerfish, SkiPole, Museum, and Severus. Preferably, the recombinase is Bxb1 and / or the Bxb1 variant.
[0011] In some embodiments, the amino acid sequence of Bxb1 and / or the Bxb1 variant is selected from amino acid sequences containing SEQ ID NO:1-4 and SEQ ID NO:152-154.
[0012] In some embodiments, the non-natural attB recognition site sequence is selected from the nucleotide sequences shown in SEQ ID NO:7-22 and SEQ ID NO:56-151.
[0013] In some embodiments, the nucleic acid molecule of the non-natural recombinase attB recognition site comprises a nucleotide sequence selected from SEQ ID NO:9, SEQ ID NO:18-22, SEQ ID NO:56-59 or SEQ ID NO:116-117.
[0014] This invention provides a site-specific recombination method, comprising i) and ii):
[0015] i) Provide a nucleic acid molecule containing the non-natural recombinase attB recognition site described above;
[0016] ii) Provide a nucleic acid molecule with a recombinase attP recognition site;
[0017] In the presence of recombinase, the nucleic acid molecules in i) and ii) are combined into paired recombinase att recognition sites.
[0018] In some embodiments, the nucleic acid molecule of the attP recognition site is selected from natural recombinase attP recognition sites or non-natural recombinase attP recognition sites.
[0019] In some embodiments, the nucleotide sequence of the natural recombinase attP recognition site is shown in SEQ ID NO:6; the nucleotide sequence of the non-natural recombinase attP recognition site is selected from the nucleotide sequences shown in SEQ ID NO:23-33.
[0020] In some embodiments, the nucleic acid molecules in i) or ii) are located in plasmids, bacteriophages, bacterial genomes, fungal genomes, plant genomes, or animal genomes.
[0021] This invention provides a site-specific recombination system, wherein the site-specific recombination system comprises:
[0022] a) Nucleic acid molecules or their expression constructs that are not natural recombinases and their attB recognition sites as described above;
[0023] b) A nucleic acid molecule of the attP recognition site or its expression construct, wherein the nucleic acid molecule of the attP recognition site is selected from a natural or non-natural recombinase attP recognition site, the nucleotide sequence of the natural attP recognition site is shown in SEQ ID NO:6, and the nucleotide sequence of the non-natural attP recognition site is selected from the nucleotide sequences shown in SEQ ID NO:23-33; and
[0024] c) Recombinases or their expression constructs.
[0025] In some embodiments, the site-specific recombination system further comprises:
[0026] d) A genome editing system that introduces the non-natural recombinase attB recognition site or recombinase attP recognition site sequence described above into the cell genome.
[0027] In some embodiments, the genome editing system is selected from a wide range of genome editing systems mediated by nucleases, zinc finger proteins, TALEN, CRISPR-Cas, topoisomerases, and recombinase systems.
[0028] The present invention provides a host cell containing a nucleic acid molecule with a non-natural recombinase attB recognition site as described above, or a site-specific recombination system as described in any one of the above.
[0029] In some embodiments, the cells are derived from prokaryotes and / or eukaryotes; the prokaryotes include bacteria; and the eukaryotes include plants, fungi, and / or vertebrates.
[0030] In some embodiments, the vertebrates include mammals, including humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and / or cats;
[0031] The plants include crop plants, which include wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, or potatoes.
[0032] This invention provides nucleic acid molecules with non-natural recombinase attB recognition sites as described above, site-specific recombination systems as described above, and the use of host cells as described above in the preparation of reagents for modifying target nucleic acids.
[0033] This invention provides the use of the non-natural recombinase attB recognition site sequence as described above, the site-specific recombination system as described above, or the host cell as described above in the preparation of medicaments for the prevention and / or treatment of diseases related to or caused by gene recombination.
[0034] Specifically, the base editor and its applications provided by this invention have the following beneficial effects:
[0035] 1) A series of new recombinase attB recognition site libraries were obtained through rational design and saturation mutagenesis;
[0036] 2) The novel non-natural attB recognition site enhances the recombinase activity;
[0037] 3) Novel non-natural attB recognition sites enhance the performance of recombinase-related gene editing tools. Attached Figure Description
[0038] To better understand the technical solutions described in this invention, the following description is provided in conjunction with the accompanying drawings.
[0039] Figure 1 This is a schematic diagram of the construct for the recombinant activity verification system 1.
[0040] Figure 2 Statistics on the recombination activity of Bxb1 recombinase variants against novel att site pairs.
[0041] Figure 3 The effect of optimized variants of the novel attB site on the efficiency of large-fragment recombination in the human genome is shown.
[0042] Figure 4 The effect of optimized variants of the novel attB site on the efficiency of large-fragment recombination in the rice genome is shown.
[0043] Figure 5 Statistics on recombination activity of novel attB site variants are shown.
[0044] Figure 6 This is a statistical graph showing the recombination activity of Bxb1 recombinase variants at the novel attB site. Detailed Implementation
[0045] I. Terminology
[0046] Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by those skilled in the art.
[0047] A numerical range includes the numbers that define the range, and explicitly includes every integer and non-integer fraction within the defined range. Unless otherwise stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0048] As used in this invention, the term "structure" or "recombinant expression structure" refers to an artificially designed DNA fragment that can be used to introduce genetic material into target cells (e.g., using a recombinant expression structure to produce a base editor or a component thereof). The term "expression" refers to the transcription and translation of a nucleic acid coding sequence to produce a coding polypeptide.
[0049] As used in this invention, the term "genetic engineering" refers to altering the genetic composition of cells through biotechnology, including intra- and inter-species gene transfer, to produce modified or non-naturally occurring cells. The term "structure-encoding base editor or a component thereof" refers to genetically engineered cells that produce base editors. Cells containing exogenous, recombinant, synthetic, and / or other modified polynucleotides are considered genetically engineered cells, and therefore are non-naturally occurring relative to any naturally occurring counterpart. In some cases, genetically engineered cells contain one or more recombinant nucleic acids. In other cases, genetically engineered cells contain one or more synthetic or genetically engineered nucleic acids (e.g., a nucleic acid containing at least one artificially inserted, deleted, reversed, or substituted sequence relative to its naturally occurring counterpart). Methods for producing genetically engineered cells are known in the art, for example, Sambrook et al., Molecular Cloning, A Laboratory Manual (Fourth Edition), Cold Spring Harbor Press, Cold Spring Harbor, NY (2012).
[0050] As used in this invention, the terms "genetically engineered cell," "genetically engineered host cell," or "recombinant expression host cell" can refer to cells modified using gene editing techniques. Gene editing refers to a form of genetic engineering that inserts, deletes, modifies, or replaces DNA in the genome of a living cell. Compared to other genetic engineering techniques that can randomly insert genetic material into the host genome, gene editing can target the insertion to a specific location (e.g., the AAVS1 allele). Examples of gene editing techniques include, but are not limited to, restriction endonucleases, zinc finger nucleases, TALENs, and CRISPR-Cas9. The base editor disclosed herein is a specific example of gene editing that allows for changes with one or more single nucleotides, particularly resulting in alterations to the cell phenotype.
[0051] In this invention, unless otherwise stated, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. For example, the standard recombinant DNA and molecular cloning techniques used in this invention are well known to those skilled in the art and are described more fully in the following literature: Sambrook, J., Fritsch, EF, and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989 (hereinafter referred to as "Sambrook"). Meanwhile, to better understand this invention, definitions and explanations of relevant terms are provided below.
[0052] As used herein, the term “and / or” covers all combinations of items connected by the term and should be regarded as if each combination had been listed separately herein. For example, “A and / or B” covers “A,” “A and B,” and “B.” For example, “A, B, and / or C” covers “A,” “B,” “C,” “A and B,” “A and C,” “B and C,” and “A and B and C.”
[0053] When the term "comprising" is used herein to describe a protein or nucleic acid sequence, the protein or nucleic acid may consist of the stated sequence, or may have additional amino acids or nucleotides at one or both ends of the protein or nucleic acid, while still possessing the activities described in this invention. Furthermore, those skilled in the art will understand that the methionine encoded by the start codon at the N-terminus of a polypeptide may be retained in certain practical situations (e.g., when expressed in a specific expression system) without substantially affecting the polypeptide's function. Therefore, when describing a specific polypeptide amino acid sequence in this specification and claims, although it may not contain the methionine encoded by the start codon at the N-terminus, the sequence containing that methionine is still included, and correspondingly, its encoding nucleotide sequence may also contain the start codon; and vice versa.
[0054] The terms “gene” and “genome” as used in this article not only encompass chromosomal DNA located in the cell nucleus, but also organelle DNA located in subcellular components of the cell (such as mitochondria and plastids).
[0055] As used herein, “organism” includes any organism suitable for genome editing, preferably eukaryotes. Examples of organisms include, but are not limited to, mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants including monocots and dicots such as rice, corn, wheat, sorghum, barley, soybeans, peanuts, and Arabidopsis thaliana.
[0056] "Genetically modified organism" or "genetically modified cell" refers to an organism or cell whose genome contains exogenous polynucleotides or modified genes or expression regulatory sequences. For example, exogenous polynucleotides can be stably integrated into the genome of an organism or cell and inherited across generations. Exogenous polynucleotides can be integrated into the genome alone or as part of a recombinant DNA construct. Modified genes or expression regulatory sequences are sequences in the genome of an organism or cell that contain single or multiple deoxynucleotide substitutions, deletions, and additions.
[0057] The terms “polynucleotide,” “nucleic acid sequence,” “nucleotide sequence,” or “nucleic acid fragment” are used interchangeably and are single-stranded or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural, or modified nucleotide bases. Nucleotides are designated by their single-letter names as follows: “A” for adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), “C” for cytidine or deoxycytidine, “G” for guanosine or deoxyguanosine, “U” for uridine, “T” for deoxythymidine, “R” for purine (A or G), “Y” for pyrimidine (C or T), “K” for G or T, “H” for A, C, or T, “I” for inosine, and “N” for any nucleotide.
[0058] "Introducing" nucleic acid molecules (such as plasmids, linear nucleic acid fragments, RNA, etc.) or proteins into an organism refers to transforming the cells of an organism with the nucleic acid or protein, enabling the nucleic acid or protein to function within the cell. The term "transformation" as used in this invention includes both stable transformation and transient transformation.
[0059] As used in this disclosure, the term "amino acid" can include natural amino acids, non-natural amino acids, amino acid analogs, and all their D and L stereoisomers. The amino acids and their abbreviations and English abbreviations in this disclosure are as follows:
[0060] Histidine (His, H); Serine (S); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Threonine (Thr, T); Phenylalanine (Phe, F); Aspartic acid (Asp, D); Tyrosine (Tyr, Y); Leucine (Leu, L); Isoleucine (Ile, I); Arginine (Arg, R); Alanine (Ala, A); Valine (Val, V); Tryptophan (Trp, W); Methionine (Met, M); Asparagine (Asn, N); Cysteine (Cys, C); Lysine (Lys, K); Proline (Pro, P).
[0061] As used herein, the terms "large serine recombinase" and "LSR" are used interchangeably and refer to a class of serine recombinases. Serine recombinases are integrase carried by bacteriophages. They integrate large DNA sequences into the bacterial genome by recognizing specific sequences on the bacterial genome and bacteriophage DNA fragments, without requiring any cell cofactors or producing harmful DNA double-strand breaks (DSBs). Many serine recombinases function to resolve translocation intermediates or regulate gene expression by reversing regulatory sequences. In the serine recombinase-catalyzed reaction, DNA first forms a complex with the recombinase. Then, a serine residue in the recombinase domain attacks the DNA phosphate backbone, causing a nick in the DNA. At the nick, a 3′-OH double-strand break terminus and a 5′-phosphoserine residue covalently linked to the DNA are formed. The synergistic complex flips the DNA, and the double-stranded DNA rejoins to form the recombinant product. Most serine recombinases have a 150-amino acid catalytic domain at their amino terminus, followed by a small HTH (Helix-Turn-Helix)-DNA binding domain. Large serine recombinases have similar amino-terminal catalytic domains, but have larger carboxyl-terminal regions, ranging in size from 300 residues in phage R4 and A118 recombinases to 550 residues in TnpX transposases. These serine recombinases are called large serine recombinases because of their larger carboxyl-terminal regions.
[0062] As used in this article, the terms "attP" and "attB" refer to the phage attachment site (attP) and the bacterial attachment site (attB), respectively, and are collectively referred to as "att sites" in this article. The original function of integrase is to enable recombination of short DNA sequences between the phage attachment site (attP) and the bacterial attachment site (attB). Therefore, in gene editing, the recognition site of the target DNA during the editing process is called "attB," and the recognition site of the DNA sequence to be edited is called "attP."
[0063] The term "recombination" refers to the excision, integration, flipping, or exchange (e.g., translocation) of DNA fragments between sequences recognized by recombinases. In some specific embodiments, recombinase recombination activity includes either integrative or flipping activity of the recombinase.
[0064] As used in this invention, "expression construct" or "construct" refers to a vector, such as a recombinant vector, suitable for expressing a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, the expression of a nucleotide sequence can refer to the transcription of the nucleotide sequence (e.g., transcription to generate mRNA or functional RNA) and / or the translation of RNA into a precursor or mature protein.
[0065] The "expression construct" of the present invention may be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, a translatable RNA (such as mRNA).
[0066] The "expression construct" of the present invention may contain regulatory sequences and nucleotide sequences of interest from different sources, or regulatory sequences and nucleotide sequences of interest from the same source but arranged in a manner different from those normally found in nature.
[0067] II. Nucleic acid molecules with non-natural recombinase attB recognition sites
[0068] This invention provides a nucleotide molecule of a non-natural recombinase attB recognition site, wherein the non-natural recombinase attB recognition site nucleic acid molecule contains a specific base sequence: [motif1-L][motif2-L] GACGGCGGWCTCC[motif3] [motif2-R][motif1-R].
[0069] In some specific implementations, [motif1-L] is selected from at least one of SEQ ID NO:34-37; [motif2-L] is selected from at least one of SEQ ID NO:38-40; [motif3] is selected from at least one of SEQ ID NO:41-46; [motif2-R] is selected from at least one of SEQ ID NO:47-51; [motif1-R] is selected from at least one of SEQ ID NO:52-55, and W is A or T.
[0070] As used in this invention, in the above-mentioned base sequence, "GW" is the central dibase of the attB recognition site, and the central dibase of the Bxb1 wild-type attB recognition site is "GT". At the same time, the prior art has proven that specific gene recombination can also be achieved when the central dibase is "GA". The central dibase does not participate in the specific recognition process of the recombinase on the site sequence, but "GA" and "GT" are orthogonal, that is, attB-GA can only recombine with attP-GA and cannot recombine with attP-GT, and vice versa. Specific references for the prior art: Jusiak, Barbara et al. "Comparison of IntegrasesIdentifies Bxb1-GA Mutant as the Most Efficient Site-Specific IntegraseSystem in Mammalian Cells." ACS synthetic biology vol. 8,1 (2019): 16-24.doi:10.1021 / acssynbio.8b00089; Yarnall, Matthew TN et al. “Drag-and-dropgenome insertion of large sequences without double-strand DNA cleavage using CRISPR-directed integrases.” Nature biotechnology vol. 41,4 (2023): 500-512.doi:10.1038 / s41587-022-01527-4.
[0071] In some alternative implementations, the central dibase of the attB recognition site is selected from "GA" and "GT". The selection of the central dibase does not affect the function of recombination at a specific location of the recombinase, and the central dibases of paired att recognition sites are consistent.
[0072] The recombinases are selected from Bxb1, Bxb1 variants, BxZ2, PhiC31, Peaches, Veracruz, Rebeuca, Theia, Benedict, PattyP, Trouble, KSSJEB, Lockley, Scowl, Switzer, Bob3, Abrogate, Doom, ConceptII, Anglerfish, SkiPole, Museum, and Severus. For specific details of the recombinases, please refer to US Patent US10731153B2.
[0073] In some preferred embodiments, the recombinase is Bxb1 and its variants. Specifically, the sequence of Bxb1 is shown in SEQ ID NO:1. In some exemplary embodiments, Bxb1 variants include eeBxb1, Bxb1(A315R), Bxb1(T84N+E69M), Bxb1-hV30, Bxb1-V27, and Bxb1-V39, whose amino acid sequences are shown in SEQ ID NO:2-4 and SEQ ID NO:152-154, respectively.
[0074] Furthermore, the nucleic acid molecule of the non-natural recombinase attB recognition site comprises nucleotide sequences selected from those shown in SEQ ID NO:7-22 and SEQ ID NO:56-151.
[0075] III. Site-Specific Recombination Methods
[0076] This invention provides a site-specific recombination method, comprising i) and ii):
[0077] i) Provide a nucleic acid molecule containing the non-natural recombinase attB recognition site described above;
[0078] ii) Provide a nucleic acid molecule with a recombinase attP recognition site.
[0079] In particular, in the presence of recombinase, the nucleic acid molecules in i) and ii) are combined into paired recombinase att recognition sites.
[0080] In some specific implementations, the nucleic acid molecule of the attP recognition site is selected from a natural recombinase attP recognition site or a non-natural recombinase attP recognition site.
[0081] In some exemplary embodiments, the nucleotide sequence of the natural recombinase attP recognition site is shown in SEQ ID NO:6; the nucleotide sequence of the non-natural recombinase attP recognition site is selected from the nucleotide sequences shown in SEQ ID NO:23-33.
[0082] In some exemplary embodiments, the paired att recognition sites include i) a combination of a nucleic acid molecule with a sequence as shown in SEQ ID NO:6; or
[0083] The paired recombinase att recognition sites are selected from the following (A1) to (A11):
[0084] (A1) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:7 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:23;
[0085] (A2) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:8 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:24;
[0086] (A3) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:9 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:25;
[0087] (A4) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:10 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:26;
[0088] (A5) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:11 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:27;
[0089] (A6) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:12 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:28;
[0090] (A7) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:13 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:29;
[0091] (A8) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:14 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:30;
[0092] (A9) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:15 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:31;
[0093] (A10) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:16 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:32; or
[0094] (A11) Nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:17 and nucleic acid molecules with nucleotide sequences as shown in SEQ ID NO:33.
[0095] In some alternative embodiments, the nucleic acid molecules in i) or ii) are located in plasmids, bacteriophages, bacterial genomes, fungal genomes, plant genomes, or animal genomes. This includes, but is not limited to, bacteriophages such as T4 bacteriophage and λ bacteriophage; bacteria such as Escherichia coli, Bacillus subtilis, Lactobacillus, and Streptococcus; fungi such as yeasts, molds, Penicillium, Aspergillus, Mucor, Rhizopus, and Agrobacterium; mammals such as humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and cats; poultry such as chickens, ducks, and geese; and plants including monocots and dicots, such as wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, kiwifruit, tobacco, cassava, and potatoes.
[0096] IV. Site-Specific Recombination System
[0097] This invention provides a site-specific recombination system comprising:
[0098] a) Nucleic acid molecules or their expression constructs that are not natural recombinases and their attB recognition sites as described above;
[0099] b) A nucleic acid molecule of the attP recognition site or its expression construct, wherein the nucleic acid molecule of the recombinase attP recognition site is selected from a natural recombinase attP recognition site or a non-natural recombinase attP recognition site, the nucleotide sequence of the natural recombinase attP recognition site is as shown in SEQ ID NO:6, and the nucleotide sequence of the non-natural recombinase attP recognition site is selected from the nucleotide sequences shown in SEQ ID NO:23-33; and
[0100] c) Recombinase or its expression construct.
[0101] In some embodiments, the site-specific recombination system further comprises:
[0102] d) A genome editing system that introduces nucleic acid molecules with the non-natural recombinase attB recognition site or the recombinase attP recognition site as described above into the cell genome.
[0103] Those skilled in the art can choose appropriate gene editing tools to insert nucleic acid molecules with the attB recognition site or the attP recognition site of the non-natural recombinase described above into the desired location in the genome.
[0104] In some alternative implementations, the genome editing system is selected from a wide range of genome editing systems mediated by nucleases, zinc finger proteins, TALENs, CRISPR-Cas, topoisomerases, or recombinases.
[0105] Furthermore, the gene editing system can be a CRISPR-Cas mediated genome editing system, such as a guided editing system. This system includes a fusion of a CRISPR nuclease (i.e., a Cas nuclease or its variants, such as Cas9-H840A or Cas9-D10A) with a reverse transcriptase (e.g., M-MLV reverse transcriptase) (guided editing fusion protein) and a guide RNA (pegRNA, prime editing RNA, guide editing gRNA) targeting the target sequence, with a 3' end containing a repair template (RT template, RTT) and a free single-stranded binding region (primer binding site, PBS). This system binds to the free single strand produced by the Cas nuclease or its variants (e.g., Cas9-H840A or Cas9-D10A) via PBS, causing it to transcribe a single-stranded DNA sequence according to the given RTT. After cellular repair, arbitrary changes to the DNA sequence located downstream of the PAM sequence at the 3' end can be achieved in the genome; for example, it can be used to alter target genomic sites, including base substitutions, insertions, and deletions.
[0106] In some preferred embodiments, the guided editing system is an ePPE-mediated dual pegRNA guided editing system. The pegRNA expression vector targets the target region and, through ePPE mediation containing Cas nuclease and reverse transcriptase, achieves integration into the genome of plant cells (e.g., rice cells). In some embodiments, the ePPE vector sequence expressing Cas enzyme and reverse transcriptase is shown in SEQ ID NO:161, and the dual pegRNA sequences are shown in SEQ ID NO:162 and SEQ ID NO:163, respectively.
[0107] In other preferred embodiments, the guided editing system is a PE2-mediated dual pegRNA guided editing system. Integration into the genome of animal cells (e.g., human cells) is achieved by targeting the target region using a pegRNA expression vector and mediating through PE2, which contains both Cas nuclease and reverse transcriptase. In some embodiments, the PE2 vector sequence expressing Cas and reverse transcriptase is shown in SEQ ID NO:168, and the dual pegRNA sequences are shown in SEQ ID NO:155 and SEQ ID NO:156, respectively.
[0108] V. Polynucleotides
[0109] In one aspect, the present invention provides a polynucleotide encoding:
[0110] (a) Nucleic acid molecules that are non-natural recombinases that recognize the attB recognition site as described above; and / or
[0111] (b) The site-specific recombination system described above.
[0112] The polynucleotides of this invention can be in DNA or RNA form. DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. DNA can be single-stranded or double-stranded. DNA can be a coding strand or a non-coding strand. RNA form includes mRNA.
[0113] The polynucleotides include: nucleic acid molecules that encode only the non-natural recombinase attB recognition site described above; the coding sequence of the nucleic acid molecule that encodes only the site-specific recombination system described above; the coding sequence of the site-specific recombination system described above and various additional coding sequences; the coding sequence of the site-specific recombination system described above (and optional additional coding sequences) and non-coding sequences.
[0114] VI. Expression constructs and host cells
[0115] In one aspect, the present invention provides an expression construct comprising the polynucleotides described above.
[0116] In some implementations, the expression construct is selected from expression constructs of viruses, bacteria, yeast, plants, and mammalian cells.
[0117] In another aspect, the present invention provides a host cell containing a nucleic acid molecule with a non-natural recombinase attB recognition site as described above, a site-specific recombination system as described above, a polynucleotide as described above, or an expression construct as described above.
[0118] VII. Methods for modifying target nucleic acids and methods for generating at least one genetically modified cell.
[0119] In one aspect, the present invention provides a method for modifying a target nucleic acid, the method comprising the step of contacting the target nucleic acid with a nucleic acid molecule having a non-natural recombinase attB recognition site as described in the present invention, a site-specific recombination system as described in the present invention, a polynucleotide as described in the present invention, or an expression construct as described in the present invention, a pharmaceutical composition as described in the present invention, or a kit as described in the present invention, thereby causing modification, for example recombination, of the nucleic acid sequence in the target nucleotide region of the at least one cell.
[0120] On the other hand, the present invention provides a method for generating at least one genetically modified cell, the method comprising introducing a nucleic acid molecule with a non-natural recombinase attB recognition site as described in the present invention, a site-specific recombination system as described in the present invention, a polynucleotide as described in the present invention, or an expression construct as described in the present invention into at least one cell, thereby causing modification, for example recombination, of the nucleic acid sequence in a target nucleotide region of the at least one cell.
[0121] In this invention, the modification can be located anywhere in the genome, such as within a functional gene like a protein-coding gene, or in a gene expression regulatory region such as a promoter or enhancer region, thereby achieving modification of gene function or gene expression. The modification in the cell genome sequence can be detected using T7EI, PCR / RE, or sequencing methods.
[0122] In some embodiments, the introduction is carried out by a variety of methods well known to those skilled in the art: calcium phosphate transfection, protoplast fusion, electroporation, liposome transfection, microinjection, viral infection (such as baculovirus, vaccinia virus, adenovirus, adeno-associated virus, lentivirus or other viruses), N-acetylgalactosamine (GalNAc) mediated, gene gun method, PEG-mediated protoplast transformation and / or Agrobacterium tumefaciens-mediated transformation.
[0123] In some exemplary embodiments, the cells are cells derived from prokaryotes and / or eukaryotes; the prokaryotes include bacteria; the eukaryotes include plants, fungi, and / or vertebrates.
[0124] In some preferred embodiments, the vertebrates include mammals, including humans, mice, rats, monkeys, dogs, pigs, sheep, cattle, and / or cats; the plants include crop plants, including wheat, rice, corn, soybeans, sunflowers, sorghum, rapeseed, alfalfa, cotton, barley, millet, sugarcane, tomatoes, tobacco, cassava, and / or potatoes.
[0125] In some embodiments, the method of the present invention is performed in vitro. For example, the cells are isolated cells, or cells in isolated tissues or organs.
[0126] In other embodiments, the method of the present invention can also be performed in vivo. For example, the cells are cells within an organism, and the system of the present invention can be introduced into the cells in vivo via, for example, a viral or Agrobacterium-mediated method.
[0127] VIII. Pharmaceutical Compositions
[0128] In one aspect, the present invention provides a pharmaceutical composition comprising a nucleic acid molecule with a non-natural recombinase attB recognition site as described in the present invention, a site-specific recombination system as described in the present invention, a polynucleotide as described in the present invention, an expression construct as described in the present invention, and a host cell as described in the present invention.
[0129] In some alternative embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier; further, the pharmaceutically acceptable carrier comprises one or more combinations of solvents, solubilizers, cosolvents, emulsifiers, flavoring agents, odorants, colorants, binders, disintegrants, fillers, lubricants, wetting agents, osmotic pressure regulators, pH regulators, stabilizers, surfactants, and preservatives.
[0130] IX. Reagent Kit
[0131] In one aspect, the present invention also includes a kit for use in the methods of the present invention, said kit comprising a nucleic acid molecule with a non-natural recombinase attB recognition site as described in the present invention, a site-specific recombination system as described in the present invention, a polynucleotide as described in the present invention, or an expression construct as described in the present invention, or a host cell as described in the present invention. The kit generally includes a label indicating the intended use and / or method of use of the kit contents. Terminology labels include any written or documented material provided on or with the kit or otherwise accompanied by the kit.
[0132] 10. Uses
[0133] In one aspect, the present invention provides the use of the non-natural recombinase attB recognition site of the present invention, the site-specific recombination system of the present invention, the polynucleotide of the present invention, or the expression construct of the present invention, the host cell of the present invention, or the pharmaceutical composition of the present invention in the preparation of reagents for modifying target nucleic acids.
[0134] On the other hand, the present invention provides the use of the non-natural recombinase attB recognition site of the present invention, the site-specific recombination system of the present invention, the polynucleotide of the present invention, or the expression construct of the present invention, or the host cell of the present invention in the preparation of a medicament for the prevention and / or treatment of diseases related to or caused by gene recombination.
[0135] Example
[0136] A further understanding of the invention can be obtained by referring to some specific embodiments given herein, which are for illustrative purposes only and are not intended to limit the scope of the invention in any way. Obviously, various modifications and variations can be made to the invention without departing from its spirit; therefore, such modifications and variations are also within the scope of protection claimed in this application.
[0137] Example 1: Design of a novel att locus
[0138] Rational design was carried out based on the shortest natural recognition sequences attB and attP (hereinafter referred to as wild-type attB / attP) of the wild-type Bxb1 recombinase.
[0139] Specifically, the rational design of attB sites is based on the shortest natural sequence of wild-type attB. (SEQ ID NO:170), where W is A or T), this sequence is 38 bp. The substituted positions are as shown in the underlined positions above. A single underline indicates motif1, where positions 1-8 are represented as motif1-L, and the mirror-symmetric positions 31-38 are represented as motif1-R; a double underline indicates motif2, where positions 9-11 are represented as motif2-L, and the mirror-symmetric positions 28-30 are represented as motif2-R.
[0140] Further rational design of attP sites, based on the shortest natural sequence of wild-type attP ( (SEQ ID NO: 171), where W is A or T, and W corresponds to the attB site), this sequence is 48 bp. The substituted positions are as shown in the underlined positions above: motif1-L (1-8), motif2-L (14-16), motif2-R (33-35), and motif1-R (41-48). The above motif segmentation method is referenced from: Friedrich Fauser et al., Systematic Development of Reprogrammed Modular Integrases Enables Precise Genomic Integration of Large DNA Sequence, bioRxiv (2024).
[0141] The combination of motif1-L + motif2-L from attB is called BP1, the combination of motif1-R + motif2-R from attB is called BP2, the combination of motif1-L + motif2-L from attP is called BP3, and the combination of motif1-R + motif2-R from attP is called BP4. If all the motif1-L and motif2-L, as well as the motif1-R and motif2-R sequences of attB and attP, are replaced with BP1, this rationally designed pair of atts is called BP1111, and the naming follows the same pattern. Due to the mirror symmetry principle of module partitioning, the sequence when motif1-R is replaced in the position of motif1-L is the reverse complementary sequence of motif1-R, and so on.
[0142] The novel, rationally designed non-natural att sites were validated. A schematic diagram of the recombinant activity validation system 1 is shown below. Figure 1 As shown in the diagram, this construct contains a site-specific recombinase (e.g., Bxb1 or a functional variant thereof), with the tobacco mosaic virus CaMV35S promoter initiating Bxb1 expression and the T-NOS terminator terminating Bxb1 expression. The construct also contains attB and attP sequences or variants thereof arranged in opposite order, with a CaMV 35S promoter inserted into attB and attP in the opposite direction to firefly luciferase (F-LUC) expression. Firefly luciferase expression is terminated using a sequence from the 3' untranslated region of Agrobacterium tumefaciens. The construct further includes Renida luciferase as an internal control to correct firefly luciferase expression; Renida luciferase is initiated using the cassava vein mosaic virus CSVMV promoter and terminated by the CaMV 35S terminator. The recombinant activity verification system 2 was transfected into cells. Bxb1 expressed and recognized the attB and attP sites or their variant sequences. Under the action of recombinase, the CaMV 35S promoter sequence in the att site underwent recombination, and the recombinant CaMV 35S promoter initiated the expression of firefly luciferase. The expression intensity of firefly luciferase was detected using a microplate reader, and Renilla luciferase with constant expression was selected as an internal control. The luciferase detection kit used was the Transdetect Double-Luciferase Reporter Assay Kit (Transgen, FR201). The average of three biological replicates of the fold increase in activity relative to the wild-type att site of Bxb1 was taken, and the activity results were normalized. The activity of the wild-type att site was taken as 1, and the activities of other rationally designed att sites were taken as folds of 1 for comparison.
[0143] The list of att sites with enhanced recombinant activity is shown in Table 2, and the activity enhancement results are as follows: Figure 2As shown. First, replace all four parts of attB and attP with the same module combination, namely BP1111, BP2222, BP3333, and BP4444. Among them, the recombination activity of BP3333 and BP4444 is significantly improved compared with WT (BP1234). Figure 2 A), In addition, the inventors tested more permutations and combinations of modules, among which the activity of att site combinations of BP4434, BP3434, BP1414, BP3334, BP1334, and BP1434 were all enhanced, indicating that most of the att combinations with enhanced activity were BPXX34 ( Figure 2 B). The above results indicate that the recombination efficiency of the existing WT_attP module has reached a relatively ideal level. Therefore, further optimization of recombination activity mainly focuses on adjusting the attB sequence.
[0144] Table 1. List of AT site sequences for recombinant activity
[0145]
[0146] Example 2: Validation of the activity of a rationally designed novel attB site in site-directed large fragment recombination.
[0147] In Example 1, wild-type attB and wild-type attP were split into eight parts: motif1-L, motif1-R, motif2-L, and motif2-R. In this example, a large-scale rational design was performed on wild-type attB, replacing the corresponding positions of the four parts with their corresponding motif sequences. The specific replacement sequences are shown in Table 2. When replacing the left and right corresponding sequences, inverse complementary sequences are required. Since the inverse complementary sequences of attB-motif2-L and attB-motif2-R are the same, the motif2-L / R module has only three motif sequences to be replaced. As shown in Table 2, after permutations and combinations, there are 144 possible combinations, resulting in 143 rationally designed attB sequences.
[0148] To verify the application effect of novel non-natural attB sites in achieving large-fragment genome recombination at specific sites, recombination activity verification system 2 was used for activity verification. First, an att site sequence (attP) recognized by Bxb1 was inserted into the genome of eukaryotic cells (e.g., rice or human cells) using a guided editing tool. Then, the attB site sequence was embedded in the donor vector. When Bxb1 or its protein variants are expressed in the cells, they recognize the att site pair, thereby initiating a recombination reaction and ultimately integrating the target gene from the donor into the cellular genome. Recombination activity verification system 2 reference: Sun, Chao et al. “Precise integration of large DNA sequences in plant genomes using PrimeRoot editors.” Nature biotechnology vol. 42,2 (2024). The insertion efficiency of guided editing and the recombination efficiency of Bxb1 were detected by next-generation sequencing. The next-generation sequencing method was referenced from Pandey S, Gao XD, Krasnow NA, et al. Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing[J]. Nature Biomedical Engineering, 2024: 1-18.
[0149] This system was used to validate rationally designed novel non-natural attB sites. The donor vector contained the novel non-natural attB site whose activity needed to be validated, and another attP site recognized by Bxb1 was inserted into the genome, which was a wild-type attP sequence (SEQ ID NO:6). The average of three biological replicates of the fold increase in activity relative to the Bxb1 wild-type attB (WT_attB; SEQ ID NO:5) was taken, and the activity results were normalized. The recombination activity of WT_attB was taken as 1, and the recombination activity of other rationally designed attB sites was taken as a fold of 1 for comparison.
[0150] The results are shown in Table 3, which summarizes all rationally designed attB sites with higher recombinant activity than WT_attB. Table 3 shows that the preferred attB variants of the rational design have 1.04-4.75 times the recombinant activity of wild-type attB in tobacco, with attB variants V1320, V1402, V1410, and V1414 showing significantly enhanced recombinant activity in tobacco.
[0151] Table 2. Motif sequences to be replaced for each module in attB
[0152]
[0153] Table 3. attB sites for screening large fragment recombination activity
[0154]
[0155]
[0156]
[0157] Example 3: Validation of rationally designed attB sites by different recombinases
[0158] The novel attB site with enhanced recombination activity, as described in Table 3 of Example 2, was used to verify recombination activity again in different genomes using different recombinases. The rice genome selected a safe harbor site, GSH1, and the integration vector was a 5.6kb donor vector (exemplary examples show vectors with the WT_attB site as shown in SEQ ID NO:164; for other experiments, the donor vectors used only had the attB sequence replaced). The PE tool used was ePPE (SEQ ID NO:161), and the double pegRNAs are shown in SEQ ID NO:162 and SEQ ID NO:163. The specific binding sequence of the F primer used to detect efficiency in rice is shown in SEQ ID NO:166, and the R primer sequence inserted into the donor vector used in rice is shown in SEQ ID NO:167. For human HEK293T cells, the TRAC site was selected to test the recombination activity of the variant against the 5.6kb donor vector (exemplary examples show vectors with the WT_attB site as shown in SEQ ID NO:157; for other experiments, the donor vectors used only had the attB sequence replaced). The PE tool used was PE2 (SEQ ID NO:168), and the double pegRNAs are shown in SEQ ID NO:155 and SEQ ID NO:166. As shown in ID NO:156, the specific binding sequence of the F primer used for efficiency detection in human cells is shown in SEQ ID NO:159, and the R primer sequence inserted into the donor vector used in human cells is shown in SEQ ID NO:160. In this embodiment, the different recombinases verified, in addition to wild-type Bxb1 (WT_Bxb1), also include eeBxb1 (SEQ ID NO:2; Pandey S, Gao XD, Krasnow NA, et al. Efficient site-specific integration of large genes in mammalian cells via continuously evolved recombinases and prime editing[J]. Nature Biomedical Engineering, 2024: 1-18.), Bxb1-hV30 (SEQ ID NO:152), Bxb1-V27 (SEQ ID NO:153), and Bxb1-V39 (SEQ ID NO:154).
[0159] Figure 3The recombination efficiency of TRAC sites in human cells was demonstrated. Rationally designed attB sites V1246, V1320, V1402, V1410, V1414, V1434 and V1435 were applicable to different recombinases such as WT_Bxb1, Bxb1-hV30 and eeBxb1, and all showed a significant improvement in recombination efficiency compared to WT_attB. Figure 4 The recombination efficiency at the GSH1 site in rice cells was demonstrated. Rationally designed variants V1246 and V1410 were able to complete donor recombination in WT_Bxb1, Bxb1-V27, Bxb1-V39, and eeBxb1, and both showed a significant improvement in recombination efficiency compared to WT_attB.
[0160] In summary, the optimally designed attB site can produce enhanced recombination activity when using different Bxb1 recombinases on different genomes.
[0161] Example 4: Further optimization of rationally designed attB sites
[0162] In Example 1 Figure 2 The results of A further showed that the recombination activity of the att combination BP2222 was almost unchanged or even slightly lower than that of the wild-type att combination. BP2 refers to the motif1-R+motif2-R combination derived from attB. Therefore, the inventors made a similar assumption based on the "barrel principle", that is, the sequence adjustment of the right half of attB will improve the overall recombination activity of attB.
[0163] To verify the above hypothesis, the inventors used a fixed sequence of positions 12-27 in WT_attB ( The rightmost positions 25-27 of (SEQ ID NO:169) are incorporated into the adjustment range (hereinafter referred to as motif3), and a novel attB site variant library is constructed by saturating mutations of six bases, including motif3 and motif2-R. The recombination activity is verified using recombination activity verification system 1, and the specific experimental method is described in Example 1. Table 4 shows the novel attB site variants with enhanced activity compared to WT_attB, and the statistical results of the activity enhancement are as follows: Figure 5 As shown in the figure. The results showed that the preferred novel attB variants exhibited 2.2-4.1 times the recombinant activity of natural attB in tobacco, with attB-v5 exhibiting the highest recombinant activity. Therefore, optimizing only the motif sequence of the right half of attB can effectively enhance the recombinant activity of attB in recombinases.
[0164] The recombination efficiency of different Bxb1 proteins in novel attB site variants obtained through screening was further verified. Wild-type Bxb1, eeBxb1, Bxb1-A315R (Rose J, Kim H, Von Stetina J, et al. Engineered Bxb1 variants improve integrase activity and fidelity[J].bioRxiv, 2024: 2024.10. 21.619419.), and Bxb1-T84N+E69M were selected, with amino acid sequences shown in SEQ ID NO:1-4, respectively. Experimental results are as follows: Figure 6 As shown, attB-v4 and attB-v5 also enhanced the recombination activity of various Bxb1 variants.
[0165] Table 4 Preferred novel attB variants
[0166]
[0167] In summary, based on all the rational designs of the attB site in Examples 1-4 above, the positions of the novel attB site and variant alternative local motifs relative to WT_attB, as well as the alternative sequences, are shown in Table 5.
[0168] Table 5 List of local motif sequences
[0169]
[0170] The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
[0171] The present invention relates to the following sequences:
[0172] >SEQ ID NO:1 wt_Bxb1
[0173] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0174] >SEQ ID NO:2 eeBxb1
[0175] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDAIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGRKWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAIELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0176] >SEQ ID NO:3 Bxb1(A315R)
[0177] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFRGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0178] >SEQ ID NO:4 Bxb1(T84N+E69M)
[0179] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEMQPFDVIVAYRVDRLNRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0180] >SEQ ID NO:5 WT_attB
[0181] GGCTTGTCGACGACGGCGGACTCCGTCGTCAGGATCAT
[0182] >SEQ ID NO:6 WT_attP
[0183] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0184] >SEQ ID NO:7 BP1111_attB
[0185] GGCTTGTCGACGACGGCGGACTCCGTCGTCGACAAGCC
[0186] >SEQ ID NO:8 BP2222_attB
[0187] ATGATCCTGACGACGGCGGACTCCGTCGTCAGGATCAT
[0188] >SEQ ID NO:9 BP3333_attB
[0189] GGTTTGTCAACGACGGCGGACTCCGTCGTTGACAAACC
[0190] >SEQ ID NO:10 BP4444_attB
[0191] GGTTTGTACACGACGGCGGACTCCGTCGTGTACAAACC
[0192] >SEQ ID NO:11 BP4434_attB
[0193] GGTTTGTACACGACGGCGGACTCCGTCGTGTACAAACC
[0194] >SEQ ID NO:12 BP3434_attB
[0195] GGTTTGTCAACGACGGCGGACTCCGTCGTGTACAAACC
[0196] >SEQ ID NO:13 BP2424_attB
[0197] ATGATCCTGACGACGGCGGACTCCGTCGTgTACAAACC
[0198] >SEQ ID NO:14 BP1414_attB
[0199] GGCTTGTCGACGACGGCGGACTCCGTCGTGTACAAACC
[0200] >SEQ ID NO:15 BP3334_attB
[0201] GGTTTGTCAACGACGGCGGACTCCGTCGTTGACAAACC
[0202] >SEQ ID NO:16 BP1334_attB
[0203] GGCTTGTCGACGACGGCGGACTCCGTCGTtGACAAACC
[0204] >SEQ ID NO:17 BP1434_attB
[0205] GGCTTGTCGACGACGGCGGACTCCGTCGTgTACAAACC
[0206] >SEQ ID NO:18 v1_attB
[0207] GGCTTGTCGACGACGGCGgaCTCCGCTGAGAGGATCAT
[0208] >SEQ ID NO:19 v2_attB
[0209] GGCTTGTCGACGACGGCGgaCTCCCCTGTGAGGATCAT
[0210] >SEQ ID NO:20 v3_attB
[0211] GGCTTGTCGACGACGGCGgaCTCCGAAGTGAGGATCAT
[0212] >SEQ ID NO:21 v4_attB
[0213] GGCTTGTCGACGACGGCGgaCTCCTTTGTGAGGATCAT
[0214] >SEQ ID NO:22 v5_attB
[0215] GGCTTGTCGACGACGGCGgaCTCCGCAGCGAGGATCAT
[0216] >SEQ ID NO:23 BP1111_attP
[0217] GGCTTGTCTGGTCGACCACCGCGGACTCAGTGGTCTACGGGACAAGCC
[0218] >SEQ ID NO:24 BP2222_attP
[0219] ATGATCCTTGGTCGACCACCGCGGACTCAGTGGTCTACGGAGGATCAT
[0220] >SEQ ID NO:25 BP3333_attP
[0221] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTTTACGGGACAAACC
[0222] >SEQ ID NO:26 BP4444_attP
[0223] GGTTTGTATGGTCCACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0224] >SEQ ID NO:27 BP4434_attP
[0225] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0226] >SEQ ID NO:28 BP3434_attP
[0227] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0228] >SEQ ID NO:29 BP2424_attP
[0229] ATGATCCTTGGTCGACCACCGCGGACTCAGTGGTgTACGGTACAAACC
[0230] >SEQ ID NO:30 BP1414_attP
[0231] GGCTTGTCTGGTCGACCACCGCGGACTCAGTGGTgTACGGTACAAACC
[0232] >SEQ ID NO:31 BP3334_attP
[0233] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTgTACGGTACAAACC
[0234] >SEQ ID NO:32 BP1334_attP
[0235] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0236] >SEQ ID NO:33 BP1434_attP
[0237] GGTTTGTCTGGTCAACCACCGCGGACTCAGTGGTGTACGGTACAAACC
[0238] >SEQ ID NO:152 Bxb1-hV30
[0239] MRALVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVICAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0240] >SEQ ID NO:153 Bxb1-V27
[0241] MRGLVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0242] >SEQ ID NO:154 Bxb1-V39
[0243] MRAMVVIRLSRVTDATTSPERQLESCQQLCAQRGWDVVGVAEDLDVSGAVDPFDRKRRPNLARWLAFEEQPFDVIVAYRVDRLTRSIRHLQQLVHWAEDHKKLVVSATEAHFDTTTPFAAVVIALMGTVAQMELEAIKERNRSAAHFNIRAGKYRGSLPPWGYLPTRVDGEWRLVPDPVQRERILEVYHRVVDNHEPLHLVAHDLNRRGVLSPKDYFAQLQGREPQGREWSATALKRSMISEAMLGYATLNGKTVRDDDGAPLVRAEPILTREQLEALRAELVKTSRAKPAVSTPSLLLRVLFCAVCGEPAYKFAGGGRKHPRYRCRSMGFPKHCGNGTVAMAEWDAFCEEQVLDLLGDAERLEKVWVAGSDSAVELAEVNAELVDLTSLIGSPAYRAGSPQREALDARIAALAARQEELEGLEARPSGWEWRETGQRFGDWWREQDTAAKNTWLRSMNVRLTFDVRGGLTRTIDFGDLQEYEQHLRLGSVVERLHTGMS
[0244] >SEQ ID NO:155 epegRNA1 for human cells
[0245]
[0246] >SEQ ID NO:156 Human cells using epeRNA2
[0247]
[0248] >SEQ ID NO:157 Human donor vector containing the WT_attB site (5.6kb)
[0249]
[0250] >SEQ ID NO:158 Vector for expressing WT_Bxb1 in human cells
[0251]
[0252] >SEQ ID NO:159 Specific binding sequence of the F primer used for efficiency detection in human cells
[0253] ACTTGCCAGCCCCACAGAG
[0254] >SEQ ID NO:160 The R primer sequence inserted into the donor vector used in human cells
[0255] TCCAGTGACAAGTCTGTCTGC
[0256] >SEQ ID NO:161 ePPE carrier
[0257]
[0258] >SEQ ID NO:162, epigRNA1, for plant cell use
[0259]
[0260] >SEQ ID NO:163, epigRNA2, for plant cell use
[0261]
[0262] >SEQ ID NO:164 Rice-5.6kb donor vector containing the WT_attB site
[0263]
[0264] >SEQ ID NO:165 Vector expressing WT_Bxb1 in rice
[0265]
[0266] >SEQ ID NO:166 Specific binding sequence of the F primer used for efficiency detection in rice
[0267] CTCATGTGCATGGAAGCATC
[0268] >SEQ ID NO:167 R primer sequence inserted into the donor vector used in rice
[0269] CTTGCTCCTTCTCACCGTG
[0270] >SEQ ID NO:168 PE2 vector, for human cell use
[0271]
[0272] >SEQ ID NO:169
[0273] GACGGCGGWCTCC
[0274] >SEQ ID NO:170
[0275] GGCTTGTCGACGACGGCGGWCTCCGTCGTCAGGATCAT
[0276] >SEQ ID NO:171
[0277] GGTTTGTCTGGTCAACCACCGCGGWCTCAGTGGTGTACGGTACAAACC。
Claims
1. A nucleic acid molecule of a non-natural recombinase attB recognition site, the nucleic acid molecule of a non-natural recombinase attB recognition site comprising the base sequence: [motifl-L][motif2-L]GACGGCGGWCTCC[motif3][motif2-R][motifl-R]; [motifl-L] is selected from at least one of SEQ ID NOs: 34-37; [motif2-L] is selected from at least one of SEQ ID NOs: 38-40; [motif3] is selected from at least one of SEQ ID NOs: 41-46; [motif2-R] is selected from at least one of SEQ ID NOs: 47-51; [motifl-R] is selected from at least one of SEQ ID NOs: 52-55; W is base A or T. wherein the recombinase is selected from BxBl, BxBl variants, BxZ2, PhiC31, Peaches, Veracruz, Rebeuca, Theia, Benedict, PattyP, Trouble, KSSJEB, Lockley, Scowl, Switzer, Bob3, Abrogate, Doom, Concept II, Anglerfish, SkiPole, Museum, and Severus, preferably the recombinase is BxBl and / or the BxBl variants.
2. The nucleic acid molecule of claim 1, wherein, the BxBl and / or the BxBl variants comprise an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-4, SEQ ID NOs: 152-154.
3. The nucleic acid molecule of claim 2, wherein, the nucleic acid molecule of a non-natural recombinase attB recognition site comprises a nucleotide sequence selected from the group consisting of SEQ ID NOs: 7-22, SEQ ID NOs: 56-151.
4. The nucleic acid molecule of claim 1, wherein, the nucleic acid molecule of a non-natural recombinase attB recognition site comprises a nucleotide sequence selected from the group consisting of SEQ ID NO: 9, SEQ ID NOs: 18-22, SEQ ID NOs: 56-59, or SEQ ID NOs: 116-117.
5. The nucleic acid molecule of claim 1, wherein the method comprises i) and ii):
6. A method of site-specific recombination, characterized in that, i) providing a nucleic acid molecule of a non-natural recombinase attB recognition site according to any one of claims 1-5; ii) providing a nucleic acid molecule of a recombinase attP recognition site; combining the nucleic acid molecules of i) and ii) into a paired recombinase att recognition site in the presence of a recombinase. the nucleic acid molecule of a recombinase attP recognition site is selected from a natural recombinase attP recognition site or a non-natural recombinase attP recognition site.
7. The method of claim 6, wherein, the nucleotide sequence of the natural recombinase attP recognition site is set forth in SEQ ID NO: 6; the nucleotide sequence of the non-natural recombinase attP recognition site is selected from the group consisting of the nucleotide sequences set forth in SEQ ID NOs: 23-33.
8. The method of claim 7, wherein, 9. The method according to any one of claims 6-8, characterized in that, The nucleic acid molecule in i) or ii) is located in a plasmid, a bacteriophage, a bacterial genome, a fungal genome, a plant genome, or an animal genome.
10. A site-specific recombination system, characterized in that, The site-specific recombination system contains: a) the nucleic acid molecule of the unnatural recombinase attB recognition site according to any one of claims 1 to 5 or an expression construct thereof; b) the nucleic acid molecule of the attP recognition site or an expression construct thereof, wherein the nucleic acid molecule of the recombinase attP recognition site is selected from the group consisting of a natural recombinase attP recognition site, the nucleotide sequence of which is shown in SEQ ID NO: 6, or an unnatural recombinase attP recognition site, the nucleotide sequence of which is selected from the group consisting of the nucleotide sequences shown in SEQ ID NOs: 23 to 33; and c) a recombinase or an expression construct thereof.
11. The site-specific recombination system of claim 10, wherein, The site-specific recombination system further contains: d) a genome editing system for introducing the nucleic acid molecule of the unnatural recombinase attB recognition site according to any one of claims 1 to 5 or the nucleic acid molecule of the recombinase attP recognition site into the genome of a cell.
12. The site-specific recombination system of claim 11, wherein, The genome editing system is selected from the group consisting of a meganuclease, a zinc finger protein, a TALEN, a CRISPR-Cas, a topoisomerase, or a recombinase system mediated genome editing system.
13. A host cell, characterized in that, The host cell contains the nucleic acid molecule of the unnatural recombinase attB recognition site according to any one of claims 1 to 5 or the site-specific recombination system according to any one of claims 10 to 12.
14. The host cell of claim 13, wherein, The cell is a cell from a prokaryote and / or a eukaryote; The prokaryote comprises a bacterium; The eukaryote comprises a plant, a fungus, and / or a vertebrate.
15. The host cell of claim 14, wherein, The vertebrate comprises a mammal, the mammal comprises a human, a mouse, a rat, a monkey, a dog, a pig, a sheep, a cow, and / or a cat; The plant comprises a crop plant, the crop plant comprises a wheat, a rice, a maize, a soybean, a sunflower, a sorghum, a rape, an alfalfa, a cotton, a barley, a millet, a sugarcane, a tomato, a tobacco, a cassava, and / or a potato.
16. Use of the nucleic acid molecule of the unnatural recombinase attB recognition site according to any one of claims 1 to 5, the site-specific recombination system according to any one of claims 10 to 12, or the host cell according to any one of claims 13 to 15 for the manufacture of a medicament for modifying a target nucleic acid.
17. Use of the nucleic acid molecule of the unnatural recombinase attB recognition site according to any one of claims 1 to 5, the site-specific recombination system according to any one of claims 10 to 12, or the host cell according to any one of claims 13 to 15 for the manufacture of a medicament for the prevention and / or treatment of a disease associated with or caused by genetic recombination.
Citation Information
Patent Citations
Recombinases and target sequences
US10731153B2