Enzymes that mediate the integration of large DNA fragments into the mammalian genome and their applications

By developing novel integrase Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 and their derivatives, the problems of low integration efficiency and multi-site integration risk of existing integrase in the human genome have been solved, achieving efficient and stable integration of large DNA fragments and enhancing the potential of gene therapy.

CN120731269BActive Publication Date: 2026-04-21YOLTECH THERAPEUTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YOLTECH THERAPEUTICS CO LTD
Filing Date
2024-02-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing integrase has low integration efficiency in the human genome and poses potential risks of multi-site integration and chromosomal rearrangement, which limits its application in gene therapy.

Method used

A series of novel integrase enzymes, such as Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 and their derivatives, have been developed. These enzymes have highly efficient DNA integration capabilities, especially in mammalian genomes, where they can mediate large-fragment integration and achieve site-specific integration by binding to specific recognition sequences.

Benefits of technology

It improves integration efficiency, reduces the risk of multi-site integration and chromosomal rearrangement, and enhances the effectiveness of gene therapy, especially by achieving stable integration of large DNA fragments in mammalian cells.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HSB0000212684250000011
    Figure HSB0000212684250000011
  • Figure HSB0000212684250000012
    Figure HSB0000212684250000012
  • Figure HSB0000212684250000021
    Figure HSB0000212684250000021
Patent Text Reader

Abstract

An enzyme capable of mediating the integration of large DNA fragments into the mammalian genome and its applications are provided. Specifically, novel enzymes capable of integrating large DNA fragments into the genome are provided, including integrases as shown in SEQ ID NO: 1-7. Further, polynucleotides containing nucleotide sequences encoding integrases and host cells containing these polynucleotides are disclosed. A method for recombining target nucleic acids into the human genome using said integrase is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene editing, and more specifically, to enzymes that mediate the integration of large DNA fragments into the mammalian genome and their applications. Background Technology

[0002] Integrase is a recombinase derived from bacteriophages that mediates recombination between the attP site on the bacteriophage genome and the attB site on the bacterial genome, thereby specifically integrating the bacteriophage genomic site into the genome. Studies have found that mammalian genomes contain pseudo attP sites homologous to attP sites. These sites can also recombine with plasmids carrying attB sites under the action of integrase, thus integrating foreign genes from the plasmid into the genome.

[0003] PhiC31 integrase is an integrase derived from PhiC31 bacteriophage that mediates recombination between the attP site on the phage's own circular genome and the attB site on the circular genome of *Streptomyces pulveratum*. Utilizing this principle, researchers first successfully integrated plasmids containing the attB site sequence into genomes containing the attP site under the catalysis of PhiC31 integrase in bacteria, yeast, and mammalian cells. This method has been used in preclinical studies of hemophilia A and hemophilia B, showing that it effectively increases clotting factor levels in hemophilia A and B mice, achieving long-term improvement in disease symptoms. Integrases possess characteristics such as unidirectional integration, relatively specific sites, no need for cofactors, and the ability to mediate long-term fragment integration, demonstrating great potential for application in gene therapy.

[0004] However, existing integrases have some drawbacks, such as low integration efficiency in the human genome, multi-site integration, and potential risks such as chromosomal rearrangement.

[0005] Therefore, there is a great need in this field to develop new integrase with advantages such as high integration efficiency. Summary of the Invention

[0006] This invention provides enzymes that can mediate the integration of large DNA fragments into the mammalian genome and their applications.

[0007] In a first aspect of the invention, an integrase is provided, said integrase being selected from the group consisting of:

[0008] (a) Selected from any of the integrases shown in SEQ ID NO:1-7;

[0009] (b) The amino acid sequence has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with any of the amino acid sequences shown in SEQ ID NO:1-7, and substantially retains the biological function of any of SEQ ID NO:1-7;

[0010] (c) A polypeptide derived from (a) having integrase function, formed by substituting, deleting or adding one or more amino acid residues of any of the amino acid sequences shown in SEQ ID NO:1-7.

[0011] In another preferred embodiment, the amino acid sequence of the integrase is as shown in any of SEQ ID No:1-7.

[0012] In another preferred embodiment, the biological function is the function of integrating foreign genes into the genome.

[0013] In another preferred embodiment, the integrase is Cyin14.

[0014] In a second aspect of the invention, a polynucleotide is provided, said polynucleotide encoding the integrase described in the first aspect of the invention.

[0015] In another preferred embodiment, the nucleotide sequence of the polynucleotide is shown in SEQ ID No:10-16.

[0016] In a third aspect of the invention, a carrier is provided, the carrier comprising the polynucleotide described in the second aspect of the invention.

[0017] In another preferred embodiment, the polynucleotide molecule is operatively linked to a regulatory sequence to allow the expression of the polynucleotide molecule.

[0018] In another preferred embodiment, the regulatory sequence includes a promoter sequence.

[0019] In a fourth aspect of the invention, a genetically engineered host cell is provided, containing a vector as described in the third aspect of the invention, or having its genome integrated with polynucleotides as described in the second aspect of the invention.

[0020] In another preferred embodiment, the host cell expresses an integrase as described in the first aspect of the invention.

[0021] In another preferred embodiment, the host cell includes eukaryotic cells and prokaryotic cells.

[0022] In another preferred embodiment, the host cell is a mammalian cell.

[0023] In a fifth aspect of the invention, a method for preparing the integrase described in the first aspect of the invention is provided, the method comprising:

[0024] (a) Culturing the host cells described in the fourth aspect of the present invention under suitable expression conditions;

[0025] (b) Isolate the integrase from the culture.

[0026] In a sixth aspect of the invention, a cell engineered using the integrase of the invention is provided, the cell comprising:

[0027] (i) an integrated sequence selected from the group consisting of: any nucleotide sequence shown in SEQ ID NO:17-42, or a nucleotide sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity with it, or a nucleotide sequence having no more than 1, 2, 3, 4, 5, 6, 7, or 8 sequence alterations (e.g., substitution, insertion, or deletion) relative to any nucleotide sequence shown in SEQ ID NO:17-42; and

[0028] (ii) Heterogeneous target nucleic acid.

[0029] In another preferred embodiment, the cell is a eukaryotic cell or a prokaryotic cell.

[0030] In another preferred embodiment, the cell is a eukaryotic cell.

[0031] In another preferred embodiment, the cell is a human or non-human mammalian cell.

[0032] In another preferred embodiment, the heterologous target nucleic acid is integrated into the genome of the cell via an integrase according to the first aspect of the invention.

[0033] In another preferred embodiment, the integrase is Cyin14.

[0034] In another preferred embodiment, when the integrase is Cyin14, the corresponding integration sequence contains the integration sequence 14-P1 (SEQ ID No:31).

[0035] In a seventh aspect of the invention, a recombination method is provided, the method comprising:

[0036] (a) Provide the target nucleic acid to be integrated, and the genome to be integrated;

[0037] (b) The presence of the integrase described in the first aspect of the invention enables the target nucleic acid to be integrated into the genome.

[0038] In another preferred embodiment, the target nucleic acid to be integrated contains one or more recognition sequences (or integration sequences).

[0039] In another preferred embodiment, the recognition sequence is located upstream (or immediately upstream), downstream (or immediately downstream) or on both sides of the target nucleic acid sequence to be integrated.

[0040] In another preferred embodiment, the recognition sequence (or integration sequence) on the target nucleic acid to be integrated and the recognition sequence (or false recognition sequence) on the genome can mediate site-specific integration under the action of the integrase described in the first aspect of the invention.

[0041] In another preferred embodiment, the false identification sequence is located in or near a genomic safe harbor site.

[0042] In another preferred embodiment, in step (b), the target nucleic acid, the genome, and the integrase are contacted.

[0043] In another preferred embodiment, in step (b), the target nucleic acid, the cell containing the genome, and the integrase are contacted.

[0044] In another preferred embodiment, the method includes in vitro methods and in vivo methods.

[0045] In another preferred embodiment, the method is non-diagnostic and non-therapeutic.

[0046] In an eighth aspect of the invention, a reaction system for gene integration is provided, the reaction system comprising:

[0047] (a) The target nucleic acid to be integrated;

[0048] (b) The genome to be integrated; and

[0049] (b) The integrase described in the first aspect of the present invention.

[0050] In a ninth aspect of the invention, a kit for detecting a target nucleic acid in a sample is provided, the kit comprising:

[0051] (a) the integrase described in the first aspect of the present invention; (b) the polynucleotide described in the second aspect of the present invention; or (c) the vector described in the third aspect of the present invention.

[0052] In a tenth aspect of the invention, the use of the integrase described in the first aspect of the invention, or the polynucleotide described in the second aspect of the invention, or the vector described in the third aspect of the invention, or the host cell described in the fourth aspect of the invention, in the preparation of a formulation or kit, said formulation or kit being used for:

[0053] (i) Gene or genome editing;

[0054] (ii) Editing target sequences in target loci to modify biological or non-human organisms.

[0055] In an eleventh aspect of the present invention, a method is provided for specifically integrating exogenous nucleic acid sites into the cellular genome or intracellular target nucleic acid, the method comprising:

[0056] (a) Introducing (i) the first construct, (ii) gRNA, (iii) the exogenous nucleic acid to be integrated, and (iv) the second construct into cells.

[0057] (i) A first construct, wherein the first construct is an expressible polynucleotide construct encoding an editing polypeptide, wherein the editing polypeptide comprises a DNA-binding nuclease domain connected via a linker to a reverse transcriptase domain, wherein the DNA-binding nuclease domain comprises nicking enzyme activity; and

[0058] (ii) gRNA, wherein the gRNA is a guide RNA containing a complementary sequence comprising a targeting sequence, a primer-binding sequence, and an integration sequence.

[0059] The gRNA interacts with the expressed editing peptide, thereby targeting the editing peptide to a specific target site in the cellular genome or intracellular target nucleic acid. The DNA-binding nuclease domain of the editing peptide cleaves one strand of the cellular genome or intracellular target nucleic acid, and the reverse transcriptase domain incorporates the integration sequence into the cleavage site, thereby incorporating at least one integration sequence into the cellular genome or target nucleic acid at a specific target site.

[0060] (iii) A foreign nucleic acid to be integrated, which is ligated to a recognition sequence homologous to the integration sequence, wherein, in the presence of both the recognition sequence and the integration sequence, the integrase mediates site-specific integration; and

[0061] (iv) A second construct, wherein the second construct is an expressible polynucleotide construct encoding the integrase described in the first aspect of the present invention.

[0062] The integrase integrates exogenous nucleic acids into specific target sites in the cell genome;

[0063] (b) In the presence of the editing polypeptide, gRNA, integrase and the exogenous nucleic acid to be integrated, the integrase integrates the exogenous nucleic acid into the cellular genome or intracellular target nucleic acid.

[0064] In another preferred embodiment, (i) the first construct, (ii) the gRNA, (iii) the exogenous nucleic acid to be integrated, and (iv) the second construct can be introduced into cells independently, in batches, or simultaneously.

[0065] In another preferred embodiment, the integrated sequence is a sequence selected from the group consisting of: SEQ ID NO:17-42, 47-48.

[0066] In another preferred embodiment, (i) a first construct, (ii) gRNA and (iv) a second construct are integrated into the cellular genome using the following methods: virus, lipid, microvesicle, gene gun, or a combination thereof.

[0067] In another preferred embodiment, the gRNA hybridizes with the complementary strand of the cell genome.

[0068] In another preferred embodiment, the exogenous nucleic acid to be integrated is integrated into the cell genome via microcircular DNA, plasmid, mRNA, linear DNA, or a combination thereof.

[0069] In another preferred embodiment, the DNA-binding nuclease domain containing nicking enzyme activity is selected from: nicking enzymes of Cas9, nicking enzymes of Cas12a, nicking enzymes of Cas12b, nicking enzymes of Cas12c, nicking enzymes of Cas12d, nicking enzymes of Cas12e, nicking enzymes of Cas12f, nicking enzymes of Cas12g, nicking enzymes of Cas12h, nicking enzymes of Cas12i, nicking enzymes of Cas12j, nicking enzymes of Cas12k, nicking enzymes of Cas12l, nicking enzymes of Cas12m, nicking enzymes of Cas12n, or combinations thereof.

[0070] In another preferred embodiment, the reverse transcriptase domain contains a mutation relative to the wild-type sequence.

[0071] In another preferred embodiment, the reverse transcriptase domain is selected from: the reverse transcriptase domain of Moloney murine leukemia virus (M-MLV), the transcriptase heteropolymerase (RTX), the avian myeloma virus reverse transcriptase (AMV-RT), the rectal eubacterium maturation enzyme RT (MarathonRT), or a combination thereof.

[0072] In another preferred embodiment, the M-MLV reverse transcriptase domain contains mutations selected from the following: D200N, T306K, W313F, T330P, L603W, or combinations thereof.

[0073] In another preferred embodiment, the exogenous nucleic acid to be integrated is selected from the group consisting of: β-hemoglobin (HBB) gene, gene causing β-thalassemia or sickle cell anemia, metabolic gene, gene of infectious disease, gene of hereditary disease, or a combination thereof.

[0074] In a twelfth aspect of the invention, a polypeptide (such as a fusion protein) is provided, comprising a DNA-binding nuclease and an integrase or an active fragment thereof as described in the first aspect of the invention.

[0075] In another preferred embodiment, the fusion protein comprises:

[0076] a) An integrase or an active fragment thereof having the first amino acid sequence as described in the first aspect of the present invention;

[0077] b) A CRISPR-associated (Cas) protein having a second amino acid sequence or an active fragment thereof; and

[0078] c) Nuclear localization signal (NLS) with a third amino acid sequence.

[0079] In another preferred embodiment, the fusion protein comprises a DNA-binding nuclease, the integrase described in the first aspect, and a reverse transcriptase.

[0080] In another preferred embodiment, the fusion protein comprises, from the N-terminus to the C-terminus, a DNA-binding nuclease, a reverse transcriptase, and the integrase described in the first aspect.

[0081] In another preferred embodiment, the reverse transcriptase is connected to the integrase described in the first aspect of the invention via a linker.

[0082] In another preferred embodiment, the DNA-binding nuclease is linked to the reverse transcriptase via a adapter.

[0083] In another preferred embodiment, the connector may be cuttable or non-cuttable.

[0084] In another preferred embodiment, the DNA-binding nuclease comprises nicking enzyme activity.

[0085] In another preferred embodiment, the DNA-binding nuclease is selected from: Cas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas12l, Cas12m, Cas12n, or combinations thereof.

[0086] In another preferred embodiment, the reverse transcriptase is selected from: the reverse transcriptase domain of Moloney murine leukemia virus (M-MLV), transcription heteropolymerase (RTX), avian myeloma virus reverse transcriptase (AMV-RT), or Marathon RT, or a combination thereof.

[0087] In another preferred embodiment, the reverse transcriptase is an M-MLV modified relative to wild-type M-MLV reverse transcriptase.

[0088] In another preferred embodiment, the M-MLV reverse transcriptase domain comprises a mutation selected from the group consisting of D200N, T306K, W313F, T330P, L603W, or a combination thereof.

[0089] In a thirteenth aspect of the present invention, a sequence integration system is provided, the integration system comprising:

[0090] (i) A first construct, wherein the first construct is an expressible polynucleotide construct encoding an editing polypeptide, wherein the editing polypeptide comprises a DNA-binding nuclease domain connected to a reverse transcriptase domain via a linker, wherein the DNA-binding nuclease domain comprises nicking enzyme activity;

[0091] (ii) gRNA, wherein the gRNA is a guide RNA containing a complementary sequence comprising a targeting sequence, a primer-binding sequence, and an integration sequence.

[0092] The gRNA interacts with the expressed editing peptide to target the editing peptide to a specific target site of the cell genome or intracellular target nucleic acid, and the DNA-binding nuclease domain of the editing peptide cleaves one strand of the cell genome or intracellular target nucleic acid, and the reverse transcriptase domain incorporates the integration sequence into the cleavage site, thereby incorporating at least one integration sequence into the cell's specific target site genome or target nucleic acid.

[0093] (iii) a foreign nucleic acid to be integrated, wherein the foreign nucleic acid to be integrated is linked to a recognition sequence homologous to the integration sequence, the recognition sequence and the integration sequence being capable of mediating site-specific integration; and

[0094] (iv) A second construct, wherein the second construct encodes the integrase described in the first aspect of the invention.

[0095] The integrase integrates exogenous nucleic acids into specific target sites in the cell genome.

[0096] In another preferred embodiment, the integration is an in vitro method.

[0097] In a fourteenth aspect of the invention, a carrier is provided comprising nucleic acid encoding the fusion protein described in the twelfth aspect of the invention.

[0098] In a fifteenth aspect of the invention, a gRNA that specifically binds to a DNA-binding nuclease is provided, comprising (a) a primer binding site for hybridizing to a nicked DNA strand; (b) an integration sequence of an integrase; and (c) a targeting sequence.

[0099] In another preferred embodiment, the integrase is Cyin14.

[0100] In another preferred embodiment, when the integrase is Cyin14, the corresponding integration sequence contains the integration sequence 14-P1 (SEQ ID No:31).

[0101] In a sixteenth aspect of the invention, a treatment method is provided, the method comprising integrating a foreign nucleic acid to be integrated into the genome of a desired object using the method described in the twelfth aspect of the invention. In another preferred embodiment, the object includes a human or a non-human mammal.

[0102] It should be understood that, within the scope of this invention, the above-described technical features of this invention and the technical features specifically described below (such as in the embodiments) can be combined with each other to form new or preferred technical solutions. Due to space limitations, they will not be described in detail here. Attached Figure Description

[0103] Figure 1 The functional domains of integrase Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 are shown.

[0104] Figure 2 The three-dimensional structures of integrase Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 are shown.

[0105] Figure 3 A flowchart of the integration experiment is shown.

[0106] Figure 4 The diagram shows how to replace the ABE8e nucleotide sequence at positions 409-5226 in the ABE8e plasmid with, for example, Cyin10. Other expression plasmids are constructed in the same way.

[0107] Figure 5 The donor plasmid expressing EGFP is shown, with the EGFP nucleotide sequence located between HindIII and KpnI restriction sites.

[0108] Figure 6 The integration efficiency of seven distinct donor plasmids, each capable of being mediated, was demonstrated on the genome of 293T cells.

[0109] Figure 7 The flow cytometry results for integrase Cyin14 and Cyin16, flow cytometry results for integrase phiC31, and flow cytometry results for the control group are shown.

[0110] Figure 8 The results show the integration efficiency of integrase Cyin14 and Cyin16 compared with the previously reported phiC31 and Bxb1.

[0111] Figure 9The phylogenetic tree of integrases Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 is shown. Detailed Implementation

[0112] Through extensive and in-depth research and numerous screenings, the inventors have developed a novel integrase for the first time. This integrase exhibits low homology to existing integrases but possesses excellent integration activity, particularly in mediating the integration of large DNA fragments into the mammalian genome. This invention is based on this foundation.

[0113] the term

[0114] The terms used in the technical solution of this invention are explained as follows:

[0115] Operable ligation: refers to the ligation of a target nucleotide sequence to a regulatory element in a manner that allows for the expression of the nucleotide sequence (e.g., when a vector is introduced into a host cell, in an in vitro transcription / translation system, or in the host).

[0116] attB / attP reaction: also known as B / P reaction, is a recombination reaction mediated between attB recognition sites and attP recognition sites.

[0117] att sites: These are attachment sites on DNA molecules used for attachment to DNA molecules or complexes.

[0118] Integrated: refers to the integration of foreign genes or nucleotide sequences into the host genome through covalent bonds formed with the host DNA.

[0119] Recognition site: also known as recombination site, refers to the nucleotide sequence that can be recognized by recombinase protein.

[0120] Target sequence: also known as target DNA, refers to a nucleotide sequence containing at least one recombinase recognition site. Target nucleotide sequences can be genes, expression cassettes, promoters, molecular markers, or parts of any of these.

[0121] Site-specific recombination: also known as sequence-specific recombination, refers to recombination between two nucleotide sequences, each containing at least one recognition site or at least one non-homologous site.

[0122] "Site specificity" refers to a specific nucleotide sequence, such as a specific location in the host cell's genome. The nucleotide sequence can be endogenous to the host cell, located in a natural location in the host genome or at another location in the genome, or it can be a heterologous nucleotide sequence that has been previously inserted into the host cell's genome by any of a variety of known methods.

[0123] Host cell: refers to a cell in which proteins and / or genetic material have been introduced. This includes not only the specific subject cell but also its progeny. Host cells can be isolated cells, cell lines, or living tissues grown in a culture. Host cells include animal or plant cells, such as human cells, non-human mammalian cells, insect cells, avian cells, rodent cells, corn cells, soybean cells, wheat cells, or rice cells.

[0124] Target nucleic acid: "Target nucleic acid" is used interchangeably with "target gene," "target sequence," and "target DNA," and includes any nucleotide sequence that imparts a desired characteristic to a cell when transferred to that cell. This desired characteristic includes (but is not limited to): viral resistance, insect resistance, antibiotic stress resistance, disease resistance, resistance to other pests, herbicide tolerance, improved nutritional value, improved performance in industrial processes, or altered reproductive capacity. Target nucleic acid can also be a sequence transferred to a cell line, mammal, or plant to produce a commercially valuable enzyme or metabolite.

[0125] As used herein, “target nucleic acid to be integrated,” “exogenous nucleic acid to be integrated,” “nucleic acid to be integrated,” “target nucleic acid to be integrated,” “target gene to be integrated,” “target sequence to be integrated,” and “target DNA to be integrated” are used interchangeably and all refer to the target nucleic acid to be integrated into the cell genome. Preferably, the nucleic acid is heterologous.

[0126] As used in this article, "heterogeneous" refers to DNA or RNA that is from a different source than the cell or cell genome to be integrated.

[0127] As used herein, the terms “recognition sequence” and “integration sequence” are used interchangeably and generally refer to a nucleic acid (e.g., DNA) sequence that is recognized (e.g., can be bound by) a recombinase polypeptide. For example, as described herein, in some cases, the recognition sequence comprises two recognition sequences, one located at the integration site (the site into which the nucleic acid is to be integrated) and the other adjacent to the target nucleic acid to which the integration site is to be introduced. The recognition sequences are generally referred to as attB and attP. The recognition sequences can be native or altered relative to the native sequence. The length of the recognition sequence can vary, but is typically from about 20 to about 200 nt, about 30 to 90 nt, and more typically 30 to 70 nucleotides. Typically, the target nucleic acid to be integrated contains one or more recognition sequences (or integration sequences). The recognition sequence can be located upstream (or immediately upstream), downstream (or immediately downstream), or flanking the target nucleic acid sequence to be integrated. The recognition sequence in the recombinant vector cooperates with the recognition sequence (or integration sequence) on the genome to integrate the target nucleic acid sequence to be integrated into the predetermined location. It should be understood that recognition sequences exist in the genomes of many organisms, and these recognition sequences do not necessarily have the same nucleotide sequence as the wild-type recognition sequence (for a given recombinase); however, such natural recognition sequences are still sufficient to promote recombinase-mediated recombination.

[0128] Integrase and its applications

[0129] In this invention, novel integrases are isolated, including Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16, and their derivative polypeptides having the same or substantially the same integrative activity. This invention also provides uses for these novel integrases, particularly in mediating the integration of large DNA fragments into the mammalian genome.

[0130] As used in this article, "isolated" means that a substance has been separated from its original environment (in the case of a natural substance, the original environment is the natural environment). For example, polynucleotides and polypeptides in their natural state within living cells are not isolated and purified, but the same polynucleotides or polypeptides are isolated and purified if they are separated from other substances present in their natural state.

[0131] As used herein, "isolated integrase or polypeptide" means an integrase polypeptide that is substantially free of other naturally occurring or associated proteins, lipids, carbohydrates, or other substances. Those skilled in the art can purify integrases using standard protein purification techniques. A substantially pure polypeptide will produce a single master band on a non-reducing polyacrylamide gel. The purity of the integrase polypeptide can be determined by amino acid sequence analysis.

[0132] The polypeptides of the present invention can be recombinant polypeptides, natural polypeptides, or synthetic polypeptides, with recombinant polypeptides being preferred. The polypeptides of the present invention can be naturally purified products, chemically synthesized products, or produced from prokaryotic or eukaryotic hosts (e.g., bacteria, yeast, higher plants, insects, and mammalian cells) using recombinant technology. Depending on the host used in the recombinant production protocol, the polypeptides of the present invention can be glycosylated or non-glycosylated. The polypeptides of the present invention may or may not include an initial methionine residue.

[0133] This invention also includes fragments, derivatives, and analogs of integrase. As used herein, the terms “fragment,” “derivative,” and “analyte” refer to polypeptides that substantially retain the same biological function or activity as the natural integrase of this invention. The polypeptide fragments, derivatives, or analogs of this invention may be (i) polypeptides in which one or more conserved or non-conserved amino acid residues (preferably conserved amino acid residues) are substituted, and such substituted amino acid residues may or may not be encoded by the genetic code; or (ii) polypeptides having substituent groups in one or more amino acid residues; or (iii) polypeptides formed by fusing a mature polypeptide with another compound (e.g., a compound that extends the half-life of the polypeptide, such as polyethylene glycol); or (iv) polypeptides formed by fusing an additional amino acid sequence to this polypeptide sequence (e.g., a leader sequence or secretion sequence, or a sequence used to purify this polypeptide, or a proteogen sequence, or a fusion protein formed with an antigen IgG fragment). Based on the teachings herein, these fragments, derivatives, and analogs are within the scope well known to those skilled in the art.

[0134] In this invention, the term "integrase polypeptide" refers to a polypeptide having an integrase activity of any of the amino acid sequences shown in SEQ ID NO:1-7. This term also includes variations of the amino acid sequences shown in SEQ ID NO:1-7 having the same function as an integrase. These variations include (but are not limited to): deletions, insertions, and / or substitutions of one or more amino acids (typically 1-50, preferably 1-30, more preferably 1-20, most preferably 1-10), and the addition of one or more amino acids (typically up to 20, preferably up to 10, more preferably up to 5) at the C-terminus and / or N-terminus. For example, in the art, substitution with amino acids of similar or comparable properties generally does not alter the function of the protein. Similarly, the addition of one or more amino acids at the C-terminus and / or N-terminus generally does not alter the function of the protein. This term also includes active fragments and active derivatives of integrase.

[0135] The variant forms of this polypeptide include: homologous sequences, conserved variants, allelic variants, natural mutants, induced mutants, proteins encoded by DNA that can hybridize with integrase DNA under high or low severity conditions, and polypeptides or proteins obtained using antiserum against integrase polypeptides. The invention also provides other polypeptides, such as fusion proteins comprising integrase polypeptides or fragments thereof. In addition to nearly full-length polypeptides, the invention also includes soluble fragments of integrase polypeptides. Typically, this fragment has at least about 10 consecutive amino acids of the integrase polypeptide sequence, typically at least about 30 consecutive amino acids, preferably at least about 50 consecutive amino acids, more preferably at least about 80 consecutive amino acids, and most preferably at least about 100 consecutive amino acids.

[0136] The invention also provides analogs of integrase or polypeptides. These analogs may differ from natural integrase polypeptides in the form of amino acid sequence differences, or differences in modifications that do not affect the sequence, or both. These polypeptides include natural or induced genetic variants. Induced variants can be obtained by various techniques, such as random mutagenesis through radiation or exposure to a mutagen, or by site-directed mutagenesis or other known molecular biology techniques. Analogs also include those having residues different from natural L-amino acids (such as D-amino acids), and those having non-naturally occurring or synthetic amino acids (such as β, γ-amino acids). It should be understood that the polypeptides of the present invention are not limited to the representative polypeptides exemplified above.

[0137] Modifications (typically without altering the primary structure) include chemically derived forms of peptides, such as acetylation or carboxylation, either in vivo or in vitro. Modifications also include glycosylation, such as those resulting from glycosylation modifications performed during peptide synthesis and processing or further processing steps. This modification can be accomplished by exposing the peptide to glycosylating enzymes (such as mammalian glycosylation or deglycosylation enzymes). Modifications also include sequences containing phosphorylated amino acid residues (such as phosphotyrosine, phosphotyserine, phosphotythreonine). Modifications also include peptides modified to improve their resistance to proteolysis or optimize their solubility.

[0138] In this invention, a "conserved variant polypeptide of integrase" refers to a polypeptide formed by replacing up to 10, more preferably up to 8, more preferably up to 5, and most preferably up to 3 amino acids with amino acids of similar or analogous properties compared to any of the amino acid sequences shown in SEQ ID NO:1-7. These conserved variant polypeptides are preferably generated by amino acid substitutions according to Table 1.

[0139] Table 1

[0140]

[0141] The polynucleotides of this invention can be in DNA or RNA form. DNA form includes cDNA, genomic DNA, or artificially synthesized DNA. The DNA can be single-stranded or double-stranded. The DNA can be a coding strand or a non-coding strand. The coding region sequence encoding the mature polypeptide can be identical to or a degenerate variant of any of the coding region sequences shown in SEQ ID NO:10-16. As used herein, taking the integrase Cyin14 as an example, "degenerate variant" in this invention refers to a nucleic acid sequence encoding the protein having SEQ ID NO:1 but differing from the coding region sequence shown in SEQ ID NO:10.

[0142] Polynucleotides encoding mature polypeptides include: coding sequences that encode only the mature polypeptide; coding sequences of the mature polypeptide and various additional coding sequences; coding sequences of the mature polypeptide (and optional additional coding sequences) and non-coding sequences.

[0143] The term "polynucleotide encoding a polypeptide" can refer to a polynucleotide that includes the polypeptide, or it can also include additional coding and / or non-coding sequences.

[0144] This invention also relates to variants of the aforementioned polynucleotides that encode polypeptides or fragments, analogs, and derivatives of polypeptides having the same amino acid sequence as those of this invention. These polynucleotide variants can be naturally occurring allelic variants or non-naturally occurring variants. These nucleotide variants include substitution variants, deletion variants, and insertion variants. As is known in the art, an allelic variant is a substitution of a polynucleotide, which may be a substitution, deletion, or insertion of one or more nucleotides, but does not substantially alter the function of the polypeptide it encodes.

[0145] The present invention also relates to polynucleotides that hybridize with the above-described sequences and have at least 50%, preferably at least 70%, and more preferably at least 80% identity between the two sequences. The present invention particularly relates to polynucleotides that hybridize with the polynucleotides described herein under stringent conditions. In the present invention, “stringent conditions” means: (1) hybridization and elution at lower ionic strength and higher temperatures, such as 0.2×SSC, 0.1% SDS, 60°C; or (2) hybridization with a denaturing agent, such as 50% (v / v) formamide, 0.1% fetal bovine serum / 0.1% Ficoll, 42°C, etc.; or (3) hybridization only occurs when the identity between the two sequences is at least 90%, preferably at least 95%. Furthermore, the polypeptide encoded by the hybridizable polynucleotide has the same biological function and activity as the mature polypeptide shown in SEQ ID NO:2.

[0146] This invention also relates to nucleic acid fragments that hybridize with the sequences described above. As used herein, a "nucleic acid fragment" is at least 15 nucleotides long, preferably at least 30 nucleotides, more preferably at least 50 nucleotides, and most preferably at least 100 nucleotides or more. The nucleic acid fragment can be used in nucleic acid amplification techniques (such as PCR) to identify and / or isolate polynucleotides encoding integrase.

[0147] The polypeptides and polynucleotides in this invention are preferably provided in isolated form and are more preferably purified to homogenization.

[0148] The full-length nucleotide sequence or fragment thereof of the integrase of the present invention can generally be obtained by PCR amplification, recombination, or artificial synthesis. For PCR amplification, primers can be designed based on the nucleotide sequences disclosed in the present invention, especially the open reading frame sequences, and the relevant sequences can be amplified using commercially available cDNA libraries or cDNA libraries prepared according to conventional methods known to those skilled in the art as templates. When the sequence is long, it is often necessary to perform two or more PCR amplifications, and then splice the fragments amplified from each amplification in the correct order.

[0149] Once the relevant sequence is obtained, it can be obtained in large quantities using recombination methods. This typically involves cloning it into a vector, transferring it into cells, and then isolating the sequence from the proliferated host cells using conventional methods.

[0150] In addition, sequences can be synthesized artificially, especially when the fragment length is short. Typically, long sequences can be obtained by first synthesizing multiple small fragments and then joining them.

[0151] Currently, the DNA sequence encoding the protein of this invention (or a fragment thereof, or a derivative thereof) can be obtained entirely through chemical synthesis. This DNA sequence can then be introduced into various existing DNA molecules (or vectors) and cells known in the art. Furthermore, mutations can be introduced into the protein sequence of this invention through chemical synthesis.

[0152] The method of amplifying DNA / RNA using PCR technology (Saiki, et al. Science 1985; 230: 1350-1354) is preferred for obtaining the gene of the present invention. Especially when it is difficult to obtain full-length cDNA from a library, the RACE method (RACE-cDNA end amplification method) is preferred. The primers used for PCR can be appropriately selected based on the sequence information of the present invention disclosed herein and can be synthesized using conventional methods. The amplified DNA / RNA fragments can be separated and purified using conventional methods such as gel electrophoresis.

[0153] The present invention also relates to vectors containing the polynucleotides of the present invention, host cells genetically engineered using the vectors or integrase encoding sequences of the present invention, and methods for generating the polypeptides of the present invention via recombinant technology.

[0154] Using conventional recombinant DNA technology (Science, 1984; 224:1431), the polynucleotide sequence of this invention can be used to express or produce recombinant integrase polypeptides. Generally, the following steps are involved:

[0155] (1) Transform or transduce suitable host cells with the polynucleotide (or variant) encoding the integrase polypeptide of the present invention, or with a recombinant expression vector containing the polynucleotide;

[0156] (2) Host cells cultured in a suitable culture medium;

[0157] (3) Isolate and purify proteins from culture media or cells.

[0158] In this invention, the integrase polynucleotide sequence can be inserted into a recombinant expression vector. The term "recombinant expression vector" refers to bacterial plasmids, bacteriophages, yeast plasmids, plant cell viruses, mammalian cell viruses such as adenoviruses, retroviruses, or other vectors well known in the art. Vectors applicable in this invention include, but are not limited to: T7-based expression vectors for expression in bacteria (Rosenberg, et al. Gene, 1987, 56: 125); pMSXND expression vectors for expression in mammalian cells (Lee and Nathans, J Bio Chem. 263: 3521, 1988); and baculovirus-derived vectors for expression in insect cells. In short, any plasmid and vector can be used as long as it can replicate and remain stable within the host. An important characteristic of expression vectors is that they typically contain an origin of replication, a promoter, a marker gene, and translational control elements.

[0159] Methods well known to those skilled in the art can be used to construct expression vectors containing integrase-encoding DNA sequences and suitable transcription / translation control signals. These methods include in vitro recombinant DNA techniques, DNA synthesis techniques, and in vivo recombination techniques (Sambroook, et al. Molecular Cloning, a Laboratory Manual, cold Spring Harbor Laboratory, New York, 1989). The DNA sequence can be efficiently ligated to an appropriate promoter in the expression vector to direct mRNA synthesis. Representative examples of these promoters include: the lac or trp promoter of *E. coli*; the PL promoter of *λ* phage; eukaryotic promoters including the CMV immediate early promoter, the HSV thymidine kinase promoter, early and late SV40 promoters, retroviral LTRs, and other known promoters that control gene expression in prokaryotic or eukaryotic cells or their viruses. The expression vector also includes a ribosome binding site for translation initiation and a transcription terminator.

[0160] In addition, the expression vector preferably contains one or more selective marker genes to provide phenotypic traits for selecting host cells for transformation, such as dihydrofolate reductase, neomycin resistance, and green fluorescent protein (GFP) for eukaryotic cell culture, or tetracycline or ampicillin resistance for Escherichia coli.

[0161] Vectors containing the appropriate DNA sequence and appropriate promoter or control sequence can be used to transform appropriate host cells so that they can express proteins.

[0162] The host cell can be a prokaryotic cell, such as a bacterial cell; a lower eukaryotic cell, such as a yeast cell; or a higher eukaryotic cell, such as a mammalian cell. Representative examples include: Escherichia coli, Streptomyces; Salmonella typhimurium bacterial cells; fungal cells such as yeast; plant cells; Drosophila S2 or Sf9 insect cells; and animal cells such as CHO, COS, 293 cells, or Bowes melanoma cells.

[0163] When the polynucleotides of this invention are expressed in higher eukaryotic cells, the insertion of an enhancer sequence into the vector will enhance transcription. Enhancers are cis-acting factors of DNA, typically approximately 10 to 300 base pairs, that act on the promoter to enhance gene transcription. Examples include the SV40 enhancer (100 to 270 base pairs) located late on the replication origin side, the polyoma enhancer located late on the replication origin side, and adenovirus enhancers.

[0164] Those skilled in the art are well aware of how to select appropriate vectors, promoters, enhancers, and host cells.

[0165] Transformation of host cells with recombinant DNA can be performed using conventional techniques well known to those skilled in the art. When the host is a prokaryote such as *E. coli*, competent cells capable of uptake DNA can be harvested after the exponential growth phase and treated with CaCl2, the steps of which are well known in the art. Another method is to use MgCl2. If desired, transformation can also be performed using electroporation. When the host is a eukaryote, the following DNA transfection methods can be used: calcium phosphate coprecipitation, conventional mechanical methods such as microinjection, electroporation, liposome packaging, etc.

[0166] The obtained transformants can be cultured using conventional methods to express the polypeptide encoded by the gene of this invention. Depending on the host cells used, the culture medium can be selected from various conventional media. Culture is carried out under conditions suitable for host cell growth. Once the host cells have grown to an appropriate cell density, the selected promoter is induced using a suitable method (such as temperature adjustment or chemical induction), and the cells are cultured for a further period.

[0167] The recombinant peptides used in the methods described above can be expressed intracellularly, on the cell membrane, or secreted extracellularly. If desired, the recombinant proteins can be separated and purified using various separation methods based on their physical, chemical, and other properties. These methods are well known to those skilled in the art. Examples of these methods include, but are not limited to: conventional refolding treatment, treatment with protein precipitants (salting out), centrifugation, permeation, ultrafiltration, ultracentrifugation, molecular sieve chromatography (gel filtration), adsorption chromatography, ion exchange chromatography, high-performance liquid chromatography (HPLC), and various other liquid chromatography techniques, as well as combinations of these methods.

[0168] Recombinant integrases or peptides have a wide range of uses. These uses include (but are not limited to) integrating target sequences into the genome of mammals to treat certain diseases or confer a specific function.

[0169] Integrating reaction methods and systems

[0170] This invention provides a gene integration method, the method comprising:

[0171] (a) Provide the target nucleic acid to be integrated, and the genome to be integrated;

[0172] (b) In the presence of the integrase of the present invention, the target nucleic acid is integrated into the genome.

[0173] The present invention also provides a reaction system for gene integration, the reaction system comprising:

[0174] (a) The target nucleic acid to be integrated;

[0175] (b) The genome to be integrated; and

[0176] (b) The integrase of the present invention.

[0177] Preferably, the integration method can be an in vivo integration method or an in vitro integration method. Furthermore, the genome can be a nucleic acid molecule located in prokaryotic cells, eukaryotic cells, or isolated from a large fragment of the genome.

[0178] Preferably, the cells to be integrated include mammalian cells, especially human and non-human mammalian cells; the cells can be somatic cells or germ cells.

[0179] The main advantages of this invention include:

[0180] 1. The integrase of the present invention has low homology with previously reported integrase and exhibits excellent integrative activity, and has broad application prospects.

[0181] 2. Using these integrases of the present invention, large fragments of DNA encoding therapeutic genes can be integrated into the genome without the aid of viral vectors, thereby greatly improving the safety of gene therapy, such as DNA fragments of CAR molecules in CAR-T therapy and genes encoding coagulation factors in hemophilia gene therapy.

[0182] 3. When combined with PE (Prime Editor) technology, it can enable the insertion of integration sequences at specific locations in the genome, achieving site-specific integration of large DNA fragments.

[0183] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Experimental methods in the following embodiments, unless otherwise specified, are generally performed under conventional conditions, such as those described in Sambrook et al., Molecular Cloning: A Laboratory Manual (New York: Cold Spring Harbor Laboratory Press, 1989), or as recommended by the manufacturer. Unless otherwise stated, percentages and parts are weight percentages and parts by weight.

[0184] Example 1. Identification of integrase

[0185] 1. Obtaining integrase

[0186] The inventors used computational programs to mine metagenomic data and analyze the metagenomics of uncultured organisms. Through redundancy removal and protein clustering analysis, they identified several new integrase enzymes, which were named Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16, respectively. Their amino acid sequences are shown in SEQ ID No. 1-7.

[0187] The amino acid sequence of the integrase is as follows:

[0188]

[0189]

[0190]

[0191] 2. Optimize integrase

[0192] The inventors optimized the obtained integrase (Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, Cyin16) with human codons for further functional experiments.

[0193] The encoded sequences for these codon optimizations are SEQ ID NO: 10-16, and the sequences are as follows:

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201] 3. Domain analysis and prediction of each enzyme were performed using multiple sequence alignment software.

[0202] The domains of integrase Cyin10–Cyin16 were predicted using the Conserved Domain Search Service online tool in the Domain & Structure section of the NCBI website.

[0203] The functional domains of integrase Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 relative to the N-terminus and C-terminus of the protein were analyzed, revealing the following distribution of each functional domain: Figure 1 As shown;

[0204] from Figure 1 As can be seen, integrases 10-13 and 15, from N-terminus to C-terminus, include a resolvevase domain, a recombinase domain, and a zinc-ribbon recombinase domain, respectively. Integrases Cyin14 and Cyin16 only contain a resolvevase domain and a zinc-ribbon recombinase domain. The predicted distribution of each integrase domain (calculated from N-terminus to C-terminus) is as follows:

[0205] Table 1

[0206]

[0207]

[0208] The evolutionary relationships of the various integrases in integrase Cyin10-16 were analyzed by sequence alignment, as follows: Figure 9 As shown, from Figure 9 As can be seen, integrases Cyin12 and 13 are more closely related to each other, as are Cyin14 and 16. Integrases Cyin10-13 belong to one large protein family, while integrases Cyin14-16 belong to another large protein family. Predictive analysis of integrase domains also confirms this.

[0209] The structures of the aforementioned proteins were predicted using the Alphafold server, and their three-dimensional structures were displayed using the molecular 3D structure visualization software PyMOL (https: / / pymol.org / 2 / ). The 3D structures of each integrase are shown below. Figure 2 As shown.

[0210] Example 2: Integration Identification Experiment of Integrase

[0211] To verify whether the newly discovered integrase can mediate the integration of large DNA fragments into the mammalian genome, the inventors designed an experiment to verify this, the specific experimental design of which is as follows: Figure 3 As shown.

[0212] The procedure for the identification experiment is as follows:

[0213] 1. Constructing integrase expression plasmids

[0214] Codon optimization was performed on seven DNA sequences encoding integrase Cyin14, Cyin10, Cyin11, Cyin12, Cyin13, Cyin15, and Cyin16, resulting in codon-optimized coding sequences SEQ ID Nos. 10-16 suitable for expression in human cells. Codon optimization was also performed on the coding sequences encoding integrase phiC31 (amino acid sequence shown in SEQ ID No. 43) and integrase Bxb1 (amino acid sequence shown in SEQ ID No. 44), resulting in codon-optimized coding sequences suitable for expression in human cells, SEQ ID Nos. 45 and SEQ ID No. 46, respectively.

[0215] Artificially synthesized DNA sequences SEQ ID No. 10-16, 45, and 46 were double-digested using NotI (Thermo Fisher, catalog number FD0593) and AgeI (Thermo Fisher, catalog number FD1464). Simultaneously, the backbone vector ABE8e (purchased from Addgene, #138489) was digested using these two restriction endonucleases, resulting in fragments of 3361 bp and 4856 bp, respectively. The 3361 bp backbone fragment was recovered using a PCR product recovery kit (purchased from Tiangen Biotech (Beijing) Co., Ltd.). Each double-digested DNA fragment was ligated into the recovered backbone fragment of the ABE8e plasmid. After ligation, the bases between positions 409 and 5226 of the ABE8e plasmid were replaced by sequences SEQ ID No. 10-16, 45, and 46, respectively, yielding nine integrase recombinant expression plasmids. The integrase recombinant expression plasmid map is shown below. Figure 4 As shown, the results were verified to be correct by Sanger sequencing.

[0216] 2. Construct a donor plasmid that expresses EGFP and has an attachment site.

[0217] The integration sequences of integrase Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 were predicted using the BLAST alignment tool on the NCBI website. The integration sequences of Cyin10, Cyin11, Cyin12, Cyin13, Cyin14, Cyin15, and Cyin16 were obtained as SEQ ID No. 17-42.

[0218] The EGFP donor plasmid (SEQ ID No. 8) and the integration sequences of the donor plasmid to be inserted (SEQ ID Nos. 17-42 and 47-48) were prepared artificially. After synthesis, the integration sequences were used to replace the sequences between the 2855th and 2868th bases of the EGFP donor plasmid, respectively, through molecular cloning, thereby obtaining different donor plasmids for the integrase. The EGFP gene is regulated by the EF1α promoter.

[0219] 3. Transformation of expression plasmids and integration site recombinant donor plasmids

[0220] The constructed integrase expression plasmid and the integration site recombinant donor plasmid were transformed into E. coli DH5α competent cells (commercially available). After plating, clone selection, and shaking, plasmid extraction was performed.

[0221] 4. Identification of integration of integrase Cyin10-16

[0222] (1) Cell transfection experiment group design

[0223] The experimental group consisted of co-transfecting 293T cells with each integrase expression plasmid and its corresponding donor plasmid. The transfection reagent was PEI (Polysciences, catalog number 23966-1). The corresponding integration sequences, the integrase expression plasmids to be transfected, and the corresponding donor plasmids containing the integration sequences are shown in the table below.

[0224] Table 2. Expression plasmids to be transfected and their corresponding donor plasmids

[0225]

[0226]

[0227] The control group consisted of 293T cells co-transfected with the control plasmid pUC19 (plasmid sequence SEQ ID No. 49) which does not express integrase and each donor plasmid.

[0228] (2) Transfection experiment

[0229] 293T cells were passaged into 96-well plates at a density of 50,000 cells / well and cultured at 37°C and 5% CO2 for 18 hours before transfection.

[0230] Negative control group transfection method: Add 0.6 μL (concentration of 1 μg / μL) of PEI (purchased from Polysciences, 23966-1) to 5 μL. Mix thoroughly in serum-depleted transfection medium (OptiMEM) (purchased from Shanghai Yuanpei Biotechnology Co., Ltd., L530KJ) to prepare solution A; take 50 ng of each donor plasmid containing the integration sequence, and add 100 ng of control plasmid (pUC19, plasmid sequence is SEQ ID No. 49) to 5 μL. Mix solution A and solution B in serum-depleted medium for transfection. After standing for 5 minutes, add solution A to solution B, mix gently, and let stand for 30 minutes to form a transfection complex. Add the entire transfection complex to the corresponding wells of a 96-well plate containing 293T cells. After 6 hours of transfection, change the medium to DMEM containing 10% fetal bovine serum and 1% penicillin antibiotic (15070063, Gibco). After 24 hours of transfection, change the medium to DMEM selection medium containing 10% fetal bovine serum and puromycin (Gibco, A1113803) at a concentration of 10 μg / ml. After 48 hours, change the medium back to DMEM medium containing 10% fetal bovine serum and continue culturing at 37°C and 5% CO2. Passage the cells every 2-3 days.

[0231] Transfection method for positive control group: Add 0.6 μL (concentration of 1 μg / μL) of PEI (purchased from Polysciences, 23966-1) to 5 μL. Mix the transfection-specific serum-reduced medium (OptiMEM) to prepare solution A. Add 100 ng of plasmid expressing integrase phiC31 and 50 ng of donor plasmid containing the attB site of the integrase phiC31 integration sequence to 5 μL of Opti-MEM serum-reduced medium (Thermo, 31985088) and mix well to prepare solution B. After standing for 5 min, add solution A to solution B, mix gently, and let stand for 30 min to form a transfection complex. Add the full amount of the transfection complex to the corresponding well of 293T cells. After 6 h of transfection, change the medium to DMEM medium containing 10% fetal bovine serum and 1% penicillin antibody (15070063, Gibco). 24 hours after transfection, the medium was replaced with 10% fetal bovine serum DMEM selection medium containing 10 μg / ml puromycin (Gibco, A1113803). After 48 hours, the medium was replaced back with 10% fetal bovine serum DMEM medium, and cultured for another 2-3 days. The plasmid expressing integrase Bxb1 and the donor plasmid containing the integration sequence of Bxb1 were transfected using the same method. After 48 hours of transfection, the medium was replaced back with 10% fetal bovine serum DMEM medium, and cultured for another 37°C at 5% CO2. The medium was then passaged for another 2-3 days.

[0232] Transfection method for the experimental group: Add 0.6 μL (concentration of 1 μg / μL) of PEI (purchased from Polysciences, 23966-1) to 5 μL. Mix thoroughly in OptiMEM (a serum-depleted medium) to prepare solution A. Add 100 ng of each integrase plasmid and 50 ng of the corresponding donor plasmid to 5 μL of OptiMEM (Thermo, 31985088) and mix thoroughly to prepare solution B. After standing for 5 min, add solution A to solution B, mix gently, and let stand for 30 min to form the transfection complex. Add the complete transfection complex to the corresponding wells of 293T cells. After 6 h of transfection, change the medium to DMEM containing 10% fetal bovine serum and 1% penicillin antibiotic (15070063, Gibco). 24 hours after transfection, the medium was replaced with 10% fetal bovine serum DMEM selection medium containing 10 μg / ml puromycin (Gibco, A1113803). 48 hours after transfection, the medium was replaced back with 10% fetal bovine serum DMEM medium and cultured at 37°C under 5% CO2 conditions. The medium was passaged every 2-3 days.

[0233] (3) Integration efficiency test

[0234] After 20 days of culture, the original culture medium in each well was discarded. 293 T cells per well were digested with 80 μL of 0.05% trypsin and incubated at 37°C for 5 min. Then, 80 μL of 10% fetal bovine serum medium was added to stop the digestion. The cells were pipetted to form a single-cell suspension, centrifuged at 400g for 5 min, the supernatant was discarded, and the cells were resuspended in 200 μL of 10% FBS medium. The proportion of EGFP-positive cells was detected by flow cytometry. This proportion represents the integration efficiency of large DNA fragments mediated on the genome.

[0235] After 20 days of cell culture, the inventors used flow cytometry to detect the percentage of EGFP-positive cells. The 293T cells successfully transfected with integrase and donor plasmids were mainly divided into two categories: one where the donor plasmid integrated into the genome, and these cells continuously expressed EGFP protein; the other where the donor plasmid did not integrate into the genome, and as the culture time increased, the donor plasmid was metabolized and therefore no longer expressed EGFP protein. Therefore, detecting whether 293T cells express EGFP protein can characterize whether the integration of large exogenous DNA fragments has occurred, and analyzing the percentage of EGFP-positive cells can determine the integration efficiency of large DNA fragments.

[0236] Figure 6 The figure shows the integration efficiency of seven integrase-mediated donor plasmids into the 293T cell genome (integrase Cyin10 corresponds to two integration sequences, while integrase Cyin11-16 correspond to three or four integration sequences, respectively). From Figure 6 As can be seen from the table, the integration efficiency of the donor plasmid genome when the integrase corresponds to the attachment sequence is shown in Table 3.

[0237] Table 3 Integration efficiency of each integrase and its corresponding attachment sequence

[0238]

[0239]

[0240] Table 3 shows that integrases Cyin10-16 exhibit integration efficiency when corresponding to specific integration sequences. In particular, integrase 14, corresponding to integration sequence 14-P1, mediates the integration of a large DNA fragment of approximately 5 kb into the mammalian genome with an integration efficiency as high as 10.12%, comparable to the integration level of the previously reported highly efficient integrase phiC31. Integrases Cyin11, Cyin12, Cyin14, Cyin15, and Cyin16 all show high integration efficiency when corresponding to attachment sequences 11-B1, 12-P1, 14-B1, 15-B2, and 16-P1, 16-P2, and 16-B2, respectively. A comparison of the integration efficiencies of each integrase and its corresponding integration sequence is provided below. Figure 6 As shown.

[0241] Figure 7 The figures show the flow cytometry results for integrase Cyin14 with the 14-P1 donor plasmid, integrase Cyin14 with a donor plasmid containing the 14-P1 integration sequence, integrase Cyin16 with a donor plasmid containing the 16-B2 integration sequence, the positive control group (integrase phiC31), and the negative control group (plastic plasmid without integrase expression, i.e., the results of random integration of donor plasmids). As can be seen from the figures, compared to the control group, the integration efficiency of integrase Cyin14, 16, and phiC31 with their associated integration sequences is significantly higher than that of the control group.

[0242] Figure 8 The results compare the integration efficiency of integrases Cyin14 and Cyin16 with that of previously reported integrases phiC31 and Bxb1. These results demonstrate that integrases Cyin14 and 16 can efficiently mediate the integration of approximately 5kb DNA fragments into the mammalian genome.

[0243] All sequence descriptions used in this invention are shown in Table 4.

[0244] SEQ ID No. 8-9,17-49 is shown below.

[0245]

[0246]

[0247]

[0248]

[0249]

[0250]

[0251]

[0252]

[0253]

[0254]

[0255] Table 4 Sequence Description

[0256]

[0257]

[0258]

[0259] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. An integrase, characterized in that, The integrase is selected from the integrase shown in SEQ ID NO:

1.

2. A polynucleotide, characterized in that, The polynucleotide encodes the integrase of claim 1.

3. A carrier, characterized in that, The vector comprises the polynucleotide of claim 2.

4. The carrier as described in claim 3, characterized in that, The regulatory sequence includes a promoter sequence.

5. A genetically engineered host cell, characterized in that, It contains the vector as described in claim 3, or its genome integrates the polynucleotide as described in claim 2.

6. A method for preparing the integrase according to claim 1, characterized in that, The method includes: (a) Culturing the host cells of claim 5 under suitable expression conditions; (b) Isolate the integrase from the culture.

7. A recombination method, characterized in that, The method includes: (a) Provide the target nucleic acid to be integrated, and the genome to be integrated; (b) In the presence of the integrase of claim 1, the target nucleic acid is integrated into the genome.

8. A reaction system for gene integration, characterized in that, The reaction system includes: (a) The target nucleic acid to be integrated; (b) The genome to be integrated; and (b) The integrase of claim 1.

9. A kit for detecting target nucleic acids in a sample, characterized in that, The kit contains: (a) the integrase of claim 1; (b) the polynucleotide of claim 2; or (c) the vector of claim 3.

10. The use of the integrase of claim 1, or the polynucleotide of claim 2, or the vector of claim 3, or the host cell of claim 5, in the preparation of a formulation or kit, wherein the formulation or kit is used for: (i) Gene or genome editing; (ii) Editing target sequences in target loci to modify biological or non-human organisms.

11. A fusion protein comprising a DNA-binding nuclease and the integrase of claim 1.

12. A method for specifically integrating exogenous nucleic acid sites into the cellular genome or intracellular target nucleic acid, characterized in that, The method includes: (a) Introduce (i) the first construct, (ii) gRNA, (iii) the exogenous nucleic acid to be integrated, and (iv) the second construct into cells. (i) A first construct, wherein the first construct is an expressible polynucleotide construct encoding an editing polypeptide, wherein the editing polypeptide comprises a DNA-binding nuclease domain connected via a linker to a reverse transcriptase domain, wherein the DNA-binding nuclease domain comprises nicking enzyme activity; and (ii) gRNA, wherein the gRNA is a guide RNA containing a complementary sequence comprising a targeting sequence, a primer-binding sequence, and an integration sequence. The gRNA interacts with the expressed editing peptide, thereby targeting the editing peptide to a specific target site in the cellular genome or intracellular target nucleic acid. The DNA-binding nuclease domain of the editing peptide cleaves one strand of the cellular genome or intracellular target nucleic acid, and the reverse transcriptase domain incorporates the integration sequence into the cleavage site, thereby incorporating at least one integration sequence into the cellular genome or target nucleic acid at a specific target site. (iii) A foreign nucleic acid to be integrated, which is ligated to a recognition sequence homologous to the integration sequence, wherein the integrase mediates site-specific integration in the presence of both the recognition sequence and the integration sequence; and (iv) A second construct, wherein the second construct is an expressible polynucleotide construct encoding the integrase of claim 1. The integrase integrates exogenous nucleic acids into specific target sites in the cell genome; (b) In the presence of the editing peptide, gRNA, integrase, and exogenous nucleic acid to be integrated, the integrase integrates the exogenous nucleic acid into the cellular genome or intracellular target nucleic acid.

13. A sequence integration system, characterized in that, The integrated system includes: (i) A first construct, wherein the first construct is an expressible polynucleotide construct encoding an editing polypeptide, wherein the editing polypeptide comprises a DNA-binding nuclease domain connected to a reverse transcriptase domain via a linker, wherein the DNA-binding nuclease domain comprises nicking enzyme activity; (ii) gRNA, wherein the gRNA is a guide RNA containing a complementary sequence comprising a targeting sequence, a primer-binding sequence, and an integration sequence. The gRNA interacts with the expressed editing peptide to target the editing peptide to a specific target site of the cell genome or intracellular target nucleic acid, and the DNA-binding nuclease domain of the editing peptide cleaves one strand of the cell genome or intracellular target nucleic acid, and the reverse transcriptase domain incorporates the integration sequence into the cleavage site, thereby incorporating at least one integration sequence into the cell's specific target site genome or target nucleic acid. (iii) a foreign nucleic acid to be integrated, wherein the foreign nucleic acid to be integrated is linked to a recognition sequence homologous to the integration sequence, the recognition sequence and the integration sequence being capable of mediating site-specific integration; and (iv) A second construct, the second construct encoding the integrase of claim 1. The integrase integrates exogenous nucleic acids into specific target sites in the cell genome.

14. The use of the sequence integration system of claim 13 in preparing a sequence integration system for integrating a foreign nucleic acid to be integrated into the genome of a desired object using the method of claim 12.

15. The application as described in claim 14, characterized in that, The objects include humans or non-human mammals.

Citation Information

Patent Citations

  • Mutants of the bacteriophage lambda integrase

    CN108064279A

  • Polypeptide variants

    CN108291212A