Structure of plasmid containing polyadenylic acid (POLYA)

By introducing a specific order of replication start sites, specific site recombination sequences and transition sequences into the polyadenylation plasmid, the plasmid aggregate problem is solved, the stability and yield of the plasmid are improved, the drug regulatory requirements are met, and the development and application of mRNA therapy are supported.

WO2025201494A1PCT designated stage Publication Date: 2025-10-02SUZHOU LEVOSTAR LIFE SCIENCES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/085583
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-28
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

In the existing technology, polyadenylation plasmids are prone to form aggregates in bacteria, resulting in low plasmid yield and low monomer ratio, affecting the research and development of mRNA therapy and drug market access. It is especially difficult to solve the problems of plasmid stability and consistency when facing the regulatory requirements of the Food and Drug Administration.

Method used

By introducing a specific order of replication start site sequence, specific site recombination sequence and transition sequence into the plasmid, the arrangement order of nucleic acid molecules is optimized, the stability and yield of plasmid PolyA are improved, and specific recombinase is used for recombination to maintain the supercoiled structure of the plasmid.

Benefits of technology

It significantly improves the stability and yield of polyadenylation plasmids, enhances the production efficiency and quality control of mRNA therapy, and meets the requirements of drug review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025085583_02102025_PF_FP_ABST
    Figure CN2025085583_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the structure and a use of a plasmid containing polyadenylic acid (polyA). The isolated nucleic acid molecule involved in the present application comprises the following components: an origin of replication sequence, a first transition sequence, and a site-specific recombination sequence. The present application further relates to a vector comprising the nucleic acid molecule, a host cell comprising the nucleic acid molecule and / or the vector, and a diagnostic or pharmaceutical composition and a kit comprising the nucleic acid molecule, the vector and / or the host cell. The present application further provides a method for expressing a target gene.
Need to check novelty before this filing date? Find Prior Art

Description

The structure of a polyadenylation PolyA plasmid Technical Field

[0001] The present application relates to the field of biomedicine, and specifically to the structure and application of a polyadenylic acid (PolyA) plasmid. Background Art

[0002] Currently, plasmids containing polyadenylation (PolyA) sequences (referred to as polyA plasmids) are important raw materials for in vitro transcription to produce mRNA. However, the PolyA sequence brings a series of problems. In addition to its low stability and prone to deletion, it also leads to a series of problems such as low plasmid yield and low monomer ratio.

[0003] Current technologies for solving the plasmid aggregate problem mainly focus on improving strains, but these technical methods have limited effectiveness when dealing with mRNA plasmids with long PolyA tails. This type of plasmid with a long PolyA tail is prone to forming aggregates in bacteria, resulting in the problem of easy aggregate formation during the production process. This phenomenon is difficult to effectively control with existing technologies. This problem is particularly prominent in the development of mRNA therapies. mRNA plasmids with long PolyA tails will encounter significant technical challenges during the research and application stages of biopharmaceutical companies, especially when facing the regulatory requirements put forward by the Drug Review Center (CDE) of the National Medical Products Administration. Existing technologies have failed to provide a stable and efficient solution to ensure the quality and consistency of such plasmids, which in turn affects the development of the entire biopharmaceutical field and the speed of market access for drugs. Therefore, new technical means are urgently needed to solve this key problem to facilitate the production and quality control of mRNA plasmids. Summary of the Invention

[0004] In response to the problems in the prior art, the present application provides a nucleic acid molecule comprising a replication initiation site sequence, a first transition sequence (screening sequence, such as KanR) and a specific site recombination sequence, so that the nucleic acid molecule maintains the supercoiled structure of the plasmid while also effectively improving the stability of the plasmid PolyA and the plasmid yield. Furthermore, when the order of the above components in the vector is the first transition sequence (screening sequence, such as KanR), the replication initiation site sequence (Ori), and the specific site recombination sequence (Cer), the stability of the plasmid PolyA is effectively improved.

[0005] In one aspect, the present application provides a polyadenylic acid (PolyA) plasmid comprising the following components: a replication initiation site sequence, a first transition sequence, and a specific site recombination sequence.

[0006] In some embodiments, the order of arrangement of the components in the nucleic acid molecule along the 5' to 3' direction can be one or more of the following:

[0007] (1) Replication origin sequence, specific site recombination sequence, first transition sequence; or

[0008] (2) replication origin sequence, first transition sequence, specific site recombination sequence; or

[0009] (3) a first transition sequence, a replication origin sequence, or a specific site recombination sequence; or

[0010] (4) a first transition sequence, a specific site recombination sequence, or a replication origin sequence; or

[0011] (5) specific site recombination sequence, replication origin sequence, first transition sequence; or

[0012] (6) Specific site recombination sequence, first transition sequence, and replication initiation site sequence.

[0013] In some embodiments, the order of the components in the nucleic acid molecule along the 5' to 3' direction may be a first transition sequence, a replication origin sequence, and a specific site recombination sequence.

[0014] In some embodiments, a second transition sequence is further included, wherein for sequence 2), it further includes: a replication origin sequence, a first transition sequence, a second transition sequence, and a specific site recombination sequence.

[0015] In some embodiments, the second transition sequence is a promoter sequence.

[0016] In some embodiments, the site-specific recombination sequence comprises a specific recombination site, wherein the specific recombination site can be combined with a specific recombinase for recombination. In some embodiments, the specific recombinase comprises one or more of XerC, XerD, CodV, RipX, Cre, Int, Xis, P22, Flp, and R1. In some embodiments, the specific recombination site comprises one or more of Ecdif, cer, psi, dif, mwr, Bsdif, loxP, FRT, and RS, or variants or functional fragments of the above enzymes.

[0017] In some embodiments, the specific recombinase is XerC and / or XerD, or a variant or functional fragment thereof.

[0018] In some embodiments, the specific recombination site is cer.

[0019] In some embodiments, the specific site recombination sequence further comprises an accessory sequence, wherein the accessory sequence comprises an accessory site capable of binding to an accessory protein. The accessory protein comprises a first accessory protein, and the accessory sequence comprises a first accessory sequence, wherein the first accessory sequence comprises a site capable of binding to the first accessory protein. In some embodiments, the accessory protein is one or more of PepA, ArgR, and ArcA. In some embodiments, the first accessory protein is PepA.

[0020] In some embodiments, the accessory protein further comprises a second accessory protein, the accessory sequence further comprises a second accessory sequence, and the second accessory sequence comprises a site capable of binding to the second accessory protein. In some embodiments, the second accessory protein is one or more of ArgR and ArcA.

[0021] In some embodiments, the site-specific recombination sequence in the nucleic acid molecule comprises, in order from 5' to 3',:

[0022] a) a first common sequence as shown in any one of SEQ ID NO: 45 or its reverse complementary sequence; and

[0023] b) a second common sequence or its reverse complementary sequence as shown in any one of SEQ ID NO: 46. The first common sequence and the second common sequence are separated by a nucleotide sequence of 6-8 bp.

[0024] In some embodiments, the first common sequence is a sequence as shown in SEQ ID: 38-40 or a reverse complementary sequence thereof.

[0025] In some embodiments, the second common sequence is the sequence shown in SEQ ID: 41 or its reverse complement.

[0026] In some embodiments, the specific site recombination sequence includes a first common sequence as shown in SEQ ID: 38-40 or its reverse complementary sequence and a second common sequence as shown in SEQ ID: 41 or its reverse complementary sequence.

[0027] In some embodiments, the site-specific recombination sequence includes a sequence as shown in any one of SEQ ID NOs.: 18-37 or a reverse complementary sequence thereof.

[0028] In some embodiments, the first transition sequence comprises a screening sequence, wherein the screening sequence comprises a prokaryotic screening sequence and / or a eukaryotic screening sequence. In some embodiments, the first transition sequence is a screening sequence.

[0029] In some embodiments, the prokaryotic screening sequence includes one or more of an ampicillin resistance gene, a kanamycin resistance gene, a chloramphenicol resistance gene, a spectinomycin resistance gene, a tetracycline resistance gene, a bleomycin resistance gene, a streptomycin resistance gene, a hygromycin resistance gene, a gentamicin resistance gene, a hexapromycin resistance gene, an erythromycin resistance gene, and a blasticidin gene.

[0030] In some embodiments, the screening sequence comprises a sequence as shown in any one of SEQ ID NOs.: 13-17 or a reverse complementary sequence thereof.

[0031] In some embodiments, the replication initiation site sequence is one or more of ColE1, pMB1, pBR322, pSC101, R6K, and P15A.

[0032] In some embodiments, the replication initiation site sequence comprises a sequence as shown in any one of SEQ ID NOs.: 11-12 or a reverse complementary sequence thereof.

[0033] In some embodiments, the nucleic acid molecule further comprises one or more of the following components: a promoter sequence, a specific enzyme cleavage site sequence, a tag sequence, and a terminator sequence.

[0034] In some embodiments, the promoter comprises one or more of a constitutive promoter, a tissue-specific promoter, or an inducible promoter.

[0035] In some embodiments, the promoter is selected from one or more of the following groups: CMV, EF1a, SV40, PGK1, Ubc, CAG, TRE, USA, Ac5, CaMkIIa, GAL1 / 10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1 and U6.

[0036] In some embodiments, the specific enzyme cleavage site sequence includes one or more sites selected from the following group: ApaI, BamHI, BglII, EcoRI, HindIII, KpnI, NcoI, NdeI, NheI, NotI, SacI, SalI, SphI, XbaI and XhoI.

[0037] In some embodiments, the tag sequence comprises a fluorescent protein sequence. In some embodiments, the fluorescent protein sequence is selected from one or more of the following groups: GFP, eGFP, eYFP, eCFP, mCherry, luc2, and hrluc. In some embodiments, the fluorescent protein sequence comprises a combination of eGFP and luc2, and / or a combination of mCherry and hrluc.

[0038] In some embodiments, the nucleic acid molecule further comprises an in vitro transcription sequence. In some embodiments, the in vitro transcription sequence is a Poly A sequence.

[0039] In some embodiments, the nucleic acid molecule comprises a sequence as shown in any one of SEQ ID NOs: 2-10, 48, 50, 51 or a reverse complementary sequence thereof.

[0040] On the other hand, the present application also provides a vector comprising the isolated nucleic acid molecule. In some embodiments, the vector is a plasmid. In some embodiments, the vector is a cosmid or a transposon.

[0041] On the other hand, the present application also provides a host cell comprising the isolated nucleic acid molecule or the vector. In some embodiments, the host cell provides a specific recombinase. In some embodiments, the host cell provides an accessory protein.

[0042] On the other hand, the present application also provides a diagnostic or pharmaceutical composition comprising the isolated nucleic acid molecule, the vector or the host cell.

[0043] On the other hand, the present application also provides a method for expressing a target gene, which comprises introducing the isolated nucleic acid molecule or the vector into a host cell, and allowing the target gene to be expressed in the host cell. In some embodiments, the method is an in vitro or ex vivo method.

[0044] On the other hand, the present application also provides a kit comprising the isolated nucleic acid molecule, the vector or the host cell.

[0045] On the other hand, the present application also provides a method for constructing a recombinant plasmid, which comprises introducing the nucleic acid molecule into a polyadenylation plasmid to obtain a recombinant plasmid.

[0046] On the other hand, the present application also provides a recombinant plasmid, wherein the nucleic acid molecule is obtained by the above-mentioned nucleic acid molecule construction method.

[0047] The present application provides a nucleic acid molecule containing a gene for maintaining supercoiling of polyadenosine plasmid monomers, which can significantly increase the supercoiling ratio of mRNA plasmid monomers with long PolyA tails (60 to 150 nucleotides in length), overcoming the limitations of the existing technology, and at the same time can improve the stability and consistency of long PolyA tail plasmids and plasmid yield, providing production technology support for the research and development and application of mRNA therapy.

[0048] Those skilled in the art can easily discern other aspects and advantages of the present application from the detailed description below. In the detailed description below, only exemplary embodiments of the present application are shown and described. As will be appreciated by those skilled in the art, the content of the present application enables those skilled in the art to make changes to the disclosed specific embodiments without departing from the spirit and scope of the invention to which the present application relates. Accordingly, the description in the drawings and specification of the present application is merely exemplary and not restrictive. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The specific features of the invention involved in this application are shown in the appended claims. The features and advantages of the invention involved in this application can be better understood by referring to the exemplary embodiments and drawings described in detail below. A brief description of the drawings is as follows:

[0050] FIG1 shows the plasmid map of pKGCT7-Fluc in the examples of this application.

[0051] FIG2 shows the plasmid map of pKGCT7-Fluc BX in the examples of this application.

[0052] FIG3 shows the plasmid map of pKGCT7-Fluc BXm in the examples of this application.

[0053] FIG4 shows the plasmid map of AX-GFP-Firefly in the examples of the present application.

[0054] FIG5 shows the plasmid map of AX-RFP-Renilla in the examples of the present application.

[0055] FIG6 shows the plasmid map of pKGCT7-Fluc-cer in the examples of the present application.

[0056] FIG7 shows the agarose gel electrophoresis diagram in Example 1.

[0057] FIG8 shows the plasmid map of AX-Fluc in the examples of the present application.

[0058] Figure 9 shows the plasmid map of pLevo-D1-Fluc in the examples of this application Specific implementation plan

[0059] The following describes the implementation scheme of the present invention through specific embodiments. People familiar with this technology can easily understand other advantages and effects of the present invention from the contents disclosed in this specification.

[0060] The technical means used in the examples are conventional means well known to those skilled in the art, and the raw materials used are all commercially available products.

[0061] Definition of terms

[0062] The term "isolated nucleic acid molecule" used in this application refers to deoxyribonucleic acid or ribonucleic acid in an isolated form. The isolated nucleic acid molecule can be a product isolated from a natural environment or artificially synthesized. The isolated nucleic acid molecule carries a target gene and can express a protein, peptide, amino acid residue or nucleic acid sequence encoded by the corresponding target gene. When the isolated nucleic acid molecule is transferred into a host cell for expression, the expression product can exert its biological function in the host cell. The isolated nucleic acid molecule not only includes nucleic acid molecules of the same sequence, but also includes nucleic acid molecules containing any base mutations, and the base mutations especially include a replacement that produces synonymous codons (different codons encoding the same amino acid residue) due to the codon degeneracy in conservative amino acid replacements. The isolated nucleic acid molecule described in the present application also includes any single-stranded nucleic acid sequence complementary to the nucleic acid molecule.

[0063] The term "vector" as used in this application generally refers to a nucleic acid delivery vehicle into which any nucleic acid molecule can be inserted. The nucleic acid molecule usually inserted can carry the target gene. When the nucleic acid molecule comprising the target gene is inserted into the vector, the target gene can be stably expressed in vivo or in vitro. The vector can be transformed, transduced or transfected into the host cell so that the target gene it carries is expressed in the host cell. For example, the vector can be derived from one or more of the following groups: plasmid, phagemid, cosmid, artificial chromosome such as yeast artificial chromosome (YAC), bacterial artificial chromosome (BAC) or artificial chromosome (PAC) from P1 and bacteriophage such as lambda phage or M13 phage and animal virus. The animal virus species used as a vector include but are not limited to retrovirus (including lentivirus), adenovirus, adeno-associated virus, herpes virus (such as herpes simplex virus), poxvirus, baculovirus, papillomavirus and / or papillomavirus (such as SV40). Generally speaking, the vector can contain one or more elements that control expression, including promoter sequence, transcription initiation sequence, enhancer sequence, screening sequence and reporter gene. In addition, the vector may also contain a replication origin sequence. The vector may also include components that assist it in entering cells, including but not limited to virus particles, liposomes and / or protein coats. In addition, the vector may also include specific site recombination sequences and / or restriction endonuclease cleavage sites.

[0064] The term "target gene" used in this application generally refers to a DNA or cDNA that can encode a gene contained in a nucleic acid molecule, vector, host cell, pharmaceutical composition or kit. For example, the "target gene" used in this application refers to a DNA sequence that is added to the vector of this application and used for the final protein expression.

[0065] The term "host cell" as used in this application generally refers to a cell that has been or can be transformed, transduced and / or transfected by a nucleic acid molecule and / or a vector. Typically, the host cell is capable of expressing the protein, peptide, amino acid residue or nucleic acid sequence encoded by the target gene on the nucleic acid molecule and / or the vector. The host cell may also include the offspring of the host cell, regardless of whether the offspring of the host cell is morphologically or genetically identical to the original host cell. The host may include host cells in an isolated form, such as host cells in culture. The host cell may also include the environmental conditions inside the host cell, such as the molecules required for various biological reactions present in the host cell internal environment, including but not limited to specific recombinases and accessory proteins. The host cell internal environment can also provide a reaction site for the biological reaction and the physical and chemical conditions required for the reaction, such as being able to achieve targeted binding and recombination of specific recombinases and specific recombination sites.

[0066] As used herein, the term "replication origin sequence" refers to a genetic element on a constructed nucleic acid molecule that enables the nucleic acid molecule to replicate autonomously in a host cell. The replication origin sequence can be any vector replicon capable of mediating autonomous replication. For example, the replication origin sequence can mediate replication of a plasmid or vector in a host cell. The presence of the replication origin sequence can also determine the copy number of the plasmid or vector in the host cell. Examples of replication origin sequences for use in E. coli hosts are one or more of ColE1, pMB1, pBR322, pSC101, R6K, and P15A.

[0067] The term "transition sequence" used in this application refers to a nucleic acid fragment positioned on a constructed nucleic acid molecule, positioned at a specific site recombination sequence downstream and a replication initiation site sequence upstream, or positioned at a replication initiation site downstream and a specific site recombination sequence upstream. In this application, a transition sequence can comprise the functional elements of one or more plasmids, such as promoter sequences, and examples include one or more of CMV, EF1a, SV40, PGK1, Ubc, CAG, TRE, USA, Ac5, CaMkIIa, GAL1 / 10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, and U6. For example, a transition sequence can be a sequence of a coded protein, a polypeptide segment, or an amino acid fragment. For example, a transition sequence can be connected to a nucleic acid fragment of a specific site recombination sequence and a replication initiation site sequence. The transition sequence can be a screening sequence, such as an antibiotic screening sequence, including but not limited to one or more of an ampicillin resistance gene, a kanamycin resistance gene, a chloramphenicol resistance gene, a spectinomycin resistance gene, a tetracycline resistance gene, a bleomycin resistance gene, a streptomycin resistance gene, a hygromycin resistance gene, a gentamicin resistance gene, a hexapromycin resistance gene, an erythromycin resistance gene, and a blasticidin gene. The term "screening sequence" as used herein refers to a sequence used to detect whether a nucleic acid molecule is present in a cell or environment to be detected. When the constructed nucleic acid molecule includes a screening sequence, the nucleic acid molecule can be present and amplified in a specific environment.

[0068] The term "specific site-recombination sequence" as used herein refers to a nucleic acid sequence located on a constructed nucleic acid molecule that can be recognized and bound by a specific recombinase. Recognition and binding of the specific site-recombination sequence by the specific recombinase allows for recombination. In this application, recombination of the specific site-recombination sequence can maintain local supercoiling of a chromosome or plasmid. The term "specific recombinase" as used herein refers to a protein that can recognize and bind to a specific recombination site on a constructed nucleic acid molecule, catalyzing and completing the exchange or excision of flanking nucleic acid segments. Typically, the specific recombinase is derived from a host cell, meaning that the host cell can express a gene encoding the specific recombinase without any genetic manipulation. Specific recombinases used in Escherichia coli include, but are not limited to, tyrosine recombinases and / or serine recombinases. Examples include one or more of XerC, XerD, CodV, RipX, Cre, Int, Xis, P22, Flp, and R1, or variants or functional fragments of the aforementioned enzymes.

[0069] The term "specific site recombination sequence" used in this application can include specific recombination sites, and specific recombination sites refer to a nucleic acid sequence in a specific site recombination sequence. Specific recombination sites or their expressed proteins can be recognized and combined by specific recombinases, which reorganize constructed nucleic acid molecules based on the specific recombination sites. For example, the combination of a specific recombinase and a specific recombination site can exchange or excise nucleic acid fragments located upstream, downstream, or inside the recombination site. In this application, specific recombination sites can be combined with specific recombinases to allow the specific site recombination sequence to carry out homologous recombination, thereby maintaining local supercoiling of a chromosome or plasmid. Examples of the specific recombination sites used in eukaryotes and prokaryotes include, but are not limited to, one or more of Ecdif, cer, psi, dif, mwr, Bsdif, loxP, FRT, and RS.

[0070] The term "common sequence" as used herein refers to a nucleic acid sequence shared by two or more nucleic acid molecules, i.e., the nucleic acid sequence is a nucleic acid fragment of two or more nucleic acid molecule sequences. A common sequence can also be a reverse complementary nucleic acid fragment of a nucleic acid fragment shared by two or more nucleic acid molecules. The nucleic acid molecules can come from different species and can encode the same protein or proteins with the same function but different structures. A common sequence need not be a completely identical nucleic acid fragment (e.g., 100%), and partial common sequences are also contemplated by the present application and are within the scope of the present application, for example, a common sequence identity of at least 95%, 94%, 93%, 92%, 91%, 90%, 89%, 88%, 87%, 86%, 85%, 84%, 83%, 82%, 81%, 80%, etc. In the present application, the common sequence can be a nucleic acid sequence shared by a specific site recombination sequence.

[0071] The term "accessory sequence" used in this application refers to the upstream and downstream nucleic acid fragment sequences of the specific site recombination sequence in the constructed nucleic acid molecule. Generally speaking, the accessory sequence can regulate the recombination of the specific site recombination sequence. In this application, the accessory sequence can regulate the recombination of the specific site recombination sequence, thereby dissociating the plasmid dimer. Usually, it is necessary to recognize and bind to the accessory sequence through the accessory protein to exert the regulatory function of the accessory sequence. The accessory protein achieves targeted binding to the accessory sequence based on the accessory site on the accessory sequence to exert the regulatory function of the accessory sequence. In this application, after the accessory protein binds to the accessory site, it can activate the specific recombinase, regulate the recombination of the specific site recombination sequence to dissociate the plasmid dimer and maintain the supercoiled structure of the plasmid. The "accessory protein" described in this application can be endogenous to the host cell, that is, the host cell can express the accessory protein without any genetic manipulation. Examples include but are not limited to one or more of PepA, ArgR and ArcA.

[0072] The term "tag sequence" as used herein refers to another molecular entity incorporated into a constructed nucleic acid molecule. Once the tag sequence completes the transcription or translation process, it can be used to screen for the tagged target product in a mixture of cellular transcription or translation. The tag sequence can be a detectable marker sequence, such as a fluorescent protein sequence that allows the tagged target product to be detected optically. Examples of fluorescent protein sequences include one or more selected from the group consisting of GFP, eGFP, eYFP, eCFP, mCherry, luc2, and hrluc.

[0073] Detailed Description of the Invention

[0074] 1. Isolated nucleic acid molecules

[0075] On the one hand, the present application provides an isolated nucleic acid molecule, which can have the characteristics of a high ratio of monomeric polyadenosine supercoiling, and at the same time, the nucleic acid molecule can also exhibit high yield and strong PolyA stability.

[0076] The isolated nucleic acid molecule may be a nucleic acid molecule isolated from nature, or may be obtained by modifying a nucleic acid molecule isolated from nature, wherein the nucleotide modification may not alter the amino acid residues encoded by the nucleic acid molecule. The isolated nucleic acid molecule may be an artificially synthesized recombinant nucleic acid.

[0077] The isolated nucleic acid molecule may comprise DNA, cDNA and / or mRNA. The isolated nucleic acid molecule may comprise single-stranded nucleic acid and / or double-stranded nucleic acid. The single-stranded DNA may be a coding strand or a non-coding strand. The isolated nucleic acid molecule may be a circular nucleic acid molecule and / or a linear nucleic acid molecule.

[0078] The isolated nucleic acid molecule described in the present application may comprise the following components: a replication origin sequence, a first transition sequence, and a specific site recombination sequence. The components are as follows from the 5' to the 3' direction:

[0079] (1) Replication origin sequence, specific site recombination sequence, first transition sequence; or

[0080] (2) replication origin sequence, first transition sequence, specific site recombination sequence; or

[0081] (3) a first transition sequence, a replication origin sequence, or a specific site recombination sequence; or

[0082] (4) a first transition sequence, a specific site recombination sequence, or a replication origin sequence; or

[0083] (5) specific site recombination sequence, replication origin sequence, first transition sequence; or

[0084] (6) Specific site recombination sequence, first transition sequence, and replication initiation site sequence.

[0085] The isolated nucleic acid molecule assembly described herein, wherein the specific site recombination sequence has the function of dissociating plasmid dimers. The specific site recombination sequence can increase nucleic acid monomer production. The specific site recombination sequence can improve the stability of the Poly A sequence in the nucleic acid molecule. The specific site recombination sequence can maintain the supercoiled structure of the plasmid, increase nucleic acid monomer production, and improve the stability of the Poly A sequence in the nucleic acid molecule.

[0086] The specific site recombination sequence can be exogenous. The specific site recombination sequence can be introduced into the host cell through genetic manipulation. The specific site recombination sequence can be endogenous, and the specific site recombination sequence can be derived from the host cell, that is, the host cell does not need any genetic manipulation to express the specific site recombination sequence. The specific site recombination sequence is derived from the same cell type as the host cell. The specific site recombination sequence can be derived from a prokaryotic organism. The specific site recombination sequence can be derived from a eukaryotic organism.

[0087] Specific site recombination sequences can be used for homologous recombination, site-specific recombination, or transposition recombination. Specific site recombination sequences can be used for site-specific recombination. Specific site recombination sequences can be used independently for recombination. Using specific site recombination sequences for recombination requires the participation of a protein factor in catalysis. The protein factor can be a specific recombinase.

[0088] The specific recombinase can recognize and target any nucleotide fragment on the specific site recombination sequence. The specific recombinase can recognize and target the specific recombination site on the specific site recombination sequence. The specific recombinase can remove the nucleic acid sequence between the specific recombination sites. The specific recombinase can remove the nucleic acid sequence between the specific recombination sites and the specific recombination site sequence. The specific recombinase can insert a new nucleic acid sequence in the forward or reverse direction while removing the nucleic acid sequence between the specific recombination sites. The new nucleic acid sequence inserted by the specific recombinase can be endogenous. The new nucleic acid sequence inserted by the specific recombinase can be exogenous. The new nucleic acid fragment added by the specific recombinase can encode a sense fragment. The new nucleic acid fragment added by the specific recombinase can encode a nonsense fragment.

[0089] Specific recombinases can be derived from prokaryotes.Specific recombinases can be derived from eukaryotes.

[0090] The specific recombinase can be exogenous. A gene encoding the specific recombinase can be introduced into the host cell through genetic manipulation. The specific recombinase can be endogenous. The specific recombinase can be derived from the host cell, that is, the host cell does not need to undergo any genetic manipulation to express the specific recombinase. The specific recombinase can be derived from the same cell type as the host cell. The specific recombinase can be inducibly or constitutively expressed. The specific recombinase can be independently expressed. The expression of the specific recombinase can depend on the regulation of host cell living environment factors, including but not limited to osmotic pressure, temperature, the presence of an inducer, the cell growth period, and the presence of a compound or analogue that can change the secondary structure or supercoil structure of DNA.

[0091] The specific recombinase may be a transposase, a site-specific recombinase, or a homologous recombinase. The specific recombinase may include tyrosine recombinases and / or serine recombinases. The specific recombinase may include one or more of XerC, XerD, CodV, RipX, Cre, Int, Xis, P22, Flp, and R1. The specific recombinase may be XerC and / or XerD, or variants or functional fragments thereof. The specific recombinase may be CodV. The specific recombinase may be RipX. The specific recombinase may be Cre. The specific recombinase may be Int. The specific recombinase may be Xis. The specific recombinase may be P22. The specific recombinase may be Flp. Among these, the specific recombinase may be R1.

[0092] A site-specific recombination sequence can include specific recombination sites, wherein the specific recombination sites can be recognized and bound by a specific recombinase. The specific recombination site can be a sequence. A nucleic acid molecule can have one or more specific recombination sites. When one or more specific recombination sites are present on a nucleic acid molecule, one or more specific recombinases can act on the sites, thereby excising the portion of nucleic acid between the sites. The specific recombination sites and recombinases can simultaneously excise the nucleic acid sequence between the sites and insert new exogenous and / or endogenous nucleic acid sequences in either the forward or reverse direction.

[0093] The specific recombination site sequence may comprise one or more specific recombination sites. The specific recombination sites may be two or more adjacent sequences. The specific recombination sites may be two or more non-adjacent sequences. The sequences of the two specific recombination sites may be identical or reverse complementary. The sequences of the two specific recombination sites may be different.

[0094] Specific recombination sites may be derived from prokaryotes. Specific recombination sites may be derived from eukaryotes. Specific recombination sites may include one or more of Ecdif, cer, psi, dif, mwr, Bsdif, loxP, FRT, and RS. Specific recombination sites may be Ecdif. Specific recombination sites may be cer. Specific recombination sites may be psi. Specific recombination sites may be dif. Specific recombination sites may be mwr. Specific recombination sites may be Bsdif. Specific recombination sites may be loxP. Specific recombination sites may be FRT. Specific recombination sites may be RS.

[0095] The specific site recombination sequence may further include an accessory sequence. The accessory sequence may direct specific recombination between specific recombination sites. The accessory sequence may direct the specific site recombination sequence to dissociate plasmid dimers. The accessory sequence may be recognized and targeted for binding by an accessory protein. The accessory protein may target and bind to the accessory sequence by recognizing any nucleic acid fragment on the accessory sequence. The accessory protein may target and bind to the accessory sequence by recognizing an accessory site on the accessory sequence. An accessory sequence may further include one or more accessory sites. The specific site recombination sequence may further include one or more accessory sequences. One or more specific recombination sites may be functionally associated with one or more accessory sites.

[0096] The accessory protein may be endogenous. The accessory protein may be inducible. The accessory protein may be one or more of PepA, ArgR, and ArcA. The accessory protein may be an inactivated form of any of PepA, ArgR, and ArcA. The accessory sequence may comprise a binding site for one or more of PepA, ArgR, and ArcA.

[0097] The isolated nucleic acid molecule assembly described in the present application, wherein the specific site recombination sequence may contain, in order from 5' to 3' direction:

[0098] a) the first common sequence as shown in SEQ ID NO: 45 or its reverse complement; or

[0099] b) a second common sequence as shown in SEQ ID NO: 46 or its reverse complement.

[0100] The specific site recombination sequence may contain, in order from 5' to 3' direction:

[0101] a) a first common sequence as shown in any one of SEQ ID NOs: 38-40 or its reverse complementary sequence; and

[0102] b) a second common sequence as shown in SEQ ID NO: 41 or its reverse complement.

[0103] Site-specific recombination sequences may include:

[0104] a) a first common sequence as shown in SEQ ID: 38 or its reverse complementary sequence and a second common sequence as shown in SEQ ID: 41 or its reverse complementary sequence; or

[0105] b) comprising a first common sequence as shown in SEQ ID: 39 or its reverse complementary sequence and a second common sequence as shown in SEQ ID: 41 or its reverse complementary sequence; or

[0106] c) comprising a first common sequence as shown in SEQ ID: 40 or its reverse complementary sequence and a second common sequence as shown in SEQ ID: 41 or its reverse complementary sequence.

[0107] There may be a 6-8 bp nucleotide sequence between the first common sequence and the second common sequence.

[0108] The specific site recombination sequence may include a sequence as shown in any one of SEQ ID NOs.: 18-37 or a variant or homolog thereof. The specific site recombination sequence may include a reverse complementary sequence of a sequence as shown in any one of SEQ ID NOs.: 18-37 or a variant or homolog thereof.

[0109] The isolated nucleic acid molecule assembly described herein, wherein the replication origin sequence can be a stringent replication origin sequence or a relaxed replication origin sequence. The replication origin sequence can participate in regulating the copy number of the nucleic acid molecule in the host cell. The replication origin sequence can participate in regulating the yield of the nucleic acid molecule. The replication origin sequence can be dependent on other regulatory components. The replication origin sequence can function independently.

[0110] The replication origin sequence may be derived from one or more of ColE1, pMB1, pBR322, pSC101, R6K and P15A.

[0111] The replication origin sequence may include a sequence as shown in any one of SEQ ID NOs.: 11-12 or a variant or homolog thereof. The replication origin sequence may include a reverse complementary sequence of a sequence as shown in any one of SEQ ID NOs.: 11-12 or a variant or homolog thereof.

[0112] The isolated nucleic acid molecule component described in the present application, wherein the first transition sequence can be a nucleic acid fragment located downstream of the specific site recombination sequence and upstream of the replication origin sequence, or upstream of the replication origin and downstream of the specific site recombination sequence.

[0113] The first transition sequence can comprise one or more functional elements of a plasmid, such as promoter sequences. Examples include one or more of CMV, EF1a, SV40, PGK1, Ubc, CAG, TRE, USA, Ac5, CaMkIIa, GAL1 / 10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, and U6. The transition sequence can be a sequence encoding a protein, polypeptide, or amino acid fragment. The first transition sequence can be a nonsense sequence.

[0114] The first transition sequence can be connected to and / or overlap with the nucleic acid fragment of the specific site recombination sequence and the replication origin sequence. The first transition sequence can be an independent nucleic acid sequence. The first transition sequence can be endogenous. The first transition sequence can be exogenous.

[0115] The first transition sequence can be a nonsense nucleic acid fragment. The first transition sequence can be a functional nucleic acid fragment. The first transition sequence can be a screening sequence, wherein the screening sequence can include a prokaryotic screening sequence and / or a eukaryotic screening sequence. The prokaryotic screening sequence can include one or more of an ampicillin resistance gene, a kanamycin resistance gene, a chloramphenicol resistance gene, a spectinomycin resistance gene, a tetracycline resistance gene, a bleomycin resistance gene, a streptomycin resistance gene, a hygromycin resistance gene, a gentamicin resistance gene, a hexaampromycin resistance gene, an erythromycin resistance gene, and a blasticidin gene.

[0116] The screening sequence may include a sequence as shown in any one of SEQ ID NOs.: 13-17 or a variant or homolog thereof. The screening sequence may include a reverse complementary sequence of a sequence as shown in any one of SEQ ID NOs.: 13-17 or a variant or homolog thereof.

[0117] The nucleic acid molecule may further include one or more of the following components: a promoter sequence, a specific enzyme cleavage site sequence, a tag sequence, and a terminator sequence.

[0118] The above nucleic acid molecule may further include a promoter sequence. The promoter sequence may be located at any position of the nucleic acid molecule, for example:

[0119] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the specific site recombination sequence, and the first transition sequence, and the promoter sequence can be located 5' upstream of the replication initiation site sequence. The promoter sequence can be located at any position between the replication initiation site sequence and the specific site recombination sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The promoter sequence can be located 3' downstream of the first transition sequence.

[0120] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence, and the promoter sequence can be located 5' upstream of the replication initiation site sequence. The promoter sequence can be located at any site between the replication initiation site sequence and the first transition sequence. The promoter sequence can be located at any site between the specific site recombination sequence and the first transition sequence. The promoter sequence can be located 3' downstream of the specific site recombination sequence.

[0121] The order of arrangement of the replication origin site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the replication origin site sequence, and the specific site recombination sequence, and the promoter sequence can be located 5' upstream of the first transition sequence. The promoter sequence can be located at any position between the replication origin site sequence and the first transition sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the replication origin site sequence. The promoter sequence can be located 3' downstream of the specific site recombination sequence.

[0122] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the specific site recombination sequence, and the replication initiation site sequence, and the promoter sequence can be located 5' upstream of the first transition sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the replication initiation site sequence. The promoter sequence can be located 3' downstream of the replication initiation site sequence.

[0123] The order of arrangement of the replication origin site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the replication origin site sequence, and the first transition sequence, and the promoter sequence can be located 5' upstream of the specific site recombination sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the replication origin site sequence. The promoter sequence can be located at any position between the first transition sequence and the replication origin site sequence. The promoter sequence can be located 3' downstream of the first transition sequence.

[0124] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the first transition sequence, and the replication initiation site sequence, and the promoter sequence can be located 5' upstream of the specific site recombination sequence. The promoter sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The promoter sequence can be located at any position between the first transition sequence and the replication initiation site sequence. The promoter sequence can be located 3' downstream of the replication initiation site sequence.

[0125] In some embodiments, a second transition sequence is further included, wherein for sequence 2), it further comprises: a replication origin sequence, a first transition sequence, a second transition sequence, and a specific site recombination sequence.

[0126] In some embodiments, the second transition sequence is a promoter sequence.

[0127] The nucleic acid molecule may further include a promoter sequence, and the promoter sequence may include one or more of a constitutive promoter, a tissue-specific promoter and / or an inducible promoter.

[0128] The promoter can be selected from one or more of the following groups: CMV, EF1a, SV40, PGK1, Ubc, CAG, TRE, USA, Ac5, CaMkIIa, GAL1 / 10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1 and U6.

[0129] The nucleic acid molecule may further include a specific enzyme cleavage site sequence. The specific enzyme cleavage site sequence may be located at any position of the nucleic acid molecule, for example:

[0130] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specificity site recombinant sequence along the 5 ' to 3 ' direction can be the replication initiation site sequence, the specificity site recombinant sequence, the first transition sequence, and the described specificity restriction enzyme cutting site sequence can be located at the 5 ' upstream of the replication initiation site sequence. The described specificity restriction enzyme cutting site sequence can be located at any site in the middle of the replication initiation site sequence and the specificity site recombinant sequence. The described specificity restriction enzyme cutting site sequence can be located at any site in the middle of the specificity site recombinant sequence and the first transition sequence. The described specificity restriction enzyme cutting site sequence can be located at the 3 ' downstream of the first transition sequence.

[0131] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombinant sequence along the 5 ' to 3 ' direction can be the replication initiation site sequence, the first transition sequence, the specific site recombinant sequence, and the described specific enzyme cutting site sequence can be located at the 5 ' upstream of the replication initiation site sequence. The described specific enzyme cutting site sequence can be located at any site in the middle of the replication initiation site sequence and the first transition sequence. The described specific enzyme cutting site sequence can be located at any site in the middle of the specific site recombinant sequence and the first transition sequence. The described specific enzyme cutting site sequence can be located at the 3 ' downstream of the specific site recombinant sequence.

[0132] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5 ' to 3 ' direction can be the first transition sequence, the replication initiation site sequence, the specific site recombination sequence, and the described specific enzyme cutting site sequence can be located at the 5 ' upstream of the first transition sequence. The described specific enzyme cutting site sequence can be located at any site in the middle of the replication initiation site sequence and the first transition sequence. The described specific enzyme cutting site sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The described specific enzyme cutting site sequence can be located at the 3 ' downstream of the specific site recombination sequence.

[0133] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the specific site recombination sequence, the replication initiation site sequence, and the specific enzyme cutting site sequence can be located at the 5' upstream of the first transition sequence. The specific enzyme cutting site sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The specific enzyme cutting site sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The specific enzyme cutting site sequence can be located at the 3' downstream of the replication initiation site sequence.

[0134] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the replication initiation site sequence, the first transition sequence, and the specific enzyme cutting site sequence can be located at the 5' upstream of the specific site recombination sequence. The specific enzyme cutting site sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The specific enzyme cutting site sequence can be located at any site in the middle of the first transition sequence and the replication initiation site sequence. The specific enzyme cutting site sequence can be located at the 3' downstream of the first transition sequence.

[0135] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specificity site recombination sequence along the 5 ' to 3 ' direction can be the specificity site recombination sequence, the first transition sequence, the replication initiation site sequence, and the specificity restriction enzyme cutting site sequence can be located at the 5 ' upstream of the specificity site recombination sequence. The specificity restriction enzyme cutting site sequence can be located at any site in the middle of the specificity site recombination sequence and the first transition sequence. The specificity restriction enzyme cutting site sequence can be located at any site in the middle of the first transition sequence and the replication initiation site sequence. The specificity restriction enzyme cutting site sequence can be located at the 3 ' downstream of the replication initiation site sequence.

[0136] The specific enzyme cleavage site sequence can be a restriction endonuclease site sequence. The nucleic acid molecule can include one or more specific enzyme cleavage site sequences. The nucleic acid molecule can include multiple identical or different specific enzyme cleavage site sequences. The nucleic acid molecule can include multiple specific enzyme cleavage site sequences that can be recognized by isozymes.

[0137] The specific enzyme cleavage site sequence includes one or more sites selected from the following group: ApaI, BamHI, BglII, EcoRI, HindIII, KpnI, NcoI, NdeI, NheI, NotI, SacI, SalI, SphI, XbaI and XhoI.

[0138] The above nucleic acid molecule may further include a tag sequence. The tag sequence may be located at any position of the nucleic acid molecule, for example:

[0139] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the specific site recombination sequence, and the first transition sequence, and the tag sequence can be located 5' upstream of the replication initiation site sequence. The tag sequence can be located at any position between the replication initiation site sequence and the specific site recombination sequence. The tag sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The tag sequence can be located 3' downstream of the first transition sequence.

[0140] The order of arrangement of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence, and the tag sequence can be located 5' upstream of the replication initiation site sequence. The tag sequence can be located at any position between the replication initiation site sequence and the first transition sequence. The tag sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The tag sequence can be located 3' downstream of the specific site recombination sequence.

[0141] The order of arrangement of the replication origin site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the replication origin site sequence, and the specific site recombination sequence, and the tag sequence can be located 5' upstream of the first transition sequence. The tag sequence can be located at any position between the replication origin site sequence and the first transition sequence. The tag sequence can be located at any position between the specific site recombination sequence and the replication origin site sequence. The tag sequence can be located 3' downstream of the specific site recombination sequence.

[0142] The order of arrangement of the replication origin site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the specific site recombination sequence, and the replication origin site sequence, and the tag sequence can be located 5' upstream of the first transition sequence. The tag sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The tag sequence can be located at any position between the specific site recombination sequence and the replication origin site sequence. The tag sequence can be located 3' downstream of the replication origin site sequence.

[0143] The order of the replication origin site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the replication origin site sequence, and the first transition sequence, and the tag sequence can be located 5' upstream of the specific site recombination sequence. The tag sequence can be located at any position between the specific site recombination sequence and the replication origin site sequence. The tag sequence can be located at any position between the first transition sequence and the replication origin site sequence. The tag sequence can be located 3' downstream of the first transition sequence.

[0144] The order of the replication initiation site sequence, the first transition sequence, and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the first transition sequence, and the replication initiation site sequence, and the tag sequence can be located 5' upstream of the specific site recombination sequence. The tag sequence can be located at any position between the specific site recombination sequence and the first transition sequence. The tag sequence can be located at any position between the first transition sequence and the replication initiation site sequence. The tag sequence can be located 3' downstream of the replication initiation site sequence.

[0145] The tag sequence can be fused to the N-terminus and / or C-terminus of the target protein after expression. After expression, the tag sequence can be connected to the target protein through hydrogen bonds or intermolecular forces.

[0146] The tag sequence can be a detectable marker, such as a radiolabeled amino acid or a biotinylated polypeptide that can be detected by labeled avidin. The tag sequence can be a fluorescent protein sequence, an enzyme tag sequence, a chemiluminescent tag sequence, a biotin group sequence, a preselected polypeptide antigen epitope sequence that can be recognized by a second receptor, or a magnetic reagent. The tag sequence can be a fluorescent protein sequence, wherein the fluorescent protein sequence is selected from one or more of the following groups: GFP, eGFP, eYFP, eCFP, mCherry, luc2, and hrluc. Fluorescent protein sequences include: a combination of eGFP and luc2, and / or a combination of mCherry and hrluc.

[0147] The above nucleic acid molecule may further include a terminator sequence. The terminator sequence may be located at any position of the nucleic acid molecule, for example:

[0148] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the specific site recombination sequence, the first transition sequence, and the terminator sequence can be located at the 5' upstream of the replication initiation site sequence. The terminator sequence can be located at any site in the middle of the replication initiation site sequence and the specific site recombination sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The terminator sequence can be located at the 3' downstream of the first transition sequence.

[0149] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5 ' to 3 ' direction can be the replication initiation site sequence, the first transition sequence, the specific site recombination sequence, and the terminator sequence can be located at the 5 ' upstream of the replication initiation site sequence. The terminator sequence can be located at any site in the middle of the replication initiation site sequence and the first transition sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The terminator sequence can be located at the 3 ' downstream of the specific site recombination sequence.

[0150] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the replication initiation site sequence, the specific site recombination sequence, and the terminator sequence can be located at the 5' upstream of the first transition sequence. The terminator sequence can be located at any site in the middle of the replication initiation site sequence and the first transition sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The terminator sequence can be located at the 3' downstream of the specific site recombination sequence.

[0151] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the specific site recombination sequence, and the replication initiation site sequence, and the terminator sequence can be located at the 5' upstream of the first transition sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The terminator sequence can be located at the 3' downstream of the replication initiation site sequence.

[0152] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the replication initiation site sequence, the first transition sequence, and the terminator sequence can be located at the 5' upstream of the specific site recombination sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the replication initiation site sequence. The terminator sequence can be located at any site in the middle of the first transition sequence and the replication initiation site sequence. The terminator sequence can be located at the 3' downstream of the first transition sequence.

[0153] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the first transition sequence, the replication initiation site sequence, and the terminator sequence can be located at the 5' upstream of the specific site recombination sequence. The terminator sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The terminator sequence can be located at any site in the middle of the first transition sequence and the replication initiation site sequence. The terminator sequence can be located at the 3' downstream of the replication initiation site sequence.

[0154] The nucleic acid molecule may further include a terminator sequence. The terminator may be Rho-dependent. The terminator may be Rho-independent. The terminator may consist of a palindromic sequence. The terminator may assist in terminating transcription with a rho factor. The terminator may independently terminate transcription.

[0155] The nucleic acid molecule may further include an in vitro transcription sequence. The in vitro transcription sequence may be located at any position of the nucleic acid molecule, for example:

[0156] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the specific site recombination sequence, and the first transition sequence, and the in vitro transcribed sequence can be located 5' upstream of the replication initiation site sequence. The in vitro transcribed sequence can be located at any site between the replication initiation site sequence and the specific site recombination sequence. The in vitro transcribed sequence can be located at any site between the specific site recombination sequence and the first transition sequence. The in vitro transcribed sequence can be located 3' downstream of the first transition sequence.

[0157] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the replication initiation site sequence, the first transition sequence, the specific site recombination sequence, and the in vitro transcribed sequence can be located at the 5' upstream of the replication initiation site sequence. The in vitro transcribed sequence can be located at any site in the middle of the replication initiation site sequence and the first transition sequence. The in vitro transcribed sequence can be located at any site in the middle of the specific site recombination sequence and the first transition sequence. The in vitro transcribed sequence can be located at the 3' downstream of the specific site recombination sequence.

[0158] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the replication initiation site sequence, the specific site recombination sequence, and the in vitro transcribed sequence can be located 5' upstream of the first transition sequence. The in vitro transcribed sequence can be located at any site between the replication initiation site sequence and the first transition sequence. The in vitro transcribed sequence can be located at any site between the specific site recombination sequence and the replication initiation site sequence. The in vitro transcribed sequence can be located 3' downstream of the specific site recombination sequence.

[0159] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the first transition sequence, the specific site recombination sequence, and the replication initiation site sequence, and the in vitro transcribed sequence can be located 5' upstream of the first transition sequence. The in vitro transcribed sequence can be located at any site between the specific site recombination sequence and the first transition sequence. The in vitro transcribed sequence can be located at any site between the specific site recombination sequence and the replication initiation site sequence. The in vitro transcribed sequence can be located 3' downstream of the replication initiation site sequence.

[0160] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the replication initiation site sequence, the first transition sequence, and the in vitro transcribed sequence can be located 5' upstream of the specific site recombination sequence. The in vitro transcribed sequence can be located at any position between the specific site recombination sequence and the replication initiation site sequence. The in vitro transcribed sequence can be located at any position between the first transition sequence and the replication initiation site sequence. The in vitro transcribed sequence can be located 3' downstream of the first transition sequence.

[0161] The order of arrangement of the replication initiation site sequence, the first transition sequence and the specific site recombination sequence along the 5' to 3' direction can be the specific site recombination sequence, the first transition sequence, and the replication initiation site sequence, and the in vitro transcribed sequence can be located at the 5' upstream of the specific site recombination sequence. The in vitro transcribed sequence can be located at any site between the specific site recombination sequence and the first transition sequence. The in vitro transcribed sequence can be located at any site between the first transition sequence and the replication initiation site sequence. The in vitro transcribed sequence can be located at the 3' downstream of the replication initiation site sequence.

[0162] The in vitro transcribed sequence can include a promoter sequence, a 5'UTR sequence, an ORF sequence, a 3'UTR sequence, and / or a PolyA sequence. The in vitro transcribed sequence can be a promoter sequence. The in vitro transcribed sequence can be a 5'UTR sequence. The in vitro transcribed sequence can be an ORF sequence. The in vitro transcribed sequence can be a 3'UTR sequence. The in vitro transcribed sequence can be a PolyA sequence. The in vitro transcribed sequence can include a precursor messenger RNA and a mature mRNA sequence. Newly synthesized in vitro transcribed sequences can be modified in several schemes to convert them into their mature functional forms.

[0163] The isolated nucleic acid molecule described herein may include a sequence as shown in any one of SEQ ID NOs: 2-10, 48, 50, 51, or a variant or homolog thereof. The isolated nucleic acid molecule may include a reverse complementary sequence of a sequence as shown in any one of SEQ ID NOs: 2-10, 48, 50, 51, or a variant or homolog thereof.

[0164] In the present application, described homology can refer to the similarity between two or more sequences or the degree of association usually.Described homologue can be to have at least 80%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100% sequence homology with described nucleotide sequence.Described variant in the present application is, compared with described nucleotide sequence, replacement, disappearance or insertion of one or more Nucleotide occur.

[0165] Nucleic acid molecules described herein can comprise changes in coding region, non-coding region or both. Nucleic acid molecule variants can comprise changes that produce silent replacement, addition or deletion, but do not change the property or activity of its coded polypeptide. Nucleic acid molecule variants can include silent replacements (due to the degeneracy of the genetic code) that cause the amino acid sequence of its coded polypeptide to not change. Nucleic acid molecule variants can be prepared to regulate or change the expression (or expression level) of the coded polypeptide. Nucleic acid molecule variants are prepared to reduce the expression of the coded polypeptide. Compared with the parent nucleic acid molecule sequence, nucleic acid molecule variants can increase the expression of the coded polypeptide. Compared with the parent nucleic acid molecule sequence, nucleic acid molecule variants can reduce the expression of the coded polypeptide.

[0166] In the present application, the nucleic acid molecules can be synthesized by conventional methods in the art. The nucleic acid molecules are amplified in vitro, for example, by polymerase chain reaction (PCR) amplification. The nucleic acid molecules can be produced by cloning and recombination. The nucleic acid molecules can be purified, for example, by enzyme digestion and gel electrophoresis fractionation. The nucleic acid molecules can be synthesized, for example, by chemical synthesis. The nucleic acid molecules can be prepared by recombinant DNA technology.

[0167] The main purpose of the present invention is to overcome the problems existing in the prior art and provide a gene for maintaining monomer supercoiling in polyadenylated plasmids. The gene can significantly increase the monomer supercoiling ratio of polyadenylated plasmids. The application of the gene is also provided.

[0168] The technical solution of the present invention to solve the technical problem is as follows:

[0169] A gene for maintaining supercoiling of polyadenylated plasmid monomers, wherein the gene sequence comprises, in order from 5' to 3' direction, a first common sequence of SEQ ID No. 45 and a second common sequence of SEQ ID No. 46; in the first common sequence and the second common sequence, r is g or a, s is g or c, n is one of a, g, c, and t, and y is t or c; and a nucleotide sequence of 6-8 bp is separated from the first common sequence and the second common sequence.

[0170] The above genes can effectively dissolve plasmid aggregates, increase the supercoiling ratio of plasmid monomers, and at the same time increase plasmid yield.

[0171] The present invention also provides:

[0172] A method for increasing the supercoil ratio of polyadenylated plasmid monomers, comprising the following steps:

[0173] The first step is to introduce the gene sequence for maintaining the supercoiling of the polyadenylation plasmid monomer or its reverse complementary sequence into the polyadenylation plasmid to obtain a recombinant plasmid;

[0174] In the second step, the recombinant plasmid obtained in the first step is transformed into competent cells, and the correct clones are obtained by screening, culturing and sequencing.

[0175] Preferably, in the first step, the polyadenylation plasmid is an mRNA plasmid with a long PolyA tail, and the length of the long PolyA tail is 60 to 150 nucleotides.

[0176] Preferably, in the first step, the polyadenylation plasmid has an ori element, an insertion sequence site, and an antibiotic resistance gene element, wherein the insertion sequence site is located between the ori element and the antibiotic resistance gene element; a gene sequence for maintaining supercoiling of the polyadenylation plasmid monomer or its reverse complementary sequence is introduced into the insertion sequence site in the polyadenylation plasmid.

[0177] Preferably, in the second step, a culture medium containing an antibiotic corresponding to the antibiotic resistance gene element in the polyadenylation plasmid is used during screening.

[0178] Preferably, in the second step, Sanger sequencing is used for sequencing.

[0179] The above method can effectively improve the supercoil ratio of polyadenylated plasmid monomers and the plasmid yield.

[0180] The present invention also provides:

[0181] A method for constructing a recombinant plasmid comprises: introducing the sequence of the gene for maintaining the supercoiling of the polyadenosine plasmid monomer or its reverse complementary sequence into the polyadenosine plasmid to obtain the recombinant plasmid.

[0182] The recombinant plasmid was constructed by the recombinant plasmid construction method described above.

[0183] The present invention also provides:

[0184] The aforementioned gene for maintaining monomer supercoiling of polyA plasmids is used to recombinant polyA plasmids to increase their monomer supercoiling ratio and plasmid yield.

[0185] Preferably, the polyadenylation plasmid is an mRNA plasmid with a long PolyA tail, and the length of the long PolyA tail is 60 to 150 nucleotides.

[0186] Compared with the existing technology, the present invention can significantly increase the supercoiling ratio of mRNA plasmid monomers with long PolyA tails (length of 60 to 150 nucleotides), reaching the highest level of supercoiling ratio of ordinary plasmid monomers (>95%), while also increasing plasmid yield. The present invention overcomes the limitations of the existing technology and can improve the stability and consistency of long PolyA tail mRNA plasmids, thereby providing key technical support for the research and development and production of mRNA therapies, greatly promoting the development of the biopharmaceutical field and the speed of market access for related drugs.

[0187] In a specific implementation, the gene for maintaining supercoiling of polyadenylated plasmid monomers of the present invention comprises, in sequence from 5' to 3', a first common sequence ggtrcsnayaa (SEQ ID No. 1) and a second common sequence ttatggtaaay (SEQ ID No. 2), wherein r is g or a, s is g or c, n is one of a, g, c, and t, and y is t or c; and a 6-8 bp nucleotide sequence is separated from the first common sequence and the second common sequence.

[0188] The above gene is one of ColE cer minimal, ColA cer minimal, and CDF cer minimal. Note: The above three sequences are collectively referred to as truncated depolymerization sequences.

[0189] Alternatively, the gene is one of ColE cer, ColA cer, and CDF cer. Note: The above three sequences are collectively referred to as depolymerization sequences.

[0190] It should be noted that ColE cer, ColA cer, and CDF cer are respectively truncated and the core regions are retained to obtain ColE cer minimal, ColA cer minimal, and CDF cer minimal.

[0191] In ColE cer or ColE cer minimal, the first common sequence is ggtgcgtacaa, the second common sequence is ttatggtaaat, and the first common sequence and the second common sequence are separated by an 8 bp nucleotide sequence ttaaggga.

[0192] In ColA cer or ColA cer minimal, the first common sequence is ggtgccgacaa, the second common sequence is ttatggtaaat, and the first common sequence and the second common sequence are separated by a 6 bp nucleotide sequence cggatg.

[0193] In CDF cer or CDF cer minimal, the first common sequence is ggtaccgataa, the second common sequence is ttatggtaaat, and a 6-bp nucleotide sequence gggatg is spaced between the first common sequence and the second common sequence.

[0194] When used, the above gene sequence or its reverse complementary sequence is introduced into any position and any direction of the Poly(A) plasmid, which can significantly increase the supercoil ratio and yield of the plasmid monomer.

[0195] 2. Recombinant vector

[0196] On the other hand, the present application also provides a recombinant vector, which may contain the isolated nucleic acid molecule.

[0197] The recombinant vector may be an expression vector or a cloning vector. The recombinant vector may be a plasmid, a cosmid, a phage or a viral vector. The recombinant vector may be a plasmid.

[0198] The recombinant vector may contain the target gene. The recombinant vector may contain expression control elements that allow the coding region to be expressed in the host cell. The expression control elements may be cis-acting elements. The recombinant vector may include, but is not limited to, one or more of the following components: a promoter sequence, a specific restriction enzyme cleavage site sequence, a tag sequence, an enhancer sequence, a silencer sequence, an insulator sequence, and a terminator sequence.

[0199] On the other hand, the present application also provides a method for constructing a recombinant vector comprising the isolated nucleic acid molecule.

[0200] In some embodiments, the method may comprise introducing the nucleic acid molecule into a polyadenylated plasmid. In some embodiments, the method may comprise introducing the nucleic acid molecule into a supercoiled plasmid.

[0201] In some embodiments, the nucleic acid molecule can be introduced into a recombinant vector by a restriction endonuclease. In some embodiments, the restriction endonuclease can target one or more of the following groups of sites for cleavage: ApaI, BamHI, BglII, EcoRI, HindIII, KpnI, NcoI, NdeI, NheI, NotI, SacI, SalI, SphI, XbaI, and / or XhoI.

[0202] In some embodiments, the nucleic acid molecule can be introduced into the recombinant vector by a homologous recombinase. In some embodiments, the nucleic acid molecule can be introduced into the recombinant vector by a site-specific recombinase.

[0203] 3. Host cells

[0204] On the other hand, the present application also provides a host cell, which may contain the isolated nucleic acid molecule and / or the recombinant vector.

[0205] The host cell can be an attenuated or non-attenuated host cell. The host cell can be a prokaryotic cell. The host cell can be a gram-negative bacterial cell. The host cell can be one or more selected from the genus Escherichia, Salmonella, Shigella, Agrobacterium, Pseudomonas and Vibrio. The host cell can be selected from Escherichia coli and Salmonella enterica, including one or more of Salmonella typhi and Salmonella typhimurium.

[0206] The host cell may be a Gram-positive bacterial cell. The host cell may be one or more selected from the genus Bacillus, Streptomyces, Listeria, Lactobacillus, Lactococcus, and Mycobacterium. The host cell may be one or more selected from the genus Bacillus subtilis or Mycobacterium bovis (e.g., BCG strain).

[0207] The host cell may be an archaeon. The host cell may be a yeast. The host cell may be one or more selected from the genera Hansenula, Pichia, Saccharomyces, and Schizosaccharomyces.

[0208] The host cell may be a non-fungal eukaryotic cell capable of replicating and / or expressing a recombinant vector. The host cell may be one or more selected from the genera Chlamydomomas, Dictyostelium, and Entamoeba. The host cell may be a mammalian cell.

[0209] Host cells can provide the protein factors required for nucleic acid recombination. Host cells can provide specific recombinases and accessory proteins. Host cells can provide nucleic acid molecules encoding specific recombinases and / or accessory proteins. Host cells can provide the reaction sites required for nucleic acid recombination.

[0210] 4. Diagnostic or pharmaceutical compositions

[0211] In another aspect, the present application further provides a diagnostic or pharmaceutical composition, which may comprise the isolated nucleic acid molecule, the recombinant vector and / or the host cell, and optionally a pharmaceutically acceptable adjuvant. The pharmaceutically acceptable adjuvant may include a buffer, an antioxidant, a preservative, a low molecular weight polypeptide, a protein, a hydrophilic polymer, an amino acid, a sugar, a chelating agent, a counterion, a metal complex, and / or a nonionic surfactant.

[0212] The diagnostic or pharmaceutical composition is useful in immunodiagnosis or therapy. The diagnostic or pharmaceutical composition may be useful in immuno-oncology. The diagnostic or pharmaceutical composition may be useful in diagnosing and / or inhibiting tumor growth in a subject. The diagnostic or pharmaceutical composition may be useful in treating cancer in a subject. The diagnostic or pharmaceutical composition may be useful in diagnosing and / or treating viral infection in a subject. The diagnostic or pharmaceutical composition may be useful in preventing disease.

[0213] The diagnostic or pharmaceutical composition can be formulated for oral administration, intravenous administration, intramuscular administration, in situ administration at the tumor site, inhalation, rectal administration, vaginal administration, transdermal administration or administration via a subcutaneous reservoir. The pharmaceutical composition can also be prepared as a solution, suspension, tablet, pill, capsule and depot.

[0214] The present application also provides a method for preparing the diagnostic or pharmaceutical composition.

[0215] 5. Methods for expressing target genes

[0216] On the other hand, the present application also provides a method for expressing a target gene, the method comprising introducing the isolated nucleic acid molecule and / or the recombinant vector into a host cell, and allowing the target gene to be expressed in the host cell.

[0217] The method can be an in vitro or ex vivo method. The method can be an in vivo method. The target gene can be recombined into the nucleic acid molecule or recombinant vector. The target gene can be derived from a nucleic acid molecule or a recombinant vector.

[0218] In another aspect, the present application also provides a method for screening target gene expression, which comprises introducing a screening sequence into the isolated nucleic acid molecule and / or the recombinant vector.

[0219] 6. Test kit

[0220] On the other hand, the present application also provides a kit comprising the isolated nucleic acid molecule, the vector, the host cell and / or the diagnostic or pharmaceutical composition.

[0221] The kit may include the isolated nucleic acid molecules, the vectors or the host cells, and / or the diagnostic or pharmaceutical compositions disclosed herein in one or more containers. The kit may include instructions for preparing and / or administering the isolated nucleic acid molecules, the vectors or the host cells and / or cell populations disclosed herein.

[0222] The kit may further include gene cloning primers. The kit may further include gene cloning tools, including but not limited to restriction endonucleases, homologous recombinases and / or site-specific recombinases.

[0223] On the other hand, the present application also provides a method for preparing the kit.

[0224] 7. Purpose

[0225] On the other hand, the present application also provides the use of the isolated nucleic acid molecule, the vector, the host cell, the diagnostic or pharmaceutical composition and / or the kit for improving the supercoiling ratio, monomer yield and PolyA stability of polyadenylated plasmid monomers.

[0226] On the other hand, the present application also provides the use of the isolated nucleic acid molecule, the vector, the host cell, the diagnostic or pharmaceutical composition and / or the kit in the preparation of mRNA drugs and / or vaccines.

[0227] On the other hand, the present application also provides the use of the isolated nucleic acid molecule, the vector, the host cell, the diagnostic or pharmaceutical composition and / or the kit in mRNA therapy.

[0228] On the other hand, the present application also provides the use of the isolated nucleic acid molecule, the vector, the host cell, the diagnostic or pharmaceutical composition and / or the kit in the treatment of tumors, viral diseases, cardiovascular diseases, metabolic diseases, and nervous system diseases.

[0229] On the other hand, the present application also provides the use of the isolated nucleic acid molecule, the vector, the host cell, the diagnostic or pharmaceutical composition and / or the kit in expressing a target gene.

[0230] Without intending to be bound by any theory, the following examples are merely intended to illustrate the isolated nucleic acid molecules, preparation methods and uses of the present application, and are not intended to limit the scope of the present invention.

[0231] Example

[0232] Example 1 Construction and Functional Verification of a Recombinant Vector Containing a Supercoiled Gene Maintaining Polyadenylation Plasmid Monomer

[0233] The gene used in this application to maintain supercoiling of polyadenylated plasmid monomers has a core region of its sequence that belongs to the XerCD binding sequence and contains several auxiliary sequences. However, not all genes containing XerCD binding sequences will achieve the same effect. In this example, a control group was established: the gene pSC101 psi (SEQ ID NO: 18) containing the XerCD binding sequence was used as a control group for comparative verification.

[0234] By recombining the gene that maintains the supercoiled polyadenylated plasmid monomer, the following five recombinant plasmids were obtained:

[0235] pKGCT7-Fluc (SEQ ID NO.: 1): This empty vector lacks the cer gene sequence. Its structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, and first transition sequence. A map of the pKGCT7-Fluc plasmid is shown in Figure 1.

[0236] pKGCT7-Fluc BW: The reverse complement sequence (SEQ ID NO.: 28) of pSC101 psi (SEQ ID NO.: 18) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. The structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, the reverse complement sequence of pSC101 psi, and the first transition sequence.

[0237] pKGCT7-Fluc BX (SEQ ID NO.: 2): The reverse complement of ColE cer (SEQ ID NO.: 19) (SEQ ID NO.: 29) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. Its structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, the reverse complement of ColE cer, and the first transition sequence.

[0238] pKGCT7-Fluc BY (SEQ ID NO.: 3): The reverse complement sequence (SEQ ID NO.: 30) of CDF cer (SEQ ID NO.: 20) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. Its structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, reverse complement sequence of CDF cer, and the first transition sequence.

[0239] pKGCT7-Fluc BZ (SEQ ID NO.: 4): The reverse complement sequence (SEQ ID NO.: 31) of ColA cer (SEQ ID NO.: 21) is inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. Its structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, the reverse complement sequence of ColA cer, and the first transition sequence.

[0240] As an example, the plasmid map of pKGCT7-Fluc BX is shown in Figure 2 .

[0241] The five plasmids were transformed into NEB Stable competent cells (NEB C3040) according to the manufacturer's instructions and cultured on kanamycin-containing LB plates at 30°C overnight. Plaques were picked and shaken in 2 mL of LB medium at 30°C for Sanger sequencing. The proportion of colonies with a poly A length greater than 110 was determined by streaking.

[0242] The sequencing primer sequence is: GTTTCGCCACCTCTGACTTG (sequence: SEQ ID NO: 42).

[0243] Four clones that were sequenced correctly were selected, and the plasmid yield and monomer supercoil ratio were detected and calculated. The monomer supercoil ratio was determined by agarose gel electrophoresis, and the electrophoretogram is shown in Figure 7. The results are shown in the following table.

[0244] Table 1 Recombinant vector plasmid yield and monomer supercoil ratio

[0245] From the above results, it can be seen that although pSC101 psi contains the XerCD binding sequence, it cannot increase the supercoiling ratio of PolyA plasmid monomers, but instead makes the supercoiling ratio of monomers worse.

[0246] ColE cer, CDF cer, and ColA cer can significantly increase the supercoiled ratio of poly(A) plasmid monomers, even to the same level as that of standard plasmid monomers (>95%). Plasmid yield is also significantly improved.

[0247] Example 2 Functional verification of different lengths of supercoiled genes maintaining polyadenylation plasmid monomers

[0248] The reverse complement of ColE cer (SEQ ID NO.: 19), CDF cer (SEQ ID NO.: 20), and ColA cer (SEQ ID NO.: 21) sequences were cloned into pKGCT7-Fluc, with the specific recombination site located between the replication origin sequence and the first transition sequence of pKGCT7-Fluc (4210-4211), to obtain the following recombinant plasmid:

[0249] pKGCT7-Fluc BX: same as Example 1.

[0250] pKGCT7-Fluc BY: same as Example 1.

[0251] pKGCT7-Fluc BZ: same as Example 1.

[0252] ColE cer (SEQ ID NO.: 19), CDF cer (SEQ ID NO.: 20), and ColA cer (SEQ ID NO.: 21) were truncated and the core regions retained to obtain ColE cer minimal (SEQ ID NO.: 25), CDF cer minimal (SEQ ID NO.: 26), and ColA cer minimal (SEQ ID NO.: 27). The reverse complement of the minimal sequences was cloned into pKGCT7-Fluc, with the specific recombination site located between the replication origin sequence and the first transition sequence (4210-4211), to obtain the following recombinant plasmids:

[0253] pKGCT7-Fluc BXm: The reverse complement of ColE cer minimal (SEQ ID NO.: 25) (SEQ ID NO.: 35) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. The pKGCT7-Fluc construct consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, the reverse complement of ColE cer minimal, and the first transition sequence.

[0254] pKGCT7-Fluc BYm: The reverse complement of CDF cer minimal (SEQ ID NO.: 26) (SEQ ID NO.: 36) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. The structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, reverse complement of CDF cer minimal, and the first transition sequence.

[0255] pKGCT7-Fluc BZm: The reverse complement of ColA cer minimal (SEQ ID NO.: 27) (SEQ ID NO.: 37) was inserted into the specific recombination site (4210-4211) of pKGCT7-Fluc. The structure consists of the CMV enhancer, CMV promoter, β-globin intron, T7 promoter AGG, luciferase, PolyA, bGH PolyA signal, replication origin sequence, reverse complement of ColA cer minimal, and the first transition sequence.

[0256] As an example, the map of pKGCT7-Fluc BXm is shown in FIG3 .

[0257] Transform the six plasmids described above onto NEB Stable competent cells (NEB C3040) and incubate overnight at 30°C on kanamycin-containing LB plates. Plaques were picked and shaken in 2 mL of LB medium at 30°C. The bacterial suspension of each clone was streaked to determine the proportion of colonies with a poly A length greater than 110. The sequencing primer sequence was: GTTTCGCCACCTCTGACTTG (SEQ ID NO: 42).

[0258] Cells transformed with the pKGCT7-Fluc plasmid were used as a control group to compare and verify the effects of the plasmids cloned with the cer sequence and the cer minimal sequence on supercoiling, monomer yield, and PolyA stability.

[0259] Table 2 Yield and supercoil ratio of recombinant plasmids containing cer sequences of different lengths

[0260] The results showed that the introduction of a minimal sequence that was a truncated version of the cer sequence increased the plasmid yield, but the effect of maintaining plasmid supercoiling was not as good as the full-length cer sequence.

[0261] Example 3: Verification of the function of the supercoiled gene in the polyadenylated plasmid monomer by the location selection of the specific recombination site

[0262] This application uses the ColE cer (SEQ ID NO.: 19) sequence, which maintains the supercoiling of polyadenylated plasmid monomers. The core region of the sequence belongs to the XerCD binding sequence and contains some auxiliary sequences. The ColE cer (SEQ ID NO.: 19) sequence was cloned into the middle of the pKGCT7-Fluc intron sequence and promoter sequence (682-683) and the middle of the pKGCT7-Fluc replication origin sequence and the first transition sequence (4210-4211), respectively, to obtain the following recombinant plasmids:

[0263] pKGCT7-Fluc BX: same as Example 1.

[0264] pKGCT7-Fluc-cer (SEQ ID NO.: 7): ColE cer (SEQ ID NO.: 19) was inserted into the specific recombination site of pKGCT7-Fluc. The plasmid map is shown in Figure 6. Its structure consists of the CMV enhancer, CMV promoter, ColE cer, β-globin intron, T7 promoter AGG, luciferase (Luc), PolyA, bGH PolyA signal, replication origin sequence, and first transition sequence.

[0265] pKGCT-Fluc-cer (SEQ ID NO.: 48): The Cer insertion position of pKGCT-Fluc-cer is behind the CMV promoter. pKGCT7-Fluc is digested with SalI and recombined with the PCR product (SEQ ID NO.: 49) using the recombination method.

[0266] AX-Fluc (SEQ ID NO.: 50): ColE cer (SEQ ID NO.: 19) is inserted downstream of pUC origin. The plasmid map is shown in Figure 8. The structure consists of the CAP binding site, lac promoter, lac operator, M13 reverse primer, T7 promoter AGG, fluorescent protein sequence, PolyA, M13 forward primer, AmpR promoter, first transition sequence (e.g., KanR), replication origin sequence, and ColE cer.

[0267] pLevo-D1-Fluc (SEQ ID NO.: 51): ColE cer (SEQ ID NO.: 19) is inserted upstream of pUC origin and downstream of the first transition sequence (e.g., KanR), as shown in the plasmid map in Figure 9. The structure consists of the T7 promoter AGG, the fluorescent protein sequence, polyA, the first transition sequence (e.g., KanR), ColE cer, and the replication origin sequence. These plasmids and the control pKGCT7-Fluc plasmid were co-transformed into NEB Stable competent cells (NEB C3040) and incubated overnight at 30°C with kanamycin in LB plates. The cells were then plated and shaken in 2 mL of LB medium at 30°C. The percentage of colonies with a polyA length greater than 110 was determined by streaking. The sequencing primer sequence for pKGCT7-Fluc is: GTTTCGCCACCTCTGACTTG (SEQ ID NO: 42). The sequencing primer sequence for AX-Fluc is: GATGTGCTGCAAGGCGATTA (SEQ ID No. 47). The sequencing primer of pLevo-D1-Fluc is CGTTTCCCGTTGAATATGGC (SEQ ID NO: 43).

[0268] Table 3 Effect of PolyA stability of recombinant plasmids containing cer sequences at different insertion positions

[0269] The effects of supercoiling, monomer yield and PolyA stability of the plasmids cloned with cer sequences at different sites were determined and compared and verified.

[0270] Table 4 Yield and supercoiling ratio of recombinant plasmids containing cer sequences at different insertion positions

[0271] When the specific recombination site is located downstream of the replication origin sequence and the first transition sequence (the order is the first transition sequence (such as the screening gene KanR), the replication origin sequence, and the specific recombination site sequence, as shown in AX-Fluc), the ColE cer sequence maintains the stability of PolyA.

[0272] Example 4 Verification of the Function of Polyadenylation Plasmid Monomer Supercoiled Genes by Different Plasmid Backbones

[0273] This application uses the ColE cer (SEQ ID NO.: 19) sequence, which maintains the supercoiling of polyadenylated plasmid monomers. The core region of the sequence belongs to the XerCD binding sequence and contains some auxiliary sequences. The ColE cer (SEQ ID NO.: 19) sequence was cloned into AX-GFP-Firefly (4994-0) and AX-RFP-Renilla (4244-0), respectively, with the specific recombination site located downstream of the replication origin sequence and the first transition sequence, to obtain the following recombinant plasmids:

[0274] AX-GFP-Firefly (SEQ ID NO.: 6): ColE cer (SEQ ID NO.: 19) was inserted into the specific recombination site of AX-GFP-Firefly. The plasmid map is shown in Figure 5. Its structure consists of the CAP binding site, lac promoter, lac operator, M13 reverse primer, T7 promoter AGG, green fluorescent protein, luciferase, PolyA, M13 forward primer, AmpR promoter, first transition sequence, replication origin sequence, and ColE cer.

[0275] AX-RFP-Renilla (SEQ ID NO.: 5): ColE cer (SEQ ID NO.: 19) was inserted into the specific recombination site of AX-RFP-Renilla. The plasmid map is shown in Figure 6. The structure consists of the CAP binding site, lac promoter, lac operator, M13 reverse primer, T7 promoter AGG, red fluorescent protein, luciferase (Rluc), PolyA, M13 forward primer, AmpR promoter, first transition sequence, replication origin sequence, and ColE cer.

[0276] The above two plasmids and the control pKGCT7-Fluc BX plasmid were co-transformed into NEB Stable competent cells (NEB C3040) and cultured overnight at 30°C with kanamycin in LB medium. The cells were then plaque-picked and shaken in 2 mL of LB medium at 30°C. The bacterial suspension of each clone was streaked to determine the proportion of colonies with a poly A length greater than 110. The sequencing primer sequence for pKGCT7-Fluc BX is: GTTTCGCCACCTCTGACTTG (SEQ ID NO: 42). The sequencing primer sequences for AX-GFP-Firefly and AX-RFP-Renilla are: GATGTGCTGCAAGGCGATTA (SEQ ID No. 47).

[0277] Table 5 Effect of PolyA stability of recombinant plasmids containing cer sequences on different plasmid backbones

[0278] The effects of supercoiling and monomer yield of the recombinant plasmids on the above-mentioned different plasmid backbones were determined and compared and verified.

[0279] Table 6 Yield and supercoil ratio of recombinant plasmids containing cer sequences on different plasmid backbones

[0280] The results showed that the PolyA stability of AX-GFP-Firefly and AX-RFP-Renilla was superior to that of pKGCT7-Fluc BX, without compromising the supercoiling ratio and monomer yield of the plasmids, maintaining a good effect. Specifically, the AX plasmid backbone (with a specific recombination site located downstream of the replication origin sequence and the first transition sequence, in the order of the first transition sequence (e.g., the selection gene KanR), the replication origin sequence, and cer) exhibited better PolyA stability than the pKGCT7-Fluc backbone, while also achieving efficient maintenance of plasmid supercoiling and high monomer yield.

[0281] Example 5: Suitability of maintaining polyadenylated plasmid monomeric supercoiled genes

[0282] The ColE cer (SEQ ID NO.: 19) sequence was cloned into GFP124A, with the specific recombination site located downstream of the replication origin sequence and the first transition sequence (2189-2190), to obtain the following recombinant plasmid:

[0283] GFP124A: does not contain the gene for maintaining supercoiling of polyadenylated plasmid monomers; its sequence is SEQ ID NO.: 44.

[0284] GFP124A-BX: ColE cer (SEQ ID NO.: 19) was inserted into the insertion sequence site (2189-2190) of GFP124A.

[0285] The two plasmids were transformed into NEB Stable competent cells (NEB C3040) according to the manufacturer's instructions and cultured on kanamycin-containing LB plates at 30°C overnight. Plaques were picked and shaken in 2 mL of LB medium at 30°C. Sanger sequencing was performed to confirm that the poly A sequence length was between 120 and 130 nucleotides.

[0286] The sequencing primer sequence is: CGTTTCCCGTTGAATATGGC (SEQ ID NO: 43).

[0287] Four clones that were sequenced correctly were selected and tested for plasmid yield and monomer supercoil ratio. The results are shown in the table below.

[0288] Table 7 Yield and supercoiling ratio of GFP124A and GFP124A-BX recombinant plasmids

[0289] From the above results, it can be seen that ColE cer can also increase the monomer supercoiling ratio and plasmid yield of other PolyA plasmids such as GFP124A.

Claims

1. An isolated nucleic acid molecule comprising the following components: Replication origin sequence, first transition sequence and specific site recombination sequence.

2. The nucleic acid molecule of claim 1, wherein the order of the modules along the 5' to 3' direction can be one or more of the following: (1) Replication origin sequence, specific site recombination sequence, first transition sequence; or (2) replication origin sequence, first transition sequence, specific site recombination sequence; or (3) a first transition sequence, a replication origin sequence, or a specific site recombination sequence; or (4) a first transition sequence, a specific site recombination sequence, or a replication origin sequence; or (5) specific site recombination sequence, replication origin sequence, first transition sequence; or (6) Specific site recombination sequence, first transition sequence, and replication initiation site sequence.

3. The nucleic acid molecule according to any one of claims 1 to 2, further comprising one or more of the following components: a promoter sequence, a specific enzyme cleavage site sequence, a tag sequence, and a terminator sequence.

4. The nucleic acid molecule according to any one of claims 1 to 3, further comprising the following component: an in vitro transcription sequence.

5. The nucleic acid molecule of any one of claims 1 to 4, wherein the in vitro transcribed sequence comprises a Poly A sequence.

6. The nucleic acid molecule according to any one of claims 1 to 5, further comprising a second transition sequence, wherein for sequence 2), further comprises: a replication initiation site sequence, a first transition sequence, a second transition sequence, and a specific site recombination sequence.

7. The nucleic acid molecule of any one of claims 1 to 6, wherein the second transition sequence comprises a promoter sequence.

8. The nucleic acid molecule according to any one of claims 1 to 7, wherein the specific site recombination sequence comprises a specific recombination site, and the specific recombinase performs recombination based on the specific recombination site.

9. The nucleic acid molecule according to any one of claims 1 to 8, wherein the specific recombinase comprises tyrosine recombinases and / or serine recombinases.

10. The nucleic acid molecule of claim 9, wherein the specific recombinase comprises one or more of XerC, XerD, CodV, RipX, Cre, Int, Xis, P22, Flp and R1, or variants or functional fragments thereof.

11. The nucleic acid molecule according to claim 10, wherein the specific recombinase is XerC and / or XerD, or a variant or functional fragment thereof.

12. The nucleic acid molecule of claim 8, wherein the specific recombination sites comprise one or more of Ecdif, cer, psi, dif, mwr, Bsdif, loxP, FRT, and RS.

13. The nucleic acid molecule of claim 12, wherein the specific recombination site is cer.

14. The nucleic acid molecule according to any one of claims 8 to 13, wherein the specific site recombination sequence further comprises an accessory sequence, wherein the accessory sequence comprises an accessory site capable of binding to an accessory protein.

15. The nucleic acid molecule according to claim 14, wherein the accessory protein is one or more of PepA, ArgR, ArcA, or a variant or functional fragment of the above accessory protein.

16. The nucleic acid molecule of claim 14, wherein the accessory protein comprises a first accessory protein, the accessory sequence comprises a first accessory sequence, and the first accessory sequence comprises a site capable of binding to the first accessory protein.

17. The nucleic acid molecule of claim 16, wherein the first accessory protein is PepA.

18. The nucleic acid molecule of claim 14, wherein the accessory protein further comprises a second accessory protein, the accessory sequence further comprises a second accessory sequence, and the second accessory sequence comprises a site capable of binding to the second accessory protein.

19. The nucleic acid molecule of claim 18, wherein the second accessory protein is one or more of ArgR and ArcA, or a variant or functional fragment of the above accessory protein.

20. The nucleic acid molecule of any one of claims 1 to 19, wherein the site-specific recombination sequence comprises a) a first common sequence as shown in any one of SEQ ID NO.: 45 or its reverse complementary sequence; and b) a second common sequence as shown in any one of SEQ ID NO.: 46 or its reverse complementary sequence. The nucleic acid molecule of claim 20 , wherein the first common sequence and the second common sequence are separated by a nucleotide sequence of 6-8 bp.

22. The nucleic acid molecule according to any one of claims 20 to 21, wherein the first common sequence is a sequence as shown in SEQ ID NO.: 38-40 or a reverse complementary sequence thereof.

23. The nucleic acid molecule according to any one of claims 20 to 22, wherein the second common sequence is a sequence as shown in SEQ ID NO.: 41 or a reverse complementary sequence thereof.

24. The nucleic acid molecule according to any one of claims 20 to 23, wherein the specific site recombination sequence is a) comprising a first common sequence as shown in SEQ ID: 38 or its reverse complementary sequence and comprising a second common sequence as shown in SEQ ID NO.: 41 or its reverse complementary sequence; or b) comprising a first common sequence as shown in SEQ ID: 39 or its reverse complementary sequence and a second common sequence as shown in SEQ ID NO.: 41 or its reverse complementary sequence; or c) comprising a first common sequence as shown in SEQ ID: 40 or its reverse complementary sequence and a second common sequence as shown in SEQ ID NO.: 41 or its reverse complementary sequence.

25. The nucleic acid molecule of claim 24, wherein the site-specific recombination sequence comprises a sequence as shown in any one of SEQ ID NOs.: 18-37 or a reverse complementary sequence thereof.

26. The nucleic acid molecule of any one of claims 1 to 25, wherein the first transition sequence comprises a screening sequence.

27. The nucleic acid molecule of claim 26, wherein the screening sequence comprises a prokaryotic screening sequence and / or a eukaryotic screening sequence.

28. The nucleic acid molecule of claim 27, wherein the screening sequence comprises one or more of an ampicillin-resistant gene, a kanamycin-resistant gene, a chloramphenicol-resistant gene, a spectinomycin-resistant gene, a tetracycline-resistant gene, a bleomycin-resistant gene, a streptomycin-resistant gene, a hygromycin-resistant gene, a gentamicin-resistant gene, a hexaampromycin-resistant gene, an erythromycin-resistant gene, and a blasticidin-resistant gene.

29. The nucleic acid molecule of any one of claims 1 to 28, wherein the screening sequence comprises a sequence as shown in any one of SEQ ID NOs.: 13 to 17 or a reverse complementary sequence thereof.

30. The nucleic acid molecule of any one of claims 1 to 29, wherein the replication origin sequence comprises one or more of ColE1, pMB1, pBR322, pSC101, R6K, and P15A.

31. The nucleic acid molecule of any one of claims 1 to 30, wherein the replication initiation site sequence comprises a sequence as shown in any one of SEQ ID NOs.: 11 to 12 or a reverse complementary sequence thereof.

32. The nucleic acid molecule of any one of claims 3 to 31, wherein the specific enzyme cleavage site sequence comprises one or more sites selected from the group consisting of ApaI, BamHI, BglII, EcoRI, HindIII, KpnI, NcoI, NdeI, NheI, NotI, SacI, SalI, SphI, XbaI, and XhoI.

33. The nucleic acid molecule of any one of claims 3 to 32, wherein the promoter comprises one or more of a constitutive promoter, a tissue-specific promoter, or an inducible promoter.

34. The nucleic acid molecule of any one of claims 1 to 33, wherein the promoter is selected from one or more of the following groups: CMV, EF1a, SV40, PGK1, Ubc, CAG, TRE, USA, Ac5, CaMkIIa, GAL1 / 10, TEF1, GDS, ADH1, CaMV35S, Ubi, H1, and U6.

35. The nucleic acid molecule of any one of claims 3 to 34, wherein the tag sequence comprises a fluorescent protein sequence.

36. The nucleic acid molecule of claim 35, wherein the fluorescent protein sequence is selected from one or more of the following: GFP, eGFP, eYFP, eCFP, mCherry, luc2, and hrluc.

37. The nucleic acid molecule of any one of claims 35-36, wherein the fluorescent protein sequence comprises: Combination of eGFP and luc2, and / or combination of mCherry and hrluc.

38. The nucleic acid molecule of any one of claims 1 to 36, comprising a sequence as shown in any one of SEQ ID NOs: 2 to 10, 48, 50, 51 or a reverse complementary sequence thereof.

39. A vector comprising the isolated nucleic acid molecule of any one of claims 1-38.

40. The vector according to claim 39, which is a plasmid. The vector according to claim 39 , which is a cosmid or a transposon.

42. A host cell comprising the isolated nucleic acid molecule according to any one of claims 1 to 38 or the vector according to any one of claims 39 to 41.

43. The host cell of claim 42, wherein the host cell provides a specific recombinase.

44. The host cell of any one of claims 42-43, wherein the host cell provides an accessory protein.

45. A diagnostic or pharmaceutical composition comprising the isolated nucleic acid molecule of any one of claims 1-38, the vector of any one of claims 39-41, or the host cell of any one of claims 42-44.

46. ​​A method for expressing a gene of interest, comprising introducing an isolated nucleic acid molecule according to any one of claims 1 to 38 or a vector according to any one of claims 39 to 41 into a host cell, and causing the gene of interest to be expressed in the host cell, the method being an in vitro or ex vivo method.

47. A kit comprising the isolated nucleic acid molecule of any one of claims 1-38, the vector of any one of claims 39-41, or the host cell of any one of claims 42-44.

48. Use of the nucleic acid molecule according to any one of claims 1 to 38 as a gene for maintaining monomer supercoiling in a polyadenylated plasmid for use in recombining a polyadenylated plasmid to increase its monomer supercoiling ratio and plasmid yield.

49. The use according to claim 48, characterized in that The polyadenylation plasmid is an mRNA plasmid with a long PolyA tail, and the length of the long PolyA tail is 10 to 1500 nucleotides.

50. A method for constructing a recombinant plasmid, characterized in that: The method comprises: introducing the nucleic acid molecule according to any one of claims 1 to 38 into a polyadenylation plasmid to obtain a recombinant plasmid.

51. A recombinant plasmid constructed by the method for constructing a recombinant plasmid according to claim 50.

Citation Information

Patent Citations

  • Gene for maintaining superhelix of polyadenosine plasmid monomer and application thereof

    CN117568348A

  • System for in vitro transportation using modified TN5 transposase

    CN1251135A