TFRE enhancer element

By constructing enhancer elements containing TFRE sequences of REL, NFκB, and other transcription factor regulatory elements in CHO cells, the problem of low site-directed integration efficiency of CMV enhancer sequences was solved, and efficient expression of recombinant proteins and antibody molecules was achieved.

WO2026158417A1PCT designated stage Publication Date: 2026-07-30JIANGSU HENGRUI MEDICINE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
JIANGSU HENGRUI MEDICINE CO LTD
Filing Date
2026-01-22
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

The plasmids constructed from existing CMV enhancer sequences in CHO cells have low site-directed integration efficiency, affecting the expression levels of recombinant proteins, especially recombinant antibody proteins.

Method used

Enhancer elements were constructed using TFRE sequences containing REL, NFκB, and transcription factor regulatory elements selected from GABPβ, Sp1, AhR/ARNT, DMP1, and CaRF, and then integrated into the host cell genome at specific sites to improve expression efficiency.

Benefits of technology

It improved the site-specific integration efficiency and expression level of recombinant proteins, especially the expression level of antibody molecules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2026074079-FTAPPB-I100001
    Figure PCTCN2026074079-FTAPPB-I100001
  • Figure PCTCN2026074079-FTAPPB-I100002
    Figure PCTCN2026074079-FTAPPB-I100002
  • Figure PCTCN2026074079-FTAPPB-I100003
    Figure PCTCN2026074079-FTAPPB-I100003
Patent Text Reader

Abstract

The present disclosure relates to a TFRE enhancer element. Specifically, the present disclosure relates to an enhancer element constructed by the sequences of transcription factor regulatory elements, and a promoter, a vector, and a cell strain comprising the enhancer element.
Need to check novelty before this filing date? Find Prior Art

Description

TFRE enhancer element Technical Field

[0001] This disclosure relates to enhancer elements constructed from transcription factor regulatory elements (TFRE) sequences, promoters and cell lines containing such enhancer elements, and their preparation methods and uses. Background Technology

[0002] Efficient expression of exogenous genes in host cells is a prerequisite for protein structure and function analysis, and for the research and development of protein or peptide drugs. Commonly used host cells include Chinese hamster ovary (CHO) cells, hamster kidney (BHK) cells, COS cells, mouse NSO thymoma cells, and mouse myeloma SP2 / 0 cells. Many factors influence the expression of exogenous proteins in mammalian cells, such as enhancers, transcription and translation control factors, gene copy number, mRNA stability, and the integration site of the exogenous gene on the chromosome. Among the factors regulating gene expression, the selection of enhancers plays a crucial role. An enhancer is a DNA sequence that increases the transcription frequency of genes linked to it. Enhancers increase transcription through promoters. Effective enhancers can be located at the 5′ end of a gene, the 3′ end, or even within introns. The effect of enhancers is significant, generally increasing gene transcription frequency by 10–200 times, and in some cases, by up to thousands of times. For example, the expression level of the human globin gene can be increased by 600 to 1000 times under the action of cytomegalovirus (CMV) enhancers. The effect of enhancers is independent of their orientation (5′→3′ or 3′→5′), and they can still have an enhancing effect even when they are thousands of kb away from the target gene.

[0003] Currently, CHO cells are the primary mammalian host cells used for expressing recombinant proteins, especially recombinant antibody proteins, and the commonly used enhancers in CHO cells are the CMV enhancer and the SV40 enhancer. The CMV enhancer is widely used in the construction of recombinant protein expression vectors to increase the expression level of recombinant proteins. However, when constructing antibody expression lines through site-directed integration, plasmids constructed using this complex CMV promoter sequence exhibit low site-directed integration efficiency. Summary of the Invention

[0004] This disclosure provides an enhancer element comprising a TFRE, said TFRE being composed of REL, NFκB and one or more (e.g., 1, 2, 3, 4 or 5) selected from GABPβ, Sp1, AhR / ARNT, DMP1 and CaRF.

[0005] This disclosure provides an enhancer element comprising a TFRE, said TFRE including REL and NFκB, and four or five selected from GABPβ, Sp1, AhR / ARNT, DMP1 and CaRF.

[0006] In some implementations, the reinforcing sub-element as described above, wherein:

[0007] REL is a sequence containing the TFRE sequence as shown in SEQ ID NO:23.

[0008] NFκB contains the TFRE sequence as shown in SEQ ID NO:22.

[0009] GABPβ contains the TFRE sequence as shown in SEQ ID NO:20.

[0010] Sp1 contains the TFRE sequence as shown in SEQ ID NO:21.

[0011] AhR / ARNT contains a TFRE sequence as shown in SEQ ID NO:18.

[0012] DMP1 contains the TFRE sequence as shown in SEQ ID NO:19.

[0013] CaRF is a sequence containing a TFRE sequence as shown in SEQ ID NO:24.

[0014] In some implementations, as described above, the number of copies of REL and NFκB in the reinforcing sub-element may be the same or different, and each independently is 6-15 (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15).

[0015] In some implementations, as described above, the REL and NFκB have the same or different copy numbers in the reinforcing sub-element, and each independently has 8-12 (e.g., 8, 9, 10, 11, 12).

[0016] In some implementations, as described above, the number of copies of REL and NFκB in the reinforcing sub-element is 8.

[0017] In some implementations, as described above, the copy numbers of GABPβ, Sp1, AhR / ARNT, and DMP1 in the reinforcing sub-element are the same or different, and each is independently 2-6 (e.g., 2, 3, 4, 5, 6).

[0018] In some implementations, as described above, the copy numbers of GABPβ, Sp1, AhR / ARNT, and DMP1 in the reinforcing sub-element may be the same or different, and each independently is 2-4.

[0019] In some implementations, as described above, the number of copies of GABPβ, Sp1, AhR / ARNT, and DMP1 in the reinforcing sub-element is 2 or 4.

[0020] In some implementations, as described above, the number of copies of the CaRF in the reinforcing sub-element is 0-2.

[0021] In some implementations, as described above, the copy number of the CaRF in the reinforcing sub-element is 0 or 2.

[0022] In some embodiments, as described above, the copy numbers of REL and NFκB in the enhancer element are the same or different, and each is independently 6-15 (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15); the copy numbers of the sequences GABPβ, Sp1, AhR / ARNT, and DMP1 in the enhancer element are the same or different, and each is independently 2-6 (e.g., 2, 3, 4, 5, 6); and the copy number of CaRF in the enhancer element is 0-2.

[0023] In some embodiments, as described above, the copy numbers of REL and NFκB in the enhancer element are the same or different, and each is independently 8-12; the copy numbers of sequences GABPβ, Sp1, AhR / ARNT and DMP1 in the enhancer element are the same or different, and each is independently 2-4; the copy number of CaRF in the enhancer element is 0-2.

[0024] In some embodiments, as described above, the number of copies of sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element is 8; the number of copies of AhR / ARNT and DMP1 in the enhancer element is 4; the number of copies of GABPβ and Sp1 in the enhancer element is 2; and the number of copies of CaRF in the enhancer element is 0 or 2.

[0025] In some embodiments, as described above, the number of copies of sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element is 8; the number of copies of AhR / ARNT and DMP1 in the enhancer element is 4; the number of copies of GABPβ and Sp1 in the enhancer element is 2; and the number of copies of CaRF in the enhancer element is 2.

[0026] In some implementations, the reinforcing sub-element as described above, wherein:

[0027] REL is the TFRE sequence shown in SEQ ID NO:23.

[0028] NFκB is the TFRE sequence shown in SEQ ID NO:22.

[0029] GABPβ is a TFRE sequence as shown in SEQ ID NO:20.

[0030] Sp1 is a TFRE sequence as shown in SEQ ID NO:21.

[0031] AhR / ARNT is a TFRE sequence as shown in SEQ ID NO:18.

[0032] DMP1 is a TFRE sequence as shown in SEQ ID NO:19.

[0033] CaRF is a TFRE sequence as shown in SEQ ID NO:24.

[0034] In some embodiments, the reinforcing sub-element as described above, wherein the TFRE comprises sequences of SEQ ID NO:22 and SEQ ID NO:23, and one or more sequences selected from SEQ ID NO:18, 19, 20 and 21.

[0035] In some embodiments, the reinforcing sub-element as described above, wherein the TFRE further comprises the sequence of SEQ ID NO:24.

[0036] In some embodiments, the reinforcing sub-element as described above, wherein the TFRE comprises sequences of SEQ ID NO:22 and SEQ ID NO:23, and one or more sequences selected from SEQ ID NO:18,19,20,21,24.

[0037] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element may be the same or different, and each independently is 6-15 (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15).

[0038] In some embodiments, as described above, the sequences SEQ ID NO:22 and SEQ ID NO:23 have the same or different copy numbers in the enhancer element, and each independently has 8-12 copies.

[0039] In some embodiments, as described above, the sequences SEQ ID NO:22 and SEQ ID NO:23 have a copy number of 8 in the enhancer element.

[0040] In some embodiments, as described above, the number of copies of the sequence SEQ ID NO:18-21 in the enhancer element may be the same or different, and each may independently be 2-6 (e.g., 2, 3, 4, 5, 6).

[0041] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:18-21 in the enhancer elements may be the same or different, and each independently ranges from 2 to 4.

[0042] In some embodiments, as described above, the number of copies of the sequence SEQ ID NO:18-21 in the enhancer element is 2 or 4;

[0043] The copy number of the TFRE shown in sequence SEQ ID NO:24 in the reinforcing sub-element is 0-2, preferably 0 or 2.

[0044] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element may be the same or different, and each independently is 6-15 (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15).

[0045] The sequences SEQ ID NO:18-21 have the same or different copy numbers in the enhancing sub-elements, and each independently has 2-6 copies (e.g., 2, 3, 4, 5, 6);

[0046] The copy number of the TFRE shown in sequence SEQ ID NO:24 in the reinforcing sub-element is 0-2.

[0047] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:22 and SEQ ID NO:23 in the reinforcing sub-element may be the same or different, and each independently is 8-12;

[0048] The sequences SEQ ID NO:18-21 have the same or different copy numbers in the reinforcing sub-elements, and each independently has 2-4 copies.

[0049] The copy number of the TFRE shown in sequence SEQ ID NO:24 in the reinforcing sub-element is 0-2.

[0050] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element is 8;

[0051] The sequences SEQ ID NO:18, 19 have a copy number of 4 in the reinforcing sub-elements; the sequences SEQ ID NO:20, 21 have a copy number of 2 in the reinforcing sub-elements.

[0052] The copy number of the TFRE, as shown in sequence SEQ ID NO:24, in the reinforcing sub-element is 0 or 2.

[0053] In some embodiments, as described above, the number of copies of the sequences SEQ ID NO:22 and SEQ ID NO:23 in the enhancer element is 8;

[0054] The sequences SEQ ID NO:18, 19 have a copy number of 4 in the reinforcing sub-elements; the sequences SEQ ID NO:20, 21 have a copy number of 2 in the reinforcing sub-elements.

[0055] The copy number of the TFRE shown in sequence SEQ ID NO:24 in the reinforcing sub-element is 2.

[0056] In some implementations, the enhancer element as described in any of the preceding claims contains a total copy number of 25-40 (e.g., 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40) of the TFRE sequence.

[0057] In some embodiments, the enhancer element as described in any of the preceding claims contains a total copy number of 25-35 (25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35) of the TFRE sequence.

[0058] In some implementations, the enhancer element as described in any of the preceding claims contains a total copy number of 28-32 TFRE sequences.

[0059] In some implementations, the enhancer element as described in any of the preceding claims contains a total of 28 or 30 copies of the TFRE sequence.

[0060] In some implementations, the enhancer element as described in any of the preceding claims contains a total of 28 copies of the TFRE sequence.

[0061] In some implementations, the enhancer element as described in any of the preceding claims contains a total of 30 copies of the TFRE sequence.

[0062] In some embodiments, as described in any of the preceding embodiments, the spacer sequence P between the TFRE sequences is selected from one or more of SEQ ID NO:25-104.

[0063] In some embodiments, the spacer sequence P between the TFRE sequences is SEQ ID NO:28, as described in any of the preceding embodiments.

[0064] In some embodiments, the reinforcing sub-element as described in any of the preceding claims is selected from a combination of the following:

[0065] Enhancer EH1-28F6:

[0066] -P-Sp1-P-REL-P-DMP1-P-GABPβ-P-REL-P-REL-P-REL-P-AhR / ARNT-P-DMP1-P-REL-P-DMP1-P-AhR / ARNT-P-NFκB-P-REL-P-N FκB-P-NFκB-P-GABPβ-P-AhR / ARNT-P-Sp1-P-NFκB-P-REL-P-AhR / ARNT-P-NFκB-P-DMP1-P-NFκB-P-REL-P-NFκB-P-NFκB-P-,

[0067] Enhancer EH2-28E7:

[0068] -P-AhR / ARNT-P-NFκB-P-DMP1-P-NFκB-P-NFκB-P-NFκB-P-Sp1-P-NFκB-P-REL-P-AhR / ARNT-P-AhR / ARNT-P-NFκB-P-REL-P-G ABPβ-P-REL-P-REL-P-DMP1-P-GABPβ-P-NFκB-P-REL-P-REL-P-NFκB-P-REL-P-Sp1-P-REL-P-DMP1-P-AhR / ARNT-P-DMP1-P-.

[0069] Pipeline EH3-28F7:

[0070] -P-AhR / ARNT-P-NFκB-P-NFκB-P-NFκB-P-REL-P-REL-P-AhR / ARNT-P-Sp1-P-NFκB-P-DMP1-P-GABPβ-P-AhR / ARNT-P-GABPβ-P -NFκB-P-Sp1-P-REL-P-NFκB-P-REL-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P-REL-P-NFκB-P-REL-P-DMP1-P-.

[0071] Design EH4-28G7:

[0072] -P-REL-P-DMP1-P-Sp1-P-REL-P-NFκB-P-Sp1-P-GABPβ-P-GABPβ-P-REL-P-AhR / ARNT-P-DMP1-P-REL-P-REL-P-NFκB-P-REL- P-NFκB-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P-NFκB-P-DMP1-P-NFκB-P-NFκB-P-AhR / ARNT-P-AhR / ARNT-P-REL-P-NFκB-P-.

[0073] Pipeline EH5-28B8:

[0074] -P-GABPβ-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-Sp1-P-GABPβ-P-REL-P-REL-P-NFκB-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P -NFκB-P-NFκB-P-NFκB-P-Sp1-P-NFκB-P-REL-P-DMP1-P-AhR / ARNT-P-AhR / ARNT-P-DMP1-P-REL-P-REL-P-NFκB-P-NFκB-P-,

[0075] Pipeline EH6-30E8:

[0076] -P-GABPβ-P-NFκB-P-AhR / ARNT-P-DMP1-P-REL-P-NFκB-P-AhR / ARNT-P-Sp1-P-REL-P-REL-P-DMP1-P-DMP1-P-REL-P-AhR / ARNT-P-NF κB-P-REL-P-REL-P-CaRF-P-NFκB-P-REL-P-AhR / ARNT-P-Sp1-P-NFκB-P-GABPβ-P-REL-P-NFκB-P-DMP1-P-NFκB-P-CaRF-P-NFκB-P-,

[0077] Pipe EH7-30F8:

[0078] -P-NFκB-P-NFκB-P-GABPβ-P-NFκB-P-NFκB-P-REL-P-DMP1-P-NFκB-P-CaRF-P-DMP1-P-DMP1-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-C aRF-P-REL-P-AhR / ARNT-P-GABPβ-P-AhR / ARNT-P-AhR / ARNT-P-Sp1-P-NFκB-P-REL-P-REL-P-Sp1-P-REL-P-NFκB-P-REL-P-NFκB-P-,

[0079] Pipeline EH8-30G8:

[0080] -P-DMP1-P-REL-P-Sp1-P-REL-P-DMP1-P-REL-P-REL-P-REL-P-AhR / ARNT-P-NFκB-P-NFκB-P-NFκB-P-NFκB-P-REL-P-DMP1-P-AhR / AR NT-P-CaRF-P-GABPβ-P-Sp1-P-NFκB-P-GABPβ-P-NFκB-P-AhR / ARNT-P-DMP1-P-REL-P-CaRF-P-NFκB-P-NFκB-P-REL-P-AhR / ARNT-P-,

[0081] Pipeline EH9-28G6:

[0082] -P-NFκB-P-NFκB-P-AhR / ARNT-P-NFκB-P-REL-P-NFκB-P-REL-P-REL-P-DMP1-P-NFκB-P-GABPβ-P-AhR / ARNT-P-REL-P-NFκB- P-DMP1-P-DMP1-P-DMP1-P-REL-P-Sp1-P-REL-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-GABPβ-P-REL-P-REL-P-NFκB-P-Sp1-P-,

[0083] Pipeline EH10-30D8:

[0084] -P-REL-P-AhR / ARNT-P-NFκB-P-CaRF-P-NFκB-P-DMP1-P-DMP1-P-GABPβ-P-NFκB-P-REL-P-REL-P-CaRF-P-DMP1-P-NFκB-P-REL-P-NF κB-P-Sp1-P-REL-P-NFκB-P-REL-P-AhR / ARNT-P-Sp1-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL-P-REL-P-NFκB-P-GABPβ-P-AhR / ARNT-P-.

[0085] Pipeline EH11-28B7 Liquid:

[0086] -P-NFκB-P-NFκB-P-GABPβ-P-REL-P-REL-P-NFκB-P-REL-P-DMP1-P-NFκB-P-AhR / ARNT-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL- P-GABPβ-P-REL-P-REL-P-DMP1-P-Sp1-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-REL-P-Sp1-P-DMP1-P-REL-P-NFκB-P-NFκB-P-.

[0087] Type EH12-28C7 plate:

[0088] -P-DMP1-P-REL-P-NFκB-P-DMP1-P-REL-P-NFκB-P-NFκB-P-AhR / ARNT-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-GABPβ-P-NFκB-P -NFκB-P-NFκB-P-REL-P-REL-P-Sp1-P-NFκB-P-GABPβ-P-Sp1-P-DMP1-P-REL-P-REL-P-AhR / ARNT-P-REL-P-REL-P-DMP1-P-.

[0089] Pipeline EH13-28D7 plate:

[0090] -P-REL-P-Sp1-P-REL-P-REL-P-DMP1-P-REL-P-REL-P-AhR / ARNT-P-REL-P-DMP1-P-DMP1-P-NFκB-P-REL-P-NFκB-P-NFκB-P- REL-P-NFκB-P-AhR / ARNT-P-NFκB-P-NFκB-P-GABPβ-P-Sp1-P-AhR / ARNT-P-DMP1-P-GABPβ-P-NFκB-P-NFκB-P-AhR / ARNT-P-.

[0091] Type EH14-28C8 plate:

[0092] -P-Sp1-P-NFκB-P-AhR / ARNT-P-NFκB-P-DMP1-P-GABPβ-P-REL-P-REL-P-NFκB-P-DMP1-P-REL-P-NFκB-P-REL-P-AhR / ARNT-P -NFκB-P-Sp1-P-NFκB-P-NFκB-P-REL-P-AhR / ARNT-P-NFκB-P-GABPβ-P-REL-P-DMP1-P-AhR / ARNT-P-REL-P-REL-P-DMP1-P-, and

[0093] Enhancer EH15-30B9 sequence:

[0094] -P-REL-P-Sp1-P-AhR / ARNT-P-NFκB-P-REL-P-NFκB-P-GABPβ-P-AhR / ARNT-P-AhR / ARNT-P-DMP1-P-REL-P-DMP1-P-DMP1-P-NFκB-P-N FκB-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL-P-REL-P-NFκB-P-Sp1-P-REL-P-CaRF-P-REL-P-GABPβca7t-NFκB-P-CaRF-P-REL-P-NFκB-P;

[0095] Where P is the interval sequence as described in the previous item;

[0096] AhR / ARNT is a TFRE sequence as shown in SEQ ID NO:18.

[0097] DMP1 is a TFRE sequence as shown in SEQ ID NO:19.

[0098] GABPβ is a TFRE sequence as shown in SEQ ID NO:20.

[0099] Sp1 is a TFRE sequence as shown in SEQ ID NO:21.

[0100] NFκB is the TFRE sequence shown in SEQ ID NO:22.

[0101] REL is the TFRE sequence shown in SEQ ID NO:23.

[0102] CaRF is a TFRE sequence as shown in SEQ ID NO:24.

[0103] In some embodiments, the reinforcing sub-element as described in any of the preceding claims is selected from one or more of SEQ ID NO:121-135.

[0104] This disclosure further provides a synthetic promoter comprising an enhancer element as described in any of the preceding claims; a restriction enzyme site sequence at the 5′ or 3′ end, the restriction enzyme site being linked to a transcription factor regulatory element sequence via a spacer sequence; and a core promoter sequence added downstream of the 3′ restriction enzyme site.

[0105] In some embodiments, the promoter as described above is synthesized, wherein the restriction enzyme site sequence is selected from one or more of SEQ ID NO:105-120.

[0106] In some embodiments, the synthetic promoter as described in any of the preceding claims, wherein the spacer sequence (P) is selected from one or more of SEQ ID NO:25-104.

[0107] In some embodiments, the synthetic promoter as described in any of the preceding claims is wherein the spacer sequence is SEQ ID NO:28.

[0108] In some embodiments, a synthetic promoter as described in any of the preceding claims is used, wherein the core promoter is selected from hCMV, hEF-1α, SV40, UbC, EF1A, PGK, and CAGG.

[0109] In some implementations, a synthetic promoter as described in any of the preceding claims is used, wherein the upper core promoter is a CMV core promoter.

[0110] In some embodiments, the synthetic promoter as described in any of the preceding claims includes an enhancer element as described in any of the preceding claims; wherein the 5′ restriction site (SpeI) sequence is shown in SEQ ID NO:117, the 3′ restriction site (NotI) sequence is shown in SEQ ID NO:114; the spacer sequence is shown in SEQ ID NO:28; and the core promoter is shown in SEQ ID NO:15.

[0111] This disclosure further provides an expression vector, wherein the expression vector contains an enhancer element as described in any of the preceding claims.

[0112] This disclosure further provides an expression vector containing a synthetic promoter as described in any of the preceding claims.

[0113] In some embodiments, the expression vector as described above contains the nucleotide sequence of the exogenous protein to be expressed; preferably, the exogenous protein nucleotide sequence encodes a non-secretory protein and a secretory protein; more preferably, the secretory protein is an antibody molecule.

[0114] In some embodiments, the antibody molecule is an anti-GITR antibody comprising a heavy chain encoded by a nucleotide sequence such as SEQ ID NO:137 and a light chain encoded by SEQ ID NO:136.

[0115] This disclosure further provides a host cell, wherein the host cell is transfected with the expression vector as described in any of the preceding claims.

[0116] In some implementations, the host cell is as described above, wherein the transfection is site-directed integration.

[0117] In some embodiments, the host cell as described above, wherein the site-specific integration site is in the Nfat5 gene or the Fer1L4 gene.

[0118] In some implementations, the host cell is as described above, wherein the site-specific integration site is in the Nfat5 gene.

[0119] In some embodiments, the host cell as described above, wherein the site-specific integration site is located within the first intron of the Nfat5 gene.

[0120] In some embodiments, the host cell as described above, wherein the site-specific integration site is located in the first intron of the Nfat5 gene, NW_003614572.1 (269232--293783).

[0121] In some embodiments, the host cell as described above, wherein the site-specific integration site is located within the first intron of the Nfat5 gene, SEQ ID NO:138 (approximately 4 kb), or within a nucleotide sequence having at least 90% sequence identity with SEQ ID NO:138.

[0122] In some embodiments, the host cell as described above, wherein the site-specific integration site is located within the first intron of the Nfat5 gene, SEQ ID NO:139 (approximately 2 kb), or within a nucleotide sequence having at least 90% sequence identity with SEQ ID NO:139.

[0123] In some embodiments, the host cell is as described above, wherein the site-specific integration site is a target sequence as shown in SEQ ID NO:2.

[0124] In some implementations, the host cell is as described above, wherein the host cell is a mammalian host cell.

[0125] In some embodiments, the host cell is as described above, wherein the host cell is a CHO, BHK, SP2 / 0, HEK293, or C127 cell.

[0126] In some implementations, the host cell is as described above, wherein the host cell is a CHO cell.

[0127] In order to improve site-specific integration efficiency and antibody expression levels, this disclosure describes the use of transcription factor regulatory elements (TFREs) to obtain enhancers with higher site-specific integration efficiency and antibody expression levels. Attached Figure Description

[0128] Figure 1: Schematic diagram of fluorescently labeled cell line construction.

[0129] Figure 2: Schematic diagram of the principle of constructing a cell line with enhanced fluorescent protein (non-secretory protein) site-specific integration using cassette exchange.

[0130] Figure 3: Schematic diagram of the principle of constructing a cell line for site-specific integration of exogenous proteins (monoclonal antibodies) using cassette exchange. Detailed Implementation

[0131] the term

[0132] The terminology used herein is for descriptive purposes only and is not intended to be limiting. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0133] Unless the context clearly requires otherwise, throughout the specification and claims, the words “comprising,” “having,” “including,” etc., should be understood as encompassing rather than exclusive or exhaustive; that is, meaning “including but not limited to.” Unless otherwise stated, “comprising” includes “consisting of.”

[0134] The term "multiple" refers to any number, including but not limited to 1, 2, 3, 4, 5, 6, 7, or 8. When the TFRE in this disclosure includes one or more selected from GABPβ, Sp1, AhR / ARNT, DMP1, and CaRF, it means that it includes 1, 2, 3, 4, or 5 of those TFREs; preferably, it means that it includes 4 or 5 of those TFREs.

[0135] The terms “polynucleotide” and “nucleotide sequence” include naturally occurring nucleic acid molecules that can be isolated from cells or recombinantly expressed nucleic acid molecules, as well as synthetic molecules that can be prepared, for example, by chemical synthesis or by enzymatic methods such as polymerase chain reaction (PCR).

[0136] The term "polynucleotide sequence encoding a polypeptide" includes DNA encoding a gene, preferably a heterologous gene expressing the polypeptide.

[0137] The term "promoter" defines a regulatory DNA sequence that mediates transcription initiation by directing RNA polymerase to bind to DNA and trigger RNA synthesis. Promoters are typically located upstream of a gene and may contain, for example, core promoters and transcription factor regulatory elements.

[0138] The term "synthetic promoter" includes the nucleotide sequence of the transcription factor regulatory element (TFRE) and the core promoter. The nucleotide sequence containing the TFRE can be located upstream of the core promoter.

[0139] The term "transcription factor regulator elements" (TFRE) refers to nucleotide sequences that serve as transcription factor binding sites (TFBS).

[0140] The term "enhancer" is defined as a nucleotide sequence that enhances gene transcription, regardless of gene identity, sequence position relative to the gene, or sequence orientation.

[0141] The term "enhancer" in this disclosure refers to an assembled double-stranded DNA molecule containing more than one type of TFRE sequence. The enhancer can be created by various means known to those skilled in the art, including linking various double-stranded TFREs together in a random or directional manner. The enhancer may also contain other nucleic acid sequences, such as spacers that do not mediate transcription factor binding but allow for the correct spatial arrangement of binding sites. Spacer regions, for example, may be common single-stranded overhangs that allow different TFREs to be easily linked together.

[0142] The enhancer may also include multiple elements, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more transcription factor regulatory elements. Many of these may be the same, or they may all be different.

[0143] The promoters or enhancers provided in this disclosure contain multiple copies of each TFRE, such as 1 to 12, 2 to 10, 2 to 8, 2 to 6, or 6 to 4 copies. In some embodiments, the TFRE has 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 copies, and any range between these point values.

[0144] The promoters or enhancers provided in this disclosure contain multiple copies of various TFREs. In some embodiments, the total number of copies of the various TFREs is 25-40, 25-38, 25-38, 25-36, 25-32, 25-30, 28-32, or 28-30. In some embodiments, the total number of copies is 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, and any range between these values.

[0145] TFREs can contain 4 to 100 nucleotides, 4 to 75 nucleotides, 4 to 50 nucleotides, 4 to 30 nucleotides, 4 to 25 nucleotides, 4 to 20 nucleotides, 4 to 15 nucleotides, or 4 to 12 nucleotides, essentially composed of or composed of them. TFREs can contain 6 to 100 nucleotides, 6 to 75 nucleotides, 6 to 50 nucleotides, 6 to 30 nucleotides, 6 to 25 nucleotides, 6 to 20 nucleotides, 6 to 15 nucleotides, or 6 to 12 nucleotides, essentially composed of or composed of them. TFREs can contain 8 to 100 nucleotides, 8 to 75 nucleotides, 8 to 50 nucleotides, 8 to 30 nucleotides, 8 to 25 nucleotides, 8 to 20 nucleotides, 8 to 15 nucleotides, or 8 to 12 nucleotides, essentially composed of or composed of them.

[0146] TFRE can be mammalian TFRE. TFRE can be CHO cell TFRE.

[0147] The TFREs in this disclosure are derived from vertebrates and range in length from 6 bp to 20 bp. Exemplary transcriptional regulatory elements (TFREs) are provided in Table 1. It should be understood that the transcriptional regulatory elements in this disclosure are not limited to the specific sequences mentioned in the specification, but also include their structural and functional analogs / homologs. Such analogs may contain truncation, deletion, insertion, and substitution of one or more nucleotides introduced directly or through random mutagenesis. For known transcriptional repressors, truncation may be introduced to delete one or more binding sites. Furthermore, such sequences may be derived from sequences naturally found in nature that exhibit high identity with the sequences of this invention. Nucleic acids of about 20 nucleotides or more will be considered to have high identity with the transcription factor regulatory elements of this invention if they hybridize to the relevant transcription factor under stringent conditions. Optionally, a nucleic acid will be considered to have a high degree of identity with the transcription factor regulatory element in this disclosure if it comprises a continuous sequence of about 20 or more nucleotides having at least 70%, 75%, 80%, 85%, 90%, 95% or more percent identity as determined by a standard alignment algorithm, such as, for example, the algorithm of the Basic Local Alignment Tool (BLAST).

[0148] The transcription factor regulatory elements disclosed herein may be selected from transcription factor regulatory elements that are known to be active in target host cells, or may be presumed regulatory elements determined by computer analysis of the upstream sequence of the core promoter using methods known to those skilled in the art.

[0149] The term "core promoter" refers to the nucleotide sequence of the smallest part of the promoter required to initiate transcription. Core promoter sequences can be derived from prokaryotic or eukaryotic genes, including, for example, the CMV immediate early gene promoter or SV40. Core promoters may contain, for example, a TATA box. Core promoters may contain, for example, a priming element. Core promoters may contain, for example, both a TATA box and a priming element.

[0150] The term "expression vector" refers to isolated and purified DNA molecules that, upon transfection into appropriate host cells, provide for the expression of recombinant gene products within the host cells. In addition to the DNA sequence encoding the recombinant or gene product, expression vectors also contain regulatory DNA sequences required for the efficient transcription of the DNA-coding sequence into mRNA and, optionally, the efficient translation of the mRNA into protein in a host cell line.

[0151] This disclosure provides for the stable integration and / or expression of recombinant proteins in eukaryotic cells. Specifically, this disclosure includes methods and compositions for improving protein expression in eukaryotic cells by employing expression-enhancing nucleotide sequences. This disclosure includes polynucleotides that facilitate recombination-mediated cassette exchange (RMCE) and modified cells. The method disclosed integrates exogenous nucleic acids into the Nfat5 locus of the Chinese hamster cell genome to facilitate enhanced and stable expression of recombinant proteins in modified cells.

[0152] DNA regions are operatively linked when they are functionally related. For example, if a promoter is capable of participating in the transcription of a coding sequence, then the promoter is operatively linked to that sequence; if a ribosome binding site is positioned to allow translation, then the ribosome binding site is operatively linked to the coding sequence. Generally, operative linking may include, but does not require, contiguity. For sequences such as secretory leader sequences, contiguity and proper placement within the reading frame are typical characteristics.

[0153] The term "exogenous nucleotide sequence" refers to any DNA sequence or gene that is not found at a locus of interest in nature. For example, an "exogenous nucleotide sequence" at the CHO locus could be a hamster gene not found at a specific CHO locus in nature (i.e., a hamster gene from another locus in the hamster genome), a gene from any other species (e.g., a human gene), a chimeric gene (e.g., human / mouse), or any other gene not found in nature present at the CHO locus of interest.

[0154] This disclosure is based, at least in part, on the discovery of unique sequences (i.e., loci) in the genome that exhibit more efficient recombination, insertion stability, and higher levels of expression compared to other regions or sequences in the genome. This disclosure is also based, at least in part, on the finding that when such expression-enhancing sequences are identified, suitable genes or constructs can be exogenously added to or near these sequences, and the exogenously added genes can be advantageously expressed or used for further genomic modifications. These sequences, referred to as expression-enhancing sequences, are considered stable and not located within coding regions of the genome. These expression-enhancing and stable regions can be engineered for future cloning or genome editing events. Therefore, reliable expression systems are constructed into the cellular genome backbone.

[0155] This disclosure also relies on exogenous gene-specific targeting of integration sites. The method disclosed herein allows for the efficient “conversion” of the cellular genome into a suitable cloning cassette, for example, by employing recombinase-mediated cassette exchange (RMCE). For this purpose, the method disclosed herein uses cellular genomic recombinase recognition sites to place the gene of interest, thereby generating high-yield cell lines for recombinant protein production.

[0156] The compositions disclosed herein can also be included in expression constructs, such as expression vectors for cloning and engineering new cell lines. Expression vectors containing the polynucleotides disclosed herein can be used for transient protein expression or can be integrated into the genome via random or targeted recombination, such as homologous recombination or recombination mediated by recombinases that recognize specific recombination sites (e.g., Cre-lox-mediated recombination). Expression vectors containing the polynucleotides disclosed herein can also be used to evaluate the efficacy of other DNA sequences, such as cis-regulatory sequences.

[0157] The synthetic promoter disclosed herein comprises two parts: an enhancer and a core promoter. The term "core promoter" refers to a short sequence encompassing the transcription start site and its vicinity, which can be linked to RNA polymerase II and drive the transcription process. The core promoter disclosed herein is derived from the smallest core promoter of the hEF-1α promoter, hCMV promoter, UbC, EF1A, PGK, CAGG, or SV40 early promoter, with the CMV core of the hCMV-IE1 promoter being the preferred choice.

[0158] The promoters mentioned are one of the following: CMV (strong mammalian expression promoter derived from human cytomegalovirus), EF-1a (strong mammalian expression promoter derived from human elongation factor 1α), SV40 (mammalian expression promoter derived from simian vacuolating virus 40), PGK1 (mammalian promoter derived from phosphoglycerate kinase gene), UBC (mammalian promoter derived from human ubiquitin C gene), human beta actin (mammalian promoter derived from β-actin gene), CAG (strong hybrid mammalian promoter), etc.

[0159] Transcriptional and translational control sequences in expression vectors suitable for transfection into vertebrate cells can be provided from viral sources. For example, commonly used promoters and enhancers are derived from viruses such as polyomavirus, adenovirus 2, simian virus 40 (SV40), and human cytomegalovirus (CMV). Viral genome promoters, control, and / or signaling sequences can be used to drive expression; these control sequences are compatible with the chosen host cell. Non-viral cell promoters (e.g., β-globulin and EF-1α promoters) can also be used, depending on the cell type expressing the recombinant protein.

[0160] DNA sequences derived from the SV40 viral genome, such as the SV40 origin, early and late promoters, enhancers, splicing sites, and polyadenylation sites, can be used to provide other genetic elements useful for the expression of heterologous DNA sequences. Early and late promoters are particularly useful because they can be readily obtained from the SV40 virus as a fragment that also contains the SV40 origin of replication. Smaller or larger SV40 fragments can also be used. Typically, this includes a sequence of approximately 250 bp extending from the Hind III site located at the SV40 origin of replication to the BglI site.

[0161] When describing a locus of interest or a segment thereof, the percentage of consistency means including homologous sequences that show consistency along adjacent homologous regions, but the presence of gaps, deletions or insertions that are not homologous in the compared sequences is not included in the calculation of the percentage of consistency.

[0162] "Integration at a specific site" refers to a gene-targeting method used to guide the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome; that is, to guide DNA to a specific site between two nucleotides in a linked polynucleotide chain. Targeted insertion of specific gene cassettes, which may include multiple genes, regulatory elements, and / or nucleic acid sequences, is also possible. "Insertion" and "integration" are used interchangeably. It should be understood that the insertion of a gene or nucleic acid sequence (e.g., a nucleic acid sequence containing an expression cassette) may result in (or may be engineered to) the substitution or deletion of one or more nucleic acids, depending on the gene-editing technology employed.

[0163] Integration sites are typically identified through random integration or by analyzing retroviral integration events. The CHO integration site described in detail in this article was identified by random integration into DNA encoding multi-stranded antibodies and by observing enhanced expression of the expressed protein.

[0164] A "recognition site," "integration site," or "recognition sequence" is a specific DNA sequence that is recognized by nucleases or other enzymes to bind to and guide site-specific cleavage of the DNA backbone. Nucleases cleave DNA within the DNA molecule. Recognition sites are also referred to as target recognition sites in their respective fields.

[0165] In the context of this disclosure, the Nfat5 gene is the wild-type Nfat5 gene, all its isoforms and all its homologues, specifically provided that the homologues have at least 40%, 50%, 60%, 70%, 80%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% sequence homology with the wild-type Nfat5 gene. Preferably, it is located in the non-coding region of NCBI accession number NW_003614572.1 (269178..352822), preferably in the wild-type CHO Nfat5 gene or a truncated form thereof, particularly within the first intron of the fat5 gene, specifically within the first intron of NW_003614572.1 (269232..293783). The efficient integration site is preferably located in sequence 2, such as SEQ ID NO:138 (approximately 4 kb). More preferably, it is located in sequence 3, such as SEQ ID NO:139 (approximately 2 kb).

[0166] As used in this disclosure, the determination of the "percentage of similarity" between, for example, the first intron of the Nfat5 gene or a fragment thereof and a species homolog will not include sequence comparisons in which no homologous sequence is compared in the alignment (i.e., where the first intron of the Nfat5 gene or a fragment thereof has an insertion at that point, or where the species homolog has a gap or deletion, as the case may be). Therefore, the "percentage of similarity" does not include penalties for gaps, deletions, and insertions.

[0167] In the context of nucleic acid sequences, a "homologous sequence" refers to a sequence that is substantially homologous to a reference nucleic acid sequence. In some embodiments, two sequences are considered substantially homologous if at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more of the corresponding nucleotides are identical in the relevant residue sequence segment. In some embodiments, the relevant sequence segment is a complete sequence.

[0168] "Integration at a specific site" refers to a gene-targeting method used to guide the insertion or integration of a gene or nucleic acid sequence into a specific location in the genome; that is, to guide DNA to a specific site between two nucleotides in a linked polynucleotide chain. Targeted insertion of specific gene cassettes, which may include multiple genes, regulatory elements, and / or nucleic acid sequences, is also possible. "Insertion" and "integration" are used interchangeably. It should be understood that the insertion of a gene or nucleic acid sequence (e.g., a nucleic acid sequence containing an expression cassette) may result in (or may be engineered to) the substitution or deletion of one or more nucleic acids, depending on the gene-editing technology employed.

[0169] Integration sites are typically identified through random integration or by analyzing retroviral integration events. The CHO integration site described in detail in this article was identified by random integration into DNA encoding multi-stranded antibodies and by observing enhanced expression of the expressed protein.

[0170] A "recombinase recognition site" is a specific DNA sequence recognized by a recombinase, such as Cre recombinase (Cre) or flipping enzyme (flp). Site-specific recombinases can perform DNA rearrangements, including deletions, inversions, and translocations, when one or more of their target recognition sequences are strategically placed in an organism's genome. In one example, Cre specifically mediates recombination events at its DNA target recognition site loxP, which consists of two 13-bp inverted repeat sequences separated by an 8-bp spacer. More than one recombinase recognition site can be used, for example, to facilitate recombination-mediated DNA exchange. Variants or mutants of recombinase recognition sites (e.g., lox sites) can also be used.

[0171] Cre-lox system

[0172] The full-length coding sequence of the Cre recombinase gene is 1029 bp (EMBL accession number X03453), encoding a 38 kDa protein. Cre recombinase is a monomeric protein composed of 343 amino acids. Belonging to the λInt enzyme superfamily, it not only possesses catalytic activity but also, similar to restriction enzymes, recognizes specific DNA sequences, namely lox sites, causing deletion or recombination of the gene sequence between lox sites.

[0173] The loxP (locus of X-over P1) sequence originates from P1 bacteriophage and consists of two 13bp inverted repeat sequences and an 8bp spacer sequence in between. The 8bp spacer sequence also determines the orientation of the loxP. Cre enzymes covalently bind to DNA during DNA strand exchange, and the 13bp inverted repeat sequences are the binding domain of Cre enzymes.

[0174] As used in this article, the term "loxP element" refers to two identical repeat sites that can be recognized by Cre recombinant proteins.

[0175] As used in this article, the term “Cre enzyme” refers to a protease that mediates specific recombination between two loxP sites, which can lead to the deletion or recombination of nucleotide sequences between loxP sites.

[0176] "Recombinase-Mediated Cassette Exchange" (RMCE) relates to a method for precisely replacing a genomic target cassette with a donor cassette. Typically, the molecular composition used in this method includes 1) a genomic target cassette with 5' and 3' side-attached sites that specifically recognize the target site for a particular recombinase, 2) a donor cassette with a matching target site side-attached site, and 3) a site-specific recombinase. Recombinase proteins are well-known in the field and are capable of precisely cleaving DNA (the DNA sequence) within a specific target site without adding or losing nucleotides. Common recombinase / site combinations include (but are not limited to) Cre / lox and Flp / frt.

[0177] As used herein, "exogenous gene" refers to an exogenous DNA molecule that performs a phased function. There are no particular limitations on the exogenous genes that can be used in this application, including various exogenous genes commonly used in the field of transgenic animals. Representative examples include (but are not limited to): antibodies, lysozyme genes, salmon calcitonin genes, or serum albumin genes.

[0178] As used herein, “screening marker gene” refers to a gene used in the transgenic process to screen transgenic cells or transgenic animals. There are no particular limitations on the screening marker genes that can be used in this application, including various screening marker genes commonly used in the transgenic field. Representative examples include (but are not limited to): neomycin resistance gene (NeoR) or puromycin resistance gene (PuroR).

[0179] As used herein, the term "expression cassette" refers to a polynucleotide sequence containing a gene to be expressed and a sequence component of the elements required for expression. For example, in this disclosure, the term "selective marker expression cassette" refers to a polynucleotide sequence containing a sequence encoding a selective marker and a sequence component of the elements required for expression. The elements required for expression include a promoter and a polyadenylation signal sequence. Furthermore, the selective marker expression cassette may contain or not contain other sequences, including (but not limited to): enhancers, secretion signal peptide sequences, etc.

[0180] In this disclosure, the promoter suitable for exogenous gene expression cassettes and selection marker gene expression cassettes can be any common promoter, which can be a constitutive promoter or an inducible promoter. Preferably, the promoter is a constitutive strong promoter, such as the bovine β-lactoglobulin promoter or other promoters suitable for eukaryotic expression.

[0181] A plasmid is a composition consisting of any polynucleotide or set of polynucleotides carrying exogenous nucleic acids, intended for introduction into a cell. It can be delivered to cells via well-known transfection methods. In one instance, the plasmid introduced into the cell may be transient and not integrated into the genome. Simultaneously, the plasmid may carry exogenous nucleic acids necessary for the integration process.

[0182] The donor plasmid of this application contains a foreign nucleotide sequence, as well as any other required elements, such as promoters, enhancers, markers, operons, ribosome binding sites, etc.

[0183] This disclosure is based, at least in part, on the discovery of unique sequences (i.e., loci) in the genome that exhibit more efficient recombination, insertion stability, and higher levels of expression compared to other regions or sequences in the genome. This disclosure is also based, at least in part, on the finding that when such expression-enhancing sequences are identified, suitable genes or constructs can be exogenously added to or near these sequences, and the exogenously added genes can be efficiently expressed or used for further genomic modifications. These sequences, referred to as expression-enhancing sequences, are considered stable and not located within coding regions of the genome. These expression-enhancing and stable regions can be engineered for future cloning or genome editing events. Therefore, reliable expression systems are constructed into the cellular genome backbone.

[0184] This disclosure also relies on exogenous gene-specific targeting of integration sites. The method disclosed herein allows for the efficient “conversion” of the cellular genome into a suitable cloning cassette, for example, by employing recombinase-mediated cassette exchange (RMCE). For this purpose, the method disclosed herein uses cellular genomic recombinase recognition sites to place the gene of interest, thereby generating high-yield cell lines for recombinant protein production.

[0185] The compositions disclosed herein can also be included in expression constructs, such as expression vectors for cloning and engineering new cell lines. Expression vectors containing the polynucleotides disclosed herein can be used for transient protein expression or can be integrated into the genome via random or targeted recombination, such as homologous recombination or recombination mediated by recombinases that recognize specific recombination sites (e.g., Cre-lox-mediated recombination). Expression vectors containing the polynucleotides disclosed herein can also be used to evaluate the efficacy of other DNA sequences, such as cis-regulatory sequences.

[0186] Physical and functional characterization of the CHO integration site included verifying the insertion of exogenous nucleotide sequences using adapter PCR and full-length PCR, verifying the gene copy number of site-directed integration of exogenous nucleotide sequences using droplet digital PCR, and determining the expression level of antibodies in the culture medium using the OCTET QK molecular interaction instrument.

[0187] The methods provided herein can be used to generate cell populations expressing enhanced levels of the protein of interest. The absolute expression level will vary for the specific protein and depends on how efficiently the cells process the protein. Cell pools generated by integrating exogenous sequences into the expression-enhancing sequences disclosed herein become stable over time and can be treated as stable cell lines for most purposes. The recombination step can also be delayed until later in the development of the cell lines disclosed herein.

[0188] CHO expression enhancement locus and its fragments

[0189] Genetic engineering to modify the cellular genome at specific locations (i.e., target loci) can be achieved in several ways. One method involves using genetic editing techniques to stably integrate nucleic acid sequences into eukaryotic cells, where these sequences are exogenous sequences not typically found in such cells. Clonal expansion is necessary to ensure that progeny cells will possess consistent genotypic and phenotypic characteristics of the engineered cell line.

[0190] In some instances, cells containing a first recombinase recognition sequence and a second recombinase recognition sequence are provided, wherein each of the first and second recombinase recognition sequences is selected from the group comprising: loxP, lox511, lox5171, lox2272, lox2372, loxm2, lox-FAS, lox71, lox66, and mutants thereof. In this case, if recombinase-mediated cassette exchange (RMCE) is required, the site-specific recombinase is Cre recombinase or a derivative thereof. In other instances, each of the first and second recombinase recognition sequences is selected from the group comprising FRT, F3, F5, FRT mutant-10, FRT mutant+10, and mutants thereof, and in this context, if RMCE is required, the site-specific recombinase is Flp recombinase or a derivative thereof. In yet another instance, each of the first and second recombinase recognition sequences is selected from the group comprising attB, attP, and mutants thereof, and in this case, if RMCE is required, the site-specific recombinase is phiC31 integrase or a derivative thereof.

[0191] Homologous recombination in eukaryotic cells can be promoted by introducing breaks at integration sites in chromosomal DNA. Model systems have demonstrated that the frequency of homologous recombination increases during gene targeting if double-strand breaks are introduced into the target chromosomal sequence. This can be achieved by targeting certain nucleases to specific integration sites. DNA-binding proteins that recognize DNA sequences at target loci are known in the field. Gene-targeting vectors are also used to promote homologous recombination. In the absence of gene-targeting vectors for homology-directed repair, cells often close double-strand breaks via non-homologous end joining (NHEJ), which may result in the deletion or insertion of multiple nucleotides at the cleavage site. Insertions or deletions (InDels) should be present, resulting in the random insertion or deletion of small amounts of nucleotides at the break site, and these InDels can shift or disrupt any open reading frames (ORFs) of the gene within the target locus.

[0192] Homology-guided repair (or homology-guided recombination) (HDR) is particularly suitable for inserting or integrating genes at target loci.

[0193] Gene targeting vector construction and nuclease selection are within the skill of a person skilled in the art to which this disclosure pertains. Common gene targeting vector construction methods (gene editing methods) include, but are not limited to, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), and clusters of regularly spaced short palindromic repeats (CRISPR). RNA-directed endonucleases (RGENs) are programmable genome engineering tools developed from bacterial adaptive immune mechanisms. In this system (CRISPR / CRISPR-associated (Cas) immune responses), the protein Cas9 forms a sequence-specific endonuclease when it complexes with two RNAs (one of which directs target selection). RGENs consist of components (Cas9 and tracrRNA) and target-specific CRISPR RNA (crRNA). The efficiency of DNA target cleavage and the location of the cleavage site vary based on the position of the pre-intercalation sequence neighbor motif (PAM), which is an additional requirement for target recognition.

[0194] In some embodiments, plasmids for introduction into the genome, i.e., exogenous nucleic acids containing sequences encoding the gene of interest, recognition sequences, or gene cassettes, may, depending on the specific context, include a vector carrying the exogenous nucleic acid and one or more additional vectors or mRNAs. In some embodiments, the one or more vectors or mRNAs may include vectors having guide RNA, tracrRNA, and nucleotide sequences encoding Cas enzymes, as well as vectors containing donor (exogenous) nucleotide sequences. Such vector sequences contain nucleotide sequences encoding the gene of interest, or recognition sequences, or gene cassettes containing any of these exogenous elements intended for targeted insertion. When using mRNA, the mRNA can be transfected into cells using common transfection methods known to those skilled in the art and may encode enzymes such as transposases or endonucleases. While the mRNA introduced into the cells may be transient and not integrated into the genome, the mRNA may carry exogenous nucleic acids necessary or beneficial for integration.

[0195] Other homologous recombination methods are available to technicians, such as BuD-derived nucleases (BuDN) with precise DNA binding specificity.

[0196] This disclosure provides a method for modifying the genome of CHO cells, comprising introducing one or more plasmids into the cells, such as a donor plasmid, a recognition plasmid, and enzyme digestion of the plasmid into CHO cells.

[0197] The donor plasmid contains a gene encoding any therapeutically or industrially applicable protein as described herein. Site-specific integration of the donor plasmid requires identification of the target sequence within the target locus; therefore, a plasmid containing the recognition sequence, such as the SgRNA plasmid disclosed herein, can be simultaneously introduced. Enzyme digestion plasmids are used to cleave the specifically recognized site, such as the Cas9 plasmid disclosed herein.

[0198] In other embodiments, the homologous arm of the donor plasmid includes a target sequence that replaces an endogenous sequence within the locus. Alternatively, in other embodiments, the homologous arm of the donor plasmid includes a target sequence integrated into or inserted into an endogenous sequence within the locus.

[0199] By modifying the genome of CHO cells, site-directed fluorescently labeled cell lines can be obtained, which can serve as convenient and stable expression systems for recombinase-mediated cassette exchange (RMCE). The nucleic acid sequence encoding the protein of interest can be readily integrated into modified cells containing the first intron of the Nfat5 gene, NW_003614572.1 (269232--293783), or an enhanced fragment thereof, and possessing at least one recombinase recognition site, for example via the RMCE method. In some embodiments disclosed herein, the integration site used is located approximately 1000 bp (NW_003614572.1(271410..273418)), 2000 bp (NW_003614572.1(270410..274418)), 2500 bp, 5000 bp, or more base pairs upstream (5') or downstream (3') of the Nfat5 gene. Preferably, the integration site used is an sgRNA targeting sequence. More preferably, the integration site used is very close to the locus of interest, for example, less than about 1000 bp, 500 base pairs (bp), 250 bp, 100 bp, 50 bp, 25 bp, 10 bp or less than about 5 bp upstream (5') or downstream (3') of SEQ ID NO:2 (NW_003614572.1, 272344..272366) on the chromosome DNA.

[0200] Recombinant expression vectors may contain synthetic or cDNA-derived DNA fragments encoding proteins, operatively linked to suitable transcriptional and / or translational regulatory elements derived from mammalian, viral, or insect genes. These regulatory elements include transcription promoters, enhancers, sequences encoding suitable mRNA ribosome binding sites, and sequences controlling transcription and translation termination, as described in detail below. Mammalian expression vectors may also contain non-transcriptional elements, such as origins of replication, other 5' or 3' flanking non-transcriptional sequences, and 5' or 3' non-translational sequences, such as splice donor and acceptor sites. Optional marker genes to help identify transfectants may also be incorporated.

[0201] Fluorescent markers are optional marker genes suitable for identifying gene cassettes that have been or have not yet been successfully inserted and / or replaced, depending on the specific circumstances. Examples of fluorescent markers are well-known in the field and include (but are not limited to) Discosoma coral (DsRed), green fluorescent protein (Zsgreen1), enhanced green fluorescent protein (eGFP), blue-green fluorescent protein (CFP), enhanced blue-green fluorescent protein (eCFP), red fluorescent protein (mcherry), yellow fluorescent protein (YFP), enhanced yellow fluorescent protein (eYFP), and far-infrared fluorescent protein.

[0202] Bicistronic expression vectors for expressing multiple transcripts have been previously described and can be used in combination with the expression-enhancing sequences or fragments thereof disclosed herein. Other types of expression vectors will also be useful, such as those described in U.S. Patent No. 4,634,665 (Axel et al.) and U.S. Patent No. 4,656,134 (Ringold et al.).

[0203] target protein

[0204] Any target protein suitable for expression in eukaryotic cells may be used. For example, target proteins include (but are not limited to) antibodies or their antigen-binding fragments, chimeric antibodies or their antigen-binding fragments, ScFv or fragments thereof, Fc fusion proteins or fragments thereof, growth factors or fragments thereof, cytokines or fragments thereof, or extracellular domains of cell surface receptors or fragments thereof. The protein of interest may be a simple polypeptide composed of a single subunit or a complex multi-subunit protein containing two or more subunits.

[0205] The nucleic acid sequence expressed by the expression vector of the present invention can encode any polypeptide or protein of interest, especially polypeptides or proteins with diagnostic or therapeutic applications, such as growth factors, cytokines (interferons, interleukins), hormones, tyrosine kinases, receptors (GPCRs), integrins, transcription factors, blood clotting factors, single-chain antibodies, antibody fragments, or antibody-like molecules (anticalins).

[0206] Host cells and transfection

[0207] The host cells used in the methods disclosed herein are mammalian host cells, including, for example, Chinese hamster ovary (CHO) cells and mouse cells. One example of a suitable integration site is the loxP site. Another example of a suitable integration site is two recombinase recognition sites, selected, for example, from a group consisting of: loxP site, lox511 site, lox2272 site, lox2372 site, loxm2 site, lox71 site, lox66 site, and lox5171 site.

[0208] This disclosure includes mammalian host cells transfected with the expression vector or mRNA disclosed herein. While any mammalian cell, such as CHO, BHK, SP2 / 0, HEK293, or C127 cells, may be used, in one particular embodiment, the host cell is a CHO cell.

[0209] In a preferred embodiment, the host cell used in this disclosure is a CHO cell or a CHO cell line.

[0210] Suitable CHO cell lines can be selected from CHO pro3-, CHO DG44, CHO P12, dhfr-negative DUK-B11, and CHOK1SV.

[0211] Several transfection methods are known in the art. The chosen transfection protocol will depend on the host cell type and the nature of the GOI, and can be selected based on routine experiments. The basic requirement of any such protocol is to first introduce DNA encoding the protein of interest into a suitable host cell, and then identify and isolate the host cell with the incorporated heterologous DNA in a relatively stable and expressible manner. The mRNA molecules encoding proteins that are suitable for integration into the host cell genome or other functions can be transient and therefore time-limited.

[0212] Transfection protocols, and protocols used to introduce peptide or polynucleotide sequences into cells, can be modified. Non-restrictive transfection methods include chemical-based methods using liposomes; nanoparticles; calcium phosphate; dendritic polymers; or cationic polymers such as DEAE-glucan or polyethyleneimine. Non-chemical methods include electroporation; acoustic perforation; and optical transfection. Particle-based transfection includes the use of gene guns and magnet-assisted transfection. Viral methods can also be used for transfection. mRNA delivery includes methods using TransMessenger™.

[0213] A common method for introducing heterologous DNA into cells is calcium phosphate precipitation. DNA introduced into host cells via this method often undergoes rearrangement, making this procedure suitable for co-transfection of single genes.

[0214] Polyethylene-induced fusion of bacterial protoplasts with mammalian cells is another suitable method for introducing heterologous DNA. Protoplast fusion protocols often produce multiple copies of plasmid DNA integrated into the mammalian host cell genome, and this technique requires selection and amplification markers to be on the same plasmid as the foreign nucleotide sequence. Electroporation can also be used to introduce DNA directly into the host cell cytoplasm. Unlike protoplast fusion, electroporation does not require selection markers and foreign nucleotide sequences to be on the same plasmid.

[0215] Other reagents suitable for introducing heterologous DNA into mammalian cells have been described, such as Lipofectin™ and Lipofectamine™ reagents (Gibco BRL, Gaithersburg, Md.). Both of these commercially available reagents are used to form lipid-nucleic acid complexes (or liposomes), which, when applied to cultured cells, facilitate nucleic acid uptake into the cells.

[0216] In one embodiment, the introduction of one or more polynucleotides into cells is mediated by electroporation, intracytoplasmic injection, viral infection, adenovirus, lentivirus, retrovirus, transfection, lipid-mediated transfection, or via Nucleofection™.

[0217] The methods used to amplify exogenous nucleotide sequences are also required for recombinant protein expression and typically involve the use of selection markers (as reviewed in Kaufman above). Resistance to cytotoxic drugs is the most commonly used trait for selection markers and can be a result of dominant traits (e.g., usable independently of host cell type) or recessive traits (e.g., applicable to specific host cell types lacking any of the selected activities). Amplifiable markers in the prior art are suitable for the expression vectors disclosed herein.

[0218] Optional markers suitable for gene amplification in drug-resistant mammalian cells are shown in Table 1 above by Kaufman, RJ, ibid., and include DHFR-MTX resistance, P-glycoprotein and multidrug resistance (MDR) – various lipophilic cytotoxic agents (e.g., adenomyomycin, colchicine, vincristine) and adenosine deaminase (ADA) – Xyl-A or adenosine and 2'-deoxycofromycin.

[0219] Other dominant selectable markers include antibiotic resistance genes derived from microorganisms, such as resistance to neomycin, kanamycin, or hygromycin. However, these selectable markers have not been shown to be amplifiable. Several suitable selection systems exist in mammalian hosts.

[0220] Useful regulatory elements previously described or known in the art may also be included in nucleic acid constructs for transfecting mammalian cells. The chosen transfection protocol and the elements selected therein will depend on the type of host cell used. Those skilled in the art are familiar with many different protocols and host cells and can select an appropriate system for expressing the desired protein based on the requirements of the cell culture system used.

[0221] Other features of this disclosure will become apparent during the following description of exemplary embodiments, which are provided for the purpose of illustrating this disclosure and are not intended to limit it.

[0222] All the components used in the structures disclosed herein are known in the art. Therefore, those skilled in the art can obtain the corresponding components using conventional methods, such as PCR, fully artificial chemical synthesis, and enzyme digestion, and then link them together using well-known DNA ligation techniques to form the plasmids disclosed herein.

[0223] The plasmid disclosed herein may be co-transformed into host cells with Cre enzyme plasmid; or the vector disclosed herein may be integrated into the host cell chromosome using TAT-Cre recombinant protein with cell penetration activity.

[0224] The term "antibody" is used in the broadest sense and encompasses a wide range of antibody structures, including but not limited to monoclonal antibodies, polyclonal antibodies; monospecific antibodies, multispecific antibodies (e.g., bispecific antibodies), full-length antibodies, and antibody fragments (or antigen-binding fragments, or antigen-binding portions), as long as they exhibit the desired antigen-binding activity. "Natural antibody" refers to a naturally occurring immunoglobulin molecule. For example, a natural IgG antibody is a heterotetraglycoprotein of approximately 150,000 Daltons, composed of two light chains and two heavy chains linked by disulfide bonds. From the N to C terminus, each heavy chain has a variable region (VH, also known as the variable heavy domain or heavy chain variable region), followed by three constant domains (CH1, CH2, and CH3). Similarly, from the N to C terminus, each light chain has a variable region (VL, also known as the variable light domain or light chain variable domain), followed by a constant light domain (light chain constant region, CL). The terms “full-length antibody,” “intact antibody,” and “all antibody” are used interchangeably in this document to refer to antibodies that have a structure substantially similar to that of natural antibodies or that have a heavy chain containing an Fc region as defined herein.

[0225] The term "antibody fragment" refers to a molecule that is distinct from the complete antibody but contains a portion of the complete antibody that retains the antigen-binding ability of the complete antibody. Examples of antibody fragments include, but are not limited to, Fv, Fab, Fab', Fab'-SH, F(ab')2, single-domain antibodies, single-chain Fab (scFab), biantibodies, linear antibodies, single-chain antibody molecules (e.g., scFv); and multispecific antibodies formed from antibody fragments.

[0226] The term "monoclonal antibody" refers to a group of substantially homogeneous antibodies, meaning that the antibody molecules contained in this group have the same amino acid sequence, except for the possible small number of naturally occurring mutations. In contrast, polyclonal antibody formulations typically contain multiple different antibodies with different amino acid sequences in their variable structural domains, and they generally specifically target different epitopes. "Monoclonal" indicates the characteristic of an antibody obtained from a substantially homogeneous group of antibodies and should not be construed as requiring the antibody to be produced by any particular method. In some embodiments, the antibodies provided in this disclosure are monoclonal antibodies.

[0227] The term “nucleic acid” is used interchangeably with the term “polynucleotide” herein and refers to deoxyribonucleotides or ribonucleotides and their polymers in single-stranded or double-stranded form. The term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or links, which are synthetic, naturally occurring, or non-naturally occurring, have similar binding properties to a reference nucleic acid, and are metabolized in a manner similar to that of a reference nucleotide. Examples of such analogs include, but are not limited to, thiophosphates, aminophosphates, methylphosphonates, chiral methylphosphonates, 2-O-methylribonucleotides, and peptide-nucleic acids (PNAs). “Isolated nucleic acid” refers to a nucleic acid molecule that has been separated from its components in its natural environment. Isolated nucleic acids include nucleic acid molecules contained in cells that typically contain such molecules but are present extrachromosomally or at chromosomal locations other than their natural chromosomal locations. Isolated nucleic acids encoding the antigen-binding molecule refer to one or more nucleic acid molecules encoding the antibody heavy and light chains (or fragments thereof), including one or more such nucleic acid molecules in a single or separate vector, and one or more such nucleic acid molecules present at one or more locations in the host cell. Unless otherwise stated, a particular nucleic acid sequence also implicitly encompasses variants of its conserved modifications (e.g., degenerate codon substitutions) and complementary sequences, as well as explicitly stated sequences. Specifically, as detailed below, degenerate codon substitutions can be obtained by generating sequences in which the third position of one or more selected (or all) codons is substituted with a degenerate base and / or a deoxyinosine residue.

[0228] The terms “peptide” and “protein” are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers, where one or more amino acid residues are artificial chemical analogs of naturally occurring amino acids, as well as to both naturally occurring and non-naturally occurring amino acid polymers. Unless otherwise stated, a particular peptide sequence also implicitly encompasses variants with conserved modifications.

[0229] The term "vector" refers to a polynucleotide molecule capable of transporting another polynucleotide linked to it. One type of vector is a "plasmid," which is a circular double-stranded DNA loop in which an additional DNA segment can be attached. Another type of vector is a viral vector, such as an adeno-associated virus vector (AAV or AAV2), in which an additional DNA segment can be attached to the viral genome. Some vectors are capable of autonomous replication in the host cells to which they are introduced (e.g., bacterial vectors with bacterial origins of replication and attachable mammalian vectors). Other vectors (e.g., non-attached mammalian vectors) can integrate into the host cell's genome after introduction into the host cell, thereby replicating along with the host genome. The term "expression vector" or "expression construct" refers to a nucleic acid sequence suitable for transforming host cells and containing a sequence that directs and / or controls (alongside the host cell) the expression of one or more heterologous coding regions operatively linked to it. Expression constructs can include, but are not limited to, sequences that affect or control transcription, translation, and, in the presence of introns, influence RNA splicing of coding regions operatively linked to them.

[0230] The term "target gene expression cassette" includes the target gene, its upstream promoter sequence containing an enhancer, and its downstream polyA sequence.

[0231] The terms “host cell,” “host cell line,” and “host cell culture” are used interchangeably and refer to cells into which exogenous nucleic acids have been introduced, including the progeny of such cells. Host cells include “transformers” and “transformed cells,” which include primary transformed cells and their derived progeny, regardless of the number of passages. Progeny may not be identical to parental cells in their nucleic acid contents and may contain mutations. In this text, the term includes mutant progeny that have the same function or biological activity as cells screened or selected in primary transformed cells. Host cells include prokaryotic and eukaryotic host cells, wherein eukaryotic host cells include, but are not limited to, mammalian cells, insect cell lines, plant cells, and fungal cells. Mammalian host cells include human, mouse, rat, dog, monkey, pig, goat, cow, horse, and hamster cells, including but not limited to Chinese hamster ovary (CHO) cells, NSO, SP2 cells, HeLa cells, young hamster kidney (BHK) cells, monkey kidney cells (COS), human hepatocellular carcinoma cells (e.g., Hep G2), A549 cells, 3T3 cells, and HEK-293 cells.

[0232] "Optional" or "optionally" means that the event or circumstances described below may, but do not have to, occur, including the circumstances in which the event or circumstances may or may not occur.

[0233] Examples and Test Cases

[0234] The present disclosure is further described below with reference to examples and test cases, but these examples and test cases are not intended to limit the scope of the disclosure. Experimental methods in the examples and test cases of this disclosure that do not specify specific conditions are generally performed under conventional conditions, such as those described in Cold Spring Harbor's Antibody Technology Manual or Molecular Cloning Manual; or under conditions recommended by the raw material or commercial manufacturer. Reagents whose specific source is not specified are commercially available, conventional reagents.

[0235] Promoters are one of the most fundamental components of synthetic biological systems. To better assess the strength of synthetic promoters, we constructed labeled cell lines expressing fluorescent proteins (Figure 1) to evaluate gene expression levels at site-directed integration of synthetic promoters. We also designed an RMCE donor plasmid containing the desired 5′ synthetic promoter (enhancer + CMV core promoter), which was then expressed by an mCherry gene (with poly(A)) and a selection marker (PuroR) cassette (without poly, this donor plasmid is RMCE-acceptable). Once the donor sequence was correctly integrated, the newly generated cells survived under the selection pressure of puromycin. This strategy facilitates the rapid integration of the desired expression cassette into the same locus, allowing for the measurement of expression differentials derived from synthetic promoters.

[0236] Example

[0237] Example 1: Establishment of site-specific labeled cell lines

[0238] This embodiment utilizes CRISPR / Cas9 site-directed integration technology to integrate exogenous nucleotide sequences, including the lox recombination recognition gene sequence and the green fluorescent protein gene sequence (ZsGreen1), into CHO cells at the target site. Specifically, after co-transfecting CHO cells with the donor plasmid, sgRNA plasmid, and Cas9 plasmid, the Cas9 protein, guided by the sgRNA, cleaves at the sgRNA target site in the genome, generating double-strand breaks (DSBs), thereby triggering homology-directed repair (HDR). This mechanism utilizes the 5′ and 3′ homology arms (HAs) upstream and downstream of the genomic break site to perform homology recombination repair with the 5′ and 3′ homology arms on the donor plasmid. This allows the exogenous nucleotide sequence carrying the green fluorescent protein (ZsGreen1) gene and the resistance gene (NeoR) to be site-directedly inserted between the homology arms, resulting in a fluorescently labeled host cell line. The target integration site is located in the non-coding region of Nfat5 gene NW_003614572.1 (269178..352822), particularly within the first intron NW_003614572.1 (269232--293783), thus not disrupting normal cellular genomic mechanisms (e.g., protein translation) or altering cell phenotype. A highly efficient integration site is preferably located in sequence 2, as shown in SEQ ID NO:138 (approximately 4 kb). More preferably, it is located in sequence 3, as shown in SEQ ID NO:139 (approximately 2 kb), and particularly preferably at sequence 4 of NW_003614572.1 (272344..272366), see SEQ ID NO:2. The contents of Chinese patent application CN 202311141326.1 are incorporated herein by reference in their entirety.

[0239] Site sequence 2:

[0240] Site sequence 3:

[0241] Host cells labeled with fluorescently integrated proteins can be conveniently and rapidly obtained through flow cytometry sorting and PCR identification. The construction of the labeled cell line is shown in Figure 1. The construction steps are as follows:

[0242] 1. Construction of plasmids

[0243] Constructing a cell line for site-directed integration requires co-transfection of CHO cells with three different plasmids: Cas9 plasmid, sgRNA plasmid, and donor plasmid.

[0244] 1.1 Cas9 plasmid

[0245] Referring to Example 6 of WO2015052231, the gene sequence of the Cas9 protein (S. pyogenes strain M1, GAS genome, SEQ ID NO:2 in WO2015052231) was cloned into pJ607-03 to obtain the Cas9 plasmid.

[0246] 1.2 sgRNA plasmid

[0247] The applicant identified CTATGGTGGT (SEQ ID NO:1) (NW_003614572.1, 272410..272418) as a target sequence for the efficient and stable expression of exogenous nucleotide proteins. This target sequence is located within the first intron of the Nfat5 gene, NW_003614572.1 (269232..293783). When using CRISPR / Cas9 technology to target and integrate the gene encoding the desired protein, this target sequence can be used to determine the integration site that the sgRNA can recognize. One preferred sgRNA target sequence is GATGTGCCATACTACACATGTGG (SEQ ID NO:2) (NW_003614572.1, 272344..272366).

[0248] The following examples demonstrate the construction of the sgRNA plasmid for this application using SEQ ID NO:2 as the target sequence:

[0249] The following sgRNA sequence for recognizing the Nfat5 gene was synthesized:

[0250] SgRNA-5-fwd:5′-TTTGGGTCAGTCTGCACTGTGAGGGT-3′

[0251] SEQ ID NO:3

[0252] SgRNA-5-rev:5′-TAAAACCCTCACAGTGCAGACTGACC-3′

[0253] SEQ ID NO:4

[0254] The U6 promoter was cloned into the pSK vector (pBluescript-SK, GenBank ID: X52324.1) to obtain the pSK-U6 vector. The above-mentioned sgRNA sequence was cloned into the pSK-U6 vector to obtain sgRNA plasmids targeting the Nfat5 gene.

[0255] 1.3 Donor plasmid 1

[0256] Synthesize exogenous nucleotide sequences containing the following gene elements:

[0257] -5′arm-loxP-EF1α-ZsGreen1-bGHpolyA-SV40-NeoR-lox2272-bGHpolyA-3′arm-SV40-tagBFP-bGHpolyA-,

[0258] The above exogenous nucleotide sequence was cloned into the pUC57(addgene,54338) vector. The exemplary cloning steps are as follows:

[0259] 1) The loxP-EF1α-ZsGreen1-bGHpolyA-SV40-NeoR-lox2272-bGHpolyA was cloned into the pUC57 vector through two restriction enzyme sites (EcoRI, KpnI);

[0260] 2) SV40-tagBFP-bGHpolyA- was cloned into the pUC57 vector via two restriction enzyme sites (KpnI, BamHI);

[0261] 3) The 5′ arm homologous arm sequence was cloned into the above vector (EcoRI restriction site);

[0262] 4) The 3′ arm homologous arm sequence was cloned into the above vector (KpnI restriction site).

[0263] Among them, bGHpolyA is Bovine Growth Hormone polyA, EF1α is a strong mammalian expression promoter derived from human elongation factor 1α, SV40 is a mammalian expression promoter derived from simian vacuolating virus 40, ZsGreen1 is a green fluorescent marker gene, NeoR is a neomycin resistance gene, and tagBFP is a blue fluorescent protein. For details, please refer to the literature "Minimizing Clonal Variation during Mammalian Cell Line Engineering for Improved Systems Biology Data Generation".

[0264] The specific sequence is as follows:

[0265] loxP recognition sites:

[0266] lox2272 identification site:

[0267] ZsGreen1 green fluorescent marker gene:

[0268] EF1α promoter:

[0269] bGHpolyA bovine growth hormone polyadenylation signaling

[0270] SV40 starter:

[0271] NeoR resistance gene:

[0272] tagBFP fluorescent protein gene:

[0273] Once the target sequence is determined, those skilled in the art can integrate the gene sequence of any industrially or therapeutically applicable protein into the target locus.

[0274] The disclosed exogenous nucleotide sequence is inserted into the Nfat5 gene, and the homologous arm sequence is as follows:

[0275] Site-5-5′arm

[0276] Site-5-3′arm

[0277] The exogenous nucleotide sequence containing the homologous arm of the Nfat5 gene was integrated into the pUC57 vector to obtain donor plasmid 1, the composition of which is shown in Figure 1.

[0278] 2. Fixed-point integration

[0279] The three plasmids (Cas9 plasmid, sgRNA plasmid, and donor plasmid 1) were mixed in a 1:1:1 ratio (molar:molar:molar) and directly electroporated into CHO cells. After 3 days of electroporation, 400 μg / mL G418 was added to the cell culture medium and maintained until the cell viability recovered to over 90%, at which point the G418 was removed from the culture medium. Using flow cytometry (BD, FACSAria Fusion) under 488 nm and 405 nm laser light, cross gates for green and blue fluorescence signals were delineated using CHO wild-type cells as a negative control. Cells containing only green fluorescence and no blue fluorescence were sorted and single-cloned to obtain the site-labeled cell line 5-92 (a cell line where the fluorescent protein gene is site-integrated into the Nfat5 gene), which is site-labeled with green fluorescent protein.

[0280] Example 2: Site-directed integration of different promoters into fluorescently labeled cell lines

[0281] 1. Construction Principles

[0282] Recombinase-mediated cassette exchange (RMCE) allows the replacement of the green fluorescent protein gene in existing labeled cells with a red fluorescent protein gene expressed by different promoters (enhancer + CMV core), thus constructing a red fluorescent cell line. In the process of constructing a red fluorescent protein expression cell line using RMCE site-specific integration technology, a plasmid containing the red fluorescent gene, the puromycin resistance gene (PuroR), and its upstream promoter, along with a Cre enzyme expression plasmid, is co-transfected into CHO cells. The red fluorescent gene contains loxP and lox 2272 elements upstream and downstream, and similarly, the green fluorescent protein (ZsGreen1) gene and the neomycin resistance gene (NeoR) in the labeled cell line contain loxP and lox 2272 elements before and after them. Cre enzymes mediate homologous recombination at loxP and lox 2272 elements in the genome of labeled cell lines and plasmids carrying red fluorescent genes. Cell populations that have completed cassette exchange are enriched through puromycin selection, ultimately yielding cell populations expressing red fluorescent protein with varying intensities. The principle of cassette exchange in constructing protein-integrated cell lines is shown in Figure 2.

[0283] Flow cytometry analysis of the fluorescence intensity of red fluorescent cell lines can conveniently and quickly reflect the strength of upstream enhancers.

[0284] 2. Constructing red fluorescent gene expression cassette vectors with different promoters

[0285] A negative control plasmid for site-directed integration into the Nfat5 site, consisting of a red fluorescent gene cassette, was synthesized. Its structure is as follows:

[0286] -loxP-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272-

[0287] A positive control plasmid for site-directed integration into the Nfat5 site was synthesized, and its structure is as follows:

[0288] -loxP-CMV-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272-

[0289] Different enhancers were synthesized and inserted into the NotI restriction site upstream of the CMV core of the negative control plasmid to construct different red fluorescent gene cassette test plasmids, the structures of which are as follows:

[0290] -loxP-enhancer-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272-

[0291] The sequence information is as follows:

[0292] CMV core

[0293] CMV:

[0294] mCherry

[0295] PuroR

[0296] The gene expression cassette was then integrated into the pUC57 vector (addgene, 54338) to form an expression cassette plasmid containing the target protein gene, the structure of which is shown in Figure 2.

[0297] Example 3: Construction of site-directed integration red fluorescent cell lines containing different promoters

[0298] The expression cassette plasmid containing the target protein (red fluorescent protein) gene (prepared according to Example 2) was transfected with the PSF-CMV-CRE plasmid (Sigma-Aldrich, OGS591) at a ratio of 3:1 (mass:mass) using the transfection reagent PEI (Yisheng Biotechnology, 40820ES) into the site-labeled cell line 5-92. The transfection conditions were as follows:

[0299] 1) Take a 96-well plate (Enzyscreen, CR1496C) and add 280 μL of 5-92 cells at a concentration of 1E6 / mL to each well.

[0300] 2) Incubate 15 μL OPTI-MEM-I (Gibco, 51985091) + 1.2 μL PEI for 5 minutes.

[0301] 3) 15 μL OPTI-MEM-I + 0.15 μg PSF-CMV-CRE plasmid + 0.45 μg promoter plasmid (expression cassette plasmids containing different promoters prepared in Example 2),

[0302] 4) After mixing, incubate for 10 minutes, then transfer the entire transfection mixture into the cells. Place the 96-well plate on a shaker at 37°C and 800 rpm for incubation.

[0303] Seven days after transfection, puromycin (Thermo, A1113802) was added to bring the final concentration in each well to 3 μg / mL to maintain cell culture. After 10 days, the puromycin concentration was gradually reduced to below 0.5 μg / mL. Cell density recovered after approximately 10-12 days.

[0304] Cells were analyzed by flow cytometry (Thermo, Attune NXT) using fluorescence intensity analysis under 488nm and 561nm laser light, and the instrument's built-in software was used.

[0305] Example 4: Selection of Transcription Factor Regulatory Elements and Determination of Optimal Copy Number

[0306] 4.1 Assessment of TFRE expression strength

[0307] Transcription factor recognition elements (TFREs) from Table 1 were selected and tandemly repeated in sets of eight copies to form isotype functional promoter sequences (TFREs in enhancers are of the same type). These sequences were then integrated into host cells along with red fluorescent protein, and the expression capacity of TFREs was evaluated using flow cytometry (Example).

[0308] Table 1. Sequence listing of TFREs, regulatory elements of transcription factors

[0309] Each TFRE sequence in the table was tandemly repeated at 8 copies, with a CAT sequence as the spacer, to form a TFRE enhancer. This enhancer was then ligated to the NotI restriction site upstream of the CMV core in the vector loxP-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272. The final vector sequence structure is -loxP-enhancer-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272-.

[0310] The synthesized tandem repeat TFRE (i.e., enhancer) together with the CMV core forms a complete promoter sequence.

[0311] The synthesized vector containing multiple promoter sequences and the vector containing the CMV core (i.e., no TFRE upstream of the CMV core) as a negative control were transfected into the 5-92 cell line using PEI. After pressurization and depressurization, the cells were analyzed by flow cytometry to determine the red fluorescence intensity of the cell population in the FITC- / mCherry+ quadrant. The results are shown in Figure 5 and Table 2 (all results have been normalized using the site-integrated cell population containing only the CMV core). The formula is as follows:

[0312] Expression fold change = Median fluorescence intensity (MFI) of each TFREs / MFI of CMV core.

[0313] Coefficient of variation = Standard deviation / Median

[0314] Table 2. Comparison of the expressive power of different TFREs

[0315] Conclusion: Based on the results shown in the table above, the expression levels of existing TFREs were preliminarily assessed. Two high-expression TFREs (fold change > 100) were identified as REL (SEQ ID NO: 23) and NFκB (SEQ ID NO: 22), and four medium-expression TFREs (10 < fold change < 100): GABPβ (SEQ ID NO: 20), Sp1 (SEQ ID NO: 21), AhR / ARNT (SEQ ID NO: 18), and DMP1 (SEQ ID NO: 19). The remaining CaRF (SEQ ID NO: 24) was identified as a low-expression TFRE (fold change < 10).

[0316] 4.2 Determining the Optimal Copy Number of TFRE

[0317] Homogeneous amplification of TFREs may easily lead to signal saturation. To avoid the stress of excessive TFRE copying on the synthesized promoter, we need to determine the optimal copy number for each TFRE. We selected seven TFREs from Table 1 for further investigation.

[0318] The TFRE sequences were tandemly repeated at copy numbers of 2, 4, 8, and 12, respectively, with CAT sequences used as spacers to form TFRE enhancers. These enhancers were then ligated to the NotI restriction site upstream of the CMV core in the vector loxP-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272. The final vector sequence structure is -loxP-enhancer-CMV core-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272-.

[0319] The synthesized tandem repeat TFRE (i.e., enhancer) together with the CMV core constitutes a complete artificial promoter. Vectors containing different artificial promoters were constructed according to Example 2.

[0320] The synthesized vector containing multiple artificial promoter sequences and the negative control vector containing the CMV core (i.e., without TFRE upstream of the CMV core) (loxP-CMVcore-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272) were transfected into 5-92 cell lines using PEI. After pressurization and depressurization, the cells were analyzed by flow cytometry to determine the red fluorescence intensity of the cell population in the FITC- / mCherry+ quadrant. The results are shown in Tables 3 and 4 (all results have been normalized using site-integrated cell populations containing only the CMV core).

[0321] Expression fold change = Mean fluorescence intensity (MFI) of each TFREs / MFI of CMV core.

[0322] Table 3. Comparison of copy number expression gene intensity among different TFRE species

[0323] Conclusion: Based on the results shown in the table above, the optimal copy numbers for AhR / ARNT, CaRF, DMP1, GABPβ, NFκB, Sp1, and REL were determined as follows:

[0324] Table 4. Determination of the optimal copy number for TFRE

[0325] 4.3 Determining the total number of TFRE copies contained in the promoter

[0326] It has been reported that heteropromoters drive stronger reporter gene expression than homopromoters. After determining the optimal copy number of different TFREs in the promoter, in order to design stronger synthetic promoters, we constructed heteropromoters (containing different types of TFREs) by combining different TFREs, and evaluated the total copy number of enhancers composed of different TFREs.

[0327] 4.3.1 Selection of TFRE spacer sequence and restriction enzyme sites

[0328] Based on the presence or absence of CaRF in the enhancer sequence and the optimal copy number for each TFRE, two types of promoters were designed: one consisting of 28 TFREs and the other of 30 TFREs. Spacer sequences selected from Table 5 can be used between the different TFREs, and restriction enzyme sites as shown in Table 6 can be added to the 5′ or 3′ end.

[0329] Table 5. Available Interval Sequences

[0330] Table 6. Common enzyme cleavage sites and their sequences

[0331] 4.3.2 Evaluation of the expression capacity of synthetic promoters containing 28-30 TFREs

[0332] Enhancers were constructed using the TFREs in Table 1, and artificially synthesized promoters containing 28 TFREs (Table 7) and 30 TFREs (Table 8) were designed.

[0333] Table 7. Copy number composition of the promoter containing 28 TFREs

[0334] Table 8. Copy number composition of promoters containing 30 TFREs

[0335] Referring to the expression capabilities of different TFREs in Tables 2 and 3, we set the copy number of high-expression TFREs (NFκB and REL) to the optimal copy number found in homologous promoters (Table 4). For TFREs from the intermediate expression group (AhR / ARNT, DMP1, GABPβ, Sp1), we set their copy number to be less than or equal to their optimal copy number. For TFREs with weak expression capabilities (CaRF), we set their copy number to 0 or 2. Twenty different enhancer sequences were synthesized by randomly arranging 28 TFRE elements (Table 7) or 30 TFRE elements (Table 8) with CAT (SEQ ID NO:28) as the spacer sequence. A SpeI restriction site ACTAGT (SEQ ID NO:117) was added to the 5′ end, and a NotI restriction site GCGGCCGC (SEQ ID NO:114) was added to the 3′ end. A CMV core sequence (SEQ ID NO:15) was then added downstream of the NotI restriction site at the 3′ end to construct the corresponding synthetic promoter. For example, the synthetic promoter 30F8 consists of a CMV core sequence and an enhancer EH1-30F8 sequence, formed by restriction sites and spacer sequences. Other promoters follow the same principle.

[0336] The sequences of the 15 enhancers contained in the 15 constructed synthetic promoters are shown below:

[0337] The synthesized promoter 28F6 contains the enhancer EH1-28F6 sequence:

[0338] The synthesized promoter 28E7 contains the enhancer sequence EH2-28E7:

[0339] The synthesized promoter 28F7 contains the enhancer EH3-28F7 sequence:

[0340] The synthesized promoter 28G7 contains the enhancer sequence EH4-28G7:

[0341] The synthesized promoter 28B8 contains the enhancer sequence EH5-28B8:

[0342] The synthesized promoter 30E8 contains the enhancer sequence EH6-30E8:

[0343] The synthesized promoter 30F8 contains the enhancer sequence EH7-30F8:

[0344] The synthesized promoter 30G8 contains the enhancer sequence EH8-30G8:

[0345] The synthesized promoter 28G6 contains the enhancer sequence EH9-28G6:

[0346] The synthesized promoter 30D8 contains the enhancer EH10-30D8 sequence:

[0347] The synthesized promoter 28B7 contains the enhancer sequence EH11-28B7:

[0348] The synthesized promoter 28C7 contains the enhancer sequence EH12-28C7:

[0349] The synthesized promoter 28D7 contains the enhancer sequence EH13-28D7:

[0350] The synthesized promoter 28C8 contains the enhancer sequence EH14-28C8:

[0351] The synthesized promoter 30B9 contains the enhancer sequence EH15-30B9:

[0352] Vectors constructed from any 10 of the above-mentioned synthetic promoter sequences (Example 2) and a vector containing the CMV promoter (loxP-CMV-mcherry-bGH polyA-SV40 promoter-PuroR-lox2272) as a positive control were transfected again with PEI at cell lines 5-92. After pressurization and depressurization, the cells were analyzed by flow cytometry again. After the fluorescence values ​​were normalized, the results showed that different synthetic promoters had good expression ability for non-secretory proteins (fluorescent proteins).

[0353] Example 5: By constructing a cell line for site-directed integration antibody expression, the expression capabilities of synthetic promoters and CMV promoters for secreted proteins were evaluated and compared.

[0354] Recombinase-mediated cassette exchange (RMCE) allows the replacement of the green fluorescent protein (non-secretory protein) gene in existing labeled cells with the antibody (secretory protein) light and heavy chain genes, thus constructing antibody-expressing cell lines. In the RMCE site-specific integration technology for constructing antibody-expressing cell lines, CHO cells are co-transfected with a plasmid carrying the antibody light and heavy chain genes and a Cre enzyme plasmid. The antibody light and heavy chain genes contain loxP and lox 2272 elements upstream and downstream, and similarly, loxP and lox 2272 elements are present before and after the green fluorescent (ZsGreen1) gene and the resistance gene (NeoR) in the labeled cell line. The Cre enzyme mediates homologous recombination at the loxP and lox 2272 elements in the genome of the labeled cell line and the plasmid carrying the antibody light and heavy chain genes, thereby replacing the green fluorescent (ZsGreen1) gene and the resistance gene (NeoR) in the site-specific labeled cell line genome with the antibody light and heavy chain genes, resulting in antibody-expressing cells without green fluorescence. The principle of constructing protein-integrated cell lines using cassette exchange is shown in Figure 2.

[0355] Flow cytometry and PCR identification can be used to quickly and easily obtain antibody light and heavy chain gene-integrated cell lines without extensive screening.

[0356] 1. Construction of antibody light and heavy chain gene expression cassette vector

[0357] A gene cassette for antibody light and heavy chains, capable of site-directed integration into the Nfat5 site, was synthesized. Its structure is: -loxP-enhancer-CMV core-LC-bGH polyA-enhancer-CMV core-HC-lox2272-. Referring to the expression cassette plasmid in Figure 3, the Promoter in the figure is the enhancer-CMV core. The relevant target gene element sequence is an anti-GITR antibody (derived from WO2019001559A1), and the nucleic acid sequence encoding the antibody is as follows:

[0358] Anti-GITR antibody light chain gene sequence:

[0359] Anti-GITR antibody heavy chain gene sequence:

[0360] The gene expression cassette was then integrated into the pUC57 vector to form an expression cassette plasmid containing the target protein gene, the structure of which is shown in Figure 3.

[0361] 2. Construction of antibody-targeted integration cell lines

[0362] The expression cassette plasmid containing the target protein gene (GOI) was electroporated onto cell line 5-92 at a ratio of 3:1 (mass:mass) with the PSF-CMV-CRE plasmid (Sigma-Aldrich, OGS591). One week after transfection, the cells were sorted by flow cytometry (BD, FACSMelody). TM Under 488nm laser light, using BD FACS Chorus software, the 0.5% cell population with the weakest green fluorescence intensity was selected. Cells within this range were sorted to obtain monoclonal antibody-integrated cell lines. These were then progressively expanded for further analysis.

[0363] Test case

[0364] Test Example 1: The effect of different promoters on the positive monoclonal gain rate of antibody site-directed integration

[0365] Seven days after transfection, flow cytometry sorting was performed to separate monoclonal cells without green fluorescence signal (the weakest 0.5% of cell populations with FITC signal value), determining the number of selected monoclonal cells. Cells were statically cultured in 96-well plates, and then the monoclonal cells were progressively expanded to 24-well and 24-well deep-dip plates with shaking culture (250 rpm). The culture medium used was CD CHO (Gibco, 10743029). Flow cytometry analysis was then performed on the monoclonal cells to confirm the number of monoclonal cells without green fluorescence signal (#of FITC-). The culture supernatant was used for titer assay, and the number of positive monoclonal cells was determined when titer > 3 μg / mL. The results showed that the positive clonal rate (integration success rate) obtained using the synthetic promoter disclosed herein was significantly better than that obtained using the CMV promoter (Table 9).

[0366] Integration success rate = number of positive monoclonal antibodies / number of selected monoclonal antibodies.

[0367] Table 9. Integration success rate of secretory proteins (monoclonal antibodies) by different promoters

[0368] Test Example 2: Effects of different promoters on antibody expression levels in site-directed integration cells

[0369] Monoclonal cells expressing anti-GITR antibodies with different promoters were subjected to 14-day fed-batch culture.

[0370] Cell culture method: Site-labeled cell lines in the exponential growth phase were cultured at 5 × 10⁻⁶ cells / year. 5Cells / mL were seeded into 20 mL of CD CHO medium containing 8 mM glutamine (Lot: 10743029, Gibco). TM Cells were cultured in 125 mL shake flasks under the following conditions: 36.5℃, 6.0% CO2, 120 rpm, and 80% relative humidity. On days 4, 6, 8, 11, and 14, appropriate cell samples were taken, stained with 0.4% trypan blue solution, and the viable cell density (VCD) and viability (Viability%) were recorded using a Countstar cell counter. On days 4, 6, 8, and 11, 5%, 5%, 7.5%, and 7.5% Efficient-Pro were added to the culture system, respectively. TM Feed 1 medium (Lot: A5208801, Gibco) TM L-glutamine at a concentration of 200 mM (Lot: 25030081, Gibco) was dispensed in 0.4 mL, 0.4 mL, 0.6 mL, and 0.6 mL solutions. TM 10 μL of cell suspension was diluted with 990 μL of PBS, and the glucose concentration in the culture medium was measured using a Biosen C-Line Glucose and Lactate analyzer (EKF Diagnostic). Based on the glucose measurement results, an appropriate amount of glucose was added to bring the final glucose concentration to 12 g / L. After 14 days of culture, 100 μL of cell supernatant was diluted 5-fold with PBS, and the antibody expression level (Titer value) in the culture medium was determined using an OCTET QK molecular interaction analyzer.

[0371] Comparing the Titer value and cell growth at 14 days, it was found that the synthetic promoter was superior to the CMV promoter.

[0372] Formula for calculating total viable cell density (IVCD):

[0373] Table 10. Effects of different promoters on antibody expression levels in site-directed integration cells.

[0374] The results showed that the expression levels and IVCD of anti-GITR antibodies expressed by 10 promoters were higher than those of the positive control CMV on day 14.

[0375] Although the invention has been described in detail with the aid of accompanying drawings and examples for clarity of understanding, these descriptions and examples should not be construed as limiting the scope of this disclosure. In fact, various modifications to this disclosure will be readily apparent to those skilled in the art from the foregoing description and drawings, in addition to those described herein. It is intended that these modifications fall within the scope of the appended claims. All patent and scientific literature disclosures cited herein are clearly and fully incorporated by reference.

Claims

1. A reinforcing sub-element comprising a TFRE, said TFRE being composed of REL, NFκB and one or more selected from GABPβ, Sp1, AhR / ARNT, DMP1 and CaRF.

2. The reinforcing sub-element as claimed in claim 1, wherein the copy number of REL and NFκB in the reinforcing sub-element is the same or different, and each independently is 6-15, preferably 8-12; more preferably 8. The copy number of GABPβ, Sp1, AhR / ARNT and DMP1 in the reinforcing sub-element may be the same or different, and each is independently 2-6, preferably 2-4, and more preferably 2 or 4. The copy number of the CaRF in the reinforcing sub-element is 0-2, preferably 0 or 2.

3. The reinforcing sub-element as claimed in claim 1 or 2, wherein the TFRE comprises sequences of SEQ ID NO:22 and SEQ ID NO:23, and one or more sequences selected from SEQ ID NO:18,19,20,21,24.

4. The reinforcing sub-element as claimed in claim 3, wherein the TFREs shown in SEQ ID NO:22 and SEQ ID NO:23 have the same or different copy numbers in the reinforcing sub-element, and each independently has 6-15 copies, preferably 8-12 copies, and more preferably 8 copies; The TFREs shown in SEQ ID NO:18-21 have the same or different copy numbers in the reinforcing sub-elements, and each independently has 2-6 copies, preferably 2-4 copies, and more preferably 2 or 4 copies; The copy number of the TFRE shown in SEQ ID NO:24 in the reinforcing sub-element is 0-2, preferably 0 or 2.

5. The reinforcing sub-element as described in any one of claims 1-4, wherein the total number of copies of the TFRE is 25-35, preferably 28-32, and most preferably 30.

6. The reinforcing sub-element as described in any one of claims 1-5, wherein there is a spacer sequence P between the TFREs, the spacer sequence P being selected from one or more of SEQ ID NO:25-104; preferably, the spacer sequence is SEQ ID NO:

28.

7. The reinforcing sub-element as claimed in any one of claims 1-6, wherein the reinforcing sub-element is selected from the following combinations: Enhancer EH1-28F6: -P-Sp1-P-REL-P-DMP1-P-GABPβ-P-REL-P-REL-P-REL-P-AhR / ARNT-P-DMP1-P-REL-P-DMP1-P-AhR / ARNT-P-NFκB-P-REL-P-N FκB-P-NFκB-P-GABPβ-P-AhR / ARNT-P-Sp1-P-NFκB-P-REL-P-AhR / ARNT-P-NFκB-P-DMP1-P-NFκB-P-REL-P-NFκB-P-NFκB-P, Pipeline EH2-28E7: -P-AhR / ARNT-P-NFκB-P-DMP1-P-NFκB-P-NFκB-P-NFκB-P-Sp1-P-NFκB-P-REL-P-AhR / ARNT-P-AhR / ARNT-P-NFκB-P-REL-P-G ABPβ-P-REL-P-REL-P-DMP1-P-GABPβ-P-NFκB-P-REL-P-REL-P-NFκB-P-REL-P-Sp1-P-REL-P-DMP1-P-AhR / ARNT-P-DMP1-P-. Pipeline EH3-28F7: -P-AhR / ARNT-P-NFκB-P-NFκB-P-NFκB-P-REL-P-REL-P-AhR / ARNT-P-Sp1-P-NFκB-P-DMP1-P-GABPβ-P-AhR / ARNT-P-GABPβ-P -NFκB-P-Sp1-P-REL-P-NFκB-P-REL-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P-REL-P-NFκB-P-REL-P-DMP1-P-. Pipeline EH4-28G7: -P-REL-P-DMP1-P-Sp1-P-REL-P-NFκB-P-Sp1-P-GABPβ-P-GABPβ-P-REL-P-AhR / ARNT-P-DMP1-P-REL-P-REL-P-NFκB-P-REL- P-NFκB-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P-NFκB-P-DMP1-P-NFκB-P-NFκB-P-AhR / ARNT-P-AhR / ARNT-P-REL-P-NFκB-P-. Pipeline EH5-28B8: -P-GABPβ-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-Sp1-P-GABPβ-P-REL-P-REL-P-NFκB-P-AhR / ARNT-P-REL-P-DMP1-P-NFκB-P -NFκB-P-NFκB-P-NFκB-P-Sp1-P-NFκB-P-REL-P-DMP1-P-AhR / ARNT-P-AhR / ARNT-P-DMP1-P-REL-P-REL-P-NFκB-P-NFκB-P-, Pipeline EH6-30E8: -P-GABPβ-P-NFκB-P-AhR / ARNT-P-DMP1-P-REL-P-NFκB-P-AhR / ARNT-P-Sp1-P-REL-P-REL-P-DMP1-P-DMP1-P-REL-P-AhR / ARNT-P-NF κB-P-REL-P-REL-P-CaRF-P-NFκB-P-REL-P-AhR / ARNT-P-Sp1-P-NFκB-P-GABPβ-P-REL-P-NFκB-P-DMP1-P-NFκB-P-CaRF-P-NFκB-P-, Pipe EH7-30F8: -P-NFκB-P-NFκB-P-GABPβ-P-NFκB-P-NFκB-P-REL-P-DMP1-P-NFκB-P-CaRF-P-DMP1-P-DMP1-P-DMP1-P-REL-P-AhR / ARNT-P-REL-P-C aRF-P-REL-P-AhR / ARNT-P-GABPβ-P-AhR / ARNT-P-AhR / ARNT-P-Sp1-P-NFκB-P-REL-P-REL-P-Sp1-P-REL-P-NFκB-P-REL-P-NFκB-P-, Pipeline EH8-30G8: -P-DMP1-P-REL-P-Sp1-P-REL-P-DMP1-P-REL-P-REL-P-REL-P-AhR / ARNT-P-NFκB-P-NFκB-P-NFκB-P-NFκB-P-REL-P-DMP1-P-AhR / AR NT-P-CaRF-P-GABPβ-P-Sp1-P-NFκB-P-GABPβ-P-NFκB-P-AhR / ARNT-P-DMP1-P-REL-P-CaRF-P-NFκB-P-NFκB-P-REL-P-AhR / ARNT-P-, Pipeline EH9-28G6: -P-NFκB-P-NFκB-P-AhR / ARNT-P-NFκB-P-REL-P-NFκB-P-REL-P-REL-P-DMP1-P-NFκB-P-GABPβ-P-AhR / ARNT-P-REL-P-NFκB- P-DMP1-P-DMP1-P-DMP1-P-REL-P-Sp1-P-REL-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-GABPβ-P-REL-P-REL-P-NFκB-P-Sp1-P-, Pipeline EH10-30D8: -P-REL-P-AhR / ARNT-P-NFκB-P-CaRF-P-NFκB-P-DMP1-P-DMP1-P-GABPβ-P-NFκB-P-REL-P-REL-P-CaRF-P-DMP1-P-NFκB-P-REL-P-NF κB-P-Sp1-P-REL-P-NFκB-P-REL-P-AhR / ARNT-P-Sp1-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL-P-REL-P-NFκB-P-GABPβ-P-AhR / ARNT-P-CH Pipeline EH11-28B7 Liquid: -P-NFκB-P-NFκB-P-GABPβ-P-REL-P-REL-P-NFκB-P-REL-P-DMP1-P-NFκB-P-AhR / ARNT-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL- P-GABPβ-P-REL-P-REL-P-DMP1-P-Sp1-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-REL-P-Sp1-P-DMP1-P-REL-P-NFκB-P-NFκB-P-. Type EH12-28C7 plate: -P-DMP1-P-REL-P-NFκB-P-DMP1-P-REL-P-NFκB-P-NFκB-P-AhR / ARNT-P-AhR / ARNT-P-NFκB-P-AhR / ARNT-P-GABPβ-P-NFκB-P -NFκB-P-NFκB-P-REL-P-REL-P-Sp1-P-NFκB-P-GABPβ-P-Sp1-P-DMP1-P-REL-P-REL-P-AhR / ARNT-P-REL-P-REL-P-DMP1-P-. Pipeline EH13-28D7 plate: -P-REL-P-Sp1-P-REL-P-REL-P-DMP1-P-REL-P-REL-P-AhR / ARNT-P-REL-P-DMP1-P-DMP1-P-NFκB-P-REL-P-NFκB-P-NFκB-P- REL-P-NFκB-P-AhR / ARNT-P-NFκB-P-NFκB-P-GABPβ-P-Sp1-P-AhR / ARNT-P-DMP1-P-GABPβ-P-NFκB-P-NFκB-P-AhR / ARNT-P-, Enhancer EH14-28C8 sequence: -P-Sp1-P-NFκB-P-AhR / ARNT-P-NFκB-P-DMP1-P-GABPβ-P-REL-P-REL-P-NFκB-P-DMP1-P-REL-P-NFκB-P-REL-P-AhR / ARNT-P -NFκB-P-Sp1-P-NFκB-P-NFκB-P-REL-P-AhR / ARNT-P-NFκB-P-GABPβ-P-REL-P-DMP1-P-AhR / ARNT-P-REL-P-REL-P-DMP1-P-, and Enhancer EH15-30B9 sequence: -P-REL-P-Sp1-P-AhR / ARNT-P-NFκB-P-REL-P-NFκB-P-GABPβ-P-AhR / ARNT-P-AhR / ARNT-P-DMP1-P-REL-P-DMP1-P-DMP1-P-NFκB-P-N FκB-P-DMP1-P-NFκB-P-AhR / ARNT-P-REL-P-REL-P-NFκB-P-Sp1-P-REL-P-CaRF-P-REL-P-GABPβca7t-NFκB-P-CaRF-P-REL-P-NFκB-P; Where REL, NFκB, GABPβ, Sp1, AhR / ARNT, DMP1, and CaRF are TFREs; P is the interval sequence as defined in claim 6.

8. The reinforcing sub-element as claimed in any one of claims 1-7, wherein the reinforcing sub-element sequence is selected from one or more of SEQ ID NO:121-135.

9. A synthetic promoter, characterized in that The device comprises an enhancer element according to any one of claims 1-8; it includes a restriction enzyme site sequence at the 5′ or 3′ end, the restriction enzyme site being linked to a transcription factor regulatory element sequence via a spacer sequence; and a core promoter sequence is added downstream of the 3′ restriction enzyme site. Preferably, the restriction enzyme site sequence is selected from SEQ ID NO:105-120, the spacer sequence (P) is selected from SEQ ID NO:25-104, and the core promoter is selected from hCMV, hEF-1α, SV40, UbC, EF1A, PGK, and CAGG; More preferably, the 5′ restriction site sequence is shown in SEQ ID NO:117, the 3′ restriction site sequence is shown in SEQ ID NO:114, the spacer sequence is shown in SEQ ID NO:28, and the core promoter is shown in SEQ ID NO:

15.

10. An expression vector comprising the nucleic acid of claim 9. The expression vector contains the synthetic promoter as described in claim 9.

11. The expression vector according to claim 10, comprising the nucleotide sequence of the exogenous protein to be expressed; preferably, the exogenous protein nucleotide sequence encodes a non-secretory protein and a secretory protein; more preferably, the secretory protein is an antibody molecule.

12. A host cell, characterized in that The host cell is transfected with the expression vector of claim 10 or 11; preferably, the transfection is site-directed integration; more preferably, the site-directed integration site is the Nfat5 gene; most preferably, the site-directed integration site is the target sequence in the Nfat5 gene as shown in SEQ ID NO:

2.

13. The host cell of claim 12, wherein The host cell is a mammalian host cell. Preferably, the host cell is a CHO, BHK, SP2 / 0, HEK293, or C127 cell; more preferably, the host cell is a CHO cell.