Truncated chromatin opening element sequence, and combination and use thereof

By combining truncated UCOE sequences with the PiggyBac transposon system, the problems of low transfection efficiency and slow cell recovery in antibody drug production using UCOE combinations were solved, resulting in shorter cell line construction time and increased protein yield, making it suitable for efficient antibody production in CHO cells.

WO2026077369A1PCT designated stage Publication Date: 2026-04-16SHANGHAI QILU PHARMACEUTICAL RESEARCH & DEVELOPMENT CENTRE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/126415
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-10
Filing Date
2025-10-09
Publication Date
2026-04-16

AI Technical Summary

Technical Problem

In existing technologies, UCOE sequence length optimization is insufficient, resulting in low transfection efficiency, slow cell recovery, and inability to perform rapid bulk pool screening. This leads to long cell line construction time and heavy workload. Furthermore, the use of UCOE in antibody drug production for stable cell line construction suffers from low integration efficiency and poor stability.

Method used

By combining truncated UCOE sequences with the PiggyBac transposon system, and integrating UCOE elements with ITR sequences into the host genome, combined with the bulkpool rapid screening process, the UCOE sequence length was optimized, improving transfection efficiency and cell line construction stability.

Benefits of technology

It significantly shortened the cell line construction cycle, improved protein production efficiency, and enabled rapid screening and efficient expression in large-scale production, with protein yield being higher than that of the control group.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PCTCN2025126415-FTAPPB-I100001
    Figure PCTCN2025126415-FTAPPB-I100001
  • Figure PCTCN2025126415-FTAPPB-I100002
    Figure PCTCN2025126415-FTAPPB-I100002
  • Figure PCTCN2025126415-FTAPPB-I100003
    Figure PCTCN2025126415-FTAPPB-I100003
Patent Text Reader

Abstract

Provided are a truncated chromatin opening element and a combination thereof, an expression system comprising the chromatin opening element or the combination thereof and applied to eukaryotic cells, and use thereof. The chromatin opening elements are truncated sequences of two chromatin opening elements from different sources, respectively, and can be used alone or in combination. The chromatin opening element significantly shortens the effective sequence, thereby facilitating the integration of a vector comprising the element or the combination thereof with other elements, such as integration with ITR, achieving semi-site-specific integration in the presence of the PiggyBac transposase, and then using a bulkpool rapid screening process to shorten the protein production and cell strain construction cycle by about 4-5 weeks while maintaining comparable yields.
Need to check novelty before this filing date? Find Prior Art

Description

Truncated chromatin open element sequences, their combinations and applications

[0001] Related applications

[0002] This application claims priority and related interests in Chinese Patent Application No. 202411412553.8, filed on October 10, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of biotechnology, specifically to a combination of truncated sequences of chromatin open elements, an expression system containing the combination of truncated sequences of chromatin open elements, and its application in the construction of CHO cell lines. Background Technology

[0004] Currently, hundreds of antibody drugs have been approved globally, and these drugs are increasingly benefiting patients across various therapeutic areas. With the development of antibody drugs, the antibody market has reached hundreds of billions of dollars, with production scales reaching tens of thousands of liters per batch. In the antibody production process, cell line construction, as the foundation of antibody production, has been extensively researched and focused on. How to achieve rapid research and development and high-yield of antibody drugs through cell line construction is a goal that the industry has consistently pursued.

[0005] In the process of constructing stable cell lines, the target gene is usually inserted into the host cell genome in a random integration manner, which has the following drawbacks: the integration efficiency is usually less than one in ten thousand; the construction of stable cell lines depends on a large number of screenings, which is labor-intensive and time-consuming; the target gene is easily affected by position effects and has a low expression level; the target gene is easily silenced by the host cell and has poor stability; the cell pool yield is low, making it difficult to obtain sufficient protein in the early stages to meet the needs of new drug development.

[0006] Studies have reported that unopen chromatin elements (UCOEs) can reduce methylation, maintain open chromatin structures, and extend this active chromatin state to other linked regulatory elements, such as promoters, thereby preventing gene transcriptional silencing and protecting the target gene from the "position effect." [1-2] This leads to high expression of the target protein in cells, thereby increasing the formation rate of high-expression clones and saving time in obtaining them. To date, four DNA sequences have been defined as UCOEs, including three human UCOEs (TBP / PSMB1, HNRPA2B1 / CBX3, and SURF1 / SURF2) and one mouse UCOE (Rps3). Studies applying UCOEs to eukaryotic host CHO cells have largely focused on HNRPA2B1 / CBX3 and Rps3. [2-6]Most of the studies involve the application of complete functional UCOE sequences or single UCOE partial functional sequences in CHO cells. However, there is a lack of research on the combined application of UCOE and CHOZN cells, as well as the combined application of truncated sequences of UCOE elements from different sources. Therefore, there is a technical need to optimize the length of UCOE sequences.

[0007] Previous studies on UCOE sequence length optimization have mostly focused on truncating the sequence based on its characteristics or functional partitions. [7-9] There are also reports of truncation based on bioinformatics analysis.

[0010] However, there are few studies that combine bioinformatics analysis with sequence features or functional partitioning for sequence optimization and functional validation, and apply this to the construction of stable cell lines for antibody drug production. Furthermore, the combined application of UCOE suffers from low transfection efficiency and slow cell recovery, and cannot utilize the blukpool rapid screening process, resulting in long cell line construction times and high workloads.

[0008] Therefore, while using UCOE to increase yield, there is a technical need to improve transfection efficiency and shorten the cell line construction cycle. Some site-directed and semi-site-directed integration technologies, such as the PiggyBac transposon expression system...

[0011] The UCOE expression system consists of a donor plasmid and a helper plasmid. The donor plasmid carries a 5' inverted repeat (5ITR) and a 3' inverted repeat (3ITR) that recognize transposase recognition sequences. The helper plasmid encodes the PiggyBac transposase, which recognizes the 5ITR and 3ITR. It can be cleaved or replicated at its original position, and then inserted into a specific location in the host genome with the assistance of circularization and transposase, achieving precise integration. This process offers significant advantages such as high integration efficiency, good stability, and uniform copy form. Integrating UCOE elements and ITR sequences can leverage both the high yield of UCOE and the high integration efficiency, good stability, and uniform copy form of transposable systems. Furthermore, the bulkpool rapid screening process can significantly shorten the cell line construction cycle, making it easier to apply in large-scale production. However, currently, there is no research on using expression systems integrating UCOE elements and ITR sequences for the construction of stable transgenic cell lines for antibody drug production, indicating a significant gap before industrial application.

[0009] Invention Overview

[0010] In a first aspect, this disclosure provides a chromatin open element (UCOE) for polypeptide or protein expression in eukaryotic cells, said chromatin open element comprising the nucleic acid sequence shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3 or SEQ ID NO:4, or comprising a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identity with any of the nucleic acid sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3 or SEQ ID NO:4.

[0011] Secondly, this disclosure also provides a combination of chromatin opening elements, comprising sequences of chromatin opening element 1 and chromatin opening element 2. The chromatin opening element 1 sequence comprises the nucleic acid sequence shown in SEQ ID NO:1 or SEQ ID NO:2, or comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with any of the nucleic acid sequences shown in SEQ ID NO:1 or SEQ ID NO:2; the chromatin opening element 2 comprises the nucleic acid sequence shown in SEQ ID NO:3 or SEQ ID NO:4, or comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity with any of the nucleic acid sequences shown in SEQ ID NO:1 or SEQ ID NO:2; The nucleic acid sequence shown in NO:4 is any nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identity.

[0012] In one embodiment, the chromatin opening element combination comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:1 and a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:3.

[0013] In some embodiments, the chromatin opening element combination is a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:1 and a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:4.

[0014] In some embodiments, the chromatin opening element combination is a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:2 and a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:3.

[0015] In some embodiments, the chromatin opening element combination is a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:2 and a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:4.

[0016] In a preferred embodiment, the combination of chromatin opening elements is any of the following combinations: a combination of nucleic acid sequences as shown in SEQ ID NO:1 and SEQ ID NO:3, a combination of nucleic acid sequences as shown in SEQ ID NO:1 and SEQ ID NO:4, a combination of nucleic acid sequences as shown in SEQ ID NO:2 and SEQ ID NO:3, or a combination of nucleic acid sequences as shown in SEQ ID NO:2 and SEQ ID NO:4.

[0017] Thirdly, this disclosure also provides a nucleic acid construct comprising the chromatin opening elements described in the first aspect or a combination of the chromatin opening elements described in the second aspect. In a preferred embodiment, the nucleic acid construct further comprises a 5' repeat terminal sequence (5ITR) recognized by PiggyBac transposase and an inverted 3' repeat terminal sequence (3ITR) and / or a selection marker gene. In some embodiments, the selection marker gene may be a GS gene.

[0018] Fourthly, this disclosure provides a transgenic vector comprising the chromatin opening elements described in the first aspect, or a combination of the chromatin opening elements described in the second aspect, or the nucleic acid constructs described in the third aspect.

[0019] In a preferred embodiment, the other elements of the transgenic vector are the original pNS03 vector sequence. Preferably, the transgenic vector further comprises a selection marker gene; more preferably, the selection marker gene is a glutamine synthase gene (GS gene) sequence or other gene sequence. Preferably, the nucleic acid construct or transgenic vector may also comprise a target gene encoding a nucleic acid sequence. In one embodiment, the target gene encoding a nucleic acid sequence is preferably an antibody or its antigen-binding fragment encoding a nucleic acid sequence, the antibody or its antigen-binding fragment comprising a heavy chain and / or a light chain sequence. Preferably, the chromatin opening element is located upstream of the promoter. In a preferred embodiment, chromatin opening element 1 or 2 is located upstream of the promoter of the light chain encoding nucleic acid sequence, and chromatin opening element 2 or 1 is located upstream of the promoter of the heavy chain encoding nucleic acid sequence. In a more preferred embodiment, chromatin opening element 1 is located upstream of the promoter of the light chain encoding nucleic acid sequence, and chromatin opening element 2 is located upstream of the promoter of the heavy chain encoding nucleic acid sequence.

[0020] Fifthly, this disclosure also provides an expression system for eukaryotic cells, comprising the chromatin opening elements described in the first aspect or a combination of the chromatin opening elements described in the second aspect, or comprising the nucleic acid constructs described in the third aspect or the transgenic vectors described in the fourth aspect, wherein the chromatin opening elements or the combination of the chromatin opening elements can improve protein yield.

[0021] In a preferred embodiment, the expression system for eukaryotic cells further comprises a 5' repeat end sequence (5ITR) recognized by the PiggyBac transposase and an inverted 3' repeat end sequence (3ITR).

[0022] In a preferred embodiment, the expression system applied to eukaryotic cells comprises the nucleic acid construct or transgenic vector.

[0023] In a preferred embodiment, the eukaryotic host cell is a CHO cell.

[0024] In a preferred embodiment, the expression system for eukaryotic cells further comprises a PiggyBac helper vector or mRNA encoding a transposase, wherein the PiggyBac helper vector contains a nucleic acid sequence encoding a transposase.

[0025] In a sixth aspect, this disclosure also provides the chromatin opening elements described in the first aspect, combinations of chromatin opening elements described in the second aspect, nucleic acid constructs comprising the chromatin opening elements or combinations of the chromatin opening elements, transgenic vectors or combinations of the chromatin opening elements, and applications of the nucleic acid constructs or transgenic vectors in eukaryotic cell expression systems.

[0026] In a preferred embodiment, the application is for gene recombination expression.

[0027] In a preferred embodiment, the gene encodes a polypeptide or protein.

[0028] In a preferred embodiment, the recombinant expression is performed in eukaryotic host cells. Preferably, the eukaryotic host cells are CHO cells.

[0029] In a seventh aspect, this disclosure also provides the application of the chromatin open element combination described in the second aspect in the construction of transgenic vectors.

[0030] Eighthly, this disclosure also provides the use of the chromatin opening elements or combinations thereof described in the first aspect in the construction of recombinant cell lines. In a specific embodiment, the steps may include:

[0031] 1) Construct transgenic vectors containing chromatin open element sequences or combinations thereof and target protein coding sequences;

[0032] 2) Construct transgenic vectors containing chromatin opening elements with added ITRs;

[0033] 3) Introduce the transgenic vector from step 1) or 2) into host cells via transfection;

[0034] 4) In step 3), the cells are placed in a culture flask to recover for a period of time;

[0035] 5) The cells from step 4) are pooled for recovery;

[0036] 6) During the recovery process of cells in step 5), the yield is detected. After the cells have fully recovered, the yield test is performed to evaluate the expression level of the target protein.

[0037] In some implementations, the target protein in step 6) can be an antibody, fusion protein, antigen, enzyme, or other types of protein or polypeptide.

[0038] In a preferred embodiment, the host cell is a CHO cell.

[0039] In a preferred embodiment, an endotoxin-free plasmid extraction kit is used in step 1) or 2).

[0040] In a preferred embodiment, the transfection method in step 3) is electrotransfection.

[0041] In the preferred embodiment, the recovery time for step 4) is 24 hours.

[0042] In a preferred embodiment, in step 5), pool is a minipool or bulkpool.

[0043] In the preferred embodiment, the yield test in step 6) is a 7-day batch, with sugar supplementation of 5 g / L during the culture process, and the expression level in the supernatant is detected during the culture process or at harvest on the 7th day.

[0044] Compared with the prior art, the advantages of the present invention are at least one of the following: the sequence truncation of UCOE elements from two different sources significantly shortens the combined UCOE sequence length while maintaining comparable yield, and the truncated UCOE sequences can be combined for application; the truncated UCOE is integrated with ITR, enabling cell line construction using a bulkpool rapid screening process, advancing protein production time by 4-5 weeks, significantly shortening cell line construction time, and maintaining comparable yield. Attached Figure Description

[0045] Figure 1 shows the pNS03 vector.

[0046] Figure 2 shows the pSU vector.

[0047] Figure 3 shows the pSU-UCOE vector.

[0048] Figure 4 shows the protein yield over 7 days as assessed by a single UCOE sequence batch.

[0049] Figure 5 shows the protein yield of a 96-well plate containing UCOE truncated sequence combinations over 12 days.

[0050] Figure 6 shows the protein yield of a 96-well plate containing UCOE truncated sequence combinations over 21 days.

[0051] Figure 7 shows the protein yield over 7 days as assessed by UCOE truncated sequence combination batch.

[0052] Figure 8 shows the pcDNA3.4-PiggyBac vector.

[0053] Figures 9A and 9B show the recovery growth curves of pNS03-5ITR-3ITR-IgG cells.

[0054] Figure 10 shows the 7-day innate protein yield of pNS03-5ITR-3ITR-IgG batch.

[0055] Figures 11A and 11B show the cell recovery growth curves of different UCOE truncated sequences and ITR integrated cells.

[0056] Figure 12 shows the protein yield over 7 days for different UCOE truncated sequences and ITR integration batches. Detailed Implementation

[0057] the term

[0058] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference as if specifically and individually indicated that each individual publication, patent or patent application is incorporated herein by reference.

[0059] Before describing the invention in detail below, it should be understood that the invention is not limited to the specific methodologies, schemes, and reagents described herein, as these can vary. It should also be understood that the terminology used herein is for describing particular embodiments only and is not intended to limit the scope of the invention. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0060] Some embodiments disclosed herein include numerical ranges, and certain aspects of the invention may be described using ranges. Unless otherwise stated, it should be understood that numerical ranges or descriptions using ranges are for purposes of brevity and convenience only and should not be considered as a strict limitation of the scope of the invention. Therefore, descriptions using ranges should be considered as specifically disclosing all possible subranges and all possible specific numerical points within those ranges, as these subranges and numerical points have been explicitly stated herein. The above principles apply equally regardless of the breadth of the numerical values. When a range description is used, the range includes the endpoints of the range.

[0061] When referring to measurable values ​​such as quantities, temporary durations, etc., the term “about” means a variation of ±20%, or in some cases ±10%, or in some cases ±5%, or in some cases ±1%, or in some cases ±0.1% of the specified value.

[0062] As used in this article, the term "antibody" typically refers to a Y-type tetrameric protein comprising two heavy (H) polypeptide chains and two light (L) polypeptide chains held together by covalent disulfide bonds and non-covalent interactions. Natural IgG antibodies possess this structure. Each light chain consists of a variable domain (VL) and a constant domain (CL). Each heavy chain contains a variable domain (VH) and a constant domain.

[0063] As used herein, the broad category of "antibody" may include polyclonal antibodies, monoclonal antibodies, chimeric antibodies, humanized antibodies and primate-derived antibodies, CDR-grafted antibodies, human antibodies (including recombinant human antibodies), recombinant antibodies, intracellular antibodies, multispecific antibodies, bifunctional fusion proteins, monovalent antibodies, multivalent antibodies, anti-individual genotype antibodies, synthetic antibodies (including mutant proteins and their variants), etc.

[0064] The term "CHO ​​platform" refers to the CHO cell line screening technology platform. After CHO (Chinese hamster ovary cells) are domesticated into the chemically defined medium CD CHO Fusion Medium, and subcloning screening is performed to establish CHO cell lines, it also includes the expression vector P3, the medium used in the clone construction stage, and the supporting reagents and processes such as the fed-batch process platform medium.

[0065] The term "nucleic acid construct" refers to a single-stranded or double-stranded nucleic acid molecule isolated from a naturally occurring gene, or modified to contain a segment of nucleic acid not found in nature, or is synthetic. "Nucleic acid construct" specifically refers to an artificial structure composed of nucleic acids that can be transcribed in target cells, preferably inserted into a vector, particularly a plasmid vector or viral vector.

[0066] The term "vector" refers to a nucleic acid molecule capable of carrying another nucleic acid linked to it. One type of vector is the "plasmid," which is a circular double-stranded DNA loop through which an additional DNA segment can be attached. Another type of vector is the viral vector, in which the additional DNA segment can be attached to the viral genome. Some vectors are capable of autonomous replication in the host cell to which they are introduced (e.g., bacterial vectors with bacterial origins of replication and attachable mammalian vectors). Other vectors (e.g., non-attached mammalian vectors) can integrate into the host cell's genome after introduction and thus replicate along with the host genome. Furthermore, some vectors are capable of directing the expression of genes operatively linked to them.

[0067] The terms "nucleic acid construct" or "vector" and similar terms should be broadly considered to include any nucleic acid (including DNA and RNA) suitable for use as a vector to transfer genetic material into cells. The term should be considered to include plasmids, viruses (including bacteriophages), granules, and artificial chromosomes. Nucleic acid constructs or vectors may include one or more regulatory elements, origins of replication, multiple cloning sites, and / or selection markers. In one embodiment, the nucleic acid construct or vector is modified to be suitable for the expression of one or more genes encoded by the construct. Nucleic acid constructs or vectors include naked nucleic acids, as well as nucleic acids formulated with one or more reagents that facilitate delivery to cells (e.g., liposome-conjugated nucleic acids, organisms containing said nucleic acids).

[0068] The term "transgenic vector" refers to an expression vector that expresses a target gene. An expression vector is a vector that adds expression elements (such as promoters, RBS, terminators, etc.) to the basic backbone of a cloning vector, enabling the expression of the target gene. Target genes include, but are not limited to, deoxynucleotide sequences encoding products such as antibodies, antigens, fusion proteins, and polypeptides in a broad sense.

[0069] The term "unopen chromatin element (UCOE)" or "unopen chromatin element (UCOE) sequence" refers to an element or sequence consisting of an extended unmethylated CpG island and a bidirectional (or unidirectional) promoter that can prevent gene transcriptional silencing, so that the transcribed target gene is not affected by the "position effect", thereby enabling the target protein to be highly expressed in the cell.

[0070] The term "PiggyBac transposon expression system" refers to the PiggyBac transposon vector system. The PiggyBac vector system mainly consists of: a helper vector or plasmid encoding a transposase; and a transposon vector (also called a donor vector) or plasmid containing optimized subterminal inverted repeat sequences at both ends, with the transposable region in the middle, into which the target gene sequence to be transposoned into the host genome can be inserted. In experiments, both the helper plasmid and the transposon plasmid are simultaneously transformed into target cells. The transposase encoded by the helper plasmid recognizes and cleaves the ITR sequences at both ends of the transposon plasmid. The released transposon region is integrated by the transposase into a site in the host genome containing the TTAA sequence, resulting in TTAA repeat sequences appearing at both ends of the transposon region.

[0071] The term "multiple cloning site" refers to a synthetically produced DNA fragment on a vector containing multiple single restriction enzyme sites. These sites serve as insertion sites for foreign DNA and are also known as multi-site adapters. They are standard configuration sequences for commonly used vector plasmids in genetic engineering. Each restriction enzyme site in a multiple cloning site is typically unique, meaning it appears only once in a specific vector plasmid. However, there may be overlap between restriction enzyme sites.

[0072] The term "transcription factors" (TFs) refers to a group of protein molecules that can specifically bind to a specific sequence upstream of the 5' end of a gene, thereby ensuring that the target gene is expressed at a specific intensity at a specific time and space.

[0073] The term "transcription factor binding site" (TFBS) refers to the region where a transcription factor binds to the gene template strand when regulating gene expression.

[0074] The term "UCSC database" refers to a database developed and maintained by a cross-departmental team at the University of California, Santa Cruz (UCSC) Genomics Institute, a multi-institutional research team based in the UK. UCSC is one of the most widely used databases in the biological field, capable of quickly and reliably displaying any desired portion of a genome of any size, along with dozens of aligned annotation tracks (known genes, predicted genes, ESTs, mRNA, CpG islands, chromatin bands, material homology, etc.). It can link to various databases such as the ENDODE Project database, JASPAR TFBS, and TFBS Conserved, facilitating user analysis.

[0075] The term "protein-protein interaction network analysis" (PPI) refers to the use of protein interaction databases, such as the STRING protein interaction database, to directly extract the interaction relationships of target gene sets (such as differentially expressed gene lists) from the database for species contained in the database and construct a network.

[0076] The terms "sequence identity," "sequence similarity," or "sequence homology" refer to the percentage of amino acid residues in a candidate sequence that are identical to those in a reference polypeptide sequence after aligning the sequences (and, where necessary, introducing gaps) to obtain the maximum percentage sequence identity, without considering any conserved substitutions as part of the sequence identity. Sequence alignment can be performed using various methods in the art to determine the percentage amino acid sequence identity, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, or MEGALIGN (DNASTAR) software. Those skilled in the art can determine suitable parameters for measuring the alignment, including any algorithm required to obtain the maximum alignment of the full length of the sequences being compared.

[0077] The term "recombinant expression" refers to the process of contacting a genetically modified oligonucleotide or polynucleotide construct with a host cell under conditions sufficient to enable the host cell to express mRNA, protein, polypeptide, or peptide, wherein the construct contains a nucleotide sequence encoding mRNA, protein, polypeptide, or peptide.

[0078] Example

[0079] The present invention will be further described in detail through the following embodiments.

[0080] The following specific embodiments are provided to illustrate the present invention. However, it should be understood that these embodiments are provided only for illustrative purposes and not for limiting the scope of the present invention.

[0081] Materials and reagents:

[0082] pNS03 (Figure 1) is a plasmid constructed in our laboratory.

[0083] pSU (Figure 2) is the plasmid constructed in our laboratory.

[0084] The sequences UCOE-L-2131, UCOE-L-2402, UCOE-S-998, and UCOE-S-1311 were synthesized by Nanjing Genscript Biotech Co., Ltd.

[0085] Example 1: Design of UCOE truncated sequences

[0086] UCOE sequences were analyzed using the BLAST function of the UCSC database, AnimalTFDB 4.0 (http: / / bioinfo.life.hust.edu.cn / AnimalTFDB4 / # / TFBS_Predict), and CpG Island Prediction (http: / / www.urogene.org / methprimer / ). Two UCOE-S sequences and two UCOE-L sequences were obtained: UCOE-S-998 (SEQ ID NO:3) and UCOE-S-1311 (SEQ ID NO:4), and UCOE-L-2131 (SEQ ID NO:1) and UCOE-L-2402 (SEQ ID NO:2).

[0087] The nucleotide sequences of SEQ ID NO:1-4 are shown below:

[0088] SEQ ID NO:1

[0089] SEQ ID NO: 2

[0090] SEQ ID NO: 3

[0091] SEQ ID NO:4

[0092] Example 2: Effect of UCOE truncated sequence on antibody expression

[0093] 1. Construction of vectors containing single UCOE sequences

[0094] By using MluI / BbsI double digestion and homologous recombination, different truncated sequences shown in SEQ ID NO:1-4 were ligated into the pSU vector to construct the vector pSU-UCOE containing a single UCOE sequence (Figure 3 exemplarily shows the pSU vector containing the full-length UCOE-L). Then, by using BsiwI and BstBI double digestion and homologous recombination, the nucleic acid sequences encoding the light chain and heavy chain of IgG antibody 904A were each ligated into pSU-UCOE to construct the recombinant expression vector pSU-UCOE-904A (a combination of a plasmid containing the light chain coding sequence and a plasmid containing the heavy chain coding sequence).

[0095] 2. List of single UCOE vectors and their abbreviations

[0096] Table 1 List of Single UCOE Vectors and Abbreviations

[0097] 3. Expression of single UCOE antibody protein

[0098] Two plasmids containing the light and heavy chain encoding genes of the 904A antibody were co-transfected into CHO cells using electroporation. The total transfection dose of the two plasmids was 50 μg. Transfection conditions: 300 V, 950 μF, cell mass 1E7, single electroporation. Cells were then cultured in culture medium. The CD CHO Fusion Medium (containing 6 mM L-glutamine) was resuspended and incubated in an incubator at 37°C and 5% CO2.

[0099] After culturing for 24 hours, centrifuge the cells to remove the supernatant, and then take an appropriate amount. Cells were resuspended in CD CHO Fusion medium and transferred to shake flasks for culture. Cells were counted and passaged every 1-5 days until pool cell viability recovered. Yield tests were then performed by seeding. The batch evaluation results are shown in Figure 4. The results showed that the protein expression levels of different truncated UCOE-L sequences were similar to those of the complete sequence. The protein expression of the truncated UCOE-S sequence SEQ ID NO:4 was similar to that of the complete sequence, while the protein expression of SEQ ID NO:3 was slightly lower.

[0100] Example 3: Effect of UCOE truncated sequence combinations on antibody expression

[0101] 1. Construction of UCOE truncated sequence combinations

[0102] An expression vector containing the original pNS03 vector sequence and two truncated UCOE sequences was constructed. The specific combinations included: SEQ ID NO:1 and SEQ ID NO:3, SEQ ID NO:1 and SEQ ID NO:4, SEQ ID NO:2 and SEQ ID NO:3, and SEQ ID NO:2 and SEQ ID NO:4. First, the pNS03 plasmid was double-digested with BbsI and BamHI. Then, using homologous recombination, the fragments shown in SEQ ID NO:1 or SEQ ID NO:2 (UCOE-L-2131 or UCOE-L-2402) were ligated into the pNS03 plasmid to replace UCOE-L, constructing plasmids pNS03-SEQ ID NO:1 and pNS03-SEQ ID NO:2 (pNS03-U2131 and pNS03-U2402). Next, using enzyme digestion-homological recombination, the fragments shown in SEQ ID NO:3 or SEQ ID NO:4 (UCOE-S-998 or UCOE-S-1311) were ligated into pNS03-SEQ ID NO:1 and pNS03-SEQ ID NO:2 (pNS03-U2131 and pNS03-U2402), replacing UCOE-S, to construct plasmid pNS03-SEQ ID NO:1. NO:1-NO:3, pNS03-SEQ ID NO:2-NO:3, pNS03-SEQ ID NO:1-NO:4 and pNS03-SEQ ID NO:2-NO:4 (pNS03-U2131-U998, pNS03-U2402-U998, pNS03-U2131-U1311 and pNS03-U2402-U1311), are collectively referred to as pNS03-UL-US for convenience. Vectors pNS03 and pNS03-UL-US were ligated by double digestion with NgoMIV / BamHI and BsiwI / BstBI, respectively. The light chain and heavy chain coding nucleic acid sequences of antibody 107K (anti-PD-L1 humanized IgG4 monoclonal antibody, which has a symmetrical structure of two light chains and two heavy chains, with the heavy chain and light chain coding nucleic acid sizes of approximately 1400bp and 700bp, respectively) were ligated to the corresponding multiple cloning sites to construct recombinant expression vectors pNS03-IgG and pNS03-UL-US-IgG. Among them, the truncated sequence shown in SEQ ID NO:1 or SEQ ID NO:2 is located upstream of the promoter of the light chain coding nucleic acid sequence, and the truncated sequence shown in SEQ ID NO:3 or SEQ ID NO:4 is located upstream of the promoter of the heavy chain coding nucleic acid sequence.

[0103] 2. List of UCOE truncated sequence combination vectors and their abbreviations

[0104] Table 2. List of Carriers and Abbreviations

[0105] 3. Stable transfection of antibody expression using UCOE truncated sequence combinations

[0106] CHO cells were transfected by electroporation, and the transfection conditions are shown in Table 3. After transfection, cells were cultured in medium... Cells were resuspended in CD CHO Fusion Medium (containing 6 mM L-glutamine) and cultured in an incubator. Cells were counted after approximately 24 hours of culture, then seeded in minipools and progressively screened and expanded in 96-well, 24-well, and TPP plates. Yield was measured at the early and late stages of 96-well plate selection. Once cells had fully recovered, batches were seeded for yield testing. Results showed that different UCOE truncated sequence combinations exhibited little difference in expression during the early recovery phase (Figure 5); some differences were observed in the later phase (Figure 6); and significant differences were observed in the 7-day batch yield test (Figure 7). The NO.1+NO.4 and NO.2+NO.4 sequence combinations both achieved good expression capabilities, with yields slightly higher or comparable to the control complete UCOE sequence. This may be because the truncated sequences not only retained the optimized expression capabilities of the control complete UCOE sequence but also improved expression efficiency due to their shorter length and smaller size.

[0107] Table 3 Stable Transfection Conditions

[0108] Example 4: Construction of the UCOE and ITR integrated vector

[0109] 1. Construction of vector pNS03-5ITR-3ITR-IgG

[0110] The sequence containing 5ITR was ligated into pNS03-IgG by double digestion with AscI / PacI to construct pNS03-5ITR-IgG. Then, the sequence containing 3ITR was ligated into pNS03-5ITR-IgG by single digestion with NotI to construct the recombinant expression vector pNS03-5ITR-3ITR-IgG.

[0111] 2. Construction of auxiliary carriers

[0112] The vector pUC57-PiggyBac and the blank vector pcDNA3.4 were double-digested with EcoRI / HindIII. The digestion products of the vector pUC57-PiggyBac and the blank vector pcDNA3.4 were then purified and recovered using the NucleoSpin Gel and PCR Clean-up kit and ligated to construct the helper plasmid pcDNA3.4-PiggyBac containing the transposase encoding gene, as shown in Figure 8.

[0113] 3. Stable transfection of antibody expression using UCOE and ITR integration vectors

[0114] CHO cells were transfected by electroporation, and the transfection conditions are shown in Table 4. After transfection, cells were cultured in... Resuspended in CD CHO Fusion Medium (containing 6 mM L-glutamine) and incubated in an incubator.

[0115] Cells were counted after approximately 24 hours of culture, then centrifuged and the supernatant was discarded. Cells were resuspended in an appropriate amount of EX-CELL CD CHO Fusion medium and transferred to shake flasks for further culture. Cells were counted and passaged every 1-5 days. The cell recovery growth curves are shown in Figures 9A and 9B. The results showed that cells recovered rapidly after ITR integration, enabling a rapid bulkpool screening process, thus shortening protein production time by 4-5 weeks. The final cell line construction cycle was significantly shorter than the minipool screening process. Once bulkpool cell viability recovered, batches were seeded for yield testing. The batch results on day 7 of stable expression are shown in Figure 10.

[0116] While the control group cells were selected and compared for growth using bulkpool, they were seeded in minipool after 24 hours of transfection recovery. The cells were then screened and expanded in 96-well, 24-well, and TPP plates. Once the cells were fully recovered (viability > 90%), they were seeded in batches for yield testing. The yield results of the top 5 minipool cells are shown in Figure 10. The results show that the yield of UCOE cells after ITR integration was comparable to that of the control group.

[0117] Table 4 Stable Transfection Conditions

[0118] Example 5: Effects of different combinations of UCOE truncated sequences and ITR integration on antibody expression

[0119] 1. Construction of different UCOE truncated sequence combinations and ITR integration vectors

[0120] Using BstBI / PacI double digestion and ligation, sequences containing 5ITR and 3ITR were ligated onto pNS03-SEQ ID NO:1-NO.4-IgG and pNS03-SEQ ID NO:2-NO.4-IgG to construct pNS03-SEQ ID NO:1-NO.4-ITR-IgG and pNS03-SEQ ID NO:2-NO.4-ITR-IgG.

[0121] 2. Stable transfection of antibodies with different UCOE truncated sequences and ITR integration vectors for antibody expression

[0122] CHO cells were transfected by electroporation, and the transfection conditions are shown in Table 5. After transfection, cells were cultured in medium... Resuspended in CD CHO Fusion Medium (containing 6 mM L-glutamine) and cultured in an incubator. Cells were counted and passaged every 1-5 days. Cell recovery growth curves are shown in Figures 11A and 11B. Once pooled cell viability recovered, batches were seeded for yield testing, and the results are shown in Figure 12. The results showed that different truncated UCOEs integrated with ITR could achieve rapid cell recovery, and the yield was significantly increased compared to without ITR. The NO.2-NO.4 sequence combination integrated with ITR yielded the highest yield.

[0123] Table 5 Stable Transfection Conditions

[0124] Note: pNS03-ITR-IgG is the same as pNS03-5ITR-3ITR-IgG.

[0125] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0126] References

[0127] [1]Rebecca E.Sizer and Robert J.White.Use of ubiquitous chromatin opening elements(UCOE) as tools to maintain transgene expression in biotechnology.Computational and Structural Biotechnology Journal 21(2023)275–283.

[0128] [2]Jonathan J.Neville,Joe Orlando,Kimberly Mann,Bethany McCloskey,Michael N.Antoniou.Ubiquitous Chromatin-opening Elements(UCOEs):Applications in biomanufacturing and gene therapy.Biotechnology Advances 35(2017)557–564.

[0129] [3]Simpson DJ,Williams SG,Irvine AS.Expression elements[Internet].US Patent.7632661,2009

[0130] [4]Suba Dharshanan,Heilly Chong,Swee Hung Cheah.Stable expression of H1C2 monoclonal antibody in NS0 and CHO cells using pFUSE and UCOE expression system.Cytotechnology.2014 Aug;66(4):625–633.

[0131] [5]Fay Saunders,Berni Sweeney,Michael N Antoniou.Chromatin function modifying elements in an industrial antibody production platform--comparison of UCOE,MAR,STAR and cHS4 elements.PLoS One.2015Apr 7;10(4):e0120096.

[0132] [6]Zeynep Betts,Alan J.Dickson.Assessment of UCOE on Recombinant EPO Production and Expression Stability in Amplified Chinese Hamster Ovary Cells.Mol Biotechnol.2015 Sep;57(9):846-58.

[0133] [7]Fang Zhang,Giorgia Santilli&Adrian J.Thrasher.Characterization of a core region in the A2UCOE that confers effective anti-silencing activity.Scientific Reports,7(1):10213.

[0134] [8]Uta Muller-Kuller,Mania Ackermann,Stephan Kolodziej,et al.A minimal ubiquitous chromatin opening element(UCOE)effectively prevents silencing of juxtaposed heterologous promoters by epigenetic remodeling in multipotent and pluripotent stem cells.Nucleic Acids Research,2015,Vol. 43, No. 3, 1577–1592.

[0135] [9] Steven Williams1, Tracey Mustoe1, Tony Mulcahy, et al. CpG-island fragments from the HNRPA2B1 / CBX3 genomic locus reduce silencing and enhance transgene expression from the hCMV promoter / enhancer in mammalian cells. BMC Biotechnology 2005, 5: 17.

[0136]

[0010] Kristian Alsbjerg Skipper, Anne Kruse Hollensen, Michael N., et al. Sustained transgene expression from sleeping beauty DNA transposons containing a core fragment of the HNRPA2B1-CBX3 ubiquitous chromatin opening element (UCOE) . BMC Biotechnology (2019) 19: 75.

[0137]

[0011] KOSUKE YUSA.PiggyBac transposon.Microbiol Spectrum 3(2):MDNA3-0028-2014.

Claims

1. A chromatin opening element, wherein the nucleic acid of the chromatin opening element comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequences shown in SEQ ID NO:1, SEQ ID NO:2, SEQ ID NO:3, or SEQ ID NO:

4.

2. A combination of chromatin opening elements, wherein the combination comprises chromatin opening element 1 and chromatin opening element 2, wherein chromatin opening element 1 comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:1 or SEQ ID NO:2, and wherein chromatin opening element 2 comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 87%, 88%, 89%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:3 or SEQ ID NO:2, and wherein chromatin opening element 2 comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 87%, 88%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:1 or SEQ ID NO:2, and wherein chromatin opening element 2 comprises a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:2 or SEQ ID NO:3, and wherein chromatin opening element 2 is identical to the nucleic acid sequence shown in SEQ ID NO:3 or SEQ ID NO:

2. The nucleic acid sequence shown in NO:4 has at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity; or, wherein The chromatin opening element 1 contains the nucleic acid sequence shown in SEQ ID NO:1 and the chromatin opening element 2 contains the nucleic acid sequence shown in SEQ ID NO:

3. The chromatin opening element 1 contains the nucleic acid sequence shown in SEQ ID NO:1 and the chromatin opening element 2 contains the nucleic acid sequence shown in SEQ ID NO:

4. The chromatin opening element 1 contains the nucleic acid sequence shown in SEQ ID NO:2 and the chromatin opening element 2 contains the nucleic acid sequence shown in SEQ ID NO:3, or The chromatin opening element 1 contains the nucleic acid sequence shown in SEQ ID NO:2 and the chromatin opening element 2 contains the nucleic acid sequence shown in SEQ ID NO:

4.

3. The combination of chromatin opening elements as described in claim 2, wherein the nucleic acid sequence of chromatin opening element 1 is a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:1 or SEQ ID NO:2, and the nucleic acid sequence of chromatin opening element 2 is a nucleic acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with the nucleic acid sequence shown in SEQ ID NO:3 or SEQ ID NO:4; or The combination of chromatin opening elements is any of the following combinations: a combination of nucleic acid sequences as shown in SEQ ID NO:1 and SEQ ID NO:3, a combination of nucleic acid sequences as shown in SEQ ID NO:1 and SEQ ID NO:4, a combination of nucleic acid sequences as shown in SEQ ID NO:2 and SEQ ID NO:3, or a combination of nucleic acid sequences as shown in SEQ ID NO:2 and SEQ ID NO:

4.

4. A nucleic acid construct comprising the chromatin opening element of claim 1, or a combination of the chromatin opening elements of claim 2 or 3.

5. The nucleic acid construct of claim 4, further comprising a 5' repeat end sequence (5ITR) and an inverted 3' repeat end sequence (3ITR) recognized by PiggyBac transposase and / or a selection marker gene, preferably, the selection marker gene is a GS gene.

6. A transgenic vector comprising the chromatin opening element of claim 1, a combination of the chromatin opening element sequences of claim 2 or 3, or the nucleic acid construct of claim 4 or 5.

7. A eukaryotic cell expression system comprising the chromatin opening element of claim 1, a combination of the chromatin opening elements of claim 2 or 3, the nucleic acid construct of claim 4 or 5, or the transgenic vector of claim 6.

8. The application of the chromatin opening element of claim 1, a combination of chromatin opening elements of claim 2 or 3, the nucleic acid construct of claim 4 or 5, the transgenic vector of claim 6, or the eukaryotic cell expression system of claim 7 in the recombinant expression of nucleic acids.

9. The application as described in claim 8, wherein the nucleic acid encodes an antibody, fusion protein, antigen, enzyme, or other type of protein or polypeptide.

10. The use of the chromatin opening element of claim 1, a combination of chromatin opening elements of claim 2 or 3, the nucleic acid construct of claims 4-5, the transgenic vector of claim 6, or the eukaryotic cell expression system of claim 7 in the preparation of proteins or peptides.

11. The application of claim 10, wherein the protein or polypeptide comprises an antibody, a fusion protein, an antigen, and an enzyme.

12. The use of the chromatin opening element of claim 1, a combination of the chromatin opening elements of claim 2 or 3, the nucleic acid construct of claim 4-5, the transgenic vector of claim 6, or the eukaryotic cell expression system of claim 7 in the construction of recombinant cell lines.

13. The application as described in claim 12, wherein the recombinant cell is a mammalian cell, preferably a CHO cell.