Spacers for gene expression constructs

The use of a modified CTCF binding sequence and spacer element in expression vectors addresses low expression and integration inefficiencies, improving the production of multiple polypeptides by enhancing integration and stability, particularly for monoclonal antibodies.

JP2026512866APending Publication Date: 2026-04-21CYTIVA SWEDEN AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CYTIVA SWEDEN AB
Filing Date
2024-03-27
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing expression vectors for protein production in mammalian cells face challenges such as low expression levels and inefficient integration, particularly when producing multiple polypeptides like the heavy and light chains of monoclonal antibodies, necessitating extensive clonal screening.

Method used

Incorporation of a modified CCCTC-binding factor (CTCF) binding sequence and a spacer element, such as a modified cHS4 core element, between expression cassettes in an expression vector to enhance polypeptide expression, allowing for stable integration and segregation of gene expression.

Benefits of technology

The modified CTCF binding sequence and spacer element improve the expression of multiple proteins, particularly in producing therapeutic protein complexes like antibodies, by enhancing integration efficiency and stability, leading to higher expression levels and cost-effective cell line development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026512866000001_ABST
    Figure 2026512866000001_ABST
Patent Text Reader

Abstract

The present invention provides a nucleic acid sequence comprising a modified CTCF binding sequence represented by the following sequence: TGCAGTACCTCCCTN1N2N3CCAGCAGGN4GGCAN5N6AGN7GAAN8GGTGAACTGGAGT (SEQ ID NO: 10), where N1, N2, N3, and N5 are each independently C or G, N4 is T or G, N6 and N7 are each independently C or T, and N8 is T or A. Each of N1 to N8 represents a nucleic acid substitution at the corresponding locus on the CTCF binding site (SEQ ID NO: 1) of the wild-type chicken high-sensitivity region 4 (cHS4) insulator. The invention also provides an expression vector and recombinant cells incorporating the modified CTCF binding sequence, as well as related methods for producing at least two polypeptides.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an expression vector construct useful for protein production in mammalian host cells. More specifically, the present invention relates to a nucleic acid sequence comprising a spacer element that can be placed between two or more expression cassettes within an expression vector and an expression vector construct. The present invention also relates to cells and cell lines containing such constructs, and methods for recombinantly producing a protein of interest.

Background Art

[0002] Plasmids are small circular extrachromosomal DNAs ranging in size from 1 to over 200 kbp and have been found to occur naturally in some bacteria. A single bacterial cell may contain several different plasmids, and each of these plasmids may be present in hundreds of copies within the cell. The most useful property of plasmids is the ability to replicate autonomously and be stably maintained within a bacterial cell line. Due to this property of plasmids, they are extremely suitable as tools for introducing foreign genetic material into bacterial host cells. Plasmids usually contain several genetic elements such as a replication origin, a replication initiation gene, and an antibiotic resistance gene. Small plasmids generally depend on the replication machinery of the bacterial host cell, while large plasmids may possess their own replication genes.

[0003] Genetically engineered artificial plasmid vectors are one of the most commonly used vectors for introducing a gene of interest into a host cell. Such engineered plasmids are widely used as cloning vectors and are designed to have genetic elements that allow the insertion of the gene of interest into the plasmid. When a bacterial host cell is transformed by such a plasmid vector, replication of the plasmid vector begins within the cell, resulting in an increase in the copy number of the plasmid, and thereby an increase in the number of copies of the inserted gene of interest. Transient expression of the gene from the plasmid can be carried out in a bacterial host cell, but this usually produces a small amount of protein.

[0004] A better approach to expressing a target gene is to use an expression vector. An expression vector may be a circular or linear DNA fragment engineered to contain the sequence of the target gene in the form of an expression cassette. Typically, an expression vector contains several other design features or components engineered within the expression vector sequence, such as regulatory genes encoding promoters and enhancers for effectively transcribing the target gene, and subsequently produces stable messenger RNA (mRNA) that can be translated into protein in mammalian host cells. However, replication of such expression vectors in mammalian cells is often observed to be insufficient.

[0005] One way to overcome the above limitations and improve the expression of the target gene is to integrate the expression vector construct into the genome of the mammalian host cell when the host cell is transfected with the expression vector. This can be done by including a restriction endonuclease binding site in the design of the expression vector to facilitate its integration into the genome of the mammalian host cell.

[0006] Several expression vectors have been developed over time with the aim of further improving protein expression by incorporating different promoter elements as well as other elements such as insulators and spacers into the vector construct. Insulators are a class of DNA sequence elements that share the common ability to protect genes from undesirable signals emanating from the surrounding environment. There are two ways insulators protect expressed genes from their surroundings. The first is by blocking the action of distal enhancers on the promoter. However, enhancer blocking only occurs when the insulator is located between the enhancer and the promoter. Such activity prevents the enhancer from activating the expression of the blocked adjacent gene, while allowing it to freely stimulate the expression of the gene located on the unblocked side. The second way insulators protect genes is by acting as a "barrier" that prevents nearby condensed chromatin domains from expanding into the transcriptionally active region, which can suppress gene expression. Some insulators can act as both enhancer blockers and barriers. For example, in the HS4 (highly sensitive region 4, HS4) insulator (hereinafter referred to as cHS4) found at the 5' end of the chicken β-globin locus, the two activities can occur together but are separable. The cHS4 insulator marks the boundary between the active euchromatin and the highly condensed, inactive upstream heterochromatin region at the chicken β-globin locus.

[0007] The incorporation of cHS4 insulators, WPREs, and SAR elements into vector constructs to improve protein expression in eukaryotic cells has been demonstrated (Ali Ramezani et al., "Performance and safety-enhanced lentiviral vectors containing the human interferon-β scaffold attachment region and the chicken β-globin insulator", Blood 2003 (101:4717-4724).

[0008] The enhancer-blocking and barrier activities of the cHS4 insulator within the 1.2kb full-length cHS4 insulator sequence were mapped to a 250bp "core" element, as shown in Figure 1A. The core element of the cHS4 insulator contains five protein binding sites / footprints (FI~FV) for three different insulator proteins: CTCF(FII), VEZF1(FI, FIII, and FV), and USF1 / USF2(FIV). The CTCF binding site or footprint II (FII) is necessary and sufficient for enhancer-blocking activity but can be deleted from the cHS4 insulator sequence without affecting barrier activity. The remaining four protein binding sites are all essential for barrier activity (FI, FIII, FIV, and FV) but unnecessary for enhancer-blocking activity. For segregation between expression cassettes and recruitment of transcription factors, it is known to use two consecutive copies of the core element, called a "double core," as shown in Figure 1B. Qin et al. demonstrated the use of a strong promoter EF1 and spacers between expression cassettes to isolate gene expression (Qin JY, Zhang L, Clift KL, Hulur I, Xiang AP et al. (2010) Systematic Comparison of Constitutive Promoters and the Doxycycline-Inducible Promoter. PLoS ONE 5(5):e10611).

[0009] An example of a commercially available plasmid expression vector is the pcDNATM 3.1(+) vector from ThermoFisher Scientific, shown in Figure 2, which has been used for decades for protein expression in mammalian cells. The pcDNATM 3.1(+) vector is a simple monocistronic expression vector containing an E. coli (E. coli) growth cassette, a selection marker cassette, and an expression cassette for the protein of interest. While this vector has been used to establish high-titer clones, good results rely on extensive clonal screening, and the vector design is not practical for producing multiple polypeptides as sometimes required, for example, when the protein is an mAb that requires two polypeptide chains to be expressed, namely the heavy chain (HC) and light chain (LC) of a monoclonal antibody (mAb).

[0010] Other known expression vectors suffer from low expression levels due to poor integration efficiency. Scientists in this field are trying to find ways to improve vector constructs so that a cost-effective and effective cell line development process can be established that produces stable clones with high expression levels.

[0011] Therefore, improved expression vector constructs are needed. [Prior art documents] [Non-patent literature]

[0012] [Non-Patent Document 1] Ali Ramezani et al., “Performance and safety-enhanced lentiviral vectors containing the human interferon-β scaffold attachment region and the chicken β-globin insulator,” Blood 2003 (101:4717-4724 [Non-Patent Document 2] Qin JY, Zhang L, Clift KL, Hulur I, Xiang AP et al. (2010) Systematic Comparison of Constitutive Promoters and the Doxycycline-Inducible Promoter. PLoS ONE 5(5):e10611 [Overview of the project] [Problems that the invention aims to solve]

[0013] The object of the present invention is to overcome, or at least partially mitigate, the drawbacks of the prior art.

[0014] Therefore, an object of the present invention is to provide improved nucleic acid sequences and expression vectors that are useful in the production of recombinant proteins in mammalian cell lines. [Means for solving the problem]

[0015] For this purpose and other purposes, the nucleic acid sequence represented by the following is: TGCAGTACCTCCCTN1N2N3CCAGCAGGN4GGCAN5N6AGN7GAAN8GGTGAACTGGAGT (Sequence ID 10) N1, N2, N3, and N5 are each independently either C or G. N4 is either T or G. N6 and N7 are independently either C or T. N8 is achieved by a nucleic acid sequence that is either T or A.

[0016] Sequence ID 10 represents a CCCTC binding factor (CTCF) binding sequence, which is modified with respect to the CTCF binding sequence of the wild-type chicken high-sensitivity region 4 (cHS4) insulator. Each of N1 to N8 in Sequence ID 10 represents a nucleic acid substitution at the corresponding locus on the CTCF binding site (Sequence ID 1) of the wild-type cHS4 insulator. Hereinafter, Sequence ID 10 will be referred to as the "modified CTCF binding sequence". Preferably, the modified CTCF binding sequence can be selected from Sequence ID 2 and Sequence ID 3.

[0017] The nucleic acid sequence may be an isolated nucleic acid sequence.

[0018] The nucleic acid sequence may include a modified cHS4 core element containing the modified CTCF-binding DNA sequence, also known as the modified footprint II (FII) of the cHS4 core element. In addition to the modified CTCF-binding DNA sequence, the modified cHS4 core element may include footprints I, III, IV, and V that provide binding sites for VEZF1 (FI, FIII, and FV) and USF1 / USF2 (FIV).

[0019] In some embodiments, the nucleic acid sequence may include two HS4 core elements, each containing a CTCF binding sequence, at least one of which is the modified CTCF binding sequence according to the present disclosure (SEQ ID NO: 10).

[0020] The present invention has the potential to improve polypeptide expression from multiple expression cassettes provided in the same expression vector. With respect to known cHS4 core elements, the present invention offers the further advantage of allowing some variability in the nucleotide sequence of the CTCF-binding motif. For example, either modified CTCF-binding DNA fragment A (SEQ ID NO: 2) or modified CTCF-binding DNA fragment B (SEQ ID NO: 3) can be used. This variability allows for the incorporation of multiple copies of CTCF-binding DNA fragments, such as in the form of a modified cHS4 double core element, with less risk of genomic instability due to the incorporation of multiple copies of the same sequence.

[0021] Thus, in the HS4 double-core element, at least one of the HS4 core elements can contain a CTCF binding sequence selected from SEQ ID NO: 2 and SEQ ID NO: 3. Optionally, in the HS4 double-core element, one of the HS4 core elements can be represented by SEQ ID NO: 4 corresponding to the wild-type cHS4 core element. For example, in the HS4 double-core element, one of the HS4 core elements can contain a CTCF binding sequence represented by SEQ ID NO: 1, and one of the HS4 core elements can contain a CTCF binding sequence selected from SEQ ID NO: 2 and SEQ ID NO: 3. Alternatively, one of the HS4 core elements can contain the modified CTCF binding sequence described in SEQ ID NO: 2, and the other of the HS4 core elements can contain the modified CTCF binding sequence described in SEQ ID NO: 3.

[0022] Thus, in the HS4 double-core element, at least one of the HS4 core elements can contain at least one CTCF binding motif described in either SEQ ID NO: 2 or 3. For example, the double-core element can have the sequence described in SEQ ID NO: 5, SEQ ID NO: 6 or SEQ ID NO: 7 (described in more detail below). Alternatively, the double-core element can have a sequence based on any one of SEQ ID NOs: 4-7, provided that at least one of the two CTCF binding motifs is a modified CTCF binding sequence of the present disclosure, such that one of the CTCF binding motifs is replaced by a different CTCF binding motif. For example, SEQ ID NO: 4 can be modified to contain one CTCF binding sequence described in SEQ ID NO: 2 or 3 while preserving one wild-type CTCF binding motif (SEQ ID NO: 1). As another example, SEQ ID NO: 5 can be modified to contain one CTCF binding motif described in SEQ ID NO: 1 or 3 while preserving one CTCF binding fragment A (SEQ ID NO: 2). As yet another example, one of SEQ ID NO: 6 or 7 can be modified to contain one CTCF binding motif described in SEQ ID NO: 1 or 2 while preserving one CTCF binding fragment B (SEQ ID NO: 3).

[0023] The nucleic acid sequence can further include a spacer sequence that is at least 100 nucleotides in length and is located downstream of the modified cHS4 core element. In the case of the HS4 double core element, the spacer sequence that is at least 100 nucleotides in length can be located downstream of the two HS4 core elements, i.e., downstream of the double core element, in the 5' to 3' direction. The spacer sequence can have a length of at least 200 nucleotides, optionally up to 500 nucleotides, for example up to 300 nucleotides. The additional spacer sequence can be a non-coding sequence or a non-functional sequence, and can also be a random sequence. The additional spacer sequence preferably should not contain a sequence motif known to interact with a DNA-binding protein.

[0024] The present invention further provides an expression vector comprising the nucleic acid sequence described herein, which is useful for the expression of at least two non-identical polypeptides. In particular, the expression vector can include a first expression cassette comprising a first coding sequence encoding a first protein of interest, a second expression cassette comprising a second coding sequence encoding a second protein of interest, and the above-described nucleic acid sequence disposed between the first expression cassette and the second expression cassette. Thus, the nucleic acid sequence containing the CTCF binding motif forms a spacer sequence. The expression vector can include further expression cassettes or sequences useful for vector propagation, genomic integration, gene expression, gene product processing, selection and / or purification. For example, the expression vector can include a third expression cassette containing a selection marker. The expression vector can further include a recombinase site for integration into the genome of the host cell. The expression vector can be provided as a circular vector, or can be linear or linearized.

[0025] The first and second target proteins typically form part of the same protein complex, which may be a therapeutic protein complex such as an antibody or antibody fragment formed from two distinct polypeptide chains. Thus, the first target protein may be the first polypeptide chain of an antibody or antibody fragment, and the second target protein may be the second polypeptide chain of an antibody or antibody fragment. As a simple example, the first target protein may represent the heavy chain of an antibody, such as a monoclonal antibody, and the second target protein may represent the light chain of an antibody. However, given the numerous types and fragments of antibodies, other variants are also conceivable. For example, the first polypeptide chain may contain a variable domain of the antibody, such as a variable domain of the antibody heavy chain, and the second polypeptide chain may contain another variable domain of the antibody, such as a variable domain of the antibody light chain.

[0026] The first expression cassette and / or the second expression cassette may include a eukaryotic promoter sequence such as the CMV, EF1α, or SV40 PGK promoter. Preferably, the first expression cassette and / or the second expression cassette may include the CMV promoter.

[0027] The present invention further provides cells, cell lines, or cell cultures comprising the expression vector described herein. The cells may be cells that proliferate in vitro. The cells may be mammalian cells, such as Chinese hamster ovary (CHO) cells, transfected with the expression vector. The expression vector may be incorporated into the cell genome, preferably stably.

[0028] The present invention further provides a method for recombinantly producing two polypeptides, the method comprising: i) transfecting mammalian cells, such as CHO cells, with an expression vector comprising an expression cassette encoding the first and second target proteins described herein; and ii) culturing the transfected cells in a cell culture medium under conditions that enable the expression of the first and second target proteins. The method may further comprise the step of purifying the polypeptides from the cell culture according to known methods. The two polypeptides may be, for example, two polypeptide chains of a monoclonal antibody or antibody fragment.

[0029] As used herein, the terms “peptide” and “polypeptide” are used synonymously and refer to compounds formed from a sequence of amino acids, without any limitation on size. “Protein” may be used to refer to larger compounds in this class.

[0030] The terms "wild type" or "wt" refer to the typical naturally occurring form.

[0031] Preferred embodiments of this disclosure are described in the following detailed description and dependent claims. It should be noted that the present invention relates to all possible combinations of the features described in the claims.

[0032] The use of the verb "includes" and its conjugations does not exclude the existence of elements or processes other than those described. The articles "a" or "an" preceding an element do not exclude the existence of multiple such elements. The mere fact that certain means are described in different dependent claims does not indicate that a combination of these means cannot be used advantageously.

[0033] This and other embodiments of the present invention will be described in more detail with reference to the accompanying drawings. [Brief explanation of the drawing]

[0034] [Figure 1A] (See background) This figure shows a schematic diagram of the core element of the wild-type chicken high-sensitivity region 4 (cHS4) insulator. [Figure 1B] (See background) This figure shows a schematic diagram of a cHS4 double core formed by arranging two core elements of a wild-type cHS4 insulator in succession. [Figure 2] (See background) This figure shows a map of known monocistronic plasmid expression vectors, pcDNATM3.1. [Figure 3] This figure shows a map of the parent / base bicistronic modular expression vector of the present invention for random integration in a circular configuration. [Figure 4] Figure 3 shows the parent / base bicistronic modular expression vector in a linear configuration. [Figure 5] This figure shows a map of the parent / base bicistronic modular expression vector of the present invention for site-specific integration in a circular arrangement. [Figure 6] Figure 5 shows the parent / base bicistronic modular expression vector in a linear configuration. [Figure 7] This figure shows flow cytometry data associated with cells transfected with the expression vector pGE0308 related to Experiment 1. The left panel shows forward scatter versus side scatter plots used to gate viable single CHO cells (N-gate). The right panel shows mTagBFP2 versus eGFP plots used to calculate the expression percentages of both eGFP and mTagBFP2 (P-gate) from CHO cells at gate N. [Figure 8] This figure shows a plot illustrating the expression rate of recombinant CHO cells from gate P associated with Experiment 1. The x-axis represents functional spacers A-D with additional spacers X of different lengths. "C" represents the control (no double core element) with only a 100 bp spacer X. [Figure 9]This figure shows flow cytometry data related to cells transfected with the expression vector pGE0359 associated with Experiment 2. The left panel shows forward scatter versus side scatter plots used to gate viable single CHO cells (F-gate). The right panel shows mTagBFP2 versus eGFP plots used to calculate the expression percentages of both eGFP and mTagBFP2 (M-gate) from CHO cells at gate F. [Figure 10] This figure plots the expression rates of recombinant CHO cells from gate M related to Experiment 2. The x-axis shows functional spacers A-C with additional spacers of different lengths. [Figure 11] This figure shows the flow cytometry plots used to calculate mTagBFP2 and eGFP expression for Experiment 3. The left panel shows the forward scatter versus side scatter plot used to gate viable cells (A-gate). The center panel shows the side scatter against the mRaspberry plot used to gate cells expressing mRaspberry, corresponding to cells integrated in LP (B-gate). The right panel shows the mTagBFP2 versus eGFP plot used to calculate the average expression value based on the fluorescence signal of the primary focal population (C-gate). [Figure 12] This figure shows the normalized expression levels when using functional spacer A, which has two different 200bp additional spacers (α and β), and when using only 200bp spacer β, which does not have a double core element. Error bars = + / - 2 standard deviations. [Modes for carrying out the invention]

[0035] As shown in the figure, some features may be exaggerated for illustrative purposes and are therefore provided to illustrate the general structure of embodiments of the present invention.

[0036] Several experiments were conducted to construct modular expression vectors for enhancing the expression of target polypeptides. Different components of such expression vectors, such as expression cassettes, promoters, and selection markers, can be readily substituted or exchanged to reconfigure the expression vector for the production of a specific target polypeptide. While improved expression is largely exemplified by the use of random insertion (RI) vectors, these findings are also applicable to site-directed insertion (SDI).

[0037] Figure 2 shows a map of known monocistronic plasmid expression vectors pcDNATM3.1 for expression in various mammalian cell lines. This vector was used as a starting point for designing the modular bicistronic expression vector of the present invention.

[0038] Figure 3 shows a map of a parent / base bicistronic modular expression vector for random incorporation in a circular configuration. A bicistronic expression vector was designed to express two polypeptides. As shown in Figure 3, the parent / base expression vector contains a standard E. coli growth cassette as shown in Figure 2 (not shown). The selection marker is represented as expression cassette 1. Expression cassette 2 contains the gene encoding eGFP, and expression cassette 3 contains the gene encoding mTagBFP2. In these examples, the two polypeptides to be expressed are eGFP and mTagBFP2. As further shown in Figure 3, a spacer element is placed between expression cassettes 2 and 3 to isolate gene expression from the cassette and enhance the expression of the target polypeptide. Figure 3 also shows promoters 2 and 3 for enhancing expression from expression cassettes 2 and 3, respectively. Figure 4 shows the same parent / base bicistronic modular expression vector in a linear configuration.

[0039] Figure 5 shows a map of parent / base bicistronic modular expression vectors for site-specific integration in a circular configuration. Bicistronic expression vectors were designed to express two polypeptides. Similar to the expression vector in Figure 3, the parent / base expression vector contains a standard E. coli growth cassette as shown in Figure 2 (not shown). The selected marker gene is represented as expression cassette 1, which has no promoter but instead has an upstream attB2 recombination site. Expression cassette 2 contains the gene encoding Fc-eGFP, and expression cassette 3 contains the gene encoding mTagBFP2. Once the vector is incorporated into a unique landing pad integrated into the CHO genome using PhiC31 recombinase, the selected marker is activated by a promoter derived from the landing pad. In these examples, the two polypeptides expressed are eGFP and mTagBFP2. As further shown in Figure 3, a spacer element is positioned between expression cassettes 2 and 3 to isolate gene expression from the cassette and enhance the expression of the target polypeptide. Figure 3 also shows promoters 2 and 3 for enhancing expression from expression cassettes 2 and 3, respectively. Figure 6 shows the same parental / base bicistronic modular expression vector in a linear configuration.

[0040] To further enhance protein expression, several different spacers were developed to be placed between expression cassettes 2 and 3, and their effects on gene expression were tested. Longer spacer elements (over 500 bp), such as the cHS4 double core element, were observed to improve gene expression. It was also observed that protein expression improved when the wild-type CTCF binding site (SEQ ID NO: 1) present in the wild-type cHS4 double core element (SEQ ID NO: 4) was modified.

[0041] For example, various modified versions of the CTCF binding site of the wild-type cHS4 double core element were developed and evaluated, including modified CTCF-binding DNA fragment A (SEQ ID NO: 2) and modified CTCF-binding DNA fragment B (SEQ ID NO: 3). The nucleic acid sequences of the wild-type and modified CTCF-binding motifs or fragments are shown in Table 1a.

[0042] [Table 1]

[0043] The sequence identity between sequence numbers 1, 2, and 3 was calculated using Geneious Prime software (Biomatters), and is shown in Table 1b (Table 2).

[0044] [Table 2]

[0045] Several modified versions of the entire cHS4 double core element, including the modified version of the CTCF binding site described above, were developed and evaluated for gene expression. These modified double core elements were used in spacers A-C, as outlined below.

[0046] Spacer A (Sequence ID 5): [ka]

[0047] Modified cHS4 double core element A (SEQ ID NO: 5) contained modified CTCF-binding DNA fragment A (SEQ ID NO: 2) as its CTCF-binding site. When modified cHS4 double core element A is used as a spacer element between two expression cassettes (here, expression cassettes 2 and 3), it is referred to as "Spacer A".

[0048] Spacer B (Sequence ID 6): [ka]

[0049] Modified cHS4 double core element B (SEQ ID NO: 6) contained modified CTCF-binding DNA fragment B (SEQ ID NO: 3) as its CTCF-binding site. When modified cHS4 double core element B is used as a spacer element between two expression cassettes (here, expression cassettes 2 and 3), it is referred to as "spacer B".

[0050] Spacer C (Sequence ID 7): [ka]

[0051] Modified cHS4 double core element C (SEQ ID NO: 7) contained modified CTCF-binding DNA fragment B (SEQ ID NO: 3) as its CTCF-binding site. Furthermore, the DNA sequences between footprints FI, FII, FIII, FIV, and FV, as well as the DNA sequences between the two core elements, differed with respect to spacer B (SEQ ID NO: 6). Modified cHS4 double core element C is referred to as "spacer C" when used as a spacer element between two expression cassettes (expression cassettes 2 and 3 in this case).

[0052] Furthermore, the wild-type cHS4 double core element (SEQ ID NO: 4) is referred to as "Spacer D" when used as a spacer element between two expression cassettes (expression cassettes 2 and 3 in this case).

[0053] Spacer D / Wild-type cHS4 double core element (SEQ ID NO: 4): [ka]

[0054] The full length of the spacer element between expression cassettes 2 and 3 was also evaluated, and it was found that gene expression from expression cassettes 2 and 3 was further increased by ligating an additional spacer sequence downstream of the modified cHS4 double core element. This combination of the additional spacer element, referred to herein as "spacer X" or "spacer element X," and the modified cHS4 double core element (spacer A, B, or C), is referred to as a functional spacer element when used as a spacer between expression cassettes 2 and 3.

[0055] For example, several different additional spacer elements X were constructed and evaluated, including nucleic acid sequences with different nucleotide orders and nucleic acid sequences of different lengths such as 100 bp, 200 bp, 300 bp, 400, and 500 bp. Their effects on gene expression were then evaluated in combination with the modified cHS4 double core elements described above.

[0056] Table 2 (Table 3) below shows various plasmid expression vectors constructed based on the parent / base expression vector designs shown in Figures 3 and 5. Expression cassette 2 was configured to encode the eGFP or Fc-eGFP protein using a human cytomegalovirus (hCMV) or chimeric CMV promoter as promoter 2. Expression cassette 3 was configured to encode the mTagBFP2 protein using a mouse cytomegalovirus (mCMV) promoter as promoter 3. The above spacers based on the cHS4 double core element were placed between expression vectors 2 and 3. The choice marker was either glutamine synthetase (SM1), TagRFP-T (SM2), or mRaspberry (SM3). For the additional spacer X sequence, 100bp, 200bp, 300bp, or 500bp stretches were used. Furthermore, two different 200bp sequences were compared to study the potential effects of different sequences.

[0057] Additional spacer array variant α (sequence number 8): [ka]

[0058] Additional spacer array variant β (sequence number 9): [ka]

[0059] Spacer D refers to the wild-type cHS4 double core element (SEQ ID NO: 4).

[0060] [Table 3] [Examples]

[0061] Abbreviation [Table 4]

[0062] (Example 1) In this experiment, plasmid expression vectors pGE0173, pGE0296-pGE0303, and pGE0307-pGE0309, having the general structures described in relation to Figure 3 and Table 2 (Table 3), were linearized and transfected into Chinese hamster ovary (CHO) cells using electroporation, randomly incorporating them into the genome of the CHO cells. The transfected CHO cells were grown with 25 μM methionine sulfoximin (MSX). After two weeks, cell samples were analyzed by flow cytometry to determine eGFP and mTagBFP2 expression levels.

[0063] Figure 7 shows flow cytometry data related to cells transfected with the expression vector pGE0308 associated with Experiment 1. The left panel shows viable single CHO cells selected for evaluation (N-gate), and the right panel shows the expression of mTagBFP2 and eGFP from the N-gate, allowing for the calculation of the percentage of CHO cells expressing both eGFP and mTagBFP2 (P-gate). All expression vectors in Experiment 1 were evaluated in the same manner as pGE0308 as described above.

[0064] The expression results from the spacers in Experiment 1 can be seen in Figure 8, where the percentage of recombinant CHO cells expressed from gate P is shown for different spacers. The results show that when there is no double core element between expression cassettes 2 and 3 ("C"), approximately 0.1% of CHO cells are expressed from both cassettes 2 and 3, compared to an average of 0.2% when a functional spacer containing a double core element is present. Functional spacer A, which has an additional 100 bp spacer X sequence, showed the most improved expression, with 0.42% of CHO cells expressed at gate P. In conclusion, functional spacers (including the double core element) show improved expression compared to the case without a functional spacer. Furthermore, spacer A with an additional 100 bp spacer is substantially improved compared to spacer D (wild-type double core).

[0065] (Example 2) In this experiment, plasmid expression vectors pGE0341, pGE0355-pGE0365, having the general structure described in relation to Figure 3, were linearized and transfected into mammalian CHO cells using electroporation, randomly incorporating them into the genome of the CHO cells. The transfected CHO cells were grown. After two weeks, cell samples were analyzed by flow cytometry to confirm the incorporation of the TagRFP-T vector and to determine the expression levels of eGFP and mTagBFP2.

[0066] Figure 9 shows a plot of flow cytometry data associated with cells transfected with the expression vector pGE0359 related to Experiment 2. The left panel shows viable single CHO cells selected for evaluation (F-gate), and the right panel shows the expression of mTagBFP2 and eGFP from the F-gate, allowing for the calculation of the percentage of CHO cells expressing both eGFP and mTagBFP2 (M-gate). All expression vectors in Experiment 2 were evaluated in the same manner as pGE0359 as described above.

[0067] The expression results for the spacers from Experiment 2 can be seen in Figure 10, where the percentage of recombinant CHO cells expressed from gate M is shown for different spacers. The results show that spacers A and B (FI~FV) with the original sequence between footprints showed good expression, averaging 0.2~0.3%. Spacer C with an additional 200bp spacer X showed the best expression of CHO cells expressed at gate M, at 0.49%. As a group, the spacer C variant generally showed improved expression compared to spacers A or B with the chimeric CMV promoter in expression cassette 2. Among the individual spacer elements, spacer B with an additional 200bp spacer showed the best improvement.

[0068] (Example 3) The effects of different additional spacer sequences were evaluated using two different 200 bp additional spacer sequences, additional spacer variant α (SEQ ID NO: 8) and additional spacer variant β (SEQ ID NO: 9), for functional spacers separating expression cassettes 2 and 3. In this experiment, site-directed integration (SDI) was used with plasmid expression vectors pUP0062, pUP0071, and pUP0073 and PhiC31 recombinase. Expression vectors were transfected into CHO cells by electroporation to integrate into the landing pad (LP) sequence pre-inserted into the CHO cell genome. Double transfection was performed for each spacer sequence element. After 1 week of proliferation, cell samples were analyzed by flow cytometry to confirm the integration of vectors containing mRaspberry and to determine Fc-eGFP and mTagBFP2 expression levels.

[0069] Figure 11 shows plots of flow cytometry data related to Experiment 3. Mean mTagBFP2 and eGFP signals were calculated based on viable cells and subpopulations of cells expressing mRaspberry, as shown in Figure 11. The data were then normalized to the donor plasmid with the lowest expression level (pUP0073). As shown in Figure 11, the left panel shows the forward scatter versus side scatter plot used to gate viable cells (Gate A). The center panel shows the side scatter against the mRaspberry plot used to gate cells expressing mRaspberry, corresponding to cells integrated in the landing pad (Gate B). The right panel shows the mTagBFP2 versus eGFP plot used to calculate mean expression values ​​based on the fluorescence signal of the primary focal population (Gate C). The mean and standard deviation for each spacer element were then calculated based on the dual data. The obtained data are plotted in Figure 12. As can be seen from this figure, Experiment 3 shows that spacer A, which has an additional 200 bp spacer, clearly improves expression compared to spacer A with only 200 bp spacer (i.e., lacking the cHS4 double core element), and that spacer A functions well with an additional 200 bp spacer of a different sequence.

[0070] Based on experiments 1, 2, and 3 described above, it was concluded that spacers A, B, and C could significantly improve gene expression, and that the specific nucleotide sequence of the additional spacer element X was not important. However, as demonstrated by performing experiment 3 using two different sequences of 200 bp length as the additional spacer element encoded by nucleic acid sequences SEQ ID NOs: 8 and 9, it was found that the length of the additional spacer affected gene expression from expression cassettes 2 and 3.

[0071] (Example 4) This embodiment demonstrates that monoclonal antibodies (Herceptin) can be produced by the expression of light chain and heavy chain expression cassettes separated using a functional spacer element according to an embodiment of the present invention.

[0072] (Example 4A) Comparison of expression vectors with different functional spacers In this experiment, we generated a minipool containing expression vectors incorporated by a random incorporation (RI) mechanism using five different expression plasmids that differed only in the design of their functional spacers. The expression vectors were constructed according to Figure 3 with the following modifications: a) The glutamine synthetase (GS) gene, driven by the mPGK promoter, was used as a selective marker (expression cassette 1). b) The Herceptin HC gene, driven by a chimeric CMV promoter, was used as expression cassette 2. c) The Herceptin LC gene, driven by the mCMV promoter, was used as expression cassette 3.

[0073] The functional spacer design of the expression vector used is summarized in Table 3 (Table 5).

[0074] [Table 5]

[0075] Different expression vectors (Table 3 (Table 5)) were linearized and used to transfect CHO-K1 cells using a lipofectamine-based method. Two days after transfection, CHO cells were sorted into 96-well plates for static culture using FACS. For each well of the plate, 2000 cells were sorted into 100 μl of growth medium supplemented with 25 μM methionine sulfoximine (MSX). The cells were then grown by static culture for a total of 26 days in the presence of 25 μM MSX. During this static culture period, only cells incorporating the expression vector (and therefore the GS gene) could divide and proliferate in the presence of 25 μM MSX and in the absence of L-glutamine. At a seeding density of 2000 cells per well, typically 0 to 5 individual cells could proliferate per well. After static culture, the wells containing the proliferating cells were transferred to a 96-deep-well plate containing 550 μl of culture medium supplemented with 37.5 μM MSX. The plates were incubated under conditions for suspension culture until stable growth was detected. The cells were then seeded onto new deep-well plates and used for 8-day batch culture. The supernatant from the batch culture plates was then used to measure Herceptin titer using a commercially available antibody titer kit (ValitaCell / Beckman Coulter).

[0076] Table 4 (Table 6) summarizes the antibody titers measured for the two minipools with the best expression for each expression vector design.

[0077] [Table 6]

[0078] As shown in Table 4 (Table 6), both expression vectors containing spacer A (pGE0349, pGE0352) and expression vectors containing spacer C (pGE0353, pGE0354) yielded desirable high antibody titers (approximately 1 g / L or higher).

[0079] (Example 4B) Comparison with reference vector The pGE0349 expression vector was compared to a reference expression vector (pGE0164) that contained spacer D as a functional spacer. The reference vector did not contain the additional spacer X.

[0080] [Table 7]

[0081] Two expression vectors (Table 5 (Table 7)) were linearized and used to transfect CHO-K1 cells using a lipofectamine-based method. Three days after transfection, CHO cells were sorted into two 96-well plates for static culture for each expression vector using FACS. For each well of the plate, 5000 cells were sorted into 100 μl of growth medium supplemented with 25 μM MSX. The cells were then grown by static culture in the presence of 25 μM MSX for a total of 25 days for pGE0164 and 28 days for pGE0349. During this static culture period, only cells incorporating the expression vector (and therefore the GS gene) could divide and proliferate in the presence of 25 μM MSX and in the absence of L-glutamine. At a seeding density of 5000 cells per well, typically 3 to 10 individual cells could proliferate per well. After static culture, wells containing growing cells were transferred to a 96-deep-well plate containing 550 μl of culture medium supplemented with 37.5 μM MSX. The plate was incubated under conditions for suspension culture until stable growth was detected. The cells were then seeded into new deep-well plates and used for 11-day batch culture. Herceptin titers were then measured using commercially available antibody titer kits with the supernatant from the batch culture plates. Table 6 (Table 8) shows the expression minipool numbers and median titers for each expression vector.

[0082] [Table 8]

[0083] From this study, it was concluded that the expression vector having the functional spacer element of the present invention produces an antibody titer at least equivalent to that of the expression vector containing spacer D (wild-type cHS4 double core).

[0084] (Example 4C) Development of established cell lines In this experiment, we conducted an established cell line development campaign using a modified expression vector (pGE0419), including evaluation of productive cell clones using fed-batch culture in a shaking flask.

[0085] The expression vector pGE0419 contained the following modifications compared to the pGE0349 vector of Examples 4A and 4B. First, the GS gene used as the selection marker was modified by a mutation from arginine to glycine at amino acid number 299 (R299G). Second, the promoter driving Herceptin HC was changed from a chimeric CMV to an mCMV sequence. Finally, the additional spacer contained 200 bp instead of 100 bp.

[0086] Transfection and minipool preparation were performed as described above for Example 4B. However, in this experiment, the best-performing minipool was selected for further growth and then used for single-cell cloning using FACS. During single-cell cloning, single cells were seeded in each well of a 96-well static culture plate containing 100 μl of cloning medium supplemented with 25 μM MSX. After growth under static culture conditions, the cells were transferred to suspension culture in deep-well plates as previously described. Following titer evaluation after 11 days of batch culture, 24 of the best-performing productive cell clones were selected for further evaluation in 24-deep-well plate fed-batch culture. Fourteen of the best-performing productive cell clones from this stage were further grown and evaluated in fed-batch culture in shaking flasks. Cell density, cell viability, and antibody titer were measured throughout the culture using commercially available equipment. Culture was terminated when cell viability fell below 20% (typically after 18-21 days). The results are summarized in Table 7 (Table 9).

[0087] [Table 9]

[0088] Therefore, Example 4C demonstrates the successful establishment of a CHO cell line capable of producing the desired antibody at very high levels.

[0089] (Example 5) This example demonstrates an established cell line development campaign for producing monoclonal antibodies (Herceptin) using site-directed integration (SDI) of antibody light and heavy chain expression cassettes separated by a functional spacer element according to an embodiment of the present invention.

[0090] In detail, using a site-directed integration (SDI) approach, sequences encoding Herceptin heavy and light antibody chains were introduced into a Chinese hamster ovary (CHO) cell line (developed in-house) that had a pre-inserted landing pad sequence in its cellular genome. Figure 13A shows the LP region of the host cell's landing pad, where two full-length HS4 sequences are adjacent.

[0091] Except for transfection and cloning, the cells were always grown in ActiPro medium containing 6 mM L-glutamine.

[0092] To facilitate integration into the landing pad, two donor vectors, pUP0295 and pUP0296 (50:50 mixture), were used for transfection with the PhiC31 plasmid. The donor vector designs are shown in Figure 13B. Each of vectors pUP0295 and pUP0296 contained either a light chain (LC) encoding sequence in the first cassette and a heavy chain (HC) encoding sequence in the second cassette, or vice versa, as shown in Figure 13B. Each cassette was under the control of the mCMV promoter. Expression cassettes were separated by spacers consisting of spacer A (SEQ ID NO: 5) fused to an additional 200 bp spacer (SEQ ID NO: 8). All donor vectors further contained a marker for SDI (eGFP), as well as markers for all insertion events, which were both random and SDI (CD8). CD8 can be stained with CD8-R667 Ab for cell sorting and flow cytometry analysis to evaluate plasmid insertion.

[0093] Transfection of host cells using donor vectors: A Herceptin donor plasmid mixture was simultaneously transfected into CHO cells with the PhiC31 recombinase plasmid by electroporation. Seven days after transfection, cells were analyzed by flow cytometry to evaluate the insertion marker (CD8) and SDI marker (eGFP). On average, 0.22% of cells showed successful plasmid SDI.

[0094] Bulk cell sorting: After confirming the success of site-directed integration of the target gene by flow cytometry, the first bulk sorting was performed using a MACSQuant Tyto cell sorter. All eGFP-expressing (SDI) cells were bulk sorted and allowed to stand for proliferation. After another flow cytometry analysis, the cells were bulk sorted again, and at this point, cells expressing both eGFP and CD8 were selected. First, a rough bulk sorting was performed to extract a large number of cells, and immediately afterward, rapid purity sorting was performed on the same cells. Analysis of the cells with a flow cytometer revealed that more than 95% of the cells expressed eGFP, indicating that they were ready for Cre-mediated excision.

[0095] Cre-mediated excision of fluorescent markers and bulk sorting: Following proliferation after the final selection marker-positive bulk sorting, cells were transfected with Cre recombinase mRNA to remove the LoxP-adjacent region of plasmids containing the selection markers. Figure 13C shows the landing pad genomic region of transfected cells after excision of the plasmid backbone and selection markers. Cells were analyzed by flow cytometry 7 days after transfection to confirm that Cre-mediated excision was sufficient. The results showed that over 97% of cells had lost the eGFP fluorescent marker. Subsequently, a final bulk sorting was performed to purify the eGFP and CD8-negative population before cloning. After 3-4 days of proliferation, cells were analyzed by flow cytometry and confirmed that the pool contained approximately 99.9% eGFP-negative cells and 99.6% CD8-negative cells, confirming that it was ready for cloning.

[0096] Cloning and Titer Assessment: Ten 96-well plates were sorted into single cells using a UP.SIGHT cell printer (Cytena). The plates were maintained in a static incubator for 16 days for growth, after which cells were prepared for titer screening. Titer screening was performed using a Biacore 8K instrument (Cytiva), and titers were normalized to confluence. Subsequently, the top 192 clones per campaign were selected for growth in 96-deep-well plates.

[0097] Growth and Titer Assessment: After two passages and 11-12 days of growth, the antibody-killing cultures were seeded onto each plate. After 9-10 days of culture, the antibody-killing cultures were ready for antibody titer assessment using a Biacore 8K instrument (Cytiva). Fourteen clones were selected for growth in shaking flasks.

[0098] Growth and cryopreservation in shaking flasks: Fourteen clones were selected for growth in shaking flasks. Two of the clones were replaced with newly selected clones due to insufficient cell count or very slow growth. Otherwise, all growth clones grew well and could be cryopreserved in smaller cell banks. When PDT was stabilized in shaking flasks, the PDT of the clones was between 19 and 21 hours, with an average PDT of 20 hours.

[0099] Flow cytometry evaluation of selected clones: All 42 proliferating clones were analyzed by flow cytometry to confirm the absence of eGFP and CD8 expression after Cre-mediated excision and selection processes. All but one clone were confirmed to be negative for both eGFP and CD8.

[0100] Herceptin-expressing clones evaluated in fed-batch culture: Ten clones with the highest titer in flow cytometry analysis and lacking fluorescent markers were selected for fed-batch experiments in shaking flasks. Cells were seeded at 0.3–0.6 M cells / ml (PDT-based) in 21 ml of ActiPro culture medium containing 6 mM L-glutamine in SF125 flasks, aiming for approximately 5 M cells / ml on day 3 when medium addition began. 3% Cell Boost 7a and 0.3% Cell Boost 7b culture medium (Cytiva) were supplied to the flasks daily. Glucose was also supplied daily to reach 4–6 g / L. Antibody titers and metabolites were measured using a Cedex Bio HT Analyzer (Roche). Clones persisted in fed-batch culture for 10–17 days. The titer in typical fed-batch culture ranged from 0.5–7.8 g / L (up to day 17). Table 8 (Table 10) shows the antibody titers on day 14 for the top 5 Herceptin-expressing clones.

[0101] [Table 10]

[0102] Overall, Examples 1-5 demonstrate that the spacer elements disclosed herein have the potential to provide cell lines that are modified to highly express the target protein, particularly when two proteins are expressed simultaneously, such as in the case of mAbs and other antibody fragments or variants incorporating both light and heavy chain components. (References) TIFF2026512866000018.tif77155

Claims

1. A nucleic acid sequence comprising a modified CTCF binding sequence represented by the following sequence, TGCAGTACCTCCCTN 1 N 2 N 3 CCAGCAGGN 4 GGCAN 5 N 6 AGN 7 GAAN 8 GGTGAACTGGAGT (Sequence ID 10) N 1 、N 2 、N 3 、and N 5 are each independently C or G, N 4 is T or G, N 6 and N 7 These are, independently, C or T, N 8 This is a nucleic acid sequence that is either T or A.

2. The nucleic acid sequence according to claim 1, wherein the modified CTCF binding sequence is selected from SEQ ID NO: 2 and SEQ ID NO:

3.

3. The nucleic acid sequence according to claim 1 or 2, comprising a modified chicken high-sensitivity region 4 (cHS4) core element containing the modified CTCF-binding DNA sequence.

4. The nucleic acid sequence according to claim 1 or 2, each comprising two highly sensitive region 4 (HS4) core elements containing a CTCF binding sequence, at least one of which is the modified CTCF binding sequence.

5. The nucleic acid sequence according to claim 4, wherein one CTCF binding sequence of the HS4 core element is represented by sequence number 1.

6. The nucleic acid sequence according to claim 4, comprising a sequence selected from sequence numbers 5 to 7.

7. The nucleic acid sequence according to claim 4, wherein one of the HS4 core elements includes the modified CTCF binding sequence described in Sequence ID No. 2, and one of the HS4 core elements includes the modified CTCF binding sequence described in Sequence ID No.

3.

8. The nucleic acid sequence according to claim 3, comprising a spacer sequence of at least 100 nucleotides in length located downstream of the modified cHS4 core element.

9. The nucleic acid sequence according to any one of claims 4 to 7, comprising a spacer sequence of at least 100 nucleotides in length located downstream of both of the two HS4 core elements.

10. The nucleic acid sequence according to claim 8 or 9, wherein the spacer sequence has a length of at least 200 nucleotides and optionally up to 500 nucleotides, for example, up to 300 nucleotides.

11. An expression vector comprising the nucleic acid sequence according to any one of claims 1 to 10.

12. A first expression cassette containing a first coding sequence encoding a first target protein, A second expression cassette containing a second coding sequence encoding a second target protein, A nucleic acid sequence according to any one of claims 1 to 10, disposed between the first expression cassette and the second expression cassette, The expression vector according to claim 11, comprising:

13. The expression vector according to claim 12, wherein the first target protein is a first polypeptide chain of an antibody or antibody fragment, and the second target protein is a second polypeptide chain of an antibody or antibody fragment.

14. The expression vector according to claim 13, wherein the first polypeptide chain comprises a variable domain of an antibody heavy chain, and the second polypeptide chain comprises a variable domain of an antibody light chain.

15. The expression vector according to any one of claims 12 to 14, wherein the first expression cassette and / or the second expression cassette comprises a eukaryotic promoter sequence such as a CMV, EF1α, or SV40 PGK promoter, and preferably a CMV promoter.

16. An expression vector according to any one of claims 11 to 15, which is a circular vector.

17. A cell comprising the expression vector according to any one of claims 11 to 16.

18. The cell according to claim 17, wherein the expression vector is incorporated into the genome of the cell.

19. The cell according to claim 17 or 18, wherein the cell is a mammalian cell, such as a Chinese hamster ovary (CHO) cell, transfected with the expression vector.

20. A step of transfecting mammalian cells such as CHO cells with an expression vector according to any one of claims 12 to 16, which includes an expression cassette encoding first and second target proteins, The steps include culturing the transfected cells in a cell culture medium under conditions that enable the expression of the first and second target proteins, A method for recombinantly producing at least two polypeptides, including [a specific compound].

21. The method according to claim 20, wherein the two polypeptides are two polypeptide chains of a monoclonal antibody or antibody fragment.

22. The method according to claim 20 or 21, further comprising the step of purifying the polypeptide.