Modified promoters for improved applications

A nucleic acid construct with an inactivated TATA box promoter and a second functional promoter enhances transcription and expression of nucleotide sequences, addressing limitations in existing methods and achieving up to 2000% increase in expression levels.

WO2026022376A1PCT designated stage Publication Date: 2026-01-29PROTEONIC BIOTECHNOLOGY IP BV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/071530
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-07-25
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing methods for regulating transcription and expression of nucleotide sequences are limited, particularly when mutations in the TATA box lead to loss of promoter activity, resulting in diseases and suboptimal expression levels.

Method used

A nucleic acid construct with a first promoter lacking a functional TATA box operably linked to a second promoter, enhancing transcription and expression of nucleotide sequences by utilizing a TATA box promoter that is inactivated, often through deletion or mutation, and optionally incorporating splice sites to further improve expression.

Benefits of technology

The method significantly increases transcription and expression of nucleotide sequences by up to 2000% compared to single-promoter systems, allowing for efficient production and purification of transcripts and polypeptides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000010_0001
    Figure IMGF000010_0001
  • Figure IMGF000026_0001
    Figure IMGF000026_0001
  • Figure 00000031_0000
    Figure 00000031_0000
Patent Text Reader

Abstract

The invention is in the field of polynucleotides. The invention relates to a method for transcription and / or expression using a nucleic acid construct that is characterized by the presence of a TATA box promoter that has been modified to disable the TATA box. This mutated promoter is followed by a second promoter. The invention further relates to that nucleic acid construct, to an expression vector or a cell comprising that construct, and to its use for improved transcription.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Modified promoters for improved applications

[0002] Field of the invention

[0003] The invention is in the field of polynucleotides. The invention relates to a method for transcription and / or expression using a nucleic acid construct that is characterized by the presence of a TATA box promoter that has been modified to disable the TATA box. This mutated promoter is followed by a second promoter. The invention further relates to that nucleic acid construct, to an expression vector or a cell comprising that construct, and to its use for improved transcription.

[0004] Background art

[0005] A TATA box (sometimes called the Goldberg-Hogness box) is a sequence of DNA that can be found in the core promoter region of genes in archaea and eukaryotes. The TATA box is considered a non-coding DNA sequence (also known as a cis- regulatory element). It was termed the "TATA box" as it contains a consensus sequence characterized by repeating T and A base pairs. Transcription is initiated in the vicinity of the TATA box in TATA-containing genes, typically the TATA box is located -25 of the site where RNA polymerase II starts transcription (TSS). The TATA box is the binding site of the TATA-binding protein (TBP) and other transcription factors in some eukaryotic genes.

[0006] Based on the sequence and mechanism of TATA box initiation, mutations such as insertions, deletions, and point mutations to this consensus sequence can result in phenotypic changes. These phenotypic changes can then turn into a disease phenotype, because the promoter with the mutated TATA box loses its original activity. Diseases such as chronic hemolytic anemia, hemophilia B Leyden, and thalassemia have all been linked to loss of promoter activity resulting from TATA box mutations (see for instance Antonarakis et al., 1984, PNAS, 81 (4): 1154-1158). Another example is Gilbert's syndrome, which is correlated with TATA box polymorphism.

[0007] WO2015 / 102487 describes nucleic acid constructs where two promoters simultaneously drive expression of a same, single nucleotide sequence of interest. The second promoter is flanked by intron splice sites. Both promoters generate different primary transcripts, but after splicing the mRNA molecules are highly comparable. The simultaneous use of two promoters improved transcription and expression.

[0008] There is an ongoing need in the art for alternative and preferably improved methods for regulating the transcription of a transcript or regulating the expression of a protein or polypeptide of interest in host cells.

[0009] Summary of the invention

[0010] The inventors identified an expression construct for increasing the transcription of mRNAs, or the expression of polypeptides of interest. The expression construct of the invention is characterized by two promoters that would both be operably linked to a coding sequence of an mRNA or polypeptide of interest, were it not that the first promoter lacks a functional TATA box. Thus, in general, from 5’ to 3’ the constructs comprise a first promoter, a second promoter, and the nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. It was found that inactivating the first promoter surprisingly increased transcription or expression. For instance, constructs similar to those described in WO2015 / 102487 (see the section on Background art) showed improved performance when the first promoter was modified to inactivate the TATA box.

[0011] Thus the invention provides a method for producing a transcript of a nucleotide sequence of interest, the method comprising the steps of: a) providing a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter, a second promoter, and the nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated; and, b) contacting a cell with the nucleic acid construct to obtain a transformed cell; and, c) allowing the transformed cell to produce the transcript of the nucleotide sequence of interest.

[0012] Optionally the method is for producing a purified transcript, wherein the method further comprises the step of: d) purifying the produced transcript. Also provided is a method for expressing a polypeptide of interest comprising step a) and step b) as above, wherein the nucleotide sequence of interest encodes the polypeptide of interest; and wherein the method further comprises the step of: d) allowing the transformed cell to express the protein or polypeptide of interest.

[0013] Preferably the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, the second promoter, and the nucleotide sequence of interest. Preferably the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, optionally a further donor splice site, an acceptor splice site, and the nucleotide sequence of interest.

[0014] In some embodiments the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein: i) the TATA box consensus sequence has been deleted; ii) part of the TATA box consensus sequence has been deleted; iii) at least 1 , 2, 3, 4, 5, 6, or 7 of the nucleotides in the TATA box consensus sequence have been mutated; iv) at most 1 , 2, 3, or 4 of the nucleotides in the TATA box consensus sequence have not been mutated.

[0015] Preferably the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein the consensus sequence TATAWAW has been mutated to VBNBNBS, preferably to CTTGATG. Preferably the first promoter and the second promoter independently are promoters that are active in mammalian cells, preferably selected from human or murine cytomegalovirus (CMV) promoter, simian virus (SV40) promoter, human, hamster, or mouse ubiquitin C (UBC) promoter, human or mouse or rat elongation factor alpha (EF1-a) promoter, apolipoprotein A1 gene promoter (apoA1), and mouse or hamster beta-actin promoter; or preferably the first promoter and the second promoter independently are promoters that are active in yeast and fungal cells, preferably selected from Leu2 promoter, galactose (Gall or Ga17) promoter, alcohol dehydrogenase I (ADH1) promoter, glucoamylase (Gia) promoter, triose phosphate isomerase (TPI) promoter, translational elongation factor EF-I alpha (TEF2) promoter, glyceraldehyde-3- phosphate dehydrogenase (gpdA) promoter, alcohol oxidase (AOX1) promoter, and glutamate dehydrogenase (gdhA) promoter. Preferably the first promoter has at least 80% sequence identity with SEQ ID NO: SEQ ID NO: 89, 90, 91 , 93, 95, 105, 107, or 1 10, preferably 91 , 93, 105, 107, or 110, more preferably 105 or 107 or 110, even more preferably 107 or 110, most preferably 107.

[0016] Also provided is a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter and a second promoter, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. Preferably it comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, and the second promoter, preferably it comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, and an acceptor splice site; more preferably it comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, a further donor splice site, and the acceptor splice site. The nucleic acid construct preferably has at least 80% sequence identity with SEQ ID NO: 109. Also provided is an expression vector comprising a nucleic acid construct as defined above, further comprising a nucleotide sequence of interest. Also provided is a cell comprising a nucleic acid construct as defined above, or comprising an expression vector as defined above.

[0017] Also provided is the use of a TATA box promoter wherein the TATA box is inactivated, for increasing transcription of a nucleotide sequence of interest when the nucleotide sequence of interest is operably linked to a second promoter that does not have an inactivated TATA box.

[0018] Description of embodiments

[0019] The inventors identified an expression construct for increasing the transcription of mRNAs, or the expression of polypeptides of interest. The expression construct of the invention is characterized by two promoters that would both be operably linked to a coding sequence of an mRNA or polypeptide of interest, were it not that the first promoter lacks a functional TATA box. Thus, in general, from 5’ to 3’ the constructs comprise a first promoter, a second promoter, and the nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. It was found that inactivating the first promoter surprisingly increased transcription or expression.

[0020] Method for improved production

[0021] The invention allows for the improved production of mRNA transcripts, and for the improved production of polypeptides. The invention thus provides a method for producing a transcript of a nucleotide sequence of interest, the method comprising the steps of: a) providing a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter, a second promoter, and the nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated; and, b) contacting a cell with the nucleic acid construct to obtain a transformed cell; and, c) allowing the transformed cell to produce the transcript of the nucleotide sequence of interest.

[0022] Preferably, the nucleotide construct of the invention comprising a first promoter and a second promoter and is capable of increasing the transcription of the nucleotide sequence of interest that is under the control of said second promoter. It would also be under the control of the first promoter, were it not that the first promoter has an inactivated TATA box. Alternatively or in combination with the increased transcription, the nucleotide construct is also preferably capable of increasing expression of a protein or polypeptide of interest encoded by the nucleotide sequence of interest. Preferably, transcription levels are assessed in an expression system using an expression construct comprising said first promoter and second promoter operably linked to a nucleotide sequence of interest using a suitable assay such as RT-qPCR. Preferably, the nucleotide construct of the invention comprising the first promoter and second promoter of the invention allows for an increase in transcription of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500% or 2000% of said nucleotide sequence of interest as compared to transcription using a construct which only differs in that the nucleotide sequence of interest is under the control of a single promoter, preferably the second promoter.

[0023] Preferably, expression levels are established in an expression system using an expression construct comprising said first promoter and second promoter operably linked to a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, said protein or polypeptide of interest is a secreted protein or polypeptide and expression of said protein or polypeptide of interest is detected by a suitable assay such as an enzyme-linked immunosorbent assay (ELISA) assay, Western blotting or, dependent on the identity of the protein or polypeptide of interest, any suitable protein identification and / or quantification assay known to the person skilled in the art. Preferably, the first promoter and second promoter of the invention allow for an increase in expression of protein or polypeptide expression of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500% or 2000% as compared to expression of said protein or polypeptide using a construct which only differs in that the encoding sequence of the protein or polypeptide of interest is under the control of a single promoter, preferably of the second promoter, preferably when tested in a system as exemplified in the Examples which are enclosed herein.

[0024] The nucleic acid construct of step a) is defined elsewhere herein. A skilled person can select a nucleotide sequence of interest depending on their intentions. Examples are nucleotide sequences that express enzymes such as demonstrated in the Examples for SeAP, that express antibodies such as demonstrated in the examples, or that express relevant RNAs when production of RNAs is desired.

[0025] The skilled person is capable of transforming cells in accordance with step b). Transformation methods as used in step b) include, but are not limited to transfer of purified DNA via cationic lipid reagents and polyethyleneimide (PEI), calcium-phosphate co-precipitation, microparticle bombardment, electroporation of protoplasts, microinjection, viral infection, or use of silicon fibers to facilitate penetration and transfer of DNA into the host cell. In step c) the transformed cell is allowed to produce a transcript of the nucleotide sequence of interest. For example, the transformed cell may be subjected to conditions leading to transcription of the nucleotide sequence of interest. The person skilled in the art is well aware of techniques to be used for transcription of the nucleotide sequence of interest. Methods in which the transformed cell does not need to be subjected to specific conditions leading to transcription of the nucleotide sequence of interest, but in which the nucleotide sequence of interest is automatically (e.g., constitutively) transcribed, are also included in the method of the present invention.

[0026] In preferred embodiments the method is for producing a purified transcript, wherein the method further comprises the step of: d) purifying the produced transcript.

[0027] In this step the produced transcript is recovered, allowing its further use or application. Purification steps depend on the transcript produced. Purification can also be seen as isolation of the produced transcript. The term "isolation" indicates that the transcript is found in a condition other than its native environment. In a preferred form, the isolated transcript is substantially free of other cellular components, particularly other homologous cellular components. It is preferred to provide the transcript in a greater than 40% pure form, more preferably greater than 60% pure form. Even more preferably it is preferred to provide the transcript in a highly purified form, i.e., greater than 80% pure, more preferably greater than 95% pure, and even more preferably greater than 99% pure, as determined by Northern blot. Preferably the purified transcript is stored in a composition further comprising water, preferably ultrapure water.

[0028] Also provided is a method for expressing a polypeptide of interest comprising step a) and step b) as described above, wherein the nucleotide sequence of interest encodes the polypeptide of interest; and wherein the method further comprises the step of: d) allowing the transformed cell to express the protein or polypeptide of interest.

[0029] This method is preferably in vitro or ex vivo. This method is preferably not a therapeutic method practiced on a human or animal body. In step c) the transformed cell is further allowed to express the protein or polypeptide of interest, and in step d) that protein or polypeptide is subsequently recovered. For example, the transformed cell may be subjected to conditions leading to expression of the protein or polypeptide of interest. The person skilled in the art is well aware of techniques to be used for expressing or overexpressing the protein or polypeptide of interest. Methods in which the transformed cell does not need to be subjected to specific conditions leading to expression of the protein or polypeptide of interest, but in which the protein or polypeptide of interest is automatically (e.g., constitutively) expressed, are also included in the method of the present invention. Purification steps and definitions related to these steps as the definition of an isolated protein or polypeptide are the same as defined elsewhere herein, or as known to a skilled person. Preferably, the method allows for an increase in expression of a protein or polypeptide of interest. Inactivated TATA box

[0030] Within the context of the invention a promoter is a promoter capable of initiating transcription in a host cell of choice. Promoters as used herein include tissue-specific, tissuepreferred, cell-type specific, inducible, and constitutive promoters as defined herein in the Definitions section. Promoters that may be comprised within said first or second promoter as defined herein are promoters that may be employed in transcription of nucleotide sequences of interest and / or expression of polypeptides of interest, preferably in mammalian cells, and include, but are not limited to, the human or murine cytomegalovirus (CMV) promoter, a simian virus (SV40) promoter, a human or mouse ubiquitin C (UBC) promoter, a human or mouse or rat elongation factor alpha (EF1-a) promoter, apolipoprotein A1 gene promoter (apoA1 , as described by Matsunaga et al., doi: 10.1161 / 01 .ATV.19.2.348), or a mouse or hamster beta-actin promoter. The Tet-Off and Tet-On responsive elements upstream of a minimal promoter such as a CMV promoter is an example of an inducible mammalian promoter. Examples of suitable yeast and fungal promoters are Leu2 promoter, the galactose (Gall or Ga17) promoter, alcohol dehydrogenase I (ADH1) promoter, glucoamylase (Gia) promoter, triose phosphate isomerase (TPI) promoter, translational elongation factor EF-I alpha (TEF2) promoter, glyceraldehyde-3-phosphate dehydrogenase (gpdA) promoter, alcohol oxidase (AOX1) promoter, or glutamate dehydrogenase (gdhA) promoter. An example of a strong ubiquitous promoter for expression in plants is cauliflower mosaic virus (CaMV) 35S promoter.

[0031] Constructs of the invention often comprise two promoters as described herein. Preferably the first promoter and the second promoter independently are promoters that are active in mammalian cells, preferably selected from human or murine cytomegalovirus (CMV) promoter, simian virus (SV40) promoter, human, hamster, or mouse ubiquitin C (UBC) promoter, human or mouse or rat elongation factor alpha (EF1-a) promoter, apolipoprotein A1 gene promoter (apoA1), and mouse or hamster beta-actin promoter; or preferably the first promoter and the second promoter independently are promoters that are active in yeast and fungal cells, preferably selected from Leu2 promoter, galactose (Gall or Ga17) promoter, alcohol dehydrogenase I (ADH1) promoter, glucoamylase (Gia) promoter, triose phosphate isomerase (TPI) promoter, translational elongation factor EF-I alpha (TEF2) promoter, glyceraldehyde-3-phosphate dehydrogenase (gpdA) promoter, alcohol oxidase (AOX1) promoter, and glutamate dehydrogenase (gdhA) promoter. Promoters named in this paragraph are TATA box promoters.

[0032] "Promoter" refers to a nucleic acid sequence located upstream or 5' to a translational start codon of an open reading frame (or protein-coding region) of a gene and that is involved in recognition and binding of RNA polymerase II and other proteins (trans-acting transcription factors) to initiate transcription. The term promoter refers to a nucleic acid fragment that functions to control the transcription of one or more genes, located upstream with respect to the direction of transcription of the transcription initiation site of the gene, and is structurally identified by the presence of a binding site for DNA-dependent RNA polymerase, transcription initiation sites and any other DNA sequences, including, but not limited to transcription factor binding sites, repressor and activator protein binding sites, and any other sequences of nucleotides known to one skilled in the art to act directly or indirectly to regulate the amount of transcription from the promoter. The promoter does not include the transcription start site (TSS) but rather ends at nucleotide -1 of the transcription site, and does not include nucleotide sequences that become untranslated regions in the transcribed mRNA such as the 5'-UTR. Promoters of the invention may be tissue-specific, tissue-preferred, cell-type specific, inducible and constitutive promoters. Tissue-specific promoters are promoters which initiate transcription only in certain tissues and refer to a sequence of DNA that provides recognition signals for RNA polymerase and / or other factors required for transcription to begin, and / or for controlling expression of the coding sequence precisely within certain tissues or within certain cells of that tissue. Expression in a tissue-specific manner may be only in individual tissues or in combinations of tissues. Tissue-preferred promoters are promoters that preferentially initiate transcription in certain tissues. Cell-type-specific promoters are promoters that primarily drive expression in certain cell types. Inducible promoters are promoters that are capable of activating transcription of one or more DNA sequences or genes in response to an inducer. The DNA sequences or genes will not be transcribed when the inducer is absent. Activation of an inducible promoter is established by application of the inducer. Constitutive promoters are promoters that are active under many environmental conditions and in many different tissue types. Preferably, capability to initiate transcription is established in an expression system using an expression construct comprising said promoter operably linked to a nucleotide sequence of interest using a suitable assay such a RT-qPCR or Northern blotting. A promoter is said to be capable to start transcription if a transcript can be detected or if an increase in a transcript level is found of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500% or 2000% as compared to transcription using a construct which only differs in that it is free of said promoter. In a further preferred embodiment, capability to initiate expression is established in an expression system using an expression construct comprising said promoter operably linked to a nucleotide sequence encoding a protein or polypeptide of interest. Preferably, said protein or polypeptide of interest is a secreted protein or polypeptide and expression of said protein or polypeptide of interest is detected by a suitable assay such as an ELISA assay, Western blotting or, dependent on the identity of the protein or polypeptide of interest, any suitable protein identification and / or quantification assay known to the person skilled in the art. A promoter is said to be capable to initiate expression if the protein or polypeptide of interest can be detected or if an increase in a expression level is found of at least 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 200%, 300%, 500%, 1000%, 1500% or 2000% as compared to expression using a construct which only differs in that it is free of said promoter. As a first and second promoter of the invention, an induced or constitutive promoter or a combination thereof may be used in the present invention.

[0033] A skilled person is familiar with the TATA box and can identify the TATA box in a promoter. Accordingly a skilled person can also identify whether a promoter is a TATA box promoter. Generally, a TATA box promoter is a promoter that comprises a TATA box or analogue thereof. Promoters generally comprise a single TATA box, if any. The TATA box is a component of the eukaryotic core promoter and generally contains the consensus sequence 5'-TATA(A / T)A(A / T)-3' which can also be denoted as TATAWAW. In yeast, for example, one study found that various Saccharomyces genomes had the consensus sequence 5'-TATA(A / T)A(A / T)(A / G)-3', or TATAWAWR. Genes containing the TATA-box tend to be involved in stress-responses and certain types of metabolism and are more highly regulated as compared to genes under the control of a promoter that does not comprise a TATA box. The TATA box is usually located 25-35 base pairs upstream of the transcription start site. Genes containing the TATA box usually comprise additional promoter elements, including an initiator site located just upstream of the transcription start site and a downstream core element (DCE). These additional promoter regions work in conjunction with the TATA box to regulate initiation of transcription.

[0034] In the invention a promoter is used wherein the TATA box is inactivated. A TATA box can be considered inactivated when the interaction with TATA-binding protein (TBP) is reduced or even absent. Savinkova et al. (PLoSONE8(2): e54626. doi:10.1371 / journal. pone.0054626) teach a simulation to predict the KD value for a given TATA box sequence, mutated or not, and TBP. This simulation was experimentally found to have a strong predictive value and can thus reliably be used to predict effects of a selected mutation on the tightness of TBP binding to the TATA box. Preferably, a TATA box is considered inactivated when TBP associates with the promoter with an affinity that is at most 60% of its original affinity, more preferably at most 50%, even more preferably at most 40%, still more preferably at most 30%, most preferably at most 20%. Generally speaking this leads to absence of promoter activity. A promoter wherein the TATA box is inactivated preferably does not comprise a functional TATA box, for instance it would not comprise a further, intact TATA box. A promoter wherein the TATA box is inactivated preferably differs from a wildtype TATA box promoter (which comprises a functional wildtype TATA box) only in this inactivation.

[0035] Preferably the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein: i) the TATA box consensus sequence has been deleted; ii) part of the TATA box consensus sequence has been deleted; iii) at least 1 , 2, 3, 4, 5, 6, or 7 of the nucleotides in the TATA box consensus sequence have been mutated; iv) at most 1 , 2, 3, or 4 of the nucleotides in the TATA box consensus sequence have not been mutated.

[0036] Thus a TATA box promoter wherein the TATA box is inactivated can be identified by comparing it to its wildtype counterpart, which comprises a functional TATA box. Insertion of single nucleotides is not preferred as this can sometimes leave the TATA box operable while shifting the transcription initiation site. For this reason, deletion or partial deletion of the TATA box is also not preferred. Most preferably, at least one of options iii) and iv) above apply. Point mutations to the TATA box have similar varying phenotypic changes depending on the gene that is being affected. Studies have consistently shown that mutations in the TATA box sequence hinders the binding of TBP. For example, a mutation from TATAAAA to CATAAAA can hinder the binding sufficiently to change transcription. As another example, a change can be seen in HeLa cells with a TATAAAA to TATACAA which leads to a 20 fold decrease in transcription. Preferably at least 3, 4, 5, 6, or 7 of the nucleotides in the TATA box consensus sequence have been mutated; more preferably at least 4, still more preferably at least 5, 6, or 7 of the nucleotides in the TATA box consensus sequence have been mutated. Preferably at most 4 of the nucleotides in the TATA box consensus sequence have not been mutated, more preferably at most 3, most preferably at most 2 have not been mutated. Preferably, positions 1 , 2, 4, 6, and 7 are mutated in an inactivated TATA box, so that these positions do not match TATAWAW. Preferably the length of an inactivated TATA box is not altered by the mutations. Preferably the promoter does not have any mutations outside of any mutations that inactivate the TATA box. Preferably the promoter does not have any deletions outside of any deletions that inactivate the TATA box.

[0037] Preferably the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein the consensus sequence TATAWAW has been mutated to VBNBNBS, preferably to CTTGATG. Examples of inactivated TATA boxes are CTCGGTG, CTCGGAG, AGGCTTC, and CTTGATG, the latter of which has shown particularly good results. The IUPAC nucleotide codes are commonly known, and the following are particularly relevant here:

[0038] Preferred promoters wherein the TATA box is inactivated are therefore promoters the wildtype of which are TATA box promoters, wherein the TATA box has been mutated to inactivate it, particularly to prevent binding by TBP. Examples of promoters inactivated this way are shown in SEQ ID NOs: 91 , 93, 105, and 110. By way of illustration, SEQ ID NO: 106 is the wildtype promoter corresponding to SEQ ID NO: 105. Preferably the first promoter has at least 80% sequence identity with SEQ ID NO: 91 or 93 or 105 or 107 or 110, preferably 105 or 107 or 110, more preferably 107 or 1 10, most preferably 107. Identity is preferably over the entire length of the SEQ ID NO. In more preferred embodiments the first promoter has at least 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity, most preferably it has 100% sequence identity. It should be noted that positions 876-882 of SEQ ID NO: 105 represent the inactivated TATA box (CTTGATG). By comparison, positions 876-882 of SEQ ID NO: 106 are TATATAA. Given how for the first promoter the TATA box is to be inactivated, the first promoter has at least 80% sequence identity with SEQ ID NO: 105 with the proviso that positions 876-882 are not TATAWAW, preferably with the proviso that positions 876-882 are VBNBNBS, thus forming SEQ ID NO: 107, which is highly preferred as a first promoter. SEQ ID NO: 110 is also highly preferred in this regard.

[0039] As demonstrated herein, a TATA box promoter wherein the TATA box is inactivated can surprisingly contribute to further increasing the transcription or expression of a nucleotide sequence of interest, particularly when the inactivated promoter is used in a construct or method as described herein. Accordingly the invention also provides the use of a TATA box promoter wherein the TATA box is inactivated for increasing transcription of a nucleotide sequence of interest when the nucleotide sequence of interest is operably linked to a second promoter that does not have an inactivated TATA box. Preferably this first and second promoter are as described elsewhere herein, and are configured as described herein.

[0040] Nucleic acid construct

[0041] The invention provides a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter and a second promoter, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. Because they are particularly useful for being transcribed, it is preferred that the nucleic acid constructs according to the invention are DNA constructs. In other embodiments the invention provides a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter and a donor splice site, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. Preferably the second promoter is not a TATA box promoter wherein the TATA box is inactivated. The second promoter can be another type of functional promoter and is most preferably a TATA box promoter with a functional TATA box. In further embodiments the nucleic acid construct comprises in the 5’ to the 3’ direction a first promoter, a donor splice site, and a second promoter, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. Preferably when an order of elements is indicated for constructs according to the invention, or for constructs for use in the invention, the order of elements is not interrupted by further elements. This does not preclude the presence of short spacer sequences that may be present, but which do not have promoter activity, or which do not have any epigenetic function.

[0042] These nucleotide constructs have been found to be useful for improving transcript production or polypeptide expression. For this, a nucleotide sequence of interest can be operably linked to the promoters. A skilled person can select a nucleotide sequence of interest depending on their intentions. Examples are nucleotide sequences that express enzymes such as demonstrated in the Examples for SeAP, that express antibodies such as demonstrated in the examples, or that express relevant RNAs when production of RNAs is desired. Accordingly the invention provides a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter, second promoter, and a nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. In other embodiments the invention provides a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter, a donor splice site, and nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. Preferably the second promoter is not a TATA box promoter wherein the TATA box is inactivated. The second promoter can be another type of functional promoter and is most preferably a TATA box promoter with a functional TATA box. In further embodiments the nucleic acid construct comprises in the 5’ to the 3’ direction a first promoter, a donor splice site, a second promoter, and a nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated. It is known from WO2015 / 102487 that so-called intronic promoters can help improve transcription or expression. In line with this, the inventors found that use of a TATA box promoter wherein the TATA box is inactivated also leads to improved results when used in a construct featuring an intronic promoter. Accordingly the invention provides a nucleic acid construct which comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, the second promoter, optionally a further donor splice site, an acceptor splice site, and optionally a nucleotide sequence of interest. In some embodiments is provided a nucleic acid construct which comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, the second promoter, a further donor splice site, an acceptor splice site, and optionally a nucleotide sequence of interest. In some embodiments is provided a nucleic acid construct which comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, the second promoter, a further donor splice site, an acceptor splice site, and a nucleotide sequence of interest. Use of this construct in the methods as described above is also provided.

[0043] In preferred embodiments the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, and the second promoter, preferably the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, and an acceptor splice site; more preferably the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, a further donor splice site, and the acceptor splice site.

[0044] Such constructs were found to enhance transcription or expression beyond the enhancement that is already offered by the use of the intronic promoter technology. It should be noted that the donor splice site in between the first promoter and the second promoter is not found in transcripts, because the first promoter does not initiate transcription due to its inactivated TATA box. Thus, when a pre-mRNA containing an intron is desired, use of a construct is preferred wherein the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, optionally a donor splice site, the second promoter, a further donor splice site, and an acceptor splice site. This ensures the presence of both a donor and acceptor splice site, and thus an intron, in the pre-mRNA.

[0045] An “intron” is a nucleotide sequence within a primary RNA transcript that is removed by RNA splicing or intron splicing while the final mature RNA product is being generated. Assessment whether intron splicing occurs can be done using any suitable method known to the person skilled in the art, such as but not limited to reverse-transcriptase polymerase chain reaction (RT-PCR) followed by size or sequence analysis of the RT-PCR product. Preferably, a nucleotide sequence is an intron if at least 2%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the primary RNA loses this sequence by RNA splicing using an assay suitable to detect intron splicing as indicated above. Preferably, an intron comprises a splice site GT at the 5’ end of the nucleotide sequence, and a splice site AG at the 3’ end of the nucleotide sequence, which splice site AG is preceded by a pyrimidine rich nucleotide sequence or polypyrimidine tract, optionally separated from splice site AG by 1-50 nucleotides. An intron may further comprise a branch site comprising the sequence Y-T-N-A-Y, at the 5’ side of the polypyrimidine tract. The branch site may have the nucleotide sequence C-Y-G-A-C. An “intronic sequence” is understood to be at least part of the nucleotide sequence of an intron.

[0046] Preferably, a donor splice site comprises at least a donor splice site or splice site GT. A donor splice site is understood herein as a splice site that, when combined with an acceptor splice site as defined herein, results in the formation of an intron as defined elsewhere herein. Preferably, a nucleotide sequence is an intron if at least 2%, 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 100% of the primary RNA loses this sequence by RNA splicing using an assay suitable to detect intron splicing, such as but not limited to reverse-transcriptase polymerase chain reaction (RT-PCR) followed by size or sequence analysis of the RT-PCR. Preferred donor splice sites of the invention are M-W-G-[cut]-G-T-R-A-G-K or M-A-R-[cut]-G-T-R-A-G-K in case the host cell is a mammalian cell, A-G-[cut]-G-T-A-W-K in case the host cell is a plant cell, [cut]-G-T-A-W-G-T-T in case the host cell is a yeast cell and R-G-[cut]-G-T-R-A-G, in case the host cell is an insect cell, “[cut]” is to be understood herein as the specific cut site where splicing will take place. Intron splicing can be assessed functionally using an assay as detailed elsewhere herein. Most preferably, the donor splice site is C-T-G-[cut]-G-T-G-A-G-G or A-A-A-[cut]-G-T-G-A-G-G.

[0047] Preferably, an acceptor splice site is understood herein as the splice site AG preferably preceded by a pyrimidine rich sequence or polypyrimidine tract nucleotide sequence, optionally separated from splice site AG by 1-50 nucleotides, and optionally further comprising a branch site comprising the sequence Y-T-N-A-Y, at the 5’ site of the polypyrimidine tract nucleotide sequence, wherein the branch site may have the nucleotide sequence C-Y-G-A-C. An acceptor splice site is understood herein as a splice site that, when combined with a donor splice site encompassed within a construct, results in the formation of an intron as defined elsewhere herein. Preferably, the acceptor splice site or splice site AG has the sequence [Y-rich]-N-Y-A-G-[cut], Preferably, the acceptor splice site or splice site AG has the sequence [Y-rich]-N-Y-A-G-[cut]-R in case the host cell is a mammalian cell, [Y-rich]-D-Y-A-G-[cut]-R or [Y-rich]-D-Y-A-G-[cut]-R-W in case the host cell is a plant cell, [Y-rich]-A-Y-A-G-[cut] in case the host cell is a yeast cell and [Y-rich]-N-Y-A-G- [cut] in case the host cell is an insect cell. “[Y-rich]” is to be understood herein as the polypyrimidine tract which is preferably defined as a consecutive sequence of at least 10 nucleotides comprising at least 6, 7, 8, 9 or preferably 10 pyrimidine nucleotides. Preferably, said acceptor splice site or splice site GT has the sequence Y-A-G-[cut]-R.

[0048] In an alternative embodiment, a donor splice site and an acceptor splice site as defined herein flank the second promoter. Preferably, the second promoter and the intronic sequences flanking the second promoter are configured to form an intronic promoter (referred is to Fig. 1 A). An intronic promoter is known to a person skilled in the art as a promoter located within an intronic sequence. Preferably, said intronic promoter is an intron as defined elsewhere herein. Preferably, the boundaries of the intronic promoter of the present invention are being formed by the donor splice site upstream of the second promoter of the invention and the acceptor splice site at the 3’ site or downstream of the second promoter of the invention. The intronic promoter of the invention can have a length that is comparable or similar to naturally occurring introns, preferably comparable or similar to naturally occurring introns in the host cell or organism as defined herein. Preferably, said intronic promoter as defined herein is at most 12,000 nucleotides in length. Preferably, said donor splice site at the 5’ site or upstream of said second promoter is located at the 3’ site or downstream of said first promoter. Preferably, the first promoter and second promoter, the intronic sequences flanking the second promoter, and a nucleotide sequence encoding a protein or polypeptide of interest are configured in such a way that the first promoter is upstream of the second promoter, wherein the second promoter is flanked by said intronic sequences to form an intronic promoter, and wherein said first promoter and second promoter are configured to be both upstream and operably linked to the nucleotide sequence encoding a protein or polypeptide of interest (Fig. 1). It should be noted that when the first promoter has an inactivated TATA-box, there is no production of a pre-mRNA where the intronic promoter is spliced from the mature mRNA.

[0049] The intronic promoter may comprise further expression enhancing elements, but preferably the intronic promoter is free of further splice sites apart from the donor and acceptor splice sites as defined herein within the first and second intronic sequences as defined herein. Preferably, one or more expression enhancing sequences are comprised within said first and / or said second promoter.

[0050] Provided is the method or construct as described herein, wherein the donor splice site comprises the donor splice site consensus sequence GT, and is preferably selected from any one of M-W-G-[cut]-G-T-R-A-G-K , M-A-R-[cut]-G-T-R-A-G-K, A-G-[cut]-G-T-A-W-K, [cut]-G-T-A-W-G- T-T, and R-G-[cut]-G-T-R-A-G, more preferably the donor splice site is C-T-G-[cut]-G-T-G-A-G-G or A-A-A-[cut]-G-T-G-A-G-G and / or wherein the acceptor splice site comprises the acceptor splice site consensus sequence AG, and is preferably selected from [Y-rich]-N-Y-A-G-[cut], [Y-rich]-N-Y-A-G-[cut]-R, [Y-rich]-D-Y- A-G-[cut]-R, [Y-rich]-D-Y-A-G-[cut]-R-W, [Y-rich]-A-Y-A-G-[cut], and [Y-rich]-N-Y-A-G-[cut], wherein “[Y-rich]” is a polypyrimidine tract, preferably defined as a consecutive sequence of at least 10 and at most 30, more preferably at most 20 nucleotides comprising at least 6, 7, 8, 9 or preferably 10 pyrimidine nucleotides.

[0051] The distance between the first promoter and the second promoter is not of particular relevance, as long as both promoters are operably linked to the nucleotide sequence of interest (taking into account that the first promoter would be operably linked were it not deactivated). In preferred embodiments the first promoter is separated from the second promoter by about 0 to about 5000 bp, preferably by about 10-4000 bp, more preferably about 50-3000 bp, still more preferably about 100-2500 bp, more preferably about 200-2000 bp, more preferably about 250-1800 bp, more preferably about 300-1600 bp, more preferably about 350-1400 bp, more preferably about 400- 1200 bp, more preferably about 450-1100 bp, more preferably about 500-1000 bp, more preferably about 520-900 bp, more preferably about 540-800 bp, more preferably about 560-700 bp, most preferably about 570-600 bp, such as about 580 bp. The nucleic acid sequence separating the first and the second promoter should preferably not comprise donor splice sites, more preferably should not comprise splice sites, except those that are explicitly described as being present.

[0052] In embodiments wherein the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, and the second promoter, the first promoter is preferably separated from the donor splice site by about 0 to about 2000 bp, preferably by about 10-1500 bp, more preferably about 20-1250 bp, still more preferably about 30-1100 bp, more preferably about 40-1000 bp, more preferably about 50-900 bp, more preferably about 60-800 bp, more preferably about 70-700 bp, more preferably about 75-600 bp, more preferably about 80-500 bp, more preferably about 85-400 bp, more preferably about 90-300 bp, more preferably about 95-200 bp, more preferably about 98-150 bp, most preferably about 100-125 bp, such as about 100 bp. The distance between the first promoter and the donor splice site is not of particular relevance.

[0053] The distance between the second promoter and the nucleotide sequence of interest is not of particular relevance, as long as they are operably linked. In preferred embodiments the second promoter is separated from the nucleotide sequence of interest by about 0 to about 2000 bp, preferably by about 10-1800 bp, more preferably about 50-1500 bp, still more preferably about 100- 1300 bp, more preferably about 150-1100 bp, more preferably about 200-1000 bp, more preferably about 250-900 bp, more preferably about 260-800 bp, more preferably about 270-700 bp, more preferably about 280-600 bp, more preferably about 290-500 bp, more preferably about 295-400 bp, more preferably about 297-360 bp, more preferably about 298-340 bp, most preferably about 300-320 bp, such as about 300 bp. The nucleic acid sequence separating the second promoter and the nucleotide sequence of interest should preferably not comprise donor splice sites, more preferably should not comprise splice sites, except those that are explicitly described as being present.

[0054] The distance between the second promoter and the further donor splice site, when present, is not of particular relevance. In preferred embodiments the second promoter is separated from the further donor splice site by about 0 to about 2000 bp, preferably by about 10-1800 bp, more preferably about 30-1500 bp, still more preferably about 50-1300 bp, more preferably about 70- 1100 bp, more preferably about 90-1000 bp, more preferably about 100-900 bp, more preferably about 110-800 bp, more preferably about 120-700 bp, more preferably about 130-600 bp, more preferably about 140-500 bp, more preferably about 150-400 bp, more preferably about 160-300 bp, more preferably about 165-250 bp, most preferably about 170-200 bp, such as about 170bp.

[0055] The distance between the further donor splice site and the acceptor splice site, when present, is not of particular relevance. In preferred embodiments the further donor splice site is separated from the acceptor splice site by about 0 to about 2000 bp, preferably by about 10-1800 bp, more preferably about 30-1500 bp, still more preferably about 50-1300 bp, more preferably about 70- 1100 bp, more preferably about 90-1000 bp, more preferably about 100-900 bp, more preferably about 110-800 bp, more preferably about 120-700 bp, more preferably about 130-600 bp, more preferably about 140-500 bp, more preferably about 150-400 bp, more preferably about 160-300 bp, more preferably about 165-250 bp, most preferably about 170-200 bp, such as about 170bp. The donor splice site and the acceptor splice site define an intron. The size of such an intron is preferably about 30-1000, more preferably about 50-750, more preferably about 70-500, still more preferably about 80-400, still more preferably about 90-300, most preferably about 100-200 bp such as about 100 bp. Such an intron preferably comprises a Kozak sequence. In preferred embodiments the nucleic acid construct according to the invention has at least 80% sequence identity with SEQ ID NO: 109. More preferably it has 81 %, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 109. Preferably the sequence identity is with the proviso that a TATA-box is inactivated in the first promoter, and that the second promoter does comprise a functional TATA box.

[0056] In some embodiments, said first and said second promoters are similar promoters. Preferably, said first promoter has at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to said second promoter, outside of the TATA box, which is inactivated in the first promoter. It is highly preferred that the second promoter does comprise a functional TATA box.

[0057] In some embodiments, said first promoter and second promoter are distinct or different promoters. Preferably, said first promoter has less than 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10% or 5% sequence identity to said second promoter.

[0058] In preferred embodiments, said first promoter sequence comprises or consists of a UBC promoter or a CCT8 promoter and said second promoter comprises or consists of a CMV promoter, or the other way around. Preferably, said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to nucleotides 1-906 of SEQ ID NO: 1 or with nucleotides 1-508 of SEQ ID NO: 2, most preferably to SEQ ID NO: 107 or SEQ ID NO: 110. Preferably, said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 58, 79, or 111 , or preferably to SEQ ID NO: 57 or 96, more preferably 57.

[0059] Preferred is a nucleic acid construct of the invention wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to nucleotides 1-906 of SEQ ID NO: 1 , most preferably to SEQ ID NO: 107 or SEQ ID NO: 110, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 57 or 96 or 111 , more preferably SEQ ID NO: 57 or 111 , even more preferably SEQ ID NO: 111.

[0060] Also preferred is a nucleotide sequence of the invention wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity with nucleotides 1-508 of SEQ ID NO: 2, or with SEQ ID NO: 89, or more preferably with SEQ ID NO: 107 or 110, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to SEQ ID NO: 57 or 96 or 111 , more preferably SEQ ID NO: 57 or 111 , even more preferably SEQ ID NO: 111.

[0061] Also preferred is a nucleotide sequence of the invention wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 58, or preferably to SEQ ID NO: 57, more preferably SEQ ID NO: 96 or 97, most preferably to SEQ ID NO: 107, 79, or 110, and wherein said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to nucleotides 1-906 of SEQ ID NO: 1 or with nucleotides 1-508 of SEQ ID NO: 2.

[0062] Also provided is a construct wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with to SEQ ID NO: 57, most preferably to SEQ ID NO: 107 or 110, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to nucleotides 1-906 of SEQ ID NO: 1. Also preferred is a nucleotide sequence of the invention wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 57, most preferably to SEQ ID NO: 107 or 110, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with nucleotides 1-508 of SEQ ID NO: 2.

[0063] Also provided is a construct wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with to SEQ ID NO: 96, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to nucleotides 1-906 of SEQ ID NO: 1 . Also preferred is a nucleotide sequence of the invention wherein said first promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity to SEQ ID NO: 96, and said second promoter comprises or consists of a sequence that has at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% identity with nucleotides 1-508 of SEQ ID NO: 2.

[0064] Expression vectors and cells

[0065] The invention provides an expression vector comprising a nucleic acid construct as defined above, further comprising a nucleotide sequence of interest. The nucleic acid construct according to the invention is preferably a vector, in particular a plasmid, cosmid, or phage nucleotide sequence, linear or circular, of a single or double stranded DNA or RNA, preferably a double stranded DNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing any one of the nucleotide sequences of the invention in sense or antisense orientation into a cell. The choice of vector is dependent on the recombinant procedures followed and the host cell used. The vector may be an autonomously replicating vector or may replicate together with the chromosome into which it has been integrated. Preferably, the vector comprises a selection marker. Useful markers are dependent on the host cell of choice and are well known to persons skilled in the art and can be selected from selection markers as defined elsewhere herein. A preferred expression vector is the pcDNA3.1 expression vector. Preferred selection markers are the neomycin resistance gene, zeocin resistance gene, and blasicidin resistance gene, as well as genomic selection markers, and a dihydrofolate reductase selection system.

[0066] The invention also provides a cell comprising a nucleic acid construct as defined herein, or comprising an expression vector as defined above. Within the context of the invention, a cell may be mammalian (including human), plant, animal, insect, fungal, yeast, or bacterial cell. A recombinant host cell, such as a mammalian, including human, plant, animal, insect, fungal or bacterial cell, containing one or more copies of a nucleic acid construct according to the invention is an additional subject of the invention. By host cell is meant a cell which contains a nucleic acid construct such as a vector and supports the replication and / or transcription and / or expression of the nucleic acid construct. Examples of suitable bacteria are species of Escherichia such as E. coli. Fungal cells include yeast cells. Expression in yeast can be achieved by using yeast strains such as Komagataella phaffii (Pichia pastoris), Saccharomyces cerevisiae, and Hansenula polymorpha. Other fungal cells of interest include filamentous fungi cells as Aspergillus niger, Trichoderma reesei, and the like. Furthermore, insect cells such as cells or cell lines from Drosophila melanogaster, Spodoptera frugiperda, and Trichoplusia ni, such as, but not limited to, S2, Sf9, Sf21 , and High Five cells can be used as host cells. Alternatively, a suitable expression system can be a baculovirus system or expression systems using mammalian cells such as CHO, COS, CPK (porcine kidney), MDCK, BHK, Sp2 / 0, NSO, and Vero cells. A suitable human cell or human cell line is an astrocyte, adipocyte, chondrocyte, endothelial, epithelial, fibroblast, hair, keratinocyte, melanocyte, osteoblast, skeletal muscle, smooth muscle, stem, synoviocyte cell or cell line. Examples of suitable human cell lines also include HEK 293 (human embryonic kidney), HeLa, Per.C6, CAP (cell lines derived from primary human amniocytes), and Bowes melanoma cells. In an embodiment a human cell is not an embryonic stem cell. In some preferred embodiments the cell is a microbial cell. In highly preferred embodiments the cell is a mammalian cell such as CHO or HEK, or primary human cells, although cell lines such as CHO or HEK are more preferred.

[0067] A nucleic acid construct preferably is stably maintained, either as an autonomously replicating element, or, more preferably, the nucleic acid construct is integrated into the host cell’s genome, in which case the construct is usually integrated at random positions in the host cell's genome, for instance by non-homologous recombination. Stably transformed host cells are produced by known methods. The term stable transformation refers to exposing cells to methods to transfer and incorporate foreign DNA into their genome. These methods include, but are not limited to transfer of purified DNA via cationic lipid reagents and polyethyleneimide (PEI), calciumphosphate co-precipitation, microparticle bombardment, electroporation of protoplasts and microinjection or use of silicon fibers to facilitate penetration and transfer of DNA into the host cel, or use of viral vectors such as lentiviral, adenoviral, or adeno associated viral vectors. A nucleic acid construct according to the invention preferably also comprises a marker gene which can provide selection or screening capability in a treated host cell. Selectable markers are generally preferred for host transformation events, but are not available for all host cells. A nucleic acid construct disclosed herein can also include a nucleotide sequence encoding a marker product. A marker product can be used to determine if the construct or portion thereof has been delivered to the cell and once delivered is being expressed. Examples of marker genes include, but are not limited to the E. coli lacZ gene, which encodes B-galactosidase, and a gene encoding the green fluorescent protein. Examples of suitable selectable markers for mammalian cells include, but are not limited to dihydrofolate reductase (DHFR), glutathione synthetase (GS), thymidine kinase, neomycin, neomycin analog G418, hygromycin, blasticidin, zeocin and puromycin.

[0068] Other suitable selectable markers include, but are not limited to antibiotic, metabolic, auxotrophic or herbicide resistant genes which, when inserted in a host cell in culture, would confer on those cells the ability to withstand exposure to an antibiotic. Metabolic or auxotrophic marker genes enable transformed cells to synthesize an essential component, usually an amino acid, which allows the cells to grow on media that lack this component. Another type of marker gene is one that can be screened by histochemical or biochemical assay, even though the gene cannot be selected for. A suitable marker gene found useful in such host cell transformation experience is a luciferase gene. Luciferase catalyzes the oxidation of luciferin, resulting in the production of oxyluciferin and light. Thus, the use of a luciferase gene provides a convenient assay for the detection of the expression of introduced DNA in host cells by histochemical analysis of the cells. In an example of a transformation process, a nucleotide sequence sought to be expressed in a host cell could be coupled in tandem with the luciferase gene. The tandem construct could be transformed into host cells, and the resulting host cells could be analyzed for expression of the luciferase enzyme. An advantage of this marker is the non-destructive procedure of application of the substrate and the subsequent detection.

[0069] When such selectable markers are successfully transferred into a host cell, the transformed host cell can survive if placed under selective pressure. There are two widely used distinct categories of selective regimes. The first category is based on a cell's metabolism and the use of a mutant cell line which lacks the ability to grow independent of a supplemented media. Two nonlimiting examples are CHO DHFR-cells and mouse LTK-cells. These cells lack the ability to grow without the addition of such nutrients as thymidine or hypoxanthine. Because these cells lack certain genes necessary for a complete nucleotide synthesis pathway, they cannot survive unless the missing nucleotides are provided in a supplemented media. An alternative to supplementing the media is to introduce an intact DHFR or TK gene into cells lacking the respective genes, thus altering their growth requirements. Individual cells which were not transformed with the DHFR or TK gene will not be capable of survival in non-supplemented media.

[0070] The second category is dominant selection which refers to a selection scheme used in any cell type and does not require the use of a mutant cell line. These schemes typically use a drug to arrest growth of a host cell. Those cells which have a novel gene would express a protein conveying drug resistance and would survive the selection. Examples of such dominant selection use the drugs neomycin, mycophenolic acid, or hygromycin. The three examples employ bacterial genes under eukaryotic control to convey resistance to the appropriate drug G418 or neomycin (geneticin), xgpt (mycophenolic acid) or hygromycin, respectively. Others include the neomycin analog G418 and puromycin. Other useful markers are dependent on the host cell of choice and are well known to persons skilled in the art.

[0071] General definitions

[0072] The phrase "nucleic acid" as used herein refers to a naturally occurring or synthetic oligonucleotide or polynucleotide, whether DNA or RNA or DNA-RNA hybrid, single-stranded or double-stranded, sense or antisense, which is capable of hybridization to a complementary nucleic acid by Watson-Crick base-pairing. A nucleic acid of the invention is preferably modified as compared to its naturally occurring counterpart by comprising at least 1 , 2, 3, 4, 5, 10, 20, 30 or 50 nucleotide mutations as compared to its naturally occurring counterpart. Preferably, a nucleic acid of the invention does not occur in nature. Nucleic acids of the invention can also include nucleotide analogs (e.g., BrdU), and nonphosphodiester internucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids can include, without limitation, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, ssRNA, dsRNA, non-coding RNAs, hnRNA, premRNA, matured mRNA or any combination thereof. The terms "nucleic acid sequence" and "nucleotide sequence" as used herein are interchangeable, and have their usual meaning in the art. The term refers to a DNA or RNA molecule in single or double stranded form. An "isolated nucleic acid sequence" refers to a nucleic acid sequence which is no longer in the natural environment from which it was isolated. A nucleic acid molecule is represented by a nucleotide sequence. Furthermore, an element such as, but not limited to an expression enhancing element and a transcription regulating element, is represented by a nucleotide sequence.

[0073] A “recombinant construct” (or chimeric construct) refers to any nucleic acid sequence or molecule, which is not normally found in nature in a species, in particular a nucleic acid sequence, molecule or gene in which one or more parts of the nucleic acid sequence are present that are not associated with each other in nature. For example, a recombinant construct comprises a promoter that is not associated in nature with part or all of the transcribed region or with another regulating region comprised within said recombinant construct. The term "recombinant construct" is understood to include expression constructs in which a promoter or expression regulating sequence is operably linked to one or more sense sequences (e.g. coding sequences) or to an antisense (reverse complement of the sense strand) or inverted repeat sequence (sense and antisense, whereby the RNA transcript forms double stranded RNA upon transcription), or to any other sequence coding for a functional RNA molecule.

[0074] A “nucleic acid construct” is defined as a polynucleotide which is isolated from a naturally occurring gene or which has been modified to contain segments of polynucleotides which are combined or juxtaposed in a manner which would not otherwise exist in nature. Optionally, a polynucleotide present in a nucleic acid construct is operably linked to one or more control sequences, which direct the production or transcription of a nucleotide sequence of interest and / or the expression or a peptide or polypeptide of interest in a cell or in a subject.

[0075] A “vector” or “plasmid” is herein understood to mean a man-made (usually circular) nucleic acid molecule resulting from the use of recombinant DNA technology and which is used to deliver exogenous DNA into a host cell. Vectors usually comprise further genetic elements to facilitate their use in molecular cloning, such as e.g. selectable markers, multiple cloning sites and the like (see below). A nucleic acid construct may also be part of a recombinant viral vector for expression of a protein in a plant or plant cell (e.g. a vector derived from cauliflower mosaic virus, CaMV, or tobacco mosaic virus, TMV) or in a mammalian organism or mammalian cell system (e.g. a vector derived from Moloney murine leukemia virus (MMLV; a Retrovirus) a Lentivirus, an Adeno-associated virus (AAV) or an adenovirus (AdV)).

[0076] A “transformed cell” are terms referring to a new individual cell (or organism), arising as a result of the introduction into said cell of at least one nucleic acid molecule, especially comprising a chimeric or recombinant construct encoding a desired protein or a nucleic acid sequence which upon transcription yields an antisense RNA for silencing of a target gene / gene family. The host cell may be a plant cell, a bacterial cell (e.g. an Agrobacterium strain), a fungal cell (including a yeast cell), an animal (including insect, mammalian) cell, etc. The transformed cell may contain the nucleic acid construct as an extra-chromosomally (episomal) replicating molecule, as a non-replicating molecule or comprises the recombinant construct integrated in the nuclear or organellar DNA of the host cell. The term “organism” as used herein, encompasses all organisms consisting of more than one cell, i.e. multicellular organisms, and includes multicellular fungi.

[0077] “Transformation” and “transformed” refers to the transfer of a nucleic acid sequence, generally a nucleic acid sequence comprising a recombinant construct or gene of interest (GOI), into the nuclear genome of a cell to create a “transgenic” cell or organism comprising a transgene. The introduced nucleic acid sequence is generally, but not always, integrated in the host genome. When the introduced nucleic acid sequence is not integrated in the host genome, one may speak of “transfection”, “transiently transfected”, and “transfected”. For the purposes of the present patent specification, the terms “transformation”, “transiently transfected”, and “transfection” are used interchangeably, and refer to stable or transient presence of a nucleic acid sequence into a cell or organism. When the cell is a bacterial cell, the term usually refers to an extrachromosomal, selfreplicating vector which harbors a selectable antibiotic resistance.

[0078] "Sequence identity" or “identity” in the context of amino acid- or nucleic acid-sequence is herein defined as a relationship between two or more amino acid (peptide, polypeptide, or protein) sequences or two or more nucleic acid (nucleotide, polynucleotide) sequences, as determined by comparing the sequences. In the art, "identity" also means the degree of sequence relatedness between amino acid or nucleotide sequences, as the case may be, as determined by the match between strings of such sequences. Within the present invention, sequence identity with a particular sequence indicated with a particular SEQ ID NO preferably means sequence identity over the entire length of said particular polypeptide or polynucleotide sequence indicated with said particular SEQ ID NO. However, sequence identity with a particular sequence indicated with a particular SEQ ID NO may also mean that sequence identity is assessed over a part of said SEQ ID NO. A part may mean at least 50%, 60%, 70%, 80%, 90% or 95% of the length of said SEQ ID NO. The sequence information as provided herein should not be so narrowly construed as to require inclusion of erroneously identified bases. The skilled person is capable of identifying such erroneously identified bases and knows how to correct for such errors.

[0079] Any nucleotide sequences capable of hybridizing to the nucleotide sequences of the invention are defined as being part of the cis-acting elements of the invention. Stringent hybridization conditions are herein defined as conditions that allow a nucleic acid sequence of at least 25, preferably 50, 75 or 100, and most preferably 150 or more nucleotides, to hybridize at a temperature of about 65°C or of 65°C in a solution comprising about 1 M salt or 1 M salt, preferably 6 x SSC or any other solution having a comparable ionic strength, and washing at 65°C in a solution comprising about 0.1 M salt, or 0.1 M salt or less, preferably 0.2 x SSC or any other solution having a comparable ionic strength. Preferably, the hybridization is performed overnight, i.e. at least for 10 hours and preferably washing is performed for at least one hour with at least two changes of the washing solution. These conditions will usually allow the specific hybridization of sequences having about 90% or more sequence identity or at least 90% sequence identity. Moderate hybridization conditions are herein defined as conditions that allow a nucleic acid sequence of at least 50, preferably 150 or more nucleotides, to hybridize at a temperature of about 45°C or of 45°C in a solution comprising about 1 M salt or 1 M salt, preferably 6 x SSC or any other solution having a comparable ionic strength, and washing at room temperature in a solution comprising about 1 M salt, or 1 M salt preferably 6 x SSC or any other solution having a comparable ionic strength. Preferably, the hybridization is performed overnight, i.e. at least for 10 hours, and preferably washing is performed for at least one hour with at least two changes of the washing solution. These conditions will usually allow the specific hybridization of sequences having up to 50% sequence identity. The person skilled in the art will be able to modify these hybridization conditions in order to specifically identify sequences varying in identity between 50% and 90%.

[0080] "Identity", particularly sequence identity, can be readily calculated by known methods, including but not limited to those described in Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heine, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991 ; and Carillo, H., and Lipman, D., SIAM J. Applied Math., 48:1073 (1988).

[0081] Preferred methods to determine identity are designed to give the largest match between the sequences tested. Methods to determine identity and similarity are codified in publicly available computer programs. Preferred computer program methods to determine identity and similarity between two sequences include e.g. the GCG program package (Devereux, J., et al., Nucleic Acids Research 12 (1): 387 (1984)), BestFit, BLASTP, BLASTN, and FASTA (Altschul, S. F. et al., J. Mol. Biol. 215:403-410 (1990). The BLAST X program is publicly available from NCBI and other sources (BLAST Manual, Altschul, S„ et al., NCBI NLM NIH Bethesda, MD 20894; Altschul, S„ et al., J. Mol. Biol. 215:403-410 (1990). The well-known Smith Waterman algorithm may also be used to determine identity.

[0082] Preferred parameters for nucleic acid comparison include the following: Algorithm: Needleman and Wunsch, J. Mol. Biol. 48:443-453 (1970); Comparison matrix: matches=+10, mismatch=0; Gap Penalty: 50; Gap Length Penalty: 3. Available as the Gap program from Genetics Computer Group, located in Madison, Wis. Given above are the default parameters for nucleic acid comparisons. Preferred program and parameter for assessing identity for nucleic acid comparison is calculated using EMBOSS Needle Nucleotide Alignment algorithm with the following parameters: DNAfull matrix with the following gap penalties: open = 10; extend = 0.5.

[0083] The term “derived from” in the context of being derived from a particular naturally occurring gene or sequence is defined herein as being chemically synthesized according to a naturally occuring gene or sequence and / or isolated and / or purified from a naturally occurring gene or sequence. Techniques for chemical synthesis, isolation and / or purification of nucleic acid molecules are well known in the art. In general, a derived sequence is a partial sequence of the naturally occurring gene or sequence or a fraction of the naturally occurring gene or sequence. Optionally, the derived sequence comprises nucleic acid substitutions or mutations, preferably resulting in a sequence being at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical over its whole length to the naturally occurring gene partial gene or sequence or partial sequence.

[0084] “Polypeptide” as used herein refers to any peptide, oligopeptide, polypeptide, gene product, expression product, or protein. A polypeptide is comprised of consecutive amino acids. The term "polypeptide" encompasses naturally occurring or synthetic molecules. A polypeptide is represented by an amino acid sequence. A polynucleotide is represented by a nucleotide sequence. A polypeptide is represented by an amino acid sequence.

[0085] The term "homologous" when used to indicate the relation between a given nucleic acid or polypeptide molecule and a given host organism or host cell, is understood to mean that in nature the nucleic acid or polypeptide molecule is produced by a host cell or organisms of the same species, preferably of the same variety or strain. If homologous to a host cell, a nucleic acid sequence of interest, preferably encoding a polypeptide will typically be operably linked to another promoter sequence or, if applicable, another secretory signal sequence and / or terminator sequence than in its natural environment.

[0086] When used to indicate the relatedness of two nucleic acid sequences the term "homologous" means that one single-stranded nucleic acid sequence may hybridize to a complementary single-stranded nucleic acid sequence. The degree of hybridization may depend on a number of factors including the extent of identity between the sequences and the hybridization conditions such as temperature and salt concentration as discussed later. Preferably, the region of identity is greater than 5 bp, more preferably the region of identity is greater than 10 bp.

[0087] The term "heterologous" when used to indicate the relation between a given (recombinant) nucleic acid or polypeptide molecule and a given host organism or host cell, is understood to mean a nucleic acid or polypeptide molecule from a foreign cell which does not occur naturally as part of the organism, cell, genome or DNA or RNA sequence in which it is present, or which is found in a cell or location or locations in the genome or DNA or RNA sequence that differ from that in which it is found in nature. Heterologous nucleic acids or proteins are not endogenous to the cell into which they are introduced, but have been obtained from another cell or synthetically or recombinantly produced.

[0088] When used to indicate the relatedness of two nucleic acid sequences, the term the term "heterologous sequence" or "heterologous nucleic acid" is one that is not naturally found operably linked as neighboring sequence of the other sequence. As used herein, the term “heterologous” may mean “recombinant”. "Recombinant" refers to a genetic entity distinct from that generally found in nature. As applied to a nucleotide sequence or nucleic acid molecule, this means that said nucleotide sequence or nucleic acid molecule is the product of various combinations of cloning, restriction and / or ligation steps, and other procedures that result in the production of a construct that is distinct from a sequence or molecule found in nature. In preferred embodiments herein, the nucleotide sequence of interest encodes a heterologous sequence.

[0089] “Operably linked” is defined herein as a configuration in which a control sequence or regulating sequence is appropriately placed at a position relative to the nucleotide sequence of interest, preferably coding for the polypeptide of interest such that the control or regulating sequence directs or affects the transcription and / or production or expression of the nucleotide sequence of interest, preferably encoding a peptide or polypeptide in a cell and / or in a subject. For instance, a promoter is operably linked to a coding sequence if the promoter is able to initiate or regulate the transcription or expression of a coding sequence, in which case the coding sequence should be understood as being “under the control of’ the promoter. When one or more nucleotide sequences and / or elements comprised within a construct are defined herein to be “configured to be operably linked to an optional nucleotide sequence of interest”, said nucleotide sequences and / or elements are understood to be configured within said construct in such a way that these nucleotide sequences and / or elements are all operably linked to said nucleotide sequence of interest once said nucleotide sequence of interest is present in said construct.

[0090] “Expression” will be understood to include any step involved in the production of the peptide or polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification and secretion.

[0091] Optionally, a promoter represented by a nucleotide sequence present in a nucleic acid construct is operably linked to another nucleotide sequence encoding a peptide or polypeptide or pre-mRNA or mRNA as identified herein. An expression vector may be any vector which can be conveniently subjected to recombinant DNA procedures and can bring about the expression of a nucleotide sequence encoding a polypeptide of the invention in a cell and / or in a subject.

[0092] As used herein, the "5'-UTR" is the sequence starting with nucleotide 1 of the mRNA and ending with nucleotide -1 of the start codon. It is possible that a regulating part of the promoter is comprised within the nucleotide sequence becoming a 5'-UTR; however, in such case, the 5'-UTR is still not part of the promoter as herein defined.

[0093] The term "control sequences" or “regulatory elements” is defined herein to include all components, which are necessary or advantageous for the expression of a polynucleotide or a polypeptide. Each control sequence may be native or foreign to the nucleic acid sequence harboring or encoding the polynucleotide or the polypeptide. Such control sequences include, but are not limited to, a leader, optimal translation initiation sequences (as described in Kozak, 1991 , J. Biol. Chem. 266:19867-19870), a polyadenylation sequence, a pro-peptide sequence, a pre-pro-peptide sequence, a promoter, a signal sequence, and a transcription terminator. At a minimum, the control sequences include a promoter, and transcriptional and translational stop signals.

[0094] The control sequences may be provided with linkers for the purpose of introducing specific restriction sites facilitating ligation of the control sequences with the coding region of the nucleic acid sequence encoding a polypeptide. The control sequence may be an appropriate promoter sequence, a nucleic acid sequence, which is recognized by a host cell for expression of the nucleic acid sequence. The promoter sequence contains transcriptional control sequences, which mediate the expression of the polypeptide. The promoter may be any nucleic acid sequence, which shows transcriptional activity in the cell including mutant, truncated, and hybrid promoters, and may be obtained from genes encoding extracellular or intracellular polypeptides either homologous or heterologous to the cell.

[0095] The control sequence may also be a suitable transcription terminator sequence, a sequence recognized by a host cell to terminate transcription. The terminator sequence is operably linked to the 3' terminus of the nucleic acid sequence of interest, preferably encoding a polypeptide of interest. Any terminator, which is functional in the cell, may be used in the present invention.

[0096] The control sequence may also be a suitable leader sequence, a non-translated region of a mRNA which is important for translation by the host cell. The leader sequence is operably linked to the 5' terminus of the nucleic acid sequence of interest, preferably encoding a polypeptide of interest. Any leader sequence, which is functional in the cell, may be used in the present invention.

[0097] The control sequence may also be a polyadenylation sequence, a sequence which is operably linked to the 3' terminus of the nucleic acid sequence and which, when transcribed, is recognized by the host cell as a signal to add adenine residues to transcribed mRNA. Any polyadenylation sequence, which is functional in the cell, may be used in the present invention.

[0098] In this document and in its claims, the verb "to comprise" and its conjugations is used in its non-limiting sense to mean that items following the word are included, but items not specifically mentioned are not excluded. In addition the verb “to consist” may be replaced by “to consist essentially of’ meaning that a product or a composition or a nucleic acid molecule or a peptide or polypeptide of a nucleic acid construct or vector or cell as defined herein may comprise additional component(s) than the ones specifically identified; said additional component(s) not altering the unique characteristic of the invention. In addition, reference to an element by the indefinite article "a" or "an" does not exclude the possibility that more than one of the elements is present, unless the context clearly requires that there be one and only one of the elements. The indefinite article "a" or "an" thus usually means "at least one". All patent and literature references cited in the present specification are hereby incorporated by reference in their entirety. Description of drawings

[0099] Fig. 1A - Schematic map of constructs. The intronic promoter construct from WO2015 / 102487 comprises 2 promoters that simultaneously drive expression of the same gene. Transcription from the 2ndPromoter results in an mRNA encoding the gene (see Transcript 1). Transcription from the 1stPromoter (often a UBC promoter) results in a primary transcript including the intron, as bordered by 5’ and 3’-splice sites, that contains the complete 2ndPromoter sequence (see Primary Transcript 2). After intron splicing, the primary transcript results in an mRNA without that intron (see mRNA of Transcript 2), and thus after splicing both mRNAs are highly similar. In WO2015 / 102487 various regions of the construct are named as indicated (TEE or EEE1).

[0100] Fig. 1 B - Constructs of the invention comprise a 1stpromoter wherein the TATA box has been inactivated. Only the 2ndpromoter (which has not been inactivated) produces a transcript.

[0101] Fig. 1 C - Application of the invention to the technology of Fig. 1A, resulting in an inactivated 1stpromoter followed by an intronic 2ndpromoter. Only the 2ndpromoter (which has not been inactivated) produces a transcript.

[0102] Fig. 2 - Comparison of SeAP activity in the exhaust media of CHO pools stably transfected with a construct of the invention (see Fig. 1 B), as compared to an analogue differing only by the fact that the TATA box had not been inactivated. Thus, results are without (left) and with (right) an altered TATA box. Bars represent average activities of 4 pools derived from 2 independent transfections, measured using the SEAP Reporter Gene Assay Kit (Abeam). Pools were grown in 125 ml shakeflasks in 30 ml CD FortiCHO selection medium.

[0103] Fig. 3 - Comparison of HuMab production level by a construct of the invention or an analogue without inactivated TATA box in stably transfected CHO pools. Thus, results are without (left) and with (right) altered TATA box. Bars represent average activities of 4 pools derived from 2 independent transfections.

[0104] Examples

[0105] Example 1 - Inactivation of the TATA box in a 1stpromoter increases transcription

[0106] In the promoter of the Chinese hamster (Cricetulus griseus) ubiquitin-C (UBC) gene the TATA box was inactivated. Inactivation was by point mutation of the TATA box as shown below, comparing fragments of SEQ ID NOs: 105 and 106 (the fragments are SEQ ID NOs: 113 and 112, respectively).

[0107] These promoters were used in a construct as shown in Fig. 1 B, thus featuring a first promoter, followed by a second promoter (CMV, SEQ ID NO: 57), followed by a gene of interest. For the wildtype promoter there was adequate transcription. Surprisingly, when the inactivated TATA box was used there was about twice as much transcription of the gene of interest. Example 2 -Increased SeAP expression

[0108] In this example the invention is compared to known technology for enhanced expression, particularly that ofWO2015 / 102487. There, the expression enhancing element represented by SEQ ID NO: 1 based on the UBC gene is used. The construct comprises the predicted promoter sequence and part of the 5’-untranslated region (see Fig. 1A) and is referred to as EEE1. In the embodiment of the invention demonstrated here, the same nucleic acid sequence was used except the TATA box had been inactivated in the UBC promoter (as schematically depicted in Fig. 1 C). Inactivation was as in Example 1 .

[0109] Expression plasmids were based on the pcDNA3.1 expression vector (SEQ ID NO:71). of which the f1 -ori was removed. In the control vector, the EEE1 sequence (SEQ ID NO:1) was inserted upstream of a CMV promoter (SEQ ID NO:57). The coding sequence for secreted Alkaline Phosphatase (SeAP; SEQ ID NO: 72) was inserted downstream of the CMV promoter (SEQ ID NO:57). The 5’UTR of the vector was replaced by SEQ ID NO:19. The EEE1 was modified to inactivate the TATA box of the UBC promoter by replacing the EcoRI-Bgll I part of EEE1 , harboring the TATA sequence, with sequence M59-S11 (SEQ ID: 92). Plasmid fragments of bacterial backbone-free EEE1 and its variant with the inactivated TATA box (SEQ ID NO: 91) were gel- purified after Sbfl digestion.

[0110] CHO-S cells (Life Technologies) were maintained per manufacturer’s instructions. Duplicate transfections were performed using 3E7 cells, 50 pg of linearized DNA and Freestyle MAX Reagent (Life Technologies). Post-transfection pools were selected in CD FortiCHO medium supplemented with 8 mM glutamine and 800 pg / ml G418. Selected pools were seeded in 30 ml of the same medium at a density of 3E5 cells / ml in 125 ml shake-flasks. SeAP activity was measured in the exhaust medium using the SEAP Reporter Gene Assay Kit, Abeam.

[0111] The inactivated TATA box pools produced 30% higher activity as compared to the wildtype TATA box pools (Fig. 2). This shows that inactivation of the TATA box surprisingly enhances the expression of the SeAP enzyme even beyond the improvement that is provided by the use of EEE1 in general (see WQ2015 / 102487). Results are also shown in Table S1.

[0112] Table S1 : Titers of SeAP

[0113] SeAP enzyme titer (mU / ml)

[0114] EEE1 (wt TATA) Production day 6 307 ± 30

[0115] EEE1 (wt TATA) Production day 7 285 ± 55

[0116] EEE1 (inactivated TATA) Production day 6 391 ± 50

[0117] EEE1 (inactivated TATA) Production day 7 397 ± 96

[0118] Example 3 - Increased antibody expression using multiple plasmids

[0119] Starting from the expression plasmids that were generated in Example 2, the sequence coding for the SeAP enzyme in the previous example was replaced by either the heavy or light chain of a biosimilar human monoclonal antibody (HuMab derived from DrugBank Accession Number DB00072). For the CMV promoter SEQ ID NO:111 was used. Variants of EEE1 with a 800 bp extension from the genomic C. griseus UBC sequence at the 5’-end (referred to in WO2015 / 102487 as EEE1-Xt, SEQ ID NO:59) were made by exchanging the Sfil-Bglll EEE1 element for the Sfil- Bglll of EEE1-Xt. Similarly, heavy chain and light chain expression plasmids with the same EEE1- Xt having an inactivated TATA sequence (as for Example 1) were cloned, resulting in SEQ ID NO: 93.

[0120] The constructs were used to generate CHO-S pools as described in Example 2. The titers at production day 6 and 7 were determined using AlphaLISA (AlphaLISA Human IgG Detection Kit AL205 I Perkin Elmer Enspire 2300). The data indicate that inactivating the TATA box surprisingly increases the titer for these complex polypeptides by about 25%, further improving the effect that could be expected for EEE1-Xt without the modification of the invention. Results are also shown in Table S2.

[0121] Table S2: Titers of HuMab DB00072

[0122] Antibody titer (pg / ml)

[0123] EEE1-Xt (wt TATA) Production day 6 382.0 ± 37.7

[0124] EEE1-Xt (wt TATA) Production day 7 455.8 ± 48.0

[0125] EEE1-Xt (inactivated TATA) Production day 6 495.9 ± 66,7

[0126] EEE1-Xt (inactivated TATA) Production day 7 552.6 ± 43.8

Claims

Claims1 . A method for producing a transcript of a nucleotide sequence of interest, the method comprising the steps of: a) providing a nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter, a second promoter, and the nucleotide sequence of interest, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated; and, b) contacting a cell with the nucleic acid construct to obtain a transformed cell; and, c) allowing the transformed cell to produce the transcript of the nucleotide sequence of interest.

2. The method according to claim 1 , wherein the method is for producing a purified transcript, wherein the method further comprises the step of: d) purifying the produced transcript.

3. A method for expressing a polypeptide of interest comprising step a) and step b) of claim 1 , wherein the nucleotide sequence of interest encodes the polypeptide of interest; and wherein the method further comprises the step of: d) allowing the transformed cell to express the protein or polypeptide of interest.

4. The method according to any one of claims 1-3, wherein the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, a donor splice site, the second promoter, and the nucleotide sequence of interest.

5. The method according to claim 4, wherein the nucleic acid construct comprises in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, optionally a further donor splice site, an acceptor splice site, and the nucleotide sequence of interest.

6. The method according to any one of claims 1-5, wherein the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein: i) the TATA box consensus sequence has been deleted; ii) part of the TATA box consensus sequence has been deleted; iii) at least 1 , 2, 3, 4, 5, 6, or 7 of the nucleotides in the TATA box consensus sequence have been mutated; iv) at most 1 , 2, 3, or 4 of the nucleotides in the TATA box consensus sequence have not been mutated.

7. The method according to any one of claims 1-6, wherein the TATA box promoter wherein the TATA box is inactivated is a TATA box promoter wherein the consensus sequence TATAWAW has been mutated to VBNBNBS, preferably to CTTGATG.

8. The method according to any one of claims 1-7, wherein the first promoter and the second promoter independently are promoters that are active in mammalian cells, preferably selected from human or murine cytomegalovirus (CMV) promoter, simian virus (SV40) promoter, human, hamster, or mouse ubiquitin C (UBC) promoter, human or mouse or rat elongation factor alpha (EF1-a) promoter, apolipoprotein A1 gene promoter (apoA1), and mouse or hamster beta-actin promoter; or wherein the first promoter and the second promoter independently are promoters that are active in yeast and fungal cells, preferably selected from Leu2 promoter, galactose (Gall or Ga17) promoter, alcohol dehydrogenase I (ADH1) promoter, glucoamylase (Gia) promoter, triose phosphate isomerase (TPI) promoter, translational elongation factor EF-I alpha (TEF2) promoter, glyceraldehyde-3-phosphate dehydrogenase (gpdA) promoter, alcohol oxidase (AOX1) promoter, and glutamate dehydrogenase (gdhA) promoter.

9. The method according to any one of claims 1-8, wherein the first promoter has at least 80% sequence identity with SEQ ID NO: 89, 90, 91 , 93, 95, 105, 107, or 110, preferably 91 , 93, 105, 107, or 110, more preferably 105 or 107 or 110, even more preferably 107 or 110, most preferably 107.

10. The method according to any one of claims 1-9, wherein the first promoter and the second promoter are both operably linked to the nucleotide sequence of interest.11 . The method according to any one of claims 1-10, wherein the first promoter is separated from the second promoter by about 0 to about 5000 bp, and / or wherein the second promoter is separated from the nucleotide sequence of interest by about 0 to about 2000 bp.

12. The method according to any one of claims 1-10, wherein the first promoter differs only from a wildtype TATA box promoter in that the TATA box has been inactivated.

13. A nucleic acid construct comprising in the 5’ to the 3’ direction a first promoter and a second promoter, wherein the first promoter is a TATA box promoter wherein the TATA box is inactivated.

14. The nucleic acid construct according to claim 13, comprising in the 5’ to the 3’ direction the first promoter, a donor splice site, and the second promoter, preferably comprising in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, and an acceptor splice site; more preferably comprising in the 5’ to the 3’ direction the first promoter, the donor splice site, the second promoter, a further donor splice site, and the acceptor splice site.

15. The nucleic acid construct according to claim 13 or 14, having at least 80% sequence identity with SEQ ID NO: 109.

16. The nucleic acid construct according to any one of claims 13-15, wherein the first promoter is separated from the second promoter by about 0 to about 5000 bp, and / or wherein the second promoter is separated from the nucleotide sequence of interest by about 0 to about 2000 bp.

17. The nucleic acid construct according to any one of claims 13-16, wherein the first promoter differs only from a wildtype TATA box promoter in that the TATA box has been inactivated.

18. An expression vector comprising a nucleic acid construct as defined in any one of claims 13- 17, further comprising a nucleotide sequence of interest.

19. The expression vector according to claim 18, wherein the first promoter and the second promoter are both operably linked to the nucleotide sequence of interest.

20. A cell comprising a nucleic acid construct as defined in any one of claims 13-16, or an expression vector as defined in claim 18 or 19.21 . Use of a TATA box promoter wherein the TATA box is inactivated for increasing transcription of a nucleotide sequence of interest when the nucleotide sequence of interest is operably linked to a second promoter that does not have an inactivated TATA box.

22. The use of claim 21 , wherein the TATA box promoter wherein the TATA box is inactivated differs only from a wildtype TATA box promoter in that the TATA box has been inactivated.

Citation Information

Patent Citations

  • Construct and sequence for enhanced gene expression

    WO2015102487A1

  • Self-immolative plasmid backbone

    WO2020174079A1

  • Improved rep genes for recombinant AAV production

    WO2024121778A1